Cross-device scheduling method, device, equipment, medium and product
By introducing the target execution device as a decision-making intermediary in cross-device collaboration scenarios, and based on user intent and basic device information, one-time intent parsing and precise device scheduling are achieved, solving the problems of cloud computing resource waste and device collaboration flexibility, and improving response efficiency and user experience.
Patent Information
- Application Number
- CN202511785351.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies waste cloud computing resources due to multiple redundant simulation processes in cross-device collaboration scenarios, and cannot flexibly handle device collaboration needs in complex scenarios, thus limiting the scope of application and scalability.
By introducing the target execution device as a decision-making intermediary, the second control device is intelligently selected based on user intent and device basic information, and structured control instructions are generated to achieve one-time intent parsing and precise device scheduling.
It reduces cloud computing load, improves response efficiency and device resource utilization, supports multi-device collaborative tasks in complex scenarios, parallel control and device status matching, and improves control success rate and user experience.
Smart Images

Figure CN121509520A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and specifically to a method, apparatus, device, medium, and product for cross-device scheduling. Background Technology
[0002] In related technologies, in cross-device collaboration scenarios, after receiving a user request from the wake-up device (the first control device), the cloud typically simulates the identities of multiple devices in parallel to process the user request.
[0003] For example, one process is performed as the wake-up device itself, while another or more processes are performed as other potential controlled devices. Then, based on the simulation results of each process, a target execution device is selected, and the corresponding instructions are returned to the wake-up device so that the wake-up device can control the target execution device to execute the user request.
[0004] However, the need to perform redundant simulations on the same user request leads to a serious waste of cloud computing resources. Summary of the Invention
[0005] In view of this, embodiments of this application provide a method, apparatus, device, medium, and product for cross-device scheduling.
[0006] On one hand, this application provides a method for cross-device scheduling, including: In response to receiving a user request from the first control device, the user request is subjected to intent recognition to obtain the recognized user intent; the user request is generated by the first control device based on user instructions issued by the user. Based on the user's intent and the device basic information of at least one user-associated device, a second control device is determined; the device basic information is used to evaluate the compatibility between the user-associated device and the user's intent; the second control device is different from the first control device and belongs to the type of central device that can be configured to drive at least one other user-associated device to perform operations. Based on the user's intent, control commands are sent to the second control device.
[0007] In this way, the user request is parsed once to generate a structured user intent; then, based on the device's basic information, a second control device different from the first control device is intelligently selected from at least one user-associated device; finally, control commands are issued to the device. This solves the problem of wasted cloud computing resources caused by multi-path parallel intent understanding in related technologies. Through architectural optimization, the computing load is reduced and the response efficiency is improved, providing a basic support for building an efficient cross-device collaborative system.
[0008] In one implementation, determining a second control device based on user intent and device basic information of at least one user-associated device includes: Based on the user intent and the device basic information of at least one user-associated device, determine the target execution device among at least one user-associated device; If the target execution device is a central device, then the target execution device is identified as the second control device; If the target execution device is not a central device type, then a second control device for controlling the target execution device is obtained from at least one user-associated device.
[0009] In this way, by introducing the target execution device as a decision-making intermediary, the device scheduling logic is refined. First, the target execution device for the final operation is determined. Then, the second control device is dynamically determined based on its device type. If the target device itself is a control device, it is directly reused; otherwise, a suitable control device is assigned to it. This two-level decision-making mechanism ensures the accuracy of device control and improves the utilization rate of device resources, enabling flexible adaptation to collaborative scenarios with different device types.
[0010] In one implementation, based on user intent, a control command is issued to a second control device, including: Based on the user intent, as well as the device identifier and device status information of the target execution device, control commands are generated; the device status information includes the device runtime status related to the execution of the user intent. Send control commands to the second control device.
[0011] In this way, by incorporating the real-time runtime status of the target execution device, precise generation of control commands is achieved. When generating commands, not only user intent and target device identifier are considered, but also dynamic contextual information such as the device's current operating status is combined to generate precise commands that match the actual state of the device. This represents an upgrade from standardized commands to contextualized commands, effectively avoiding execution failures or result deviations caused by mismatched device states, and significantly improving the success rate of cross-device control and user experience.
[0012] In one implementation, intent recognition of a user request includes: Input the user's request into the intent recognition model, and obtain the user's intent output by the intent recognition model; The user intent includes the target action to be performed and the device conditions on which the target action depends for its execution.
[0013] In this way, by using structured parsing methods, natural language requests are transformed into explicit semantic units containing target actions and device conditions. This structured output provides accurate criteria for subsequent device screening, enabling the direct conversion of explicit requirements and implicit constraints in user instructions into executable device screening standards, ensuring the accuracy and executability of intent understanding from the source.
[0014] In one implementation, the method further includes, before performing intent recognition on the user request: Based on the device location information of the first control device, the user request is updated to obtain a new user request.
[0015] In this way, by introducing a location context-aware mechanism, the user request is semantically enhanced before intent recognition. By utilizing the real-time location information of the first control device, the user request is supplemented with context and semantically disambiguated, enabling subsequent intent recognition to make more accurate judgments based on richer environmental information, effectively improving the accuracy of intent understanding and scene adaptability in complex scenarios.
[0016] In one embodiment, the first control device is a first speaker, and at least one user-associated device includes a second speaker and a camera device, wherein the second speaker is associated with the camera device; Based on user intent and device basic information of at least one user-associated device, a second control device is determined, including: Based on the user's intent and the device basic information of at least one user-associated device, the camera device is identified as the target execution device. In response to the fact that the camera device is not a central device type, the second speaker associated with the camera device is identified as a second control device.
[0017] In one implementation, determining a second control device based on user intent and device basic information of at least one user-associated device includes: If the user's intent contains at least two target actions, then for each target action, a second control device is determined to perform that target action; Send control commands to the second control device, including: If there are multiple second control devices, then the corresponding control command is issued to each of the second control devices.
[0018] This enables single-instruction, multi-device parallel control capabilities. When a user instruction contains multiple target actions, the task can be automatically decomposed, each action can be independently assigned a suitable control device, and corresponding control instructions can be issued in parallel. This mechanism allows cross-device collaboration in complex scenarios to be completed through a single interaction, maintaining the simplicity of user operation while significantly improving the execution efficiency of multi-device collaborative tasks.
[0019] In one implementation, based on user intent and device basic information of at least one user-associated device, a target execution device is determined from at least one user-associated device, including: Based on the device conditions contained in the user intent and the basic device information of each user's associated devices, at least one candidate device that meets the device conditions and can execute the user intent is selected from each user's associated devices. Determine the target execution device based on at least one candidate device.
[0020] In this way, a two-stage screening mechanism is used to accurately determine the target execution equipment. First, an initial screening is conducted based on equipment conditions and basic equipment information to obtain a set of candidate equipment that meets the conditions. Then, the final execution equipment is determined from the candidate equipment through a subsequent decision-making process. This step-by-step screening method not only ensures the comprehensiveness of equipment screening but also provides an operational basis for subsequent optimization selection (such as status checks, user confirmation, etc.), ensuring the accuracy and flexibility of equipment decision-making.
[0021] In one embodiment, determining a target execution device based on at least one candidate device includes: Perform anomaly detection on at least one candidate device; Candidate devices whose test results indicate that the device is functioning normally are selected from at least one candidate device; If the selected candidate devices are empty, control commands are not allowed to be issued, and an abnormal prompt message is sent to the first control device. If at least one candidate device is selected, the target execution device is determined based on the selected candidate devices.
[0022] This abnormal state pre-detection mechanism improves reliability. Before determining the final execution device, all candidate devices undergo availability checks, automatically filtering out offline, faulty, or other abnormal devices. This approach effectively avoids execution failures caused by issuing commands to unavailable devices, and by proactively providing abnormal information to guide user intervention, it significantly improves the system's success rate and user experience.
[0023] In one embodiment, determining a target execution device includes: When there is only one candidate device, the candidate device is determined as the target execution device; When there are multiple candidate devices, a device query request is sent to the first control device, and a first user response is received from the first control device. Based on the first user response, the target execution device selected by the user is determined. The device query request contains device information of each candidate device.
[0024] In this way, intelligent device selection is achieved through a human-machine collaborative decision-making mechanism. It automatically confirms the device when there is only one candidate, and proactively prompts the user when multiple candidates appear, obtaining a clear choice through interactive dialogue. This approach effectively solves the problem of decision ambiguity in multi-device scenarios while ensuring automation efficiency. It can fully utilize contextual information for automatic filtering, and also rely on user judgment for precise targeting when necessary, thus balancing execution efficiency and decision accuracy.
[0025] In one embodiment, the equipment conditions include the target equipment type and / or the target equipment location; Basic equipment information includes equipment attribute information, equipment functions, and equipment connection status, which correspond to the equipment conditions.
[0026] In this way, a multi-dimensional device matching mechanism enables precise device selection. By comparing the device conditions specified in the user's intent with the basic device information from multiple dimensions, it ensures that the selected devices not only meet the user's explicit needs but also possess execution capabilities and usability. This matching mechanism represents an upgrade from simple device discovery to intelligent device selection, effectively improving the accuracy and scenario adaptability of device selection.
[0027] In one embodiment, before issuing control commands to the second control device based on user intent, the method further includes: When a confirmation event is detected, an execution authorization request is sent to the first control device, so that the first control device outputs the execution authorization request to the user; The system receives a second user response from the first control device; the second user response is used to indicate whether the user agrees to or refuses to execute the command; the second user response is generated by the first control device based on the user's input response command. Based on the second user's response, determine whether to allow the issuance of control commands to the second control device; When the second user replies indicating agreement to execution, it is determined that the issuance of control commands is permitted; when the second user replies indicating refusal to execution, it is determined that the issuance of control commands is not permitted.
[0028] In this way, the critical operation confirmation mechanism enhances the system's security and controllability. In specific scenarios (such as initial cross-device control or high-risk operations), the system pauses the execution process, proactively requests authorization from the user, and decides whether to continue based on the user's explicit feedback. This approach maintains automation efficiency while providing users with necessary decision-making intervention points, effectively preventing misoperation or unintended execution, and enhancing users' trust and sense of security in the intelligent system.
[0029] In one embodiment, the method further includes: If an abnormal event is detected in the control device, control commands are sent to the target device based on the user's intent via the push server.
[0030] In one implementation, the abnormal event of the control device includes at least one of the following: The communication connection between the second control device and the target execution device was detected to be disconnected. The network status of the second control device was detected as offline. The response time after issuing a command to the second control device exceeds a preset threshold. An anomaly was detected in the second control device.
[0031] In this way, the abnormal degradation mechanism ensures the high availability of the system. When an abnormal condition such as connection loss, offline status, or response timeout is detected in the second control device, the system automatically switches to the cloud push channel and directly sends commands to the target execution device. This automatic fault switching capability ensures that commands can still be delivered when the main control path fails, greatly improving the robustness and task completion rate of the system and providing users with a continuous and reliable cross-device control experience.
[0032] In one embodiment, after issuing a control command to the second control device based on user intent, the method further includes: The execution result is returned to the first control device so that the first control device can output the execution result to the user.
[0033] On one hand, this application provides a cross-device scheduling apparatus, including: The identification unit is used to respond to a user request sent by the first control device, identify the intent of the user request, and obtain the identified user intent; the user request is generated by the first control device based on the user command issued by the user. The determining unit is used to determine a second control device based on the user's intent and the device basic information of at least one user-associated device; the device basic information is used to evaluate the compatibility between the user-associated device and the user's intent; the second control device is different from the first control device and belongs to the type of central device that can be configured to drive at least one other user-associated device to perform operations. The sending unit is used to send control commands to the second control device based on the user's intent.
[0034] In one implementation, the determining unit is used to: Based on the user intent and the device basic information of at least one user-associated device, determine the target execution device among at least one user-associated device; If the target execution device is a central device, then the target execution device is identified as the second control device; If the target execution device is not a central device type, then a second control device for controlling the target execution device is obtained from at least one user-associated device.
[0035] In one embodiment, the transmitting unit is used to: Based on the user intent, as well as the device identifier and device status information of the target execution device, control commands are generated; the device status information includes the device runtime status related to the execution of the user intent. Send control commands to the second control device.
[0036] In one embodiment, the identification unit is used for: Input the user's request into the intent recognition model, and obtain the user's intent output by the intent recognition model; The user intent includes the target action to be performed and the device conditions on which the target action depends for its execution.
[0037] In one embodiment, the identification unit is further configured to: Based on the device location information of the first control device, the user request is updated to obtain a new user request.
[0038] In one embodiment, the first control device is a first speaker, and at least one user-associated device includes a second speaker and a camera device, wherein the second speaker is associated with the camera device; The identification unit is used for: Based on the user's intent and the device basic information of at least one user-associated device, the camera device is identified as the target execution device. In response to the fact that the camera device is not a central device type, the second speaker associated with the camera device is identified as a second control device.
[0039] In one implementation, the determining unit is used to: If the user's intent contains at least two target actions, then for each target action, a second control device is determined to perform that target action; The transmitting unit is used for: If there are multiple second control devices, then the corresponding control command is issued to each of the second control devices.
[0040] In one implementation, the determining unit is used to: Based on the device conditions contained in the user intent and the basic device information of each user's associated devices, at least one candidate device that meets the device conditions and can execute the user intent is selected from each user's associated devices. Determine the target execution device based on at least one candidate device.
[0041] In one implementation, the determining unit is used to: Perform anomaly detection on at least one candidate device; Candidate devices whose test results indicate that the device is functioning normally are selected from at least one candidate device; If the selected candidate devices are empty, control commands are not allowed to be issued, and an abnormal prompt message is sent to the first control device. If at least one candidate device is selected, the target execution device is determined based on the selected candidate devices.
[0042] In one implementation, the determining unit is used to: When there is only one candidate device, the candidate device is determined as the target execution device; When there are multiple candidate devices, a device query request is sent to the first control device, and a first user response is received from the first control device. Based on the first user response, the target execution device selected by the user is determined. The device query request contains device information of each candidate device.
[0043] In one embodiment, the equipment conditions include the target equipment type and / or the target equipment location; Basic equipment information includes equipment attribute information, equipment functions, and equipment connection status, which correspond to the equipment conditions.
[0044] In one embodiment, the transmitting unit is further configured to: When a confirmation event is detected, an execution authorization request is sent to the first control device, so that the first control device outputs the execution authorization request to the user; The system receives a second user response from the first control device; the second user response is used to indicate whether the user agrees to or refuses to execute the command; the second user response is generated by the first control device based on the user's input response command. Based on the second user's response, determine whether to allow the issuance of control commands to the second control device; When the second user replies indicating agreement to execution, it is determined that the issuance of control commands is permitted; when the second user replies indicating refusal to execution, it is determined that the issuance of control commands is not permitted.
[0045] In one embodiment, the transmitting unit is further configured to: If an abnormal event is detected in the control device, control commands are sent to the target device based on the user's intent via the push server.
[0046] In one implementation, the abnormal event of the control device includes at least one of the following: The communication connection between the second control device and the target execution device was detected to be disconnected. The network status of the second control device was detected as offline. The response time after issuing a command to the second control device exceeds a preset threshold. An anomaly was detected in the second control device.
[0047] In one embodiment, the transmitting unit is further configured to: The execution result is returned to the first control device so that the first control device can output the execution result to the user.
[0048] On one hand, this application provides an electronic device, including: Processor; and The memory stores computer instructions that cause the processor to perform steps of the methods provided in the various alternative implementations of any of the cross-device scheduling methods described above.
[0049] On one hand, embodiments of this application provide a computer-readable storage medium storing computer instructions for causing a computer to perform steps of the methods provided in various alternative implementations of any of the above-described cross-device scheduling methods.
[0050] On one hand, this application provides a computer program product including computer-readable code or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in the processor of an electronic device, the processor in the electronic device performs the steps of the method provided in any of the various optional implementations of cross-device scheduling described above. Attached Figure Description
[0051] Figure 1 This is a flowchart of a cross-device scheduling method according to an embodiment of this application.
[0052] Figure 2 This is an example architecture diagram of a cross-device scheduling system in an embodiment of this application.
[0053] Figure 3 This is a flowchart of a device decision-making method according to an embodiment of this application.
[0054] Figure 4 This is a flowchart of a device screening method according to an embodiment of this application.
[0055] Figure 5 This is a structural block diagram of a cross-device scheduling device according to an embodiment of this application.
[0056] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0057] The technical solution of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.
[0058] In today's interconnected smart device ecosystem, users typically own multiple smart devices (such as mobile phones, smart speakers, TVs, and in-vehicle systems). Users expect to be able to interact with a voice assistant through any device (i.e., the "wake-up device" or "first control device") to intelligently control the best device in the ecosystem to perform tasks; this is the core scenario of "cross-device collaboration."
[0059] To achieve this goal, the industry generally adopts a "command transfer" collaborative solution. In this solution, after receiving a user request from the first control device, the cloud simulates the identities of multiple devices in parallel and performs redundant simulation processing on the user request.
[0060] Specifically, one processing flow interprets the intent as the first control device itself, while another or more flows interpret the intent as other potential controlled devices. Subsequently, the arbitration module selects the target execution device based on the simulation results of each flow and routes the corresponding instructions back to the first control device, which then controls the final target execution device.
[0061] However, this existing technical solution has a series of interconnected serious flaws: On the one hand, resource efficiency is low and costs are high: because multiple redundant intent understanding and simulation processing are required for the same user request, the computing load on the cloud server is greatly increased, resulting in a serious waste of cloud computing resources and driving up service operation costs.
[0062] On the other hand, its application scope is limited and its adaptability to different scenarios is poor: this solution requires all controlled devices to report their full context information to the cloud upon request for multi-channel simulation processing. This is unsuitable for devices with limited computing power or those unsuitable for full reporting due to power consumption and privacy considerations (such as smartwatches, wristbands, and TVs), thus limiting the solution's application scope. Furthermore, its execution path is fixed at "the instruction must be sent to the wake-up device," failing to support flexible methods via other control devices, making complex collaborative scenarios such as "screen mirroring" and "audio mirroring" impossible to achieve within this framework.
[0063] Thirdly, the architecture is rigid and lacks scalability: the decision-making logic of this solution (i.e., how to select the target device) is scattered across multiple independent service modules in the cloud, such as pre-processing, arbitration, and post-processing. Each new cross-device control function (such as "glasses controlling a mobile phone") requires adaptation and configuration in these scattered modules, resulting in poor system scalability and high development and maintenance costs.
[0064] Fourthly, the control granularity is coarse and the intelligence is insufficient: the framework basically only supports routing instructions to a single device for execution, and cannot flexibly handle complex scenarios that are implicit in a user instruction and require multiple devices to execute different parts of the function respectively, and cannot meet the growing demand for fine-grained collaboration.
[0065] Based on the deficiencies of the aforementioned related technologies, this application provides a method, apparatus, device, medium, and product for cross-device scheduling, aiming to reduce the resources consumed by cross-device calls.
[0066] This application provides a method for cross-device scheduling. This method can be applied to electronic devices. This application does not limit the type of electronic device, which can be any suitable type of device, such as terminal devices and servers. This application will not elaborate further on this.
[0067] See Figure 1 The diagram shown is a flowchart of a cross-device scheduling method according to an embodiment of this application, applied to a server (e.g., in the cloud). The following is a combination of... Figure 1 The method is described below, and the specific implementation process is as follows: Step 101: In response to receiving a user request sent by the first control device, perform intent recognition on the user request to obtain the recognized user intent.
[0068] The user request is generated by the first control device based on the user instructions issued by the user.
[0069] In one implementation, the following steps may be used when performing intent recognition on a user request: The user request is input into the intent recognition model to obtain the user intent output by the intent recognition model; wherein, the user intent includes the target action to be performed and the device conditions on which the target action depends for execution. Optionally, the device conditions may include the target device type and / or the target device location.
[0070] The intent recognition model is a deep learning-based natural language processing model, such as a language model trained on a large-scale corpus, used to parse unstructured natural language user requests and transform them into structured, machine-understandable, and processable user intents. User intent, as the core input to subsequent processes, is designed to include two key parts: the target action and the device conditions. The target action defines the specific operation to be performed, such as playing a song, setting the air conditioner temperature to 25 degrees Celsius, or turning on the bedroom light. The device conditions define the device attributes on which this action depends or is specified, providing clear constraints for subsequent device selection. Specifically, device conditions can be reflected in the target device type, such as a television, speakers, or lights, and / or the target device location, such as in the living room or bedroom. This design accurately captures the user's explicit or implicit device selection preferences in the instructions.
[0071] For example, when a user tells the living room speaker (the primary control device) to play the news on the bedroom TV, the intent recognition model will parse the target action as playing the news, the device conditions as target device type: TV, and target device location: bedroom. If the user's instruction is more simply to play music, the model might parse the target action as playing music, and the device conditions might only include target device type: speaker. In this case, the device location information might be supplemented by the system based on the location of the primary control device or other context, or used as a default selection.
[0072] In this way, complex and ambiguous natural language commands are transformed into a clear, structured, and complete semantic representation through a one-time, centralized intent parsing. This provides a unified and clear input for all subsequent processing modules, fundamentally avoiding the redundant parsing and wasted computing resources caused by simulating multiple device identities in related technologies, thus laying the foundation for the efficiency and intelligence of the entire solution.
[0073] Furthermore, before performing intent recognition on the user request, the user request can be updated based on the device location information of the first control device to obtain a new user request; the device basic information includes device location information.
[0074] In one implementation, based on the device location information of the first control device, device-related information in the user request is extracted, and the user request is updated based on the device-related information to obtain a new user request.
[0075] The device location information of the first control device provides crucial contextual information for the intent recognition model, significantly improving the accuracy and intelligence of intent parsing. The device location information acts as a context enhancer. The physical or logical location information of the first control device (e.g., "living room," "master bedroom," "kitchen") is used as an important contextual feature and fused with the original user request. This fusion is not a simple string concatenation, but rather feature enhancement at the model input level. Techniques can include embedding location feature vectors into the request data or constructing a structured request object containing location information, thereby forming a new user request with more complete information.
[0076] By incorporating device location, more accurate semantic disambiguation and device association can be achieved. For example, when a user says "turn on the light" to a speaker in the kitchen, the new request, which incorporates the location information of "kitchen," allows the intent recognition model to more accurately understand that the user's true intent is to turn on the "kitchen light," thus outputting the user's intent with "location: kitchen" included in the device condition. Conversely, without this location information, the model might only be able to parse the action of "turning on the light" without being able to associate it with the specific location context.
[0077] The introduction of device location information of the first control device improves the accuracy of intent understanding, making the output of user intent more precise, reducing ambiguity and errors in the subsequent device selection stage, and allowing users to implicitly convey key information simply by waking up the device location without explicitly repeating the device location in the command (i.e., without saying "turn on the kitchen light"), thus achieving more natural and convenient implicit control and optimizing the human-computer interaction experience as a whole.
[0078] As an example, if the first control device is a first speaker, and at least one user-associated device includes a second speaker and a camera device, and the second speaker is associated with the camera device, then the camera device can be determined as the target execution device based on the user's intent and the device basic information of at least one user-associated device; in response to the camera device not being a central device type, the second speaker associated with the camera device is determined as the second control device.
[0079] Step 102: Determine the second control device based on the user's intent and the device basic information of at least one user-associated device.
[0080] Among them, the basic device information is used to assess the compatibility between the user's associated device and the user's intent; the second control device is different from the first control device and belongs to the type of central device that can be configured to drive at least one other user-associated device to perform operations.
[0081] The central device type referred to in this application embodiment is a type of user-associated device designed or configured in a device network or system to possess core coordination and control capabilities. This type of device is fundamentally different from single-function terminal execution devices, such as light bulbs, curtain motors, or switches, whose initial design purpose is merely to receive and execute simple on / off or adjustment commands. The essential characteristic of the central device type lies in its cross-device transaction scheduling capability. It can act as a control node, responding to relatively abstract user intentions from the cloud, parsing them into a series of specific, operable control commands, and issuing them to one or more other associated devices to collaboratively complete complex tasks. This means that the central device type is upstream of the terminal execution devices in the system's logical control hierarchy.
[0082] Specifically, devices of this type typically have embedded device control services or integrated control protocol stacks, such as common smart home gateway protocols or control modules with infrared learning and transmission capabilities. These hardware and software foundations enable them to directly establish communication connections with a specific set of terminal devices and manage their status. Functionally, hub devices not only perform their native functions but, more importantly, act as secondary command distribution centers. For example, a smart speaker defined as a hub device, upon receiving a command to activate cinema mode from the cloud, can simultaneously play background music and, through its built-in local network connection, send commands to the smart light strip in the living room to dim the lights and to the smart curtain motor to close the curtains. Therefore, determining whether a device belongs to the hub device type is not solely based on whether it is controlling other devices, but rather on whether it has been endowed with such an architectural role and inherent capabilities, thus allowing it to be configured to assume the function of such a control hub.
[0083] In one implementation, based on the user's intent and the device basic information of at least one user-associated device, a target execution device is determined among at least one user-associated device; if the device type of the target execution device is a central device type, then the target execution device is determined as a second control device; if the device type of the target execution device is not a central device type, then a second control device for controlling the target execution device is obtained from at least one user-associated device.
[0084] In this context, the control device is a device capable of controlling other devices. For example, a control device could be a speaker equipped with the Xiaoai Software Development Kit (SDK). The first control device can be a central device or not. The target execution device can be a second control device or a user-associated device controlled by the second control device. The aforementioned at least one user-associated device refers to all devices under a user account, including at least one control device, and may also include other devices that each control device has control capabilities over. The second control device is uniformly abstracted as the final bearer and coordinator of instruction execution. Control refers to the ability of the second control device to operate and drive the target execution device after receiving control instructions from the cloud. The realization of this control relationship can rely on a pre-established connection perception layer between devices, including establishing a control channel between devices through a unified connection perception layer that is compatible with local area network protocols and infrared remote control, among other methods. For example, a smart speaker (i.e., the second control device) can directly send playback instructions to a smart TV (i.e., the target execution device) on the same network via Wi-Fi.
[0085] The two relationships described above aim to cover all practical scenarios of cross-device collaboration. The first relationship occurs when the target execution device is a user-associated device controlled by the second control device. For example, a user tells the car's infotainment system (the first control device) to turn on the home air conditioner. After cloud-based decision-making, the instruction is sent to the home gateway or smart speaker (the second control device), which then controls the air conditioner (the target execution device). The second relationship occurs when the target execution device and the second control device are the same device. For example, a user tells the first speaker (the first control device) to play music. After cloud-based decision-making, the second speaker is deemed the better execution device. In this case, the second speaker simultaneously acts as both the second control device and the target execution device, and the control instruction will be directly sent to the second speaker for execution.
[0086] This approach unifies and simplifies complex device relationships, eliminating the need to concern oneself with whether instructions are ultimately executed indirectly or directly by the controlling device. Its core responsibility is to intelligently select the optimal executor (the target device) and the intermediary capable of driving that executor (the second control device), and then issue instructions to the second control device through a unified interface. This significantly enhances the system's flexibility and scalability, enabling efficient handling of both single-device tasks and complex cross-device collaborations within a unified architecture.
[0087] Furthermore, if the user's intent includes at least two target actions, a second control device for performing the target action can be determined for each target action.
[0088] In one implementation, if the user intent includes at least two target actions, then for each target action, a target execution device and its associated second control device are determined to perform the target action.
[0089] This enables complex multi-device collaboration under a single user command. By automatically breaking down a complex command into multiple independent subtasks and intelligently assigning the most suitable execution device to each subtask, the system can control multiple devices to complete different operations in parallel or sequentially. This greatly improves the automation and efficiency of control, allowing users to achieve complex scenario-based linkages without issuing multiple commands, thus providing a truly intelligent and integrated cross-device interactive experience.
[0090] Furthermore, if only the first control device exists and no other control devices exist, then control commands are sent directly to the first control device, and the target execution device is controlled by the first control device to perform the corresponding operation.
[0091] In one implementation, when determining a target execution device among at least one user-associated device, the following steps may be taken: Based on the device conditions contained in the user intent and the basic device information of each user's associated devices, at least one candidate device that meets the device conditions and can execute the user intent is selected from each user's associated devices; based on at least one candidate device, the target execution device is determined. Specifically, if there is only one candidate device, that candidate device is determined as the target execution device; otherwise, the target execution device is selected from the candidate devices.
[0092] The core of this implementation lies in standardizing the process of determining the target execution device into a standardized two-stage screening process. First, based on the device conditions parsed from the user's intent, such as the target device type being a television or the target device location being the living room, and combined with the device attributes, functions, and real-time connection status recorded in the device's basic information, all candidate devices meeting the criteria are initially screened from all user-related devices. This stage ensures that the candidate devices not only meet the user's explicit or implicit requirements but also possess the basic capability and usability to perform the task.
[0093] When there is only one candidate device, the system directly identifies it as the target execution device, which is an efficient optimization path. When there are multiple candidate devices, a second round of screening is triggered, selecting a unique target execution device from multiple candidates based on a more refined strategy (such as device priority, signal strength, user historical preferences, or user selection). After the target execution device is determined, the system then determines the second control device that will ultimately be responsible for driving the target execution device to execute instructions, based on the control relationships between the devices.
[0094] In this way, an intelligent decision-making mechanism with both breadth and depth is constructed. It not only ensures that the final selected equipment is an optimal match between user intent and equipment capabilities, but also takes into account decision-making efficiency and accuracy through a hierarchical screening design. This ensures that the system can reliably and automatically output the optimal equipment scheduling scheme when facing equipment environments of varying complexity, thereby realizing intelligent and automated cross-device collaboration.
[0095] Furthermore, anomaly detection and device screening can also be performed. In one embodiment, anomaly detection is performed on at least one candidate device; candidate devices with normal detection results are selected from the at least one candidate device; if no candidate devices are selected, control commands are not allowed to be issued, and an anomaly prompt message is sent to the first control device; if at least one candidate device is selected, the target execution device is determined based on the selected candidate devices.
[0096] This anomaly detection aims to identify potential problems with the device that could affect the successful execution of commands, such as the device being offline, in a busy state, or having lost connection with a secondary control device. This establishes a preventative execution assurance mechanism. By performing status verification in advance in the cloud, it effectively avoids negative experiences such as "command unresponsive" or "execution failure" caused by sending commands to a device that cannot execute normally. This not only saves network and computing resources consumed by invalid command issuance and failed retries, but also provides immediate feedback to users, guiding them to troubleshoot and intervene, thereby significantly improving the system's robustness, interaction efficiency, and user satisfaction.
[0097] Furthermore, if there are multiple candidate devices, the target execution device selected by the user can be obtained through interaction with the user.
[0098] In one embodiment, a device query request is sent to a first control device, causing the first control device to output a device query request to a user; the device query request includes device information of each candidate device; a first user response is received from the first control device; the first user response is generated by the first control device based on the device selection instruction input by the user; based on the first user response, the target execution device selected by the user is determined.
[0099] In this way, when the target device cannot be determined automatically, a human-machine collaborative decision-making mechanism can be adopted. This involves presenting the user with a clear list of candidate devices (e.g., by asking via voice prompt, "Do you want to control the TV in the living room or the TV in the bedroom?"), thus returning the final choice to the user. The device query request includes necessary device information, such as device name and location, to enable the user to make a clear distinction. The user inputs a device selection command via voice or an application interface. This command is received by the first control device and returned to the cloud as the first user response, based on which the system ultimately determines the target device.
[0100] This approach cleverly solves the decision-making ambiguity problem in scenarios with multiple candidate devices, combining machine computing power with human judgment. This ensures the system can accurately execute user intent even in complex situations, avoiding errors that might arise from automatic selection, while also respecting user preferences. Thus, it guarantees the final accuracy of task execution and user satisfaction amidst uncertainty, improving system reliability and user experience.
[0101] Step 103: Based on the user's intent, send control commands to the second control device.
[0102] In one implementation, a control command is generated based on the user intent, the device identifier of the target execution device, and the device status information; the device status information includes the device runtime status related to the execution of the user intent; and the control command is sent to a second control device.
[0103] Optionally, if the target execution device is a second control device, control commands can be generated based solely on the user intent and the device status information of the target execution device.
[0104] This approach doesn't simply forward the original user intent; instead, it integrates information from three sources to construct a precise and executable control command. This information includes: the user intent defining what to do, the device identifier specifying the target device, and device status information reflecting the current state of the target device. Specifically, device status information refers to the device's runtime state closely related to executing the current intent. For example, for the intent to "play video," the relevant state might be the TV's current input source; for the intent to "adjust the air conditioner," the relevant state is the air conditioner's current operating mode and set temperature. By incorporating device status information, the system can generate more context-aware intelligent commands. For instance, if the user command is "continue playing," after obtaining the target TV's device status information (such as the currently playing episode and its progress), the system can generate a precise playback command containing the specific episode identifier and start time, thus achieving seamless continuation of playback, rather than simply issuing a general "play" command. This significantly improves the accuracy of command execution and the consistency of the user experience. It allows control commands to match the latest real-time state of the device, avoiding execution failures or results that don't meet user expectations due to state asynchrony. This marks the upgrade of the system from simple command transmission to an intelligent agent capable of understanding context and generating optimal control strategies accordingly, thus achieving truly intelligent device control.
[0105] Furthermore, if there are multiple second control devices, corresponding control commands can be issued to each second control device separately.
[0106] Furthermore, if there are multiple target execution devices, corresponding control commands can be issued to the second control device associated with each target execution device.
[0107] This enables parallel and collaborative control of multiple devices via a single command. When a user command contains multiple tasks that can be executed in parallel, the system can automatically split the tasks and assign them to devices, generating multiple independent control links, each issued to its corresponding target execution device. This simplifies complex scenario-based operations into a single interaction, eliminating the need for users to issue commands to each device separately. This significantly improves the processing efficiency and automation of complex tasks, achieving a leap from "single-point control" to "overall collaboration," and providing users with a seamless, integrated intelligent experience.
[0108] Furthermore, it is possible to determine whether the user has authorized the execution through interaction with the user.
[0109] In one implementation, upon detecting a confirmation event, an execution authorization request is sent to a first control device, causing the first control device to output the execution authorization request to the user; a second user response is received from the first control device; the second user response indicates whether the user agrees to or refuses execution; the second user response is generated by the first control device based on the user's input response instruction; based on the second user response, it is determined whether to allow the issuance of control instructions to the second control device; wherein, if the second user response indicates agreement to execution, it is determined that the issuance of control instructions is allowed; if the second user response indicates refusal to execute, it is determined that the issuance of control instructions is not allowed. For example, the confirmation event can be the first cross-device execution.
[0110] Thus, a user authorization confirmation step is introduced before performing critical operations or initial cross-device control, aiming to ensure operational security and user controllability. The specific process is as follows: when the system detects a preset confirmation event (e.g., initial cross-device execution, high-risk operation such as door lock control, or scenarios requiring confirmation according to product policy), it pauses the command issuance process and instead sends an execution authorization request to the user through the first control device. This request clearly informs the user of the upcoming operation, and the user can give a clear instruction of agreement or refusal via voice or interface interaction. After receiving and parsing this second user response, the system only continues the command issuance process if the user explicitly agrees; if the user refuses, the request is terminated.
[0111] By embedding key human decision-making points into automated processes, a balance between intelligence and security is achieved. On the one hand, it respects the user's ultimate control, preventing the system from performing sensitive or unwanted operations without authorization, thus enhancing user trust in the intelligent system. On the other hand, through standardized interaction processes, it provides a security barrier for high-risk or uncertain operations, significantly improving system security and the controllability of the user experience.
[0112] Furthermore, control commands can be directly executed by the target device through push notifications.
[0113] In one implementation, if an abnormal event is detected in the control device, a control command is sent to the target execution device based on the user's intent via a push server.
[0114] Among them, abnormal events of control equipment include at least one of the following: The following events were detected: the communication connection between the second control device and the target execution device was disconnected; the network status of the second control device was detected as offline; the response time after issuing a command to the second control device exceeded a preset threshold; and an abnormality was detected in the second control device.
[0115] This constructs a highly reliable alternative execution path to handle anomalies when the ideal control link determined by the system fails. For example, cloud-to-cloud execution methods such as Mi Cloud or Mi DigitalCar Cloud (MDCD) can be used. The inability to control via a second control device typically refers to scenarios where the local connection between the second control device and the target execution device is broken, the second control device itself is offline, or command forwarding times out or fails. In this case, instead of relying on local relay, a direct control channel from the cloud to the target execution device is established through a universal push server. Technically, the cloud sends the encapsulated control commands directly to the target execution device via the IoT cloud platform (such as Mi Cloud) it is connected to, using a server push mechanism. For example, a user at work issues a command to "turn on the air conditioner at home" to their mobile phone. The system determines that the home smart gateway should act as the second control device to operate the air conditioner, but if the gateway is found to be offline, this alternative path is triggered, directly pushing the "turn on" command to the home air conditioner via the Mi Cloud service.
[0116] This greatly enhances the robustness and ultimate success rate of cross-device control systems. It ensures that even in the event of a local network failure or intermediate device malfunction, as long as the target execution device itself remains connected to the Internet, the user's control intentions can still be reliably executed, thus providing the user with a seamless, highly available, and coherent intelligent experience.
[0117] Furthermore, the execution results can be fed back to the user.
[0118] In one implementation, the execution result is returned to the first control device so that the first control device outputs the execution result to the user.
[0119] This constitutes the final closed loop of the cross-device scheduling process, designed to ensure that users are clearly aware of the final execution status of their commands. Specifically, after the cloud sends a control command to the second control device, regardless of whether the command executes successfully or not, the cloud receives feedback from the execution chain and then returns the execution result to the first control device that initially interacted with the user. The first control device clearly outputs this result to the user through voice announcements, screen displays, or specific light prompts.
[0120] Its technical effect is to establish a complete user interaction loop, providing clear operational feedback, thereby significantly improving system reliability and the integrity of the user experience. It eliminates the waiting and uncertainty after a user issues a command, enabling the user to understand the control status in a timely manner. Whether it's confirmation after successful execution or reason prompts after failure (such as "Device is on" or "Operation failed, device not responding"), it enhances the user's sense of control and trust in the intelligent system, ensuring that every interaction has a beginning and an end.
[0121] The above embodiments are illustrated below with two scenarios. The first scenario is: cockpit-controlled mobile phone. When driving, the user tells the in-vehicle speaker (i.e., the first control device) to open WeChat on their phone. The cloud receives this voice message forwarded by the in-vehicle speaker, recognizes the user's intention to open WeChat, and determines that the target execution device is the user's mobile phone. Since the mobile phone itself is not a control device, the in-vehicle system paired with the mobile phone is selected as the second control device. The in-vehicle system sends commands to the mobile phone, and finally, WeChat is automatically launched on the mobile phone, realizing seamless control of mobile devices in the driving scenario.
[0122] The second scenario is: using smart glasses to check phone status. When a user wears smart glasses, they say how much battery their phone has. The cloud receives this voice message from the smart glasses, recognizes the query intent, and identifies the target device as the phone. A smart speaker paired with the phone is selected as the second control device. The smart speaker queries the phone's battery status and returns the result to the glasses display, enabling convenient status checks for the phone via wearable devices.
[0123] The following is combined with Figure 2 The above embodiments are illustrated by examples; please refer to [reference needed]. Figure 2 The diagram shown is an example of the architecture of a cross-device scheduling system, which mainly consists of two parts: the cloud and the user-associated devices.
[0124] The user-associated devices (i.e., the client) mainly include a first speaker, a second speaker, and several other devices. The first speaker (i.e., the first control device) serves as the entry point for user interaction and is responsible for receiving user commands. The second speaker (i.e., the second control device) can receive and execute control commands from the cloud. The cloud is the intelligent hub of the system, containing several core modules, each with its own function: a pre-processing module, an agent-parser (intent understanding) module, and an agent-skill (skill service) module.
[0125] The preprocessing module, a common intent module within the preprocessing service, is responsible for extracting device slots from user requests. For example, it can rewrite user requests based on the location information of the first speaker (i.e., the first control device). The agent-parser module is responsible for deeply parsing the semantics of user requests and transforming them into structured, executable instruction code.
[0126] The specific workflow is as follows: The user issues a voice command (i.e., a user command) to the first speaker, which acts as the first control device, such as "play Peppa Pig". The first speaker generates a user request based on this and sends it to the cloud. The cloud's preprocessing module obtains the device location information of the first speaker (i.e., the first control device). Then, the agent-parser module rewrites the user request based on the device location information of the first speaker to obtain a new user request. Based on the semantic understanding of the new user request, it generates structured code such as play(Media=Peppa Pig, Device=TV), which clarifies the target action and device conditions.
[0127] The process then proceeds to the agent-skill module. The agent-skill module is the command execution planning center, which includes the device decision module, the customization fulfillment module, and the device execution module.
[0128] The device decision module invokes the device information caching service to obtain basic device information for each user's associated devices. Based on the user's intent and the basic device information of each user's associated devices, it makes a comprehensive judgment and decides to execute the user request on the TV in the living room. The customization module obtains the current context (device status information) of the TV in the living room and generates an instruction set (control commands), such as detailed information like which video platform "Peppa Pig" is played on. The device execution module determines that the second control device associated with the TV is the second speaker, encapsulates the issued instructions into control commands, and sends the control commands to the second speaker. The control commands include the TV's device identifier (deviceId). The second speaker forwards the control commands to the TV via a local connection, controlling the TV to play "Peppa Pig".
[0129] See Figure 3 The diagram shows a flowchart of a device decision-making method applied to a server. The specific process includes the following steps: Step 301: Perform intent analysis on the user request from the first control device to obtain the user intent.
[0130] The user intent can take the form of intent function code.
[0131] Step 302: Based on the user's intent, determine whether it is a cross-device call. If yes, proceed to step 303; otherwise, proceed to step 309.
[0132] For example, there is only the first speaker, and no other speakers.
[0133] Step 303: Obtain the basic device information of each user's associated devices.
[0134] Step 304: Based on the user's intent and the basic device information of each user's associated devices, obtain at least one candidate device.
[0135] Furthermore, at least one candidate device can be further filtered based on set filtering criteria. Optionally, the filtering criteria can be set according to the actual application scenario, and there are no restrictions here.
[0136] Step 305: Determine if there are multiple candidate devices. If yes, proceed to step 306; otherwise, proceed to step 307.
[0137] Step 306: Send a device query request to the first control device and receive the first user response returned by the first control device.
[0138] Optionally, interaction can be achieved through a protocol agreed upon with the first control device.
[0139] Step 307: Determine the target execution device.
[0140] Step 308: Send a control command to the second control device associated with the target execution device to control the target execution device to perform the corresponding operation.
[0141] Optionally, different control commands can be issued to different target execution devices.
[0142] Step 309: Send control commands to the first control device and execute the control commands through the first control device.
[0143] See Figure 4 The diagram shown is a flowchart of a device screening method, which includes the following steps: Step 401: Perform intent analysis on the user request from the first control device to obtain the user intent.
[0144] Step 402: Obtain the basic device information of each user's associated devices.
[0145] Step 403: Based on the user's intent and the basic device information of each user's associated devices, select candidate devices from each user's associated devices.
[0146] Step 404: Perform anomaly detection on candidate devices.
[0147] Step 405: Select candidate devices with normal test results.
[0148] Step 406: Determine whether there is at least one candidate device. If yes, proceed to step 407; otherwise, proceed to step 408.
[0149] Step 407: Determine the target execution device based on the candidate devices.
[0150] Step 408: Send an abnormality alert to the first control device.
[0151] The error message can include error codes and / or technical suggestions to guide users or upstream systems in troubleshooting.
[0152] In practical applications, when performing steps 301-311 and steps 401-47, please refer to steps 101-103 above for specific steps, which will not be repeated here.
[0153] This application aims to construct a unified and efficient cross-device collaborative processing framework. Its core lies in treating multiple user devices as a whole through centralized intelligent decision-making, intelligently selecting the optimal device to execute user commands. First, the user request is subjected to intent recognition. The intent recognition model is used to parse the user command, outputting a structured user intent, including the target action and the device conditions required to execute that action. Optionally, the request can be semantically enhanced by incorporating the location information of the first control device to improve recognition accuracy.
[0154] The process then proceeds to a centralized decision-making phase. Based on user intent and basic device information (including device attributes, functions, and connection status), a two-stage screening process determines the second control device to execute the command: first, candidate devices are selected based on device conditions; then, the target execution device is determined using mechanisms such as anomaly detection and multi-turn dialogue. If the target execution device is a control device, it directly becomes the second control device; otherwise, a separate second control device is determined to control it. This centralized decision-making mechanism addresses the poor scalability problem caused by the decentralized decision-making logic in existing technologies.
[0155] Before issuing commands, anomaly detection and user authorization confirmation can be performed to improve system reliability. Command generation is customized based on the real-time status information of the target execution device, enabling on-demand context acquisition and reducing data transmission and cloud load. Finally, commands are issued to the second control device, supporting execution via local device forwarding or cloud push to ensure high availability. Execution results are returned to the first control device to complete the interaction loop.
[0156] In this way, by unifying decision-making, obtaining context on demand, and executing multiple paths, cross-device scheduling with high resource efficiency, strong scalability, and consistent user experience is achieved, overcoming the shortcomings of existing technologies such as resource waste, poor scalability, and coarse control granularity.
[0157] Based on the same inventive concept, this application also provides a cross-device scheduling apparatus. Since the principle of the above-mentioned apparatus and device in solving the problem is similar to that of a cross-device scheduling method, the implementation of the above-mentioned apparatus can refer to the implementation of the method, and repeated details will not be elaborated further. This apparatus can be applied to electronic devices. This application does not limit the type of electronic device; it can be any suitable type of device, such as terminal devices and servers, etc., which will not be elaborated further in this application. The apparatus embodiment can be implemented by software, or by hardware, or a combination of software and hardware. Taking software implementation as an example, as a logically defined apparatus, it is formed by the processor of the electronic device reading the corresponding computer program instructions from the non-volatile memory into memory and running them.
[0158] See Figure 5 The diagram shown is a structural block diagram of a cross-device scheduling apparatus according to an embodiment of this application. In some embodiments, the cross-device scheduling apparatus exemplified in this application includes: The identification unit 501 is used to respond to a user request sent by the first control device, perform intent identification on the user request, and obtain the identified user intent; the user request is generated by the first control device based on user instructions issued by the user. The determining unit 502 is used to determine a second control device based on the user intent and device basic information of at least one user-associated device; the device basic information is used to evaluate the compatibility between the user-associated device and the user intent; the second control device is different from the first control device and belongs to the type of central device that can be configured to drive at least one other user-associated device to perform operations. The sending unit 503 is used to send control commands to the second control device based on the user's intent.
[0159] In one embodiment, the determining unit 502 is used to: Based on the user intent and the device basic information of at least one user-associated device, determine the target execution device among at least one user-associated device; If the target execution device is a central device, then the target execution device is identified as the second control device; If the target execution device is not a central device type, then a second control device for controlling the target execution device is obtained from at least one user-associated device.
[0160] In one embodiment, the sending unit 503 is used for: Based on the user intent, as well as the device identifier and device status information of the target execution device, control commands are generated; the device status information includes the device runtime status related to the execution of the user intent. Send control commands to the second control device.
[0161] In one embodiment, the identification unit 501 is used for: Input the user's request into the intent recognition model, and obtain the user's intent output by the intent recognition model; The user intent includes the target action to be performed and the device conditions on which the target action depends for its execution.
[0162] In one embodiment, the identification unit 501 is further configured to: Based on the device location information of the first control device, the user request is updated to obtain a new user request.
[0163] In one embodiment, the first control device is a first speaker, and at least one user-associated device includes a second speaker and a camera device, wherein the second speaker is associated with the camera device; The identification unit 501 is used for: Based on the user's intent and the device basic information of at least one user-associated device, the camera device is identified as the target execution device. In response to the fact that the camera device is not a central device type, the second speaker associated with the camera device is identified as a second control device.
[0164] In one embodiment, the determining unit 502 is used to: If the user's intent contains at least two target actions, then for each target action, a second control device is determined to perform that target action; The transmitting unit 503 is used for: If there are multiple second control devices, then the corresponding control command is issued to each of the second control devices.
[0165] In one embodiment, the determining unit 502 is used to: Based on the device conditions contained in the user intent and the basic device information of each user's associated devices, at least one candidate device that meets the device conditions and can execute the user intent is selected from each user's associated devices. Determine the target execution device based on at least one candidate device.
[0166] In one embodiment, the determining unit 502 is used to: Perform anomaly detection on at least one candidate device; Candidate devices whose test results indicate that the device is functioning normally are selected from at least one candidate device; If the selected candidate devices are empty, control commands are not allowed to be issued, and an abnormal prompt message is sent to the first control device. If at least one candidate device is selected, the target execution device is determined based on the selected candidate devices.
[0167] In one embodiment, the determining unit 502 is used to: When there is only one candidate device, the candidate device is determined as the target execution device; When there are multiple candidate devices, a device query request is sent to the first control device, and a first user response is received from the first control device. Based on the first user response, the target execution device selected by the user is determined. The device query request contains device information of each candidate device.
[0168] In one embodiment, the equipment conditions include the target equipment type and / or the target equipment location; Basic equipment information includes equipment attribute information, equipment functions, and equipment connection status, which correspond to the equipment conditions.
[0169] In one embodiment, the sending unit 503 is further configured to: When a confirmation event is detected, an execution authorization request is sent to the first control device, so that the first control device outputs the execution authorization request to the user; The system receives a second user response from the first control device; the second user response is used to indicate whether the user agrees to or refuses to execute the command; the second user response is generated by the first control device based on the user's input response command. Based on the second user's response, determine whether to allow the issuance of control commands to the second control device; When the second user replies indicating agreement to execution, it is determined that the issuance of control commands is permitted; when the second user replies indicating refusal to execution, it is determined that the issuance of control commands is not permitted.
[0170] In one embodiment, the sending unit 503 is further configured to: If an abnormal event is detected in the control device, control commands are sent to the target device based on the user's intent via the push server.
[0171] In one implementation, the abnormal event of the control device includes at least one of the following: The communication connection between the second control device and the target execution device was detected to be disconnected. The network status of the second control device was detected as offline. The response time after issuing a command to the second control device exceeds a preset threshold. An anomaly was detected in the second control device.
[0172] In one embodiment, the sending unit 503 is further configured to: The execution result is returned to the first control device so that the first control device can output the execution result to the user.
[0173] The cross-device scheduling method in this embodiment includes, in response to receiving a user request sent by a first control device, performing intent recognition on the user request to obtain the recognized user intent; the user request is generated by the first control device based on user instructions issued by the user; determining a second control device based on the user intent and device basic information of at least one user-associated device; the device basic information is used to evaluate the compatibility between the user-associated device and the user intent; the second control device is different from the first control device and belongs to a central device type that can be configured to drive at least one other user-associated device to perform operations; and issuing control instructions to the second control device based on the user intent. In this way, the user request is parsed once to generate a structured user intent; subsequently, a second control device different from the first control device is intelligently selected from at least one user-associated device based on the device basic information; and finally, control instructions are issued to that device. This solves the problem of wasted cloud computing resources caused by multi-path parallel intent understanding in related technologies, and achieves reduced computing load and improved response efficiency through architectural optimization, providing a foundation for building an efficient cross-device collaborative system.
[0174] In this embodiment of the application, an electronic device is also provided, including: Processor; and The memory stores computer instructions that cause the processor to execute the methods of any of the above-described embodiments.
[0175] In this application embodiment, a computer-readable storage medium is provided, storing computer instructions for causing a computer to perform the methods of any of the above embodiments.
[0176] This application also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in the processor of an electronic device, the processor in the electronic device performs the method of any of the above-described embodiments.
[0177] Figure 6 A schematic diagram of the structure of an electronic device 6000 is shown. (See also...) Figure 6 As shown, the electronic device 6000 includes a processor 6010 and a memory 6020, and optionally may also include a power supply 6030, a display unit 6040, and an input unit 6050.
[0178] The processor 6010 is the control center of the electronic device 6000. It connects various components through various interfaces and lines, and performs various functions of the electronic device 6000 by running or executing software programs and / or data stored in the memory 6020, thereby performing overall monitoring of the electronic device 6000.
[0179] In this embodiment, when the processor 6010 calls the computer program stored in the memory 6020, it executes the steps in the above embodiments.
[0180] Optionally, the processor 6010 may include one or more processing units; preferably, the processor 6010 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 6010.
[0181] The memory 6020 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, various applications, etc.; the data storage area may store data created based on the use of the electronic device 6000, etc. In addition, the memory 6020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device, etc.
[0182] The electronic device 6000 also includes a power supply 6030 (such as a battery) that supplies power to various components. The power supply can be logically connected to the processor 6010 through a power management system, thereby enabling the management of charging, discharging, and power consumption.
[0183] The display unit 6040 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 6000. In this embodiment, it is mainly used to display the display interfaces of various applications in the electronic device 6000, as well as text, images, and other objects displayed on the display interfaces. The display unit 6040 may include a display panel 6041. The display panel 6041 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0184] The input unit 6050 can be used to receive information such as numbers or characters input by the user. The input unit 6050 may include a touch panel 6051 and other input devices 6052. The touch panel 6051, also known as a touch screen, can collect touch operations on or near the touch panel 6051 (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 6051).
[0185] Specifically, the touch panel 6051 can detect user touch operations and the signals generated by these operations, convert them into touch point coordinates, send them to the processor 6010, and receive and execute commands from the processor 6010. Furthermore, the touch panel 6051 can be implemented using various types of touch technologies, including resistive, capacitive, infrared, and surface acoustic wave. Other input devices 6052 can include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0186] Of course, the touch panel 6051 can cover the display panel 6041. When the touch panel 6051 detects a touch operation on or near it, it transmits the information to the processor 6010 to determine the type of touch event. Subsequently, the processor 6010 provides corresponding visual output on the display panel 6041 based on the type of touch event. Although in Figure 6 In this embodiment, the touch panel 6051 and the display panel 6041 are two separate components to realize the input and output functions of the electronic device 6000. However, in some embodiments, the touch panel 6051 and the display panel 6041 can be integrated to realize the input and output functions of the electronic device 6000.
[0187] The electronic device 6000 may also include one or more sensors, such as a pressure sensor, a gravity acceleration sensor, a proximity sensor, etc. Of course, depending on the specific application, the electronic device 6000 may also include other components such as a camera. Since these components are not the focus of this application embodiment, therefore... Figure 6 It is not shown in the text and will not be described in detail here.
[0188] Those skilled in the art will understand that Figure 6 This is merely an example of an electronic device and does not constitute a limitation on the electronic device. It may include more or fewer components than shown, or a combination of certain components, or different components.
[0189] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.
Claims
1. A method for cross-device scheduling, characterized in that, The method includes: In response to receiving a user request sent by a first control device, the user request is subjected to intent recognition to obtain the recognized user intent; the user request is generated by the first control device based on user instructions issued by the user. Based on the user intent and the device basic information of at least one user-associated device, a second control device is determined; the device basic information is used to evaluate the compatibility between the user-associated device and the user intent; the second control device is different from the first control device and belongs to a type of central device that can be configured to drive at least one other user-associated device to perform operations. Based on the user's intent, a control command is sent to the second control device.
2. The method according to claim 1, characterized in that, The step of determining the second control device based on the user intent and device basic information of at least one user-associated device includes: Based on the user intent and the device basic information of at least one user-associated device, determine the target execution device among the at least one user-associated device; If the device type of the target execution device is the central device type, then the target execution device is determined to be the second control device; If the device type of the target execution device is not the central device type, then the second control device for controlling the target execution device is obtained from the at least one user-associated device.
3. The method according to claim 2, characterized in that, The step of issuing control commands to the second control device based on the user intent includes: Based on the user intent, and the device identifier and device status information of the target execution device, the control command is generated; the device status information includes the device runtime status related to the execution of the user intent; The control command is sent to the second control device.
4. The method according to claim 2, characterized in that, The intent recognition of the user request includes: The user request is input into the intent recognition model to obtain the user intent output by the intent recognition model; The user intent includes the target action to be performed and the device conditions on which the target action depends for execution.
5. The method according to claim 4, characterized in that, Before performing intent recognition on the user request, the method further includes: Based on the device location information of the first control device, the user request is updated to obtain a new user request.
6. The method according to any one of claims 2-5, characterized in that, The first control device is a first speaker, and the at least one user-associated device includes a second speaker and a camera device, wherein the second speaker is associated with the camera device; The step of determining the second control device based on the user intent and device basic information of at least one user-associated device includes: Based on the user intent and the device basic information of the at least one user-associated device, the camera device is determined to be the target execution device; In response to the camera device not being the central device type, the second speaker associated with the camera device is identified as the second control device.
7. The method according to any one of claims 1-5, characterized in that, The step of determining the second control device based on the user intent and device basic information of at least one user-associated device includes: If the user intent contains at least two target actions, then for each target action, a second control device for performing that target action is determined. The step of issuing control commands to the second control device includes: If there are multiple second control devices, then corresponding control commands are issued to each second control device.
8. The method according to claim 2, characterized in that, The step of determining the target execution device among the at least one user-associated devices based on the user intent and the device basic information of at least one user-associated device includes: Based on the device conditions contained in the user intent and the basic device information of each user's associated devices, at least one candidate device that meets the device conditions and can execute the user intent is selected from each user's associated devices. The target execution device is determined based on the at least one candidate device.
9. The method according to claim 8, characterized in that, The step of determining the target execution device based on the at least one candidate device includes: Anomaly detection is performed on each of the at least one candidate device; Candidate devices whose test results indicate that the device is functioning normally are selected from the at least one candidate device; If the selected candidate devices are empty, the control command is not allowed to be issued, and an abnormal prompt message is sent to the first control device. If at least one candidate device is selected, the target execution device is determined based on the selected candidate devices.
10. The method according to claim 8 or 9, characterized in that, The determination of the target execution device includes: When there is only one candidate device, the candidate device is determined as the target execution device; When there are multiple candidate devices, a device query request is sent to the first control device, and a first user response is received from the first control device. Based on the first user response, the target execution device selected by the user is determined. The device query request includes device information of each candidate device.
11. The method according to claim 8 or 9, characterized in that, The equipment conditions include the target equipment type and / or the target equipment location; The basic device information includes device attribute information, device functions, and device connection status corresponding to the device conditions.
12. The method according to any one of claims 1-5, characterized in that, Before issuing control commands to the second control device based on the user intent, the method further includes: Upon detecting a confirmation event, an execution authorization request is sent to the first control device, so that the first control device outputs the execution authorization request to the user; The system receives a second user response returned by the first control device; the second user response is used to indicate whether the user agrees to or refuses to execute the command; the second user response is generated by the first control device based on the user's input response command. Based on the second user's response, determine whether to allow the control command to be sent to the second control device; When the second user replies indicating agreement to execution, it is determined that the control command can be issued; when the second user replies indicating refusal to execution, it is determined that the control command cannot be issued.
13. The method according to any one of claims 2-5, characterized in that, The method further includes: If an abnormal event is detected in the control device, the control command is sent to the target execution device via the push server based on the user's intent.
14. The method according to claim 13, characterized in that, The abnormal event of the control device includes at least one of the following: The communication connection between the second control device and the target execution device is detected to be disconnected. The network status of the second control device was detected as offline; The response time after sending a command to the second control device exceeds a preset threshold. An abnormality was detected in the second control device.
15. The method according to any one of claims 1-5, characterized in that, After issuing a control command to the second control device based on the user intent, the method further includes: The execution result is returned to the first control device so that the first control device outputs the execution result to the user.
16. A device for cross-device scheduling, characterized in that, The device includes: The identification unit is configured to, in response to receiving a user request sent by the first control device, identify the intent of the user request and obtain the identified user intent; the user request is generated by the first control device based on user instructions issued by the user. The determining unit is configured to determine a second control device based on the user intent and device basic information of at least one user-associated device; the device basic information is used to evaluate the compatibility between the user-associated device and the user intent; the second control device is different from the first control device and belongs to a type of central device that can be configured to drive at least one other user-associated device to perform operations. The sending unit is used to send control commands to the second control device based on the user's intent.
17. The apparatus according to claim 16, characterized in that, The determining unit is used for: Based on the user intent and the device basic information of at least one user-associated device, determine the target execution device among the at least one user-associated device; If the target execution device is a central device, then the target execution device is identified as the second control device; If the device type of the target execution device is not a central device type, then the second control device for controlling the target execution device is obtained from the at least one user-associated device.
18. The apparatus according to claim 17, characterized in that, The sending unit is used for: Based on the user intent, and the device identifier and device status information of the target execution device, the control command is generated; the device status information includes the device runtime status related to the execution of the user intent; The control command is sent to the second control device.
19. An electronic device, characterized in that, include: processor; as well as A memory storing computer instructions for causing the processor to perform the method according to any one of claims 1 to 15.
20. A computer-readable storage medium, characterized in that, The computer contains computer instructions for causing the computer to perform the method described in any one of claims 1 to 15.
21. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is executed in a processor of an electronic device, the processor in the electronic device performs the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Information processing method and center control equipment
CN104615456A
Smart home equipment voice control method and system based on earphone
CN109949801A
Voice interaction device and system, device control method, computing device and medium
CN111696534A
Smart home vehicle-mounted control method, device and equipment and storage medium
CN112051748A
Server, intelligent device and intelligent voice control method
CN114067798A