Intelligent equipment control method and device, equipment and storage medium

The target task is disassembled through the language planning model, and the execution meta actions are mapped to the basic skills and underlying control skills of the intelligent device, and the target execution instructions are generated, which solves the problems of high cost of training data acquisition and high computing power requirements in the existing technology, and achieves low-cost and efficient intelligent device control.

CN120029127APending Publication Date: 2025-05-23PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510073270.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the end-to-end control of embodied robots, the cost of collecting training data is high, and the demand for hardware computing power increases sharply when controlling multiple robots, resulting in a high control cost.

Method used

A control method for intelligent devices is proposed, and the target tasks are disassembled through a large language planning model, subdivided into multiple execution meta actions, and map these execution meta actions to the basic skill sequence and underlying control skills of the intelligent device to generate target execution instructions.

Benefits of technology

It reduces the cost of collecting training data and computing power requirements, simplifies the control process of smart devices, improves the generalization adaptability to different smart devices, and expands the application scenarios of the control process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029127A_ABST
    Figure CN120029127A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control method and device of intelligent equipment, equipment and a storage medium, and relates to the technical field of intelligent control. The method comprises the following steps: inputting a target task into a language planning large model for task planning to obtain an execution action sequence, and mapping execution meta-actions to corresponding basic execution skills in a basic skill sequence one by one to obtain a first-layer skill mapping result; and mapping the basic execution skills to at least one bottom layer control skill one by one to obtain a bottom layer skill mapping result corresponding to the first layer skill mapping result, and finally generating a target execution instruction of the intelligent device for the target task according to the bottom layer skill mapping result. Through the layer-by-layer progressive disassembly mode, decoupling of an action planning process based on a large model and a specific control process of the intelligent equipment is realized, all planning processes can be completed without depending on a complex model architecture, a subsequent mapping and decomposition process can be realized based on a low-computing-power matching algorithm, the computing power demand of the intelligent equipment is simplified, and the implementation efficiency of the intelligent equipment is improved. The training cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent control technology, and in particular to control methods, devices, equipment and storage media for intelligent devices. Background Art

[0002] End-to-end control of an embodied robot means that the entire process from the robot sensing environmental information to the robot making actions and decisions is handled by a unified model or system. In this process, there is no need to manually design complex intermediate steps and rules. The trained model can directly control the embodied robot end-to-end based on the input environmental information, allowing the robot to perform the corresponding task.

[0003] In related technologies, training models are mainly achieved by teaching and collecting a large amount of data. Specifically, in the teaching process, humans demonstrate the operation of the robot to obtain data samples corresponding to different scenarios, so the cost of collecting training data is relatively high. In addition, if multiple robots need to be controlled, this control method will increase the demand for hardware computing power sharply, resulting in higher control costs. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to propose a control method, device, equipment and storage medium for intelligent devices, reduce the cost of collecting training data, and reduce the computing power requirements during the control process.

[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a control method for a smart device, comprising:

[0006] Acquire a target task, and input the target task into a language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes a plurality of execution meta-actions;

[0007] Acquire a basic skill sequence of the smart device, map the execution meta-actions to corresponding basic execution skills in the basic skill sequence one by one, and obtain a first-level skill mapping result;

[0008] Acquire the underlying control skills of the smart device, map the basic execution skills to at least one of the underlying control skills one by one, and obtain the underlying skill mapping result corresponding to the first-layer skill mapping result, wherein the underlying control skill is used to call one or more API control units of the smart device;

[0009] Generate a target execution instruction for the target task for the smart device according to the underlying skill mapping result.

[0010] In some embodiments, mapping the execution meta-actions to corresponding basic execution skills in the basic skill sequence one by one to obtain a first-level skill mapping result includes:

[0011] Acquire functional features and timing information corresponding to the execution meta-actions in the execution action sequence, and acquire at least one skill feature corresponding to each of the basic execution skills;

[0012] Selecting a target skill feature corresponding to each of the functional features from the skill features, and obtaining a basic target execution skill corresponding to the execution meta-action according to the basic execution skill corresponding to the target skill feature;

[0013] The first-level skill mapping result is obtained according to the timing information and the basic target execution skill.

[0014] In some embodiments, selecting a target skill feature corresponding to each of the functional features from the skill features includes:

[0015] Calculating a similarity value between the functional feature and each of the skill features;

[0016] The skill feature with the largest similarity value is selected as the target skill feature.

[0017] In some embodiments, the step of comparing the functional features with the skill features and selecting a target skill feature corresponding to each functional feature from the skill features includes:

[0018] If at least one of the functional features corresponds to one of the skill features, at least one of the functional features is merged to obtain a functional group, and the skill feature is used as the target skill feature corresponding to each of the functional features in the functional group;

[0019] If the functional feature corresponds to at least one sub-skill feature, obtain the skill feature corresponding to each sub-skill feature, sort the skill features according to the sub-time series information of the sub-skill features, and obtain at least one target skill feature corresponding to the functional feature.

[0020] In some embodiments, obtaining the first-level skill mapping result according to the timing information and the basic target execution skill includes:

[0021] Sorting the basic target execution skills according to the time sequence information to obtain an initial sorting result;

[0022] Acquire input and output data of the basic target execution skill, and if it is determined according to the input and output data that the basic target execution skill includes a pre-action, and the pre-action is another basic target execution skill, adjust the initial sorting result according to the pre-action to obtain a target sorting result;

[0023] The first-level skill mapping result is obtained based on the target sorting result.

[0024] In some embodiments, before obtaining the first-level skill mapping result based on the target ranking result, the method further includes:

[0025] If a plurality of said basic target execution skills include a synchronous execution flag;

[0026] The target sorting result is adjusted according to the synchronous execution identifier.

[0027] In some embodiments, mapping the basic execution skills to at least one of the bottom-level control skills one by one to obtain a bottom-level skill mapping result corresponding to the first-level skill mapping result includes:

[0028] Acquire at least one of the underlying control skills corresponding to each of the basic execution skills;

[0029] According to the order of the basic execution skills in the first-layer skill mapping result, the underlying skill mapping result is obtained according to the corresponding underlying control skills.

[0030] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a control device for a smart device, comprising:

[0031] Task decomposition module: used to obtain the target task, and input the target task into the language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes multiple execution meta-actions;

[0032] The first-layer mapping module is used to obtain the basic skill sequence of the smart device, map the execution element actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtain the first-layer skill mapping result;

[0033] Bottom-layer mapping module: used to obtain the bottom-layer control skills of the smart device, map the basic execution skills to at least one of the bottom-layer control skills one by one, and obtain the bottom-layer skill mapping result corresponding to the first-layer skill mapping result, wherein the bottom-layer control skill is used to call one or more API control units of the smart device;

[0034] Instruction generation module: used to generate a target execution instruction for the smart device for the target task according to the underlying skill mapping result.

[0035] To achieve the above objectives, a third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0036] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0037] The control method, device, equipment and storage medium of the intelligent device proposed in the embodiment of the present application obtains the target task and inputs the target task into the language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes multiple execution meta-actions; then obtains the basic skill sequence of the intelligent device, maps the execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtains the first-level skill mapping result; then obtains the underlying control skills of the intelligent device, maps the basic execution skills to at least one underlying control skill one by one, obtains the underlying skill mapping result corresponding to the first-level skill mapping result, and finally generates the target execution instruction of the intelligent device for the target task according to the underlying skill mapping result. In the embodiment of the present application, the target task is disassembled with the help of the language planning model and subdivided into multiple execution meta-actions. Then these execution meta-actions are mapped to the basic execution skills of the specific intelligent device, thereby obtaining the first-level skill mapping result composed of the basic functions of the intelligent device. Subsequently, the first-level skill mapping result is further decomposed and further refined using the underlying control skills that can call the API control unit. Finally, the target execution instruction adapted to the intelligent device is obtained. Through this progressive disassembly method, the action planning process based on the large model and the specific control process of the smart device can be decoupled. This means that there is no need to rely on a complex model architecture to complete all planning processes from start to finish, and the subsequent mapping and decomposition processes can be implemented based on a low-computing matching algorithm. In this way, it not only simplifies the computing power requirements of smart devices, but also reduces the training cost of large models. In addition, this mechanism of hierarchical decoupling of planning and control enables the control process to have good generalization adaptability to different smart devices, thereby expanding the application scenarios of the control process. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flow chart of a control method of a smart device provided in an embodiment of the present application.

[0039] Figure 2 It is a schematic diagram of a control method for a smart device provided in an embodiment of the present application.

[0040] Figure 3 It is a flowchart provided in an embodiment of the present application for mapping execution meta-actions to corresponding basic execution skills in a basic skill sequence one by one to obtain a first-level skill mapping result.

[0041] Figure 4 It is a flowchart of selecting the target skill feature corresponding to each functional feature from the skill features provided in an embodiment of the present application.

[0042] Figure 5 It is a flowchart provided in an embodiment of the present application for selecting a target skill corresponding to each functional feature from the skill features according to the skill features of the functional features.

[0043] Figure 6 It is a flowchart of obtaining the first-level skill mapping result according to the timing information and basic target execution skills provided in the embodiment of the present application.

[0044] Figure 7 It is a flowchart provided in an embodiment of the present application for mapping basic execution skills to at least one underlying control skill one by one to obtain underlying skill mapping results corresponding to the first-level skill mapping results.

[0045] Figure 8 It is a logical schematic diagram of the control method of the smart device provided in the embodiment of the present application.

[0046] Fig. 9 This is a structural block diagram of a control device for an intelligent device provided in yet another embodiment of the present application.

[0047] Fig.10 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments characterized herein are only used to explain the present application and are not used to limit the present application.

[0049] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or features may be executed in a different order than the module division in the device or the order in the flowchart.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of characterizing the embodiments of this application and are not intended to limit this application.

[0051] First, some nouns involved in this application are analyzed:

[0052] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. AI is a branch of computer science. AI attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. AI can simulate the information process of human consciousness and thinking. AI is also a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0053] End-to-end control of embodied robots means that the entire process from the robot's perception of environmental information to the robot's actions and decisions is handled by a unified model or system, and then deployed to a large computing power platform to output direct control of the robot / robotic arm's joint movements. In this process, there is no need to manually design complex intermediate steps and rules. The trained model can directly control the embodied robot end-to-end based on the input environmental information, allowing the robot to perform the corresponding tasks. This method relies on a large amount of computing power training resources and a large computing power terminal control unit, which increases the R&D cost and hardware configuration cost.

[0054] In related technologies, training models are mainly achieved by teaching and collecting large amounts of data. Specifically, in the teaching process, humans demonstrate the operation of the robot to obtain data samples corresponding to different scenarios. Therefore, the cost of collecting training data is high, and because the performance and functional range of robots / robotic arms vary, the models trained with heterogeneous data cannot be effectively unified, and there is an inherent lack of generalization ability. In addition, if multiple robots need to be controlled, this control method will increase the demand for hardware computing power sharply, resulting in higher control costs.

[0055] Based on this, the embodiment of the present application provides a control method, device, equipment and storage medium for an intelligent device, which uses a language planning large model to disassemble the target task and subdivide it into multiple execution meta-actions. Then, these execution meta-actions are mapped to the basic execution skills of the specific intelligent device, thereby obtaining the first-level skill mapping result consisting of the basic functions of the intelligent device. The first-level skill mapping result is then further decomposed and further refined using the underlying control skills that can call the API control unit. Finally, the target execution instruction adapted to the intelligent device is obtained. Through this progressive disassembly method, the decoupling of the action planning process based on the large model and the specific control process of the intelligent device is achieved. This means that there is no need to rely on a complex model architecture to complete all planning processes from beginning to end, and the subsequent mapping and decomposition process can be achieved based on a low-computing matching algorithm. In this way, not only the computing power requirements of intelligent devices are simplified, but also the training cost of large models is reduced. In addition, this mechanism of hierarchical decoupling of planning and control enables the control process to have good generalization adaptability to different intelligent devices, thereby expanding the application scenarios of the control process.

[0056] The embodiments of the present application provide a control method, apparatus, device and storage medium for a smart device, which are specifically described through the following embodiments. First, the control method for a smart device in the embodiments of the present application is characterized.

[0057] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.

[0058] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0059] The control method of the intelligent device provided in the embodiment of the present application relates to the field of intelligent control technology. The control method of the intelligent device provided in the embodiment of the present application can be applied to the intelligent device, can also be applied to the server side, and can also be a computer program running in the intelligent device or the server side. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports the control of the intelligent device, that is, a program that can be run only by downloading it to the browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plug-in. Among them, the intelligent device communicates with the server via a network. The control method of the intelligent device can be executed by the intelligent device or the server, or by the intelligent device and the server in collaboration.

[0060] In some embodiments, the intelligent device may be an embodied robot or a robotic arm. An embodied robot is a robot system with a physical entity that can sense the environment through sensors and use its own actuators to physically interact with the environment. The server may be an independent server, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms; or a service node in a blockchain system, where each service node in the blockchain system forms a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The intelligent device and the server may be connected via Bluetooth, Universal Serial Bus (USB) or a network, etc., which is not limited in this embodiment.

[0061] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be featured in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0062] The following is a method for controlling a smart device in an embodiment of the present application.

[0063] Figure 1 is an optional flow chart of the control method of the smart device provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps 110 to 140. It can also be understood that this embodiment Figure 1 The order of step 110 to step 140 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0064] Step 110: Obtain the target task, and input the target task into the language planning model for task planning to obtain an execution action sequence.

[0065] In one embodiment, the target task is a task that needs to be completed by the smart device, such as "pick up the apples on the table and put them in the fruit basket", "clean the floor of the living room and bedroom", etc. This embodiment does not limit the specific content of the target task.

[0066] In addition, there are several ways to obtain the target task. The first is that the user interacts with the embodied robot through voice or text and directly conveys the task to the robot. For example, the user says "clean the living room" to the home service robot, or enters "please wash the dishes in the kitchen sink and put them in the disinfection cabinet" in the control interface. The second is that the smart device directly receives external system instructions. For example, in an automated factory, the production management system will send tasks to the robot, such as "take part Y from shelf X in the warehouse and deliver it to workstation Z on the production line." Or the embodied robot autonomously identifies and triggers the target task through its own sensors and intelligent algorithms. For example, when the cleaning robot detects a large area of ​​stains on the ground through sensors, it automatically sets "cleaning the stained area" as the target task. This embodiment does not limit the method of obtaining the target task.

[0067] Next, after obtaining the target task, the target task will be input into the language planning model to carry out task planning. The language planning model of the embodiment of the present application can be a generalized conventional language model or a multimodal model. Before actual application, the selected large language model is subjected to scene adaptation operation so that the model can effectively handle task inputs in specific fields and formats. For example, the model can be trained and fine-tuned using a task-related data set to enhance the model's understanding and planning capabilities for specific tasks.

[0068] In one embodiment, the language planning model will perform task planning reasoning based on the type of smart device and the target task, breaking down the target task into multiple execution meta-actions. These execution meta-actions together constitute an execution action sequence, where the execution meta-actions can be regarded as a series of sub-tasks necessary to complete the target task.

[0069] In addition, there are corresponding timing relationships between the execution meta-actions in the execution action sequence. These timing relationships can ensure the accuracy of the execution logic of the target task and ensure the coherence and feasibility of different actions. If some actions in the execution action sequence cannot be realized by the corresponding smart device, then these actions need to be adjusted before output. For example, if the smart device is a sweeping robot, it cannot complete actions such as "throwing" and "picking up".

[0070] The disassembly process of the language planning model in the embodiment of the present application has a certain degree of abstract generalization and does not depend on the specific operating details of a specific smart device. Therefore, it can be applied to various types of target tasks, whether it is a task in a home scene or a task in an industrial production, office, etc. scene, it can be effectively disassembled. This versatility enables the obtained execution action sequence to adapt to different smart devices, without the cumbersome process of designing a planning model separately for each specific target task and smart device.

[0071] In one embodiment, referring to Figure 2 , Figure 2 It is a schematic diagram of a control method for a smart device provided in an embodiment of the present application. Figure 2 The intelligent device is a robotic arm, and the target task is: "Pick up the apple on the table and put it in the fruit basket". After entering it into the language planning model, the execution action sequences are: 1. Find the apple on the table and locate A; 2. Run the end gripper of the robotic arm to A according to the current real-time posture of the apple; 3. Execute the clamping action to complete the clamping of the apple; 4. Find the posture of the fruit basket on the table, and find the placement point B where the apple can be placed; 5. Plan the path s1 for the gripper to run to point B after clamping the apple; 6. Plan the path for the gripper to run to point B after clamping the apple; 7. Execute the path to run the end of the robotic arm to point B; 8. Release the gripper to place the apple; 9. The robotic arm returns to the original posture state and completes the ready waiting state again. It can be seen that Figure 2 The language planning model generates 9 corresponding execution meta-actions according to the target task, and there is a logical order between these actions.

[0072] Step 120: Obtain the basic skill sequence of the smart device, map the execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtain the first-level skill mapping result.

[0073] In one embodiment, for a smart device, it usually has several to more than a dozen basic execution skills. These basic execution skills belong to the category of coarse-grained functional division. They are relatively independent and functionally clear operation units that can cover all the functions of the smart device. At the same time, multiple basic execution skills together constitute the basic skill sequence of the smart device. Smart devices with different hardware or software configurations have different basic skill sequences. It is stated here that in the embodiments of this application, whether it is execution meta-actions, basic execution skills, or underlying control skills, they are only for illustrative purposes and do not represent limitations on them in the embodiments of this application.

[0074] Taking the robotic arm as an intelligent device as an example, its basic skill sequence generally includes the following four skills: (1) Object positioning skill: It can also be called "positioning the target posture state". This skill is used to determine the position and posture of the target object (such as an apple, fruit basket, etc.). (2) Path planning skill: Its main function is to plan the movement path of the robotic arm or gripper from one position to another. (3) Clamping and releasing skill: It is responsible for executing the actions of clamping and releasing objects. (4) Reset skill: It aims to return the intelligent device to the initial state or the specified ready state. Figure 2"Find target object 1 (apple) position A" and "Find target object 2 (fruit basket) position B" are both object positioning skills, "Grab (apple) A and put it in B" is a clamping and releasing skill, and "reset" is a resetting skill.

[0075] Taking the sweeping robot as an example of an intelligent device, its basic skill sequence generally includes the following four skills: (1) Object positioning skill: Use a variety of sensor technologies to achieve object positioning. (2) Path planning skill: Based on the environmental information obtained by the object positioning skill, the sweeping robot performs path planning. Common path planning algorithms include random collision and planned cleaning. (3) Cleaning and garbage collection skills: The side brush gathers garbage near the suction port, and the high-power vacuum cleaner generates strong suction to suck the garbage into the dust box. (4) Recharging and resetting skills: When the sweeping robot is low on power or completes the cleaning task, the recharging and resetting skill will be activated. Using the built-in navigation system, it plans the path back to the charging base based on the previously constructed map and positioning information.

[0076] In one embodiment, after obtaining the basic skill sequence of the smart device, it is necessary to map the execution action sequence obtained by the language planning model, and use the basic execution skills to express the execution meta-actions in the execution action sequence. Figure 3 , Figure 3 The flowchart of the embodiment of the present application is to map the execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one to obtain the first-level skill mapping result, which specifically includes the following steps:

[0077] Step 310: Obtain functional features and timing information corresponding to the execution meta-actions in the execution action sequence, and obtain at least one skill feature corresponding to each basic execution skill.

[0078] In one embodiment, each execution meta-action is first subjected to feature, keyword or similar word extraction to obtain the corresponding functional feature. At the same time, it is also necessary to obtain the timing information of the execution meta-action and the execution action sequence. It is understandable that there may be more than one functional feature corresponding to the execution meta-action. For example, the functional features corresponding to the execution meta-action "find the apple on the table and locate A" may be "locating" and "target object", and the functional features corresponding to the execution meta-action "perform clamping action, complete clamping of the apple" may be "clamping", etc. This can be achieved using a simple keyword matching process, which is not limited in this embodiment.

[0079] In one embodiment, according to the same extraction method, the basic execution skills are also extracted by features, keywords or similar words to obtain at least one corresponding skill feature. For example, the skill features corresponding to the "object positioning skill" may be "object" and "positioning", while the skill features corresponding to the "path planning skill" may be "path", "position", "movement", etc., and the skill features corresponding to the "clamping and releasing skill" may be "clamping", "clamping", "releasing", etc.

[0080] Step 320: Select the target skill feature corresponding to each functional feature from the skill features, and obtain the basic target execution skill corresponding to the execution meta-action according to the basic execution skill corresponding to the target skill feature.

[0081] In one embodiment, referring to Figure 4 , Figure 4 This is a flow chart of selecting a target skill feature corresponding to each functional feature from skill features provided by an embodiment of the present application, which specifically includes the following steps:

[0082] Step 410: Calculate the similarity value between the functional feature and each skill feature.

[0083] Step 420: Select the skill feature with the largest similarity value as the target skill feature.

[0084] In one embodiment, both the function feature and the skill feature can be represented as word vectors with the same dimension, and then the cosine similarity or Euclidean distance between the function feature and each skill feature is calculated as the corresponding similarity value, and finally the skill feature with the largest similarity value is used as the target skill feature. The purpose of this is to map each execution meta-action to the basic execution skill corresponding to the smart device.

[0085] Next, the basic target execution skill corresponding to the execution meta-action is obtained according to the basic execution skill corresponding to the target skill feature. Here, the basic target execution skill corresponding to the execution meta-action is a basic execution skill that sets the operation object as an input parameter. Therefore, the description of the basic target execution skill may add the corresponding input parameters, so the surface description may be different from the basic execution skill. For example, the input parameter of the basic target execution skill "find target object 1 (apple) posture A" is "target object 1 (apple)", and the input parameter of the basic target execution skill "find target object 2 (fruit basket) posture B" is "target object 2 (fruit basket)". It can be understood that even if some basic target execution skills have different display contents due to different input parameters, the corresponding basic execution skills are essentially the same.

[0086] In one embodiment, referring to Figure 5 , Figure 5The flowchart of selecting a target skill corresponding to each functional feature from the skill features according to the functional features provided by the embodiment of the present application specifically includes the following steps:

[0087] Step 510: If at least one functional feature corresponds to a skill feature, at least one functional feature is merged to obtain a functional group, and the skill feature is used as a target skill feature corresponding to each functional feature in the functional group.

[0088] In one embodiment, in actual applications, it is possible that multiple execution meta-actions are mapped to the same basic execution skill. For example, for an industrial robot, the two execution meta-actions of "moving the screw above the screw hole" and "moving the nut near the screw rod" are highly similar to the basic execution skill of "path planning skill" in terms of operation logic and functional realization.

[0089] Therefore, in the embodiment of the present application, if multiple functional features correspond to one skill feature, these functional features are merged to obtain a functional group, and then the skill feature is used as the target skill feature corresponding to each functional feature in the functional group, which is similar to a many-to-one mapping method.

[0090] Step 520: If the functional feature corresponds to at least one sub-skill feature, obtain the skill feature corresponding to each sub-skill feature, sort the skill features according to the sub-time series information of the sub-skill features, and obtain at least one target skill feature corresponding to the functional feature.

[0091] In one embodiment, there is also a situation where an execution meta-action requires the help of multiple basic execution skills to be realized. Taking an industrial robot completing a complex assembly task as an example, the functional feature may be "assembly of component A", so it can be divided into multiple sub-skill features, such as "grabbing component A", "moving to the assembly position", "assembly", etc. Through detailed disassembly and analysis of the task, it is determined that the functional feature "assembly of component A" corresponds to these sub-skill features.

[0092] Next, after identifying the sub-skill features corresponding to the functional features, it is necessary to clarify the skill features corresponding to each sub-skill feature. For example, the sub-skill feature of "grasping component A" may correspond to "gripping and releasing skills", "moving to the assembly position" may correspond to "path planning skills", and "performing assembly operations" may correspond to "assembly operation skills".

[0093] Through the above method, the skill feature corresponding to each sub-skill feature is obtained, and then the mapped skill features are sorted according to the sub-time series information corresponding to the sub-skill feature, so as to obtain at least one target skill feature corresponding to the functional feature. This is similar to a one-to-many mapping method.

[0094] It can be seen that the above two different mapping methods can accurately break down complex tasks into specific skill operation steps, thereby effectively improving the efficiency and accuracy of smart devices in executing tasks. This mapping method makes the control logic of smart devices clearer and simpler. Each execution element action corresponds to a specific basic execution skill, avoiding the direct coupling between complex tasks and device operations. Smart devices only need to focus on executing their basic execution skills without having to worry about the complex logic of the entire task. This not only reduces the difficulty of controlling smart devices, but also improves the maintainability and scalability of the device. This embodiment does not specifically limit the mapping method.

[0095] Step 330: Execute skills according to the timing information and basic objectives to obtain the first-level skill mapping result.

[0096] In one embodiment, after the basic target execution skills corresponding to each execution meta-action are obtained, the basic target execution skills are arranged according to the timing information of the execution meta-action, and the same basic target execution skills located in adjacent positions in the arrangement result are merged to obtain the first-level skill mapping result. It can be understood that even if the input parameters corresponding to two basic target execution skills are different, the corresponding basic execution skills can be consistent, that is, the first-level skill mapping result is likely to show that the same basic execution skill appears repeatedly at different timings.

[0097] In one embodiment, referring to Figure 2 , the first-level skill mapping results can be: "Find target object 1 (apple) position A [object positioning skill]", "Find target object 2 (fruit basket) position B [object positioning skill]", "Grip and release skill [execute (apple) A to grab B and put it in]" and "Reset skill".

[0098] In one embodiment, in the process of generating the first-level skill mapping result, the logical particularity of the operation of some smart devices also needs to be considered. Figure 6 , Figure 6 This is a flowchart of obtaining a first-level skill mapping result according to timing information and basic target execution skills provided by an embodiment of the present application, which specifically includes the following steps:

[0099] Step 610: Sort the basic target execution skills according to the timing information to obtain an initial sorting result.

[0100] In one embodiment, the basic target execution skills are directly sorted according to the timing information to obtain an initial sorting result.

[0101] Step 620: Obtain input and output data of the basic target execution skill. If it is determined based on the input and output data that the basic target execution skill includes a preceding action, and the preceding action is other basic target execution skills, adjust the initial sorting result based on the preceding action to obtain the target sorting result.

[0102] Step 630: Obtain the first-level skill mapping result based on the target sorting result.

[0103] In one embodiment, it is necessary to obtain the input and output data corresponding to each basic target execution skill. The input and output data are used to determine which execution meta-actions need to be executed before other execution meta-actions. For example, when a robotic arm clamps an object, in this time route, the "route planning skill" needs to be before the "clamping and releasing skill". Therefore, for the "route planning skill", the corresponding preceding action at this moment is the "clamping and releasing skill". At this time, if the preceding action is after the basic target execution skill, it is necessary to add the corresponding preceding action in front of the basic target execution skill, adjust the initial sorting result, and obtain the target sorting result. If the preceding action is before the basic target execution skill, no adjustment is required. Finally, the first-level skill mapping result is obtained based on the updated target sorting result.

[0104] It is understandable that the pre-action is determined according to the actual operation process and operation scenario, and this embodiment does not limit this.

[0105] In one embodiment, the synchronous execution relationship between the basic target execution skills also needs to be considered. In the specific process, if multiple basic target execution skills include synchronous execution identifiers, the target sorting result is adjusted according to the synchronous execution identifier. For example, for data collection related actions, different sensors can collect data synchronously to reduce the operation control time. Therefore, the basic target execution skills that can be executed synchronously are set in parallel in the target sorting result.

[0106] It is understandable that the above mapping process can be implemented according to the actual smart device programming. And because the number of basic execution skills in the basic skill sequence is limited, even if similarity calculation is required during task allocation or optimization, the computing power required is relatively small. Taking the execution of tasks by robotic arms as an example, its basic skills such as object positioning, path planning, clamping and release are fixed. When a new task is received, the system compares the similarity of the current task with past tasks, and only needs to analyze within a limited range of basic skills. For example, to judge the tasks of grabbing apples and grabbing oranges, it is only necessary to compare the differences between the two tasks in basic skills such as object positioning and clamping strength, without the need for large-scale complex calculations.

[0107] In addition, after multiple mapping processes, a mapping result table can be constructed. This table stores the mapping results of the execution meta-actions and the corresponding basic skill sequences. The next time you encounter the same or similar execution meta-actions, there is no need to re-perform complex mapping and calculations. You can directly obtain and reuse the previous mapping results from the table. For example, every time the sweeping robot receives a "clean the bedroom" task, the system first queries the mapping result table. If there is a match, the corresponding basic skill sequence is directly called, which greatly reduces computing power consumption and improves task execution efficiency, allowing smart devices to respond to task instructions more quickly and improve user experience.

[0108] Step 130: Acquire the underlying control skills of the smart device, map the basic execution skills to at least one underlying control skill one by one, and obtain the underlying skill mapping result corresponding to the first-level skill mapping result.

[0109] In one embodiment, after the first-level skill mapping result is obtained, it is also necessary to associate it with the API control unit used by the specific smart device to perform specific operations. Therefore, it is necessary to obtain the bottom-level control skills of each smart device, where the bottom-level control skills are used to call one or more API control units of the smart device. The API control unit is the lowest-level hardware unit control unit of the smart device, such as the sensor control unit, the rotation operation control unit, etc., which is obtained according to the actual situation of the smart device.

[0110] In one embodiment, for each smart device, its underlying control skills can be pre-stored in a table for easy calling. Figure 7 , Figure 7 The flowchart of mapping basic execution skills to at least one underlying control skill one by one and obtaining the underlying skill mapping result corresponding to the first-layer skill mapping result provided by the embodiment of the present application specifically includes the following steps:

[0111] Step 710: Obtain at least one underlying control skill corresponding to each basic execution skill.

[0112] In one embodiment, at least one underlying control skill corresponding to each basic execution skill may be stored in a table, and the underlying control skill corresponding to the basic execution skill may be obtained by simply querying the basic execution skill.

[0113] For example, for a robot arm, if its basic execution skill is "grasping", the corresponding underlying control skills may be: "control the rotation angle and speed of the robot arm joint motor", "move the gripper at the end of the robot arm to the object position", "control the closing force of the gripper", etc. For example, the following Table 1 also gives two possible examples.

[0114]

[0115] Table 1 Examples of basic execution skills and underlying control skills

[0116] The matching process of the above table can be implemented by regular matching or querying. For example, the names of basic execution skills are named according to certain naming specifications, such as "operation_object". For basic execution skills such as "grab_part" and "place_tool", a regular expression "grab_.*" can be constructed to match all basic execution skills starting with "grab".

[0117] Step 720: Obtain the underlying skill mapping result according to the order of the basic execution skills in the first-level skill mapping result and the corresponding underlying control skills.

[0118] In one embodiment, according to the order of the basic execution skills in the first-level skill mapping result, the corresponding underlying control skills are queried one by one for the basic execution skills, thereby obtaining the underlying skill mapping result.

[0119] It is understandable that different smart devices may have different implementation methods and control accuracy requirements for the same basic execution skills. Therefore, the operation can be refined through the underlying control skills to accurately control the characteristics of different smart devices. For example, different brands of sweeping robots may have differences in cleaning speed, brush head rotation method, etc. By adjusting the underlying control skills, the target execution instructions can be adapted to different smart devices, improving the accuracy and reliability of control.

[0120] In one embodiment, the decomposition of the first-layer skill mapping result may have multiple layers, that is, the bottom-layer skill mapping result may be further decomposed, and this embodiment does not limit the number of decomposition layers.

[0121] In one embodiment, referring to Figure 2 , decompose the first-level skill mapping results to obtain the second-level skill mapping results. The second-level skill mapping results obtained by mapping the first-level skill mapping results of "finding the target object 1 (apple) posture A" include: "the robotic arm looks around the scene to find the target object (apple)", "image RGB online recognition of the target object (apple)", "RGB segmentation of apple image and generation of 3D model, evaluation of its graspable posture", "output the graspable posture parameters of the target (apple)". It can be seen that by decomposing the first-level skill mapping results, the second-level skill mapping results can already go deep into the operation content at the API level.

[0122] Step 140: Generate target execution instructions for the smart device for the target task based on the underlying skill mapping results.

[0123] In one embodiment, after obtaining the underlying skill mapping results, target execution instructions for the target task can be generated according to the specific operations of the smart device.

[0124] by Figure 2 Take the partial second-layer skill mapping results corresponding to "find the target object 1 (apple) posture A" as an example.

[0125] The target execution instruction generated for "robotic arm scanning the scene to find the target object (apple)" can be:

[0126] "Control the robot arm joint to move along the trajectory" generates the command "MOVE_JOINT:A,θ1,v1; MOVE_JOINT:B,θ2,v2;..." (the specific angle and speed are determined by the surround trajectory). "Turn on the visual sensor" generates the command "START_VISION:0,f1". "Set the image acquisition frequency to f1", if it has been set in the sensor turn on command, there is no need to generate a separate command.

[0127] The target execution instruction generated for "image RGB online recognition of target object (apple)" can be:

[0128] "Calling image recognition algorithm" generates the instruction "RUN_ALGORITHM:1". "Extracting image RGB features" assumes the corresponding instruction "EXTRACT_RGB_FEATURE:2", where 2 represents the feature extraction method number. "Compare with Apple feature database" generates the instruction "COMPARE_WITH_APPLE_DB:3", where 3 represents the comparison mode number.

[0129] The target execution instructions generated for "RGB segmentation of apple image and generation of 3D model, evaluation of its graspable pose" can be:

[0130] "Execute image segmentation algorithm" generates the instruction "RUN_IMAGE_SEGMENTATION:4", where 4 represents the segmentation algorithm number. "Build 3D model algorithm" generates the instruction "BUILD_3D_MODEL:5", where 5 represents the 3D model building algorithm number. "Evaluate grasp pose algorithm" generates the instruction "EVALUATE_GRASP_POSE:6", where 6 represents the pose evaluation algorithm number.

[0131] The target execution instructions generated for "output the graspable pose parameters of the target (apple)" can be:

[0132] "Encode the pose parameters" assumes that the corresponding instruction is "ENCODE_POSE_PARAMETER:7", where 7 represents the encoding method number. "Send the pose parameters to the robot control system" generates the instruction "SEND_POSE_TO_ARM_CONTROL:8", where 8 represents the sending method number.

[0133] Then the above executions are integrated, and the integrated target execution instructions are expressed as:

[0134] "MOVE_JOINT:A,θ1,v1;MOVE_JOINT:B,θ2,v2;...;START_VISION:0,f1;RU N_ALGORITHM:1;EXTRACT_RGB_FEATURE:2;COMPARE_WITH_APPLE_DB:3;RUN_IMAGE_SEGMENTATION:4;BUILD_3D_MODEL:5;EVALUATE_GR ASP_POSE:6;ENCODE_POSE_PARAMETER:7;SEND_POSE_TO_ARM_CONT ROL:8".

[0135] In one embodiment, the process of mapping execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one to obtain the first-level skill mapping results, and mapping the basic execution skills to at least one underlying control skill one by one to obtain the underlying skill mapping results corresponding to the first-level skill mapping results can also be implemented by other large models or neural network models. This mapping process is simple, and even the neural network model does not require much training data and computing power. In addition, the process of generating target execution instructions for target tasks for intelligent devices based on the underlying skill mapping results can also be implemented by other large models. This embodiment does not limit the specific implementation process.

[0136] In one embodiment, referring to Figure 8 , Figure 8 Schematic diagram of the logic of the control method of the smart device provided in the embodiment of the present application. Figure 8 It can be known that in the technical implementation framework of the control method of the intelligent device of the embodiment of the present application, the top layer uses the "language planning big model" to perform execution meta-action planning, plans the execution action sequence corresponding to the completion of the target task, and realizes the target task decomposition of the target task with sequential execution steps. Next, the first layer is decomposed, and the execution action sequence is mapped in the first layer skill space, and the first layer skill mapping result is obtained as the sequential planning corresponding to the execution action sequence, so that the action instruction corresponds to the basic operation layer of the intelligent device. Next, the first layer skill space is layered again to obtain the second layer skill space. At this time, according to the content of the first layer skill space, the basic execution skills in each first layer skill mapping result are decomposed again, and the bottom-level control skills obtained by the decomposition constitute the bottom-level skill mapping result, so that the action instruction corresponds to the API bottom-level operation layer of the intelligent device, making the action more detailed and closer to the hardware system. Finally, the target execution instruction of the intelligent device for the target task can be generated according to the bottom-level skill mapping result. From Figure 8It can also be seen that the first-level skill space is a coarse-grained stratification, while the second-level skill space is a finer-grained stratification process.

[0137] The technical solution provided in the embodiment of the present application obtains the target task and inputs the target task into the language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes multiple execution meta-actions; then obtains the basic skill sequence of the intelligent device, maps the execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtains the first-level skill mapping result; then obtains the underlying control skills of the intelligent device, maps the basic execution skills to at least one underlying control skill one by one, obtains the underlying skill mapping result corresponding to the first-level skill mapping result, and finally generates the target execution instruction of the intelligent device for the target task according to the underlying skill mapping result. In the embodiment of the present application, the target task is disassembled with the help of the language planning model and subdivided into multiple execution meta-actions. Then these execution meta-actions are mapped to the basic execution skills of the specific intelligent device, thereby obtaining the first-level skill mapping result composed of the basic functions of the intelligent device. Subsequently, the first-level skill mapping result is further decomposed and further refined using the underlying control skills that can call the API control unit. Finally, the target execution instruction adapted to the intelligent device is obtained. Through this progressive disassembly method, the decoupling of the action planning process based on the large model and the specific control process of the intelligent device is achieved. This means that there is no need to rely on complex model architectures to complete the entire planning process from start to finish, and the subsequent mapping and decomposition processes can be implemented based on low-computing matching algorithms. Smart devices usually have limited resources and cannot carry out complex model operations. Through this hierarchical decoupling, smart devices only need to execute simple matching algorithms to receive and execute target execution instructions, while complex task planning is completed by large models in the cloud or on high-performance servers, reducing the requirements for the hardware performance of smart devices.

[0138] This not only simplifies the computing power requirements of smart devices, but also reduces the training cost of large models. The hierarchical decoupling mechanism also reduces the training cost of large language planning models. The large language planning model does not need to consider the specific control details of smart devices, but only focuses on the decomposition and abstract planning of tasks. This makes the training data of the large language planning model more universal and the training process more efficient. At the same time, it reduces the learning burden of the large language planning model on specific information of different smart devices, avoids the increase in training complexity caused by device diversity, and thus reduces training costs.

[0139] In addition, this mechanism of hierarchical decoupling of planning and control enables the control process to have good generalization and adaptability to different intelligent devices. Since the task decomposition and first-level skill mapping of the large model are universal, and the refinement of the underlying control skills can be adjusted for different devices, the same set of planning and control processes can be applied to many different types of intelligent devices, expanding the application scenarios of the control process and improving the versatility and practicality of the system.

[0140] The present application also provides a control device for a smart device, which can implement the control method for the smart device. Fig. 9 , the device comprises:

[0141] Task decomposition module 910: used to obtain the target task, and input the target task into the language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes multiple execution meta-actions.

[0142] The first-level mapping module 920 is used to obtain the basic skill sequence of the smart device, map the execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtain the first-level skill mapping result.

[0143] Bottom-level mapping module 930: used to obtain the bottom-level control skills of the smart device, map the basic execution skills to at least one bottom-level control skill one by one, and obtain the bottom-level skill mapping result corresponding to the first-level skill mapping result, wherein the bottom-level control skill is used to call one or more API control units of the smart device.

[0144] Instruction generation module 940: used to generate target execution instructions for the smart device for the target task based on the underlying skill mapping results.

[0145] The specific implementation of the control device of the smart device in this embodiment is basically the same as the specific implementation of the control method of the smart device described above, and will not be repeated here.

[0146] The present application also provides an electronic device, including:

[0147] at least one memory;

[0148] at least one processor;

[0149] at least one program;

[0150] The program is stored in the memory, and the processor executes the at least one program to implement the control method of the smart device implemented in the present application. The electronic device can be any smart terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.

[0151] See also Fig.10 , Fig.10 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0152] The processor 1001 may be implemented by a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0153] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other applications. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1002, and the processor 1001 calls and executes the control method of the smart device in the embodiments of this application;

[0154] Input / output interface 1003, used to implement information input and output;

[0155] Communication interface 1004, used to realize communication interaction between the device and other devices, which can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and

[0156] A bus 1005 , which transmits information between various components of the device (e.g., the processor 1001 , the memory 1002 , the input / output interface 1003 , and the communication interface 1004 );

[0157] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .

[0158] An embodiment of the present application further provides a storage medium, which is a storage medium storing a computer program, and when the computer program is executed by a processor, the control method of the above-mentioned smart device is implemented.

[0159] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0160] The control method, device, equipment and storage medium of the intelligent device proposed in the embodiment of the present application obtains the target task and inputs the target task into the language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes multiple execution meta-actions; then obtains the basic skill sequence of the intelligent device, maps the execution meta-actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtains the first-level skill mapping result; then obtains the underlying control skills of the intelligent device, maps the basic execution skills to at least one underlying control skill one by one, obtains the underlying skill mapping result corresponding to the first-level skill mapping result, and finally generates the target execution instruction of the intelligent device for the target task according to the underlying skill mapping result. In the embodiment of the present application, the target task is disassembled with the help of the language planning model and subdivided into multiple execution meta-actions. Then these execution meta-actions are mapped to the basic execution skills of the specific intelligent device, thereby obtaining the first-level skill mapping result composed of the basic functions of the intelligent device. Subsequently, the first-level skill mapping result is further decomposed and further refined using the underlying control skills that can call the API control unit. Finally, the target execution instruction adapted to the intelligent device is obtained. Through this progressive disassembly method, the action planning process based on the large model and the specific control process of the smart device can be decoupled. This means that there is no need to rely on a complex model architecture to complete all planning processes from start to finish, and the subsequent mapping and decomposition processes can be implemented based on a low-computing matching algorithm. In this way, it not only simplifies the computing power requirements of smart devices, but also reduces the training cost of large models. In addition, this mechanism of hierarchical decoupling of planning and control enables the control process to have good generalization adaptability to different smart devices, thereby expanding the application scenarios of the control process.

[0161] The embodiments of the features of the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0162] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0163] The device embodiments characterized above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0164] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0165] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used for the order or precedence of the features. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application characterized here can be implemented in an order other than those illustrated or characterized here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0166] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used for the association relationship of feature-associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0167] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments characterized above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0168] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0170] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store programs.

[0171] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A control method for an intelligent device, characterized in that: include: Acquire a target task, and input the target task into a language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes a plurality of execution meta-actions; Acquire a basic skill sequence of the smart device, map the execution meta-actions to corresponding basic execution skills in the basic skill sequence one by one, and obtain a first-level skill mapping result; Acquire the underlying control skills of the smart device, map the basic execution skills to at least one of the underlying control skills one by one, and obtain the underlying skill mapping result corresponding to the first-layer skill mapping result, wherein the underlying control skill is used to call one or more API control units of the smart device; Generate a target execution instruction for the target task for the smart device according to the underlying skill mapping result.

2. The control method of the intelligent device according to claim 1, characterized in that: The step of mapping the execution meta-actions to corresponding basic execution skills in the basic skill sequence one by one to obtain a first-level skill mapping result includes: Acquire functional features and timing information corresponding to the execution meta-actions in the execution action sequence, and acquire at least one skill feature corresponding to each of the basic execution skills; Selecting a target skill feature corresponding to each of the functional features from the skill features, and obtaining a basic target execution skill corresponding to the execution meta-action according to the basic execution skill corresponding to the target skill feature; The first-level skill mapping result is obtained according to the timing information and the basic target execution skill.

3. The control method of the intelligent device according to claim 2, characterized in that: The step of selecting a target skill feature corresponding to each of the functional features from the skill features includes: Calculating a similarity value between the functional feature and each of the skill features; The skill feature with the largest similarity value is selected as the target skill feature.

4. The control method of the intelligent device according to claim 2, characterized in that: The step of comparing the functional features with the skill features and selecting a target skill feature corresponding to each functional feature from the skill features includes: If at least one of the functional features corresponds to one of the skill features, at least one of the functional features is merged to obtain a functional group, and the skill feature is used as the target skill feature corresponding to each of the functional features in the functional group; If the functional feature corresponds to at least one sub-skill feature, obtain the skill feature corresponding to each sub-skill feature, sort the skill features according to the sub-time series information of the sub-skill features, and obtain at least one target skill feature corresponding to the functional feature.

5. The control method of the intelligent device according to claim 2, characterized in that: The obtaining the first-layer skill mapping result according to the timing information and the basic target execution skill includes: Sorting the basic target execution skills according to the time sequence information to obtain an initial sorting result; Acquire input and output data of the basic target execution skill, and if it is determined according to the input and output data that the basic target execution skill includes a pre-action, and the pre-action is another basic target execution skill, adjust the initial sorting result according to the pre-action to obtain a target sorting result; The first-level skill mapping result is obtained based on the target sorting result.

6. The control method of the intelligent device according to claim 5, characterized in that: Before obtaining the first-level skill mapping result based on the target sorting result, the method further includes: If a plurality of said basic target execution skills include a synchronous execution flag; The target sorting result is adjusted according to the synchronous execution identifier.

7. The control method of the intelligent device according to any one of claims 1 to 6, characterized in that: The mapping of the basic execution skills to at least one of the bottom-level control skills one by one to obtain a bottom-level skill mapping result corresponding to the first-level skill mapping result includes: Acquire at least one of the underlying control skills corresponding to each of the basic execution skills; According to the order of the basic execution skills in the first-layer skill mapping result, the underlying skill mapping result is obtained according to the corresponding underlying control skills.

8. A control device for an intelligent device, characterized in that: include: Task decomposition module: used to obtain the target task, and input the target task into the language planning model for task planning to obtain an execution action sequence, wherein the execution action sequence includes multiple execution meta-actions; The first-layer mapping module is used to obtain the basic skill sequence of the smart device, map the execution element actions to the corresponding basic execution skills in the basic skill sequence one by one, and obtain the first-layer skill mapping result; Bottom-layer mapping module: used to obtain the bottom-layer control skills of the smart device, map the basic execution skills to at least one of the bottom-layer control skills one by one, and obtain the bottom-layer skill mapping result corresponding to the first-layer skill mapping result, wherein the bottom-layer control skill is used to call one or more API control units of the smart device; Instruction generation module: used to generate a target execution instruction for the smart device for the target task according to the underlying skill mapping result.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the control method of the smart device according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the control method of the smart device according to any one of claims 1 to 7 is implemented.