Training data generation methods, electronic devices, media, and products

CN122575341APending Publication Date: 2026-08-14MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-08-14

Smart Images

  • Figure CN122575341A_ABST
    Figure CN122575341A_ABST
Patent Text Reader

Abstract

This application provides a method, electronic device, medium, and product for generating training data. The method, applied in the field of smart home control technology, includes: constructing a virtual control scenario and attribute features of a target user based on preset training requirements; determining multiple device control tasks based on the attribute features of the virtual control scenario and the target user; and simulating the multiple device control tasks to generate a dialogue training dataset for training a device control model. This method provides reliable scenario constraints for training data generation by pre-constructing a virtual home scenario, preventing training data from deviating from actual home scenarios. Furthermore, the training data undergoes logical verification during the simulation environment phase, ensuring its rationality and executability. Therefore, this method can significantly reduce the proportion of invalid training data generated, improving the accuracy and reliability of smart device control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home control technology, and more specifically, to a method for generating training data, electronic devices, media, and products in the field of smart home control technology. Background Technology

[0002] Currently, with the widespread application of Artificial Intelligence of Things (AIoT) technology in the smart home industry, users are provided with a convenient, fast, and intelligent way to control devices. The implementation of this intelligent device control method mainly relies on intelligent agents within the smart home scenario.

[0003] Specifically, when a user wants to control a smart device, they can output a corresponding control command (e.g., a voice command). The Agent receives and parses the user's control intention, then generates a reasonable control command to achieve precise control of the smart device.

[0004] Therefore, ensuring the accuracy of the agent's understanding and guaranteeing the quality of training data during the agent training phase has become an urgent problem to be solved. Summary of the Invention

[0005] This application provides a method, electronic device, medium, and product for generating training data. The method provides reliable scene constraints for training data generation by pre-constructing a virtual home scenario, preventing training data from deviating from the actual home environment. Furthermore, the training data undergoes logical verification during the simulation phase, ensuring its rationality and executability. Therefore, this method can significantly reduce the proportion of invalid training data generated, improving the accuracy and reliability of intelligent device control.

[0006] Firstly, a method for generating training data is provided. This method includes: constructing a virtual control scenario and attribute features of a target user based on preset training requirements; determining multiple device control tasks of the target user in the virtual control scenario based on the virtual control scenario and the attribute features of the target user, wherein each device control task includes the target user's control intention in the virtual control scenario, the adjustment action of the device corresponding to the control intention, and the execution verification conditions; simulating the multiple device control tasks to generate a dialogue training dataset, which is used to train a device control model, and the device control model is used to enable the user to control the smart device through control commands.

[0007] In the aforementioned technical solution, this application proposes a method for generating dialogue training data for intelligent device control models. This method first constructs a virtual control scenario and the attribute features of the target user, providing spatial environment constraints and user behavior style constraints for the generation of training samples, thus preventing training data from deviating from real interaction scenarios and user habits. Furthermore, based on the target user's control intentions, device adjustment actions, and device execution verification conditions in the virtual scenario, the device control task is determined and task scenario simulation is performed to obtain dialogue training data. Here, the control intention reflects the user's adjustment needs for the device, the device adjustment actions reflect the specific content that the device needs to adjust, and the execution verification conditions provide verifiable quantitative evaluation criteria for task scenario simulation. This ensures that the generated dialogue training data not only meets the interaction needs of real users but also possesses rationality and executability, providing reliable training samples for subsequent training of the device control model and significantly reducing the proportion of invalid training data generated.

[0008] In conjunction with the first aspect, in some possible implementations, determining multiple device control tasks for the target user within the virtual control scenario, based on the virtual control scenario and the attribute characteristics of the target user, includes: generating descriptive text for N scene segments of the target user within the virtual control scenario, where N is a positive integer greater than or equal to 1, based on the virtual control scenario and the attribute characteristics of the target user; the descriptive text for each scene segment is used to represent the spatiotemporal information associated with the target user's behavioral trajectory, device status, and device adjustment requirements within the virtual control scenario; and determining the multiple device control tasks based on the descriptive text for the N scene segments and the attribute characteristics of the target user.

[0009] In the above technical solution, before generating device control tasks, scene fragments are generated first, which can provide complete spatiotemporal information of the target user's behavioral trajectory in the virtual control scenario. This makes the device control tasks obtained based on the scene fragments more in line with the real user's living habits and adjustment needs, and avoids training data from deviating from the user's real needs.

[0010] Combining the first aspect and the above implementation methods, in some possible implementation methods, the virtual control scenario includes M virtual intelligent devices, where M is a positive integer greater than or equal to 1; determining the multiple device control tasks based on the description text of the N scene fragments and the attribute characteristics of the target user includes: for any scene fragment among the N scene fragments, performing intent parsing on the description text of the scene fragment to obtain the control intent corresponding to the scene fragment; determining the target intelligent device from the M virtual intelligent devices based on the control intent corresponding to the scene fragment, and determining the adjustment action of the target intelligent device; determining the execution verification condition of the target intelligent device based on the attribute characteristics of the target intelligent device and the target user; generating the device control task corresponding to the scene fragment based on the control intent, the target intelligent device, the adjustment action of the target intelligent device, and the execution verification condition of the target intelligent device; and, after the device control tasks corresponding to the N scene fragments have been generated, determining the device control tasks corresponding to the N scene fragments as the multiple device control tasks.

[0011] In the above technical solution, scene fragments are parsed to obtain multiple device control tasks. Since the generation of scene fragments conforms to the virtual control scenario and user behavior habits, the resulting device control tasks are entirely dependent on the actual capabilities of the devices in the virtual control scenario and the user needs within that space, thus avoiding tasks exceeding the device's capabilities and generating invalid training data. Furthermore, the device control tasks include execution verification conditions, which are specifically generated based on the target user's attribute characteristics, ensuring that the generation of device control tasks simultaneously considers the personalized usage needs of different users.

[0012] Combining the first aspect and the above implementation methods, in some possible implementation methods, the task scenario simulation of the multiple device control tasks to generate a dialogue training dataset includes: combining the multiple device control tasks to generate at least one task set, which contains at least one device control task; and performing task scenario simulation on the at least one task set to obtain the dialogue training dataset.

[0013] In the above technical solution, the equipment control tasks are combined to obtain a task set containing at least one equipment control task. This generates a diversified task set from a limited number of equipment control tasks, thereby simulating the continuous behavior of users in the same virtual control scenario and improving the model's ability to understand continuous tasks.

[0014] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, the task scenario simulation of the at least one task set to obtain the dialogue training dataset includes: for any task set in the at least one task set, performing task scenario simulation on the task set and determining the execution result of the task set; if the execution result of the task set is successful, determining all dialogue data of the task set during the task scenario simulation process as at least one round of dialogue training data corresponding to the task set; if the at least one task set has completed the task scenario simulation, generating the dialogue training dataset based on at least one round of dialogue training data corresponding to at least one target task set, wherein the target task set is the task set in the at least one task set whose execution result is successful.

[0015] Combining the first aspect and the above implementation methods, in some possible implementation methods, the task set is simulated to determine the execution result of the task set, including: obtaining the virtual control scenario, the attribute characteristics of the target user, and the execution verification conditions contained in the task set; for the m-th dialogue round of the task set in the task scenario simulation process, the question text for the m-th dialogue round is generated based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; Based on the question text of the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data, the device control command for the m-th dialogue round is determined; the device adjustment command for the m-th dialogue round is used to simulate device actions to obtain the device state for the m-th dialogue round; and based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set, the task simulation result for the m-th dialogue round is determined; if the task simulation result for the m-th dialogue round is successful, the execution result of the task set is determined to be successful.

[0016] In the aforementioned technical solution, when simulating task scenarios for the task set, both the question generation and instruction generation processes incorporate the virtual control scenario, the target user's attribute characteristics, the task set, and historical dialogue data. This ensures that both the question text and device control instructions are strictly constrained by the current virtual control scenario and user profile, guaranteeing that the question text matches the user's speaking style and behavioral habits, and that the device control instructions are generated around the devices in the current virtual control scenario while also taking into account the target user's usage habits. Furthermore, based on the execution verification conditions of the task set, the device state in the m-th dialogue round is verified after the task scenario simulation, ensuring that the training data is executable and conforms to the device adjustment logic, thus guaranteeing the reliability and effectiveness of the training data.

[0017] In combination with the first aspect and the above implementation methods, in some possible implementation methods, determining all dialogue data of the task set during the task scenario simulation process as at least one round of dialogue training data corresponding to the task set includes: generating at least one round of dialogue training data corresponding to the task set based on the historical dialogue data, the question text of the m-th dialogue round, and the device control command of the m-th dialogue round.

[0018] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the generation method further includes: if the task simulation result of the m-th dialogue round is a failure, determining whether the m-th dialogue round is a preset maximum dialogue round; if the m-th dialogue round is not the preset maximum dialogue round, letting m+1=m, and repeatedly executing the generation of the question text for the m-th dialogue round based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; Based on the question text of the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data, the following steps are taken: determining the device control command for the m-th dialogue round; simulating device actions for the device adjustment command of the m-th dialogue round to obtain the device state for the m-th dialogue round; determining the task simulation result for the m-th dialogue round based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set; and determining the execution result of the task set as a failure if the m-th dialogue round is the preset maximum dialogue round.

[0019] Secondly, a training data generation device is provided, comprising: a construction module for constructing a virtual control scenario and attribute features of a target user based on preset training requirements; a determination module for determining multiple device control tasks of the target user in the virtual control scenario based on the virtual control scenario and the attribute features of the target user, wherein the device control tasks include the control intention of the target user in the virtual control scenario, the adjustment actions of the device corresponding to the control intention, and the execution verification conditions; and a generation module for simulating the multiple device control tasks to generate a dialogue training dataset, wherein the dialogue training dataset is used to train a device control model, and the device control model is used to enable the user to control the smart device through control commands.

[0020] In conjunction with the second aspect, in some possible implementations, the determining module is specifically used to: generate descriptive text for N scene segments of the target user in the virtual control scenario based on the virtual control scenario and the attribute characteristics of the target user, wherein the descriptive text of the scene segments is used to represent the spatiotemporal information, device status and device adjustment requirements associated with the behavioral trajectory of the target user in the virtual control scenario; and determine the multiple device control tasks based on the descriptive text of the N scene segments and the attribute characteristics of the target user.

[0021] Combining the second aspect and the above implementation methods, in some possible implementation methods, the virtual control scenario includes M virtual intelligent devices, where M is a positive integer greater than or equal to 1; the determining module is specifically used for: for any scene segment among the N scene segments, performing intent parsing on the description text of the scene segment to obtain the control intent corresponding to the scene segment; determining the target intelligent device from the M virtual intelligent devices based on the control intent corresponding to the scene segment, and determining the adjustment action of the target intelligent device; determining the execution verification condition of the target intelligent device based on the attribute characteristics of the target intelligent device and the target user; generating the device control task corresponding to the scene segment based on the control intent, the target intelligent device, the adjustment action of the target intelligent device, and the execution verification condition of the target intelligent device; and, after the device control tasks corresponding to the N scene segments have been generated, determining the device control tasks corresponding to the N scene segments as the plurality of device control tasks.

[0022] In combination with the second aspect and the above implementation methods, in some possible implementation methods, the generation module is specifically used to: combine the multiple device control tasks to generate at least one task set, which contains at least one device control task; and perform task scenario simulation on the at least one task set to obtain the dialogue training dataset.

[0023] In conjunction with the second aspect and the above implementation methods, in some possible implementation methods, the generation module is specifically used for: for any task set in the at least one task set, performing task scenario simulation on the task set and determining the execution result of the task set; if the execution result of the task set is successful, determining all dialogue data of the task set during the task scenario simulation process as at least one round of dialogue training data corresponding to the task set; if the at least one task set has completed all task scenario simulations, generating the dialogue training dataset based on at least one round of dialogue training data corresponding to at least one target task set, wherein the target task set is the task set in the at least one task set whose execution result is successful.

[0024] In conjunction with the second aspect and the above implementation methods, in some possible implementation methods, the generation module is specifically used to: obtain the virtual control scenario, the attribute characteristics of the target user, and the execution verification conditions contained in the task set; for the m-th dialogue round of the task set in the task scenario simulation process, generate the question text for the m-th dialogue round based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; determine the device control command for the m-th dialogue round based on the question text for the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data; simulate device actions for the device adjustment command for the m-th dialogue round to obtain the device state for the m-th dialogue round, and determine the task simulation result for the m-th dialogue round based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set; if the task simulation result for the m-th dialogue round is successful, determine that the execution result of the task set is successful.

[0025] In combination with the second aspect and the above implementation methods, in some possible implementation methods, the generation module is specifically used to: generate at least one round of dialogue training data corresponding to the task set based on the historical dialogue data, the question text of the m-th dialogue round, and the device control command of the m-th dialogue round.

[0026] In conjunction with the second aspect and the above implementation methods, in some possible implementation methods, the generation module is further used to: if the task simulation result of the m-th dialogue round is a failure, determine whether the m-th dialogue round is the preset maximum dialogue round; if the m-th dialogue round is not the preset maximum dialogue round, let m+1=m, and repeatedly execute the generation of the question text for the m-th dialogue round based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; Based on the question text of the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data, the following steps are taken: determining the device control command for the m-th dialogue round; simulating device actions for the device adjustment command of the m-th dialogue round to obtain the device state for the m-th dialogue round; determining the task simulation result for the m-th dialogue round based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set; and determining the execution result of the task set as a failure if the m-th dialogue round is the preset maximum dialogue round.

[0027] Thirdly, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to perform the generation method in the first aspect or any possible implementation thereof.

[0028] Fourthly, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to execute the generation method described in the first aspect or any possible implementation thereof.

[0029] Fifthly, a computer-readable storage medium is provided that stores computer program code, which, when run on a computer, causes the computer to perform the generation method described in the first aspect or any possible implementation thereof. Attached Figure Description

[0030] Figure 1 This is a structural diagram of an intelligent device control system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a training data generation system provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the implementation of a training data generation method provided in an embodiment of this application. Figure 4 This is a schematic flowchart of a training data generation method provided in an embodiment of this application; Figure 5 This is a schematic diagram of a layered system architecture provided in an embodiment of this application; Figure 6 This is a schematic flowchart illustrating another method for generating training data provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a training data generation device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0032] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0033] Before introducing the methods of the embodiments of this application, the technical terms that may be involved in the embodiments of this application will be explained first.

[0034] AIoT refers to the integration of artificial intelligence technology into the Internet of Things (IoT) system, enabling IoT devices to not only connect and collect data, but also possess intelligent sensing, autonomous decision-making, learning, and adaptive capabilities.

[0035] Large models refer to large-scale parameter deep learning models pre-trained based on massive amounts of text and multimodal data, which have powerful capabilities in natural language understanding, semantic reasoning, text generation, intent recognition, and contextual multi-turn dialogue.

[0036] Agent: Refers to a computer entity that can perceive its environment, make autonomous decisions, and execute actions to achieve specific goals. With a large model as its core, an agent is essentially an intelligent software entity capable of perceiving its environment, making decisions, and executing actions to achieve specific objectives.

[0037] After explaining the technical terms used in the embodiments of this application, the application scenarios of the embodiments of this application are described below.

[0038] Currently, with the widespread application of Artificial Intelligence of Things (AIoT) technology in the smart home industry, various smart devices based on IoT and AI technologies are gradually changing and updating people's lifestyles and quality of life, providing users with a convenient, fast, and intelligent way to control devices. For example, common smart devices in smart home scenarios include, but are not limited to, smart speakers, smart air conditioners, smart refrigerators, and smart curtains.

[0039] The aforementioned smart devices, such as smart speakers, smart air conditioners, smart refrigerators, and smart curtains, differ primarily from traditional home devices in their interactivity. Specifically, traditional devices only support manual operation via remote control or physical buttons, lacking network connectivity. Smart devices, however, leverage network connectivity to integrate into the Internet of Things (IoT) ecosystem and support contactless control methods such as voice commands, mobile applications (APPs), or central control screens. They become a key component of smart home control systems, possessing proactive responsiveness, perception, and decision-making capabilities.

[0040] It should be understood that the above-mentioned user control process of smart devices mainly relies on the Agent in the smart home scenario. The core component of the Agent is a large model, which is used to receive the control commands output by the user (e.g., voice commands), parse the user's control intention, and then generate reasonable control commands to achieve precise control of smart devices.

[0041] The following will combine Figure 1 This application provides a detailed description of an exemplary agent-based smart device control system.

[0042] Figure 1 This is a structural diagram of an intelligent device control system provided in an embodiment of this application.

[0043] For example, such as Figure 1 As shown, the intelligent device control system (hereinafter referred to as the "control system") 100 mainly consists of the following parts: voice commands 101, dialogue terminal 102, and multiple intelligent devices supporting voice control, such as... Figure 1 The system includes smart lights 103, smart curtains 104, smart air conditioners 105, and smart refrigerators 106. The main functions of each component are as follows: Voice command 101 is the user's voice input via natural language. It is the original interaction entry point of the control system 100 and is used to carry the user's control intentions and needs. For example, voice command 101 can be "Turn on the master bedroom light for me" or "I'm a little hot, adjust the temperature for me".

[0044] As the name suggests, the dialogue terminal 102 refers to a terminal device used for voice dialogue or voice interaction with a user, responsible for receiving voice commands 101 issued by the user. Optionally, such as Figure 1 As shown, the dialogue terminal 102 can be a smart speaker, or other types of devices, such as a central control screen, a user's smartphone, a smart tablet, or a wearable device. The following example illustrates the dialogue terminal 102 as a smart speaker. Depending on the computing power of the dialogue terminal 102, its role in the control system 100 will also vary.

[0045] Optionally, when the computing power of the dialogue terminal 102 is limited, the dialogue terminal 102 acts as a "voice relay station" in the control system 100. Specifically, the dialogue terminal 102 communicates with the smart home cloud platform (… Figure 1 (Not shown in the image) A communication connection is established, responsible for receiving and recognizing voice commands 101, obtaining the recognized text, and sending the recognized text to the cloud platform for processing. At this time, the cloud platform is equivalent to the core processing hub, on which an Agent is deployed as the processing terminal for voice commands 101.

[0046] When the dialogue terminal 102 has sufficient computing power, it functions as a core processing hub, on which an Agent is deployed as the processing terminal for voice commands 101. Regardless of whether the processing terminal is the dialogue terminal 102 or a cloud platform, it can communicate and connect with various controlled smart devices. The processing terminal receives and processes the recognized text through the Agent to obtain specific control commands, which are then sent to the corresponding smart devices, such as… Figure 1 The smart air conditioner 105.

[0047] The entire control process of the above-mentioned control system 100 can be summarized as follows: user outputs voice command 101 → dialogue terminal 102 receives voice command 101 → processing terminal calls Agent (specifically large model) to process voice command 101 to obtain device control command → Agent sends device control command 101 to smart device.

[0048] Therefore, it is evident that Agen's understanding of voice control commands directly determines the reliability of its output device control commands, and thus the accuracy of intelligent device control, throughout the entire control process.

[0049] Therefore, in order to ensure that the Agent has reliable intent understanding capabilities in practical applications, how to construct high-quality and highly realistic training data (i.e., training dialogue data) during the Agent training phase (i.e., the large model training phase) has become an urgent problem to be solved.

[0050] There are two main approaches used in related technologies to solve the above problems. The two approaches and their respective drawbacks are introduced below.

[0051] The first method: generating training data based on templates or rules.

[0052] Specifically, this approach typically begins by defining a template for generating training data. For example, the template could be "Action" + "Device" + "Parameter". Next, slots are defined; for instance, the action slot could be "on", "off", "raise", "lower", etc.; the device slot could be "living room light", "master bedroom curtains", etc.; and the parameter slot could be "brightness", "temperature", "humidity", etc. Finally, a complete voice command is obtained by filling in the slots. Keywords are then recognized from the voice command, and mapped to the corresponding slots to obtain device control commands. These voice commands and device control commands together form the training dialogue data.

[0053] The problem with the above generation method is that the generation process lacks awareness of the real physical environment, causing the training data to deviate from real-life home scenarios and resulting in the large model outputting invalid commands. Training the large model with such invalid commands leads to the model learning incorrect control logic, reducing the reliability and accuracy of control in real-world home scenarios. For example, if the generated voice command is "turn on the living room light," the resulting device control command will be for the living room light. Due to the lack of constraints from the real physical environment, training the large model with this data will result in the model outputting the living room light control command even if a living room light does not exist in a real home scenario. This renders the control command unexecutable, reducing the accuracy of device control.

[0054] The second method: generate training data using a large model.

[0055] Specifically, this approach allows a large model to generate a large number of voice commands and corresponding device control commands based on prior knowledge learned through pre-training.

[0056] The problems with the above generation method are similar to those of the first method, namely, the lack of constraints from the real physical environment during the generation process. Because large models are unaware of the device states in the real physical environment during generation, they may generate training data that cannot be executed in a real physical environment.

[0057] In view of this, embodiments of this application propose a method for generating training data. This method can provide reliable scene constraints for the generation of training data by pre-constructing a virtual home scene, thus preventing the training data from deviating from the actual home scene. Furthermore, the training data undergoes logical verification in advance during the simulation environment phase, ensuring the rationality and executability of the training data. Therefore, this method can significantly reduce the proportion of invalid training data generated, improving the accuracy and reliability of intelligent device control.

[0058] It should be understood that the above application scenarios are illustrated using a smart home scenario as an example, but the application scenarios of this application embodiment are not limited to smart home scenarios. That is to say, the Agent trained using training data in this application embodiment is not limited to controlling smart devices in a smart home scenario, but can also be other smart devices that can be controlled through control commands, such as smart health devices, smart office equipment, smart building equipment, and other IoT devices. The following description of this application embodiment will use smart devices in a smart home scenario as an example.

[0059] After introducing the application scenarios of the embodiments of this application, the control method of the embodiments of this application will be introduced below.

[0060] It should be understood that the training data generation method provided in this application embodiment is applied to a smart terminal (i.e., an electronic device), which is used to generate training data. Optionally, the smart terminal can be a cloud platform server or a local smart device (e.g., a computer device), and this application embodiment does not limit this. When implementing this method, the smart terminal mainly relies on a training data generation system deployed on it. The following will first demonstrate... Figure 2 The structure and working principle of this generation system are introduced.

[0061] Figure 2 This is a schematic diagram of the structure of a training data generation system provided in an embodiment of this application.

[0062] For example, such as Figure 2 As shown, the training data generation system (hereinafter referred to as the "generation system") 200 is divided into two functional modules, specifically including two main generation modules: the task generation module and the dialogue generation module. The functions of each generation module are as follows: The task generation module is used to build the basic background environment for training data generation, construct virtual home scenarios and virtual users, and further determine various device control tasks that virtual users may need in this virtual home scenario, so that the device control behavior in the virtual home scenario is closer to the device control logic in the real home scenario.

[0063] It should be understood that the aforementioned virtual home scenarios are not real-world physical (or physical) home scenarios, but rather digital virtual home environments constructed through software simulation. Correspondingly, the virtual user is an interactive subject generated by software simulation, capable of mimicking the language style, usage habits, and behavioral patterns of a real user.

[0064] This application's embodiments generate training data by constructing virtual home scenarios and virtual users. This is because the types, number, and layout of smart devices in physical home scenarios are very limited, making it difficult to cover the various home scenarios required for large-scale model training and providing sufficiently rich environmental constraint data. Furthermore, collecting training data in real home scenarios requires large-scale collection of user interaction behavior and device operation data, resulting in high collection costs, long cycles, and overall low efficiency.

[0065] This application, by independently constructing virtual home scenes and virtual users to mimic user interaction behaviors and device control actions, can quickly generate simulated home environments with various house structures, device configurations, and device states in batches. It can cover diverse home layouts, smart device combinations, and different user habits. It can provide sufficient and scenario-rich environmental constraint data for large model training without relying on real users and physical devices for data collection, significantly reducing the cost of generating training data and shortening the overall cycle of data generation and model training.

[0066] Based on this, the task generation module specifically includes a scene generator (Home maker), a user generator (Usermaker), and a task generator.

[0067] The scene generator is used to construct virtual home scenes, providing physical constraints for dialogue generation. The virtual home scene includes the spatial layout of the virtual home, device (i.e., virtual device) configuration information, and the initial state of each device.

[0068] The user generator is used to generate profile information of virtual users, including their language style, lifestyle habits, and usage preferences.

[0069] The task generator is used to generate multiple device control tasks for virtual users in a virtual home scenario, in order to simulate the real control needs of users in a real home environment.

[0070] The dialogue generation module is used to generate a complete dialogue and instruction loop based on the results generated by the task generation module, thus obtaining the final dialogue training data.

[0071] The dialogue generation module specifically includes a question generator, an instruction generator, and a task scenario simulator.

[0072] The question generator is used to generate natural language questions that match the user profile based on the multi-task sequence obtained by the task generation module, and input them as input-side samples to the instruction generator.

[0073] The instruction generator is used to receive user queries and generate corresponding device control instructions, which are then input into the task scenario simulator.

[0074] The task scenario simulator is used to simulate and execute the simulation results and verify their executability, and to provide feedback on the execution results and update the device status, ensuring that the generated training data is authentic and effective.

[0075] Next, we will proceed through... Figure 3 This paper introduces the specific process of the training data generation system in implementing the training data generation method.

[0076] Figure 3 This is a flowchart illustrating the implementation of a training data generation method provided in an embodiment of this application.

[0077] For example, such as Figure 3 As shown, the training data generation method can be divided into two stages in its implementation: the first stage is the task generation stage, and the second stage is the dialogue generation stage. The task generation stage specifically consists of... Figure 2 The task generation module is implemented in the middle, and the dialogue generation stage is specifically handled by Figure 2 The dialogue generation module is implemented in [the framework]. The following section combines [the implementation details]. Figure 2 The specific implementation steps for these two stages will be introduced separately.

[0078] Phase 1: Task Generation Phase.

[0079] During the task generation phase, the scene generator generates a virtual home scene, specifically by generating the basic metadata within the virtual home scene. Figure 3 The scene metadata shown defines the physical rules of the entire virtual home scene. Metadata refers to the set of instantiated parameters output by the generator, which characterizes the inherent attributes, configuration parameters, and behavioral rules of virtual entities.

[0080] For example, such as Figure 3 As shown, the scene metadata includes three categories: device manuals, tool manuals, and home information. Device manuals define the callable interfaces of smart devices in the virtual home scene, such as turning on lights, playing media, and adjusting the air conditioner temperature; these serve as the basis for generating subsequent device control commands. Tool manuals define callable tools that are not device-related, such as calculators, weather forecasts, and code execution tools. Home information defines the virtual house structure, indoor environmental parameters (e.g., device list, indoor temperature, indoor humidity), and outdoor environmental information, providing a complete spatial and environmental background for the virtual home scene. The device list may include the device's Identity Document (ID), device location (e.g., the room where the device is located), preset parameter ranges for the device, the device's current operating status, and device type (e.g., lights, air conditioner, smart TV, etc.).

[0081] The user generator is used to generate profile information of virtual users, specifically to generate user profile metadata, which is used to define the behavior and preference rules of virtual users.

[0082] For example, such as Figure 3 As shown, the user profile metadata includes the virtual user's basic information, physiological characteristics, health information, lifestyle habits, and preference settings. Basic information defines the virtual user's fundamental identity characteristics, such as name, age, and gender. Physiological characteristics define the user's physiological data, such as height, weight, and physical fitness. Health information defines the user's health data, such as major illnesses, chronic diseases, and sleep quality, which influence the user's need for control over the environment and devices. Lifestyle habits define the user's daily behavioral patterns, such as wake-up time on weekdays, lunch break duration, and shower time. Preference settings define the user's personalized comfort requirements for the home environment, such as temperature preferences, humidity preferences, and cooking environment humidity preferences.

[0083] See Figure 2 The task generation module includes a task generation and verification unit. The task generation and verification unit includes, in addition to... Figure 2 In addition to the task generator, it also includes a task validator.

[0084] like Figure 3 As shown, scene metadata and user profile metadata are input into the task generator in the task generation and verification unit to generate multiple device control tasks for the virtual user in a virtual home scenario. The task verifier is used to verify and correct the multiple device control tasks generated by the task generator to ensure the rationality of the generated device control tasks.

[0085] Specifically, the task verifier verifies the following aspects, including but not limited to: whether the device configuration in the virtual home scenario matches the virtual user, whether the device control actions are within a reasonable range, and whether the device control tasks are consistent with the virtual user's preferences.

[0086] After successful verification, the task generation and verification unit outputs the final multiple device control tasks, and simultaneously outputs scene metadata and user profile metadata. These three types of information constitute the scene blueprint of the virtual home scene. In this embodiment, the scene blueprint refers to a structured digital description of the virtual home scene, the virtual user profile, and the device control tasks. Essentially, it is a complete task description text for the virtual user in the virtual home, providing environmental constraints and character behavior constraints for subsequent dialogues and commands.

[0087] Phase Two: Dialogue Generation Phase.

[0088] Specifically, in the dialogue generation stage, in order to ensure the diversity of questions asked by virtual users, the question type generator first generates question metadata so that the question generator can generate question text in different styles.

[0089] For example, such as Figure 3 As shown, the question metadata includes question style, question environment, and question device type. Question style defines the clarity of intent expressed by the virtual user when asking a question, such as vague intent, missing parameters, or explicit intent, to adapt to different users' expression habits. Question environment defines the external environment state when the user expresses their question, such as quiet environment, noisy environment, etc., simulating a real home conversation scenario. Question device type defines the device control range and quantity corresponding to the command, such as single device, multiple devices, batch devices, etc., covering various control needs from simple to complex.

[0090] The question generator uses the question style, question environment, and question device type in the question metadata, combined with the current device control task, to generate the user's question text through simulation and input it into the command generator.

[0091] In addition to receiving the user's question text, the instruction generator also receives the scene blueprint and the historical dialogue of the current device control task during the scene simulation process (the simulation process of a device control task may involve multiple rounds of dialogue), generates the corresponding device control instructions, and sends them to the task scene simulator for simulation execution.

[0092] The task scenario simulator simulates and controls virtual smart devices in a virtual home scenario to complete device actions according to device control commands. It then returns the simulated device status to the command generator, which determines whether the current device control task was successfully executed. After successful execution, the dialogue generation module integrates all historical dialogue data and various metadata from the multi-turn dialogue process, ultimately generating dialogue training data containing complete interaction information—that is, dialogue training data under the current scenario blueprint.

[0093] In passing Figure 2 and Figure 3 After the general implementation flow of the method in the embodiments of this application has been described, the following will be conducted through... Figure 4 The detailed implementation process of a training data generation method provided in the embodiments of this application is described below.

[0094] Figure 4 This is a schematic flowchart of a training data generation method provided in an embodiment of this application.

[0095] For example, such as Figure 4 As shown, the generation method 400 includes the following steps 401 to 403.

[0096] Step 401: Based on the preset training requirements, construct the attribute characteristics of the virtual control scenario and the target user.

[0097] Among them, virtual control scenarios refer to digital virtual space environments constructed through software simulation for the control of intelligent devices.

[0098] Optionally, virtual control scenarios include, but are not limited to, virtual home scenarios, virtual building scenarios, and virtual office scenarios, each corresponding one-to-one with a real control scenario. For example, when the real control scenario is a smart home scenario, the aforementioned virtual control scenario is a virtual home scenario; when the real control scenario is a smart building scenario, the virtual control scenario is a virtual building scenario. In the following embodiments of this application, the virtual control scenario is illustrated using a virtual home scenario.

[0099] It should be understood that, as described above, in the process of generating training data in this application embodiment, in order to ensure that the generated training data is closer to the real home scenario, it is necessary to first construct the attribute characteristics of the virtual control scenario and the target user based on preset training requirements. The attribute characteristics of the target user are equivalent to the virtual user profile data mentioned above.

[0100] Pre-defined training requirements refer to the constraints on training data generation set before constructing virtual control scenarios and virtual users, based on the training objectives and data coverage requirements of the large model. For example, pre-defined training requirements may include specific family room structures and home appliances for which the training data is targeted, or specific users for which the training data is targeted.

[0101] See Figure 2 and Figure 3 When a smart terminal constructs a virtual control scene based on preset training requirements, it first generates various instantiated scene metadata through a scene generator, and then assembles these scene metadata to obtain a complete virtual home scene. The virtual home scene is generated in a semi-automatic manner. Semi-automatic means that constraints are preset by the user, and the device intelligently and automatically fills in the specific parameters. This method combines both manual rules and AI-generated automatic generation.

[0102] For example, in generating virtual home scenes, AI can construct virtual spatial structures based on preset training requirements, automatically assign virtual devices to corresponding rooms, and generate the current and environmental states of each virtual device based on common sense about home life. Users can pre-set conditions such as scene type (e.g., home scene), apartment layout, range of device types, and seasonal environmental constraints. Specifically, the apartment layout of the virtual home scene is derived from a large number of common apartment layouts collected from real-life scenarios, such as one-bedroom apartments, two-bedroom apartments, etc. Under these constraints, the scene generator automatically generates multiple types of instantiated scene metadata, including outdoor environmental parameters, indoor temperature, humidity, and lighting parameters, room distribution information, device list, and functional interfaces. Subsequently, the task generation module integrates these fragmented scene metadata, associating and binding device interfaces with room locations and environmental parameters to form a virtual home scene where space, devices, and environmental states are matched.

[0103] Similarly, when constructing the attribute features of a target user based on preset training requirements, the smart terminal first generates various instantiated user metadata through a user generator, and then assembles these user metadata to obtain the complete attribute features of the target user. Specifically, the smart terminal can automatically generate multiple types of instantiated user metadata according to the specific group of people in the preset training requirements and a large amount of user feature data from pre-collected real-life scenarios. This includes user basic information (name, age, gender, etc.), physiological characteristics (height, weight, physical fitness level), health status (underlying diseases, sleep quality, etc.), lifestyle habits (work and rest time, daily behavior patterns), and environmental preferences (temperature and humidity preferences, brightness preferences), etc. The task generation module then integrates the fragmented user metadata to obtain the attribute features of the target user, in order to construct and depict a complete character image in the virtual family scene.

[0104] Step 402: Based on the virtual control scenario and the attribute characteristics of the target user, determine multiple device control tasks of the target user in the virtual control scenario. The device control tasks include the target user's control intention in the virtual control scenario, the adjustment actions of the device corresponding to the control intention, and the execution verification conditions.

[0105] After obtaining the profile data of the virtual home scene and the virtual user through step 401, the smart terminal needs to further associate the virtual user and the virtual home scene to determine the multiple device control tasks that the virtual user may need in the virtual home scene, thereby simulating the various control intentions that the virtual user may have in the virtual home scene.

[0106] In this application embodiment, the device control task (Task) refers to the executable instruction execution target generated by the virtual user in a virtual home scenario based on user attribute characteristics and scenario environment status.

[0107] In this embodiment of the application, the device control task is specifically implemented through a dialogue interaction process. Therefore, the device control task can also be understood as the device interaction task.

[0108] It should be understood that a device control task can be considered the smallest unit of execution, and a device control task typically corresponds to a user's control intent. Specifically, a device control task includes the target user's control intent in a virtual home scenario, the device's adjustment action (adjustment effect) corresponding to the control intent, and execution verification conditions. The adjustment action or effect reflects the target state that the device needs to be adjusted to in order to achieve the control intent. The execution verification conditions reflect the adjustment rules that the device needs to follow in achieving the control intent.

[0109] Combination Figure 3 As can be seen, multiple device control tasks are part of the scenario blueprint, specifically generated by the task generation and verification unit.

[0110] The following details the specific generation process of multiple device control tasks.

[0111] In one possible implementation, based on the virtual control scenario and the attribute characteristics of the target user, multiple device control tasks for the target user in the virtual control scenario are determined, including: Based on the attributes and characteristics of the virtual control scenario and the target user, generate descriptive text for N scenario segments of the target user in the virtual control scenario. The descriptive text of the scenario segments is used to represent the spatiotemporal information, device status and device adjustment requirements associated with the target user's behavior trajectory in the virtual control scenario. Based on the descriptive text of N scene fragments and the attribute characteristics of the target user, multiple device control tasks are determined.

[0112] Specifically, the generation of multiple device control tasks includes two stages: The first stage involves obtaining descriptive text for N scene fragments of the target user within the virtual control scenario, based on the virtual control scenario and the attribute characteristics of the target user. The second stage involves determining multiple device control tasks based on the descriptive text of the N scene fragments and the attribute characteristics of the target user. This step essentially involves decomposing or extracting executable device control tasks from the scene fragments.

[0113] The descriptive text of the aforementioned scene fragments represents the spatiotemporal information, device status, and device adjustment needs associated with the target user's behavioral trajectory within a virtual control scenario. Essentially, it is a natural language description of the target user's behavioral process within the virtual control scenario. Spatiotemporal information includes time and space information (e.g., the room where the device is located). A scene fragment can be understood as a segment describing a user's life scenario. For example, the descriptive text of a scene fragment could be: when, where, and what the target user A is doing, there is a certain device adjustment need.

[0114] It should be understood that the reason for generating scene fragments before generating device control tasks is that scene fragments can describe the target user's state, environmental changes, and user needs and motivations as story fragments that are closer to real life, providing a trigger and reasonable basis for device control tasks, thereby ensuring that device control tasks conform to the behavioral logic of real users.

[0115] When generating descriptive text for N scene fragments, the smart terminal can use the spatial structure, device configuration, and environmental status of the virtual home scene as a basis, driven by the attribute characteristics of the target user, and combined with the user behavior logic in a large number of real-life scenarios to generate multiple scene fragments that the target user may experience in the current virtual home scene.

[0116] Specifically, during the generation process, the smart terminal can set reasonable behavioral paths for users based on the spatial layout and device distribution of the virtual home scene. For example, the behavioral path could be a continuous path from the reading room to the bedroom and then to the kitchen. Simultaneously, combining the initial device states and environmental parameters, it constructs trigger scenarios where the user discovers a mismatch between the device state and the environment under a certain behavior, such as the lights being turned off, making the room too dark, or the air conditioning temperature being too high, affecting the afternoon nap experience. Building on this, based on the target user's age, physiological characteristics, and lifestyle preferences, it matches the target user's behavior with device adjustment motivations consistent with the user's identity. For example, an elderly user might want to brighten the reading light due to declining eyesight, or an obese user might want to lower the air conditioning temperature because it's too hot during an afternoon nap. Finally, by combining the scene, devices, environment, user behavior, and common sense from real life, the smart terminal ultimately generates a series of coherent, natural, and realistic descriptive text fragments of the scene.

[0117] For example, the illustrative text describing the N scene fragments in this application embodiment can be as follows: Name: Scene Fragment Description Description: The description of scene fragment 1: In the morning, Wang Zhicheng went into the reading room to read the newspaper and practice calligraphy. He found that the lights in the reading room were off and the room was too dark, which was not conducive to reading for the elderly. He hoped that the brightness of the reading light could be adjusted to a comfortable level.

[0118] The description of scene fragment 2: After reading, he went to the master bedroom to prepare for a nap. The air conditioner in the master bedroom was running, but the target temperature was set to 30℃, which was too hot and did not meet his nap preference. He hoped to adjust the target temperature to be cooler so that he could rest better.

[0119] The description of scene fragment 3: After lunch break, I got up to go to the kitchen to heat up milk. When I passed through the corridor, I found that the brightness of the corridor fan light was only 5%, which was not safe to walk on. I hope that the corridor light can be turned up so that I can see the road clearly.

[0120] After obtaining the descriptive text of the above N scene fragments, the smart terminal can obtain multiple device control tasks by breaking down the tasks.

[0121] In the above technical solution, before generating device control tasks, scene fragments are generated first, which can provide complete spatiotemporal information of the target user's behavioral trajectory in the virtual control scenario. This makes the device control tasks obtained based on the scene fragments more in line with the real user's living habits and adjustment needs, and avoids training data from deviating from the user's real needs.

[0122] Based on the descriptive text of N scene fragments and the attribute characteristics of the target user, the specific process for determining multiple device control tasks is as follows.

[0123] In one possible implementation, the virtual control scenario includes M virtual intelligent devices, where M is a positive integer greater than or equal to 1; based on the descriptive text of N scenario fragments and the attribute characteristics of the target user, multiple device control tasks are determined, including: For any scene segment among N scene segments, perform intent parsing on the description text of the scene segment to obtain the control intent corresponding to the scene segment; Based on the control intent corresponding to the scene segment, identify the target intelligent device from M virtual intelligent devices and determine the adjustment action of the target intelligent device; Based on the attribute characteristics of the target smart device and the target user, determine the execution verification conditions of the target smart device; Based on the control intent, the target intelligent device, the adjustment actions of the target intelligent device, and the execution verification conditions of the target intelligent device, generate the device control task corresponding to the scene segment; Once the device control tasks corresponding to N scene segments have been generated, the device control tasks corresponding to the N scene segments will be determined as multiple device control tasks.

[0124] Based on the aforementioned construction process of the virtual control scenario, it can be seen that the virtual control scenario includes M predefined virtual intelligent devices, and each of the M virtual intelligent devices has its own device location (i.e., room) within the virtual control scenario. Furthermore, the device control task includes control intent, device adjustment actions, and execution verification conditions. The adjustment action can be understood as the trend of device state adjustment, such as raising or lowering the air conditioner temperature, raising or lowering the volume, increasing or decreasing the brightness, etc. The execution verification conditions can be understood as the parameter constraints that the device can adjust, facilitating logical verification during subsequent task scenario simulation. To ensure the accuracy of the device control task execution, the intelligent terminal needs to first determine the target intelligent device to be controlled based on the descriptive text of the scenario fragment.

[0125] For any scene segment among N scene segments, the smart terminal can obtain the corresponding control intent by parsing the descriptive text of the scene segment. Optionally, the intent recognition model can be a Natural Language Understanding (NLU) model.

[0126] Specifically, the smart terminal uses an intent recognition model to perform word segmentation, entity recognition, and semantic understanding on the descriptive text of scene fragments, obtaining multiple keywords. These keywords include location keywords, user action keywords, device keywords, and user need keywords. For example, the descriptive text of scene fragment 1 is: In the morning, Wang Zhicheng entered the reading room to read the newspaper and practice calligraphy, only to find that the lights in the reading room were off, making the room too dim for elderly people to read. He hoped to adjust the reading light brightness to a comfortable level. The intent recognition model, through semantic understanding, obtains the location keyword "reading room," the user action keyword "reading," the device keyword "light," and the user need keyword "adjust to a comfortable level."

[0127] Based on the various keywords identified above, the user's control intent is extracted and constructed. Taking the description text of scenario fragment 1 above as an example, the user's control intent is "to brighten the light in the reading room." Furthermore, based on the device and location keywords contained in the control intent, the smart terminal can identify the target smart device as "the light in the reading room" from M virtual smart devices. Simultaneously, based on the user's need keywords, the smart terminal can determine the adjustment action of the target smart device as "to adjust to a comfortable brightness level." It is evident that the difference between the control intent and the adjustment action of the target smart device is that the control intent only represents the user's need and adjustment purpose, without involving the specific degree and effect of adjusting the smart device, while the adjustment action of the target smart device clearly defines the effect that the target smart device can achieve after adjustment.

[0128] In addition, in order to ensure that the adjustment of the target smart device meets the user's personalized adjustment needs, the smart terminal can determine the execution verification conditions of the target smart device based on the attribute characteristics of the target smart device and the target user, which can be used as verifiable logical conditions during subsequent task simulation.

[0129] Specifically, after obtaining the target smart device, the smart terminal can first acquire the legal parameter range supported by the target smart device based on a virtual home scenario. Based on this, the smart terminal extracts the target user's physiological characteristics, basic information, health information, and preference settings from the target user's attribute characteristics. Further, using the legal parameter range of the target smart device as a hard constraint, the extracted target user's physiological characteristics, basic information, health information, and preference settings are mapped to parameters, mapping personalized user information to a standard device parameter range. This results in execution verification conditions that satisfy both the legal parameter range constraint and the user's adjustment preferences.

[0130] Therefore, based on the control intent, the target intelligent device, the adjustment actions of the target intelligent device, and the execution verification conditions, the intelligent terminal can generate the device control task for the current scene segment.

[0131] The descriptions of the other two scene segments in the aforementioned illustrative text can also be used to obtain their corresponding device control tasks in the same way. The device control tasks corresponding to the N scene segments constitute multiple device control tasks.

[0132] For example, illustrative text representing multiple device control tasks is as follows: Task list: [3 items] 0: { Control Intent: Increase the brightness of the reading room lights Verification conditions: The lights in the reading room are on, and the brightness is between 90% and 94%. Equipment adjustment: Adjust the brightness of the reading room lights to a comfortable, slightly bright level. Quantity of equipment: [1 item] Device ID: 0x1 ] Task ID: 1 } 1: { Control Intent: Lower the target temperature of the master bedroom air conditioner. Verification conditions: The temperature range of the master bedroom air conditioner is 20~22℃. Equipment adjustment: Lower the target temperature of the master bedroom air conditioner to a temperature suitable for a midday nap. Quantity of equipment: [1 item] Device ID: 0x2 ] Task ID: 2 } 2: { Control Intent: Increase the brightness of the corridor lights Verification conditions: The brightness range of the corridor lights is 60%~75%. Equipment adjustment: Adjust the brightness of the corridor fan light to a level that provides good illumination while walking. Quantity of equipment: [1 item] Device ID: 0x3 ] Task ID: 3 } In the above illustrative text, the value of the "Task List" field is 3 items, indicating that there are 3 device control tasks; the following "0" field indicates that the current device control task is the first device control task in the task list; the value of the "Number of Devices" field is 1 item, indicating that there is 1 smart device to be controlled under the current device control task; the "Device ID" field represents the device identifier of the smart device, such as 0x1, 0x2, 0x3; and the "Task ID" field represents the task identifier of the device control task.

[0133] It should be understood that in the examples listed above, the number of descriptive texts for scene segments corresponds one-to-one with the number of device control tasks; both are the same. For example, if there are three scene segments, there are also three device control tasks. In other scenarios, the number of descriptive texts for scene segments may differ from the number of device control tasks. For instance, the descriptive text for a scene segment may be broken down into two or more device control tasks. However, each device control task corresponds to only one control intent.

[0134] For example, if the description of scenario fragment 1 is: In the morning, Wang Zhicheng entered the reading room to read the newspaper and practice calligraphy, and found that the lights in the reading room were off, making the room too dim, which is not conducive to reading for the elderly. He hoped to adjust the brightness of the reading light to a comfortable level. In addition, Wang Zhicheng found that the air conditioner temperature in the reading room was too low, and hoped to raise the air conditioner temperature to a comfortable level.

[0135] When the smart terminal performs intent recognition on the descriptive text of the scene segment, it can identify two control intents: "turn on the lights in the reading room" and "turn up the air conditioning temperature in the reading room". Based on this, there are two device control tasks corresponding to the scene segment.

[0136] Further, see Figure 3 As shown, the task verifier can verify multiple device control tasks and finally output multiple device control tasks that have been successfully verified.

[0137] Thus, the smart terminal can obtain multiple device control tasks through the above step 402.

[0138] In the above technical solution, scene fragments are parsed to obtain multiple device control tasks. Since the generation of scene fragments conforms to the virtual control scenario and user behavior habits, the resulting device control tasks are entirely dependent on the actual capabilities of the devices in the virtual control scenario and the user needs within that space, thus avoiding tasks exceeding the device's capabilities and generating invalid training data. Furthermore, the device control tasks include execution verification conditions, which are specifically generated based on the target user's attribute characteristics, ensuring that the generation of device control tasks simultaneously considers the personalized usage needs of different users.

[0139] Step 403: Simulate task scenarios for multiple device control tasks to generate a dialogue training dataset. The dialogue training dataset is used to train the device control model, which enables users to control smart devices through control commands.

[0140] The aforementioned multiple device control tasks specifically refer to the multiple device control tasks that were successfully verified in step 402, hereinafter referred to as "multiple device control tasks".

[0141] After obtaining multiple device control tasks, each task includes the user's control intent, the corresponding device adjustment action, and execution verification conditions. Based on this, the smart terminal can obtain a dialogue training dataset by simulating the device control tasks.

[0142] One possible implementation involves simulating task scenarios for multiple device control tasks to generate a dialogue training dataset, including: Multiple device control tasks are combined to generate at least one task set, which contains at least one device control task. Task scenario simulation is performed on at least one task set to obtain a dialogue training dataset.

[0143] For example, in combination Figure 3 and Figure 2 As shown, after obtaining multiple device control tasks, before inputting these tasks into the dialogue generation module, the smart terminal can combine the multiple device control tasks through a task planning unit (not shown in the figure) to obtain at least one task set.

[0144] It should be understood that any task set in the above-mentioned at least one task set contains at least one device control task. The purpose of combining device control tasks into task sets is to make the tasks ultimately used for scenario simulation more closely resemble real-life scenarios. In real-life scenarios, a user's control intent at any given moment is not singular. If a single device control task is directly used for task scenario simulation, it cannot adapt to the diverse task execution needs in real-life scenarios.

[0145] Optionally, multiple device control tasks can be combined in ways including but not limited to random combination, combination according to possible temporal order, and combination according to scene relevance. Combining according to possible temporal order refers to chaining two device control tasks that may have a temporal sequence. For example, device control task 1 is to turn on the master bedroom light to a non-glaring brightness, and device control task 2 is to turn on the bathroom light to a suitable brightness. Based on common sense, device control task 1 usually precedes device control task 2; therefore, these two device control tasks can be combined into a task set.

[0146] Grouping based on scenario relevance refers to combining device control tasks within the same spatial area into a task set. For example, if device control task 1 and device control task 2 both correspond to smart devices in the master bedroom, then these two device control tasks can be combined into a task set.

[0147] Therefore, after obtaining at least one combined task set, the smart terminal can simulate task scenarios based on at least one task set to obtain a dialogue training dataset.

[0148] In the above technical solution, the equipment control tasks are combined to obtain a task set containing at least one equipment control task. This generates a diversified task set from a limited number of equipment control tasks, thereby simulating the continuous behavior of users in the same virtual control scenario and improving the model's ability to understand continuous tasks.

[0149] Specifically, the process of simulating task scenarios for at least one task set to obtain a dialogue training dataset is as follows.

[0150] In one possible implementation, task scenario simulation is performed on at least one task set to obtain a dialogue training dataset, including: For any task set in at least one task set, perform task scenario simulation on the task set and determine the execution result of the task set; If the execution result of the task set is successful, all dialogue data of the task set during the task scenario simulation process will be determined as the dialogue training data for at least one round corresponding to the task set. If at least one task set has completed the task scenario simulation, a dialogue training dataset is generated based on at least one round of dialogue training data corresponding to at least one target task set. The target task set is the task set in at least one task set whose execution result is successful.

[0151] In this embodiment, task scenario simulations for multiple task sets are performed sequentially; that is, at any given time, only one task set is being simulated. Furthermore, the task scenario simulations for each task set are completely decoupled and have no correlation.

[0152] For example, if at least one task set includes task set 1, task set 2, and task set 3. During the task scenario simulation process, task set 1 is first simulated to obtain the execution result of task set 1, then task set 2 is simulated to obtain the execution result of task set 2, and finally task set 3 is simulated to obtain the execution result of task set 3.

[0153] The following is combined with Figure 3 The process of obtaining the execution result of any task set by simulating a task scenario is first introduced for each component in the document.

[0154] One possible implementation involves simulating task scenarios for the task set and determining the execution results of the task set, including: Obtain the attribute characteristics of the virtual control scenario and target user, as well as the execution verification conditions contained in the task set; For the m-th dialogue round in the task scenario simulation process, the question text for the m-th dialogue round is generated based on the preset questioning style, virtual control scenario, target user attribute characteristics, task set and historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1. Based on the question text of the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and historical dialogue data, determine the device control command for the m-th dialogue round. The device action simulation is performed on the device adjustment command in the m-th dialogue round to obtain the device state in the m-th dialogue round. Based on the device state in the m-th dialogue round and the execution verification conditions contained in the task set, the task simulation result of the m-th dialogue round is determined. If the task simulation result in the m-th dialogue round is successful, the execution result of the task set is determined to be successful.

[0155] For example, such as Figure 3 As shown, the task scenario simulation process of the task set is mainly achieved through the interaction between the question generator, the task scenario simulator, and the instruction generator.

[0156] The question generator is used to generate natural language questions that match the user profile based on the multi-task sequence obtained by the task generation module, and input them as input-side samples to the instruction generator.

[0157] The instruction generator is used to receive user queries and generate corresponding device control instructions, which are then input into the task scenario simulator.

[0158] The task scenario simulator is used to simulate and execute the simulation results and verify their executability, and to provide feedback on the execution results and update the device status, ensuring that the generated training data is authentic and effective.

[0159] Specifically, in this embodiment, the question generator is a question generation model based on behavioral prompting (BP), such as the fifth-generation generative pre-trained transformer model (GPT-5 model). Behavioral prompting or behavioral guidance refers to the core technology of generating dialogue by simulating user behavior trajectories, scene context, and user profiles.

[0160] The instruction generator is specifically an instruction generation model used to simulate the response behavior of agents in smart home scenarios.

[0161] The task scenario simulator is specifically a multi-home agent sandbox. It supports customizable device operation, function parameter settings, and method calls. It allows for the input of structured action objects or Pythonic code for engine control, supports an indexing system for recalling necessary devices within the home environment, and supports a task system that can automatically trigger tasks through arbitrary manipulation. This task scenario simulator serves as a simulation environment for dialogue interaction, hosting the scenario-based execution of task sets and maintaining device states and environmental parameters within the virtual home environment.

[0162] like Figure 3 As shown, in the dialogue generation phase, the data source for the question generator is the instantiated question metadata generated by the question type generator, which includes question style, question environment, and question device type. The combination of question style, question environment, and question device type yields the question rules used to constrain the question text generated by the question generator.

[0163] For example, if the instantiated question metadata generated by the question type generator includes the following: question style is ambiguous intent, question environment is a quiet environment, and question device type is a single device, then the question rule is an ambiguous instruction controlled by a single device in a quiet environment.

[0164] It should be understood that, for a given task set, during task scenario simulation, the question type generator can generate multiple different question rules according to a preset generation method. In this embodiment, under the constraint of any question rule, the question generator and the instruction generator can generate at least one round of dialogue training data. That is to say, in this embodiment, a task set can generate at least one round of dialogue training data under each question rule through task scenario simulation.

[0165] The following embodiment of this application describes the process of generating at least one round of dialogue training data for a task set under one of the questioning rules.

[0166] Specifically, the preset question style is the question rule generated by the question type generator, which can also be understood as a question template.

[0167] Taking the m-th dialogue round in the task scenario simulation process of any task set under any questioning rule as an example, combined with Figure 3 As shown, in addition to receiving question rules, the question generator also receives the virtual control scene from the scene blueprint, the attribute characteristics of the target user, the currently simulated task set, and dialogue data from the historical dialogue process, i.e., the historical dialogue data of the first m-1 dialogue rounds. Based on this data, it generates the question text for the m-th dialogue round. The entire interaction process from the generation of the question text to the completion of the device control command is called a dialogue round.

[0168] When m=1, the historical dialogue data for the first m-1 dialogue rounds is empty, meaning that the question generator only generates the question text for the m-th dialogue round based on the virtual control scenario, the target user's attribute characteristics, and the currently simulated task set.

[0169] It should be understood that when generating question text, the virtual control scenario provides the physical boundaries of the question text generator, ensuring that the question text does not deviate from the virtual home environment. The target user's attribute characteristics provide the question generator with the user's expression habits and behavioral styles, allowing the question text to retain the user's personalized features. The currently simulated task set provides the question generator with specific device control tasks, ensuring that the dialogue completely corresponds to the current device control task. Historical dialogue data provides the question generator with the context of the dialogue process, guaranteeing the correlation between the question text and the context, and enabling better integration of the question text with historical dialogues.

[0170] Specifically, when generating question text, the question generator integrates the aforementioned data to obtain a complete input text, which serves as the input to the GPT-5 model. The GPT-5 model parses the input text, constructs the target user's role characteristics based on the user profile, and clarifies the device control intent to be triggered by the question in conjunction with the current task set. Finally, based on the role characteristics, task objectives, and question rules, the model generates a natural language question text that meets the constraints of intent clarity, expression style, and device interaction type, i.e., the question text for the m-th dialogue round, and inputs it into the instruction generator.

[0171] In addition to receiving the question text from the m-th dialogue round, the instruction generator also receives the virtual control scene from the scene blueprint, the target user's attribute characteristics, the currently simulated task set, and dialogue data from historical dialogues. Similar to the logic of the question generator generating the question text, the virtual control scene and the target user's attribute characteristics provide the instruction generator with background information for this simulation (e.g., scene information and the user's physiological characteristics and behavioral habits). The currently simulated task set provides the instruction generator with specific device control tasks. Historical dialogue data provides the question generator with the context of the dialogue process, ensuring the correlation between device control instructions and the context.

[0172] Specifically, the instruction generator first performs semantic understanding, entity extraction, and intent recognition on the question text of the m-th dialogue round, extracting key information from the question text, such as intent, device location, etc. Then, it uses the attribute features of the virtual control scenario and the target user as global constraints to ensure that the device in the device control instruction is the device in the virtual control scenario, and that the adjustment value conforms to the user's inherent preferences and physiological characteristics. Finally, it combines the historical dialogue data of the previous m-1 rounds to perform contextual association and completion, thereby obtaining the device control instruction for the m-th dialogue round.

[0173] The device control commands are input into the task scenario simulator, which then simulates the device actions to obtain the action simulation results for the m-th dialogue round.

[0174] Specifically, MHA has a complete virtual home scene pre-set inside. After receiving the device control command, MHA uses its own device indexing system and device simulation engine to complete the device location, drive the corresponding virtual device to perform simulation operation according to the command requirements, and finally update the status of the virtual device as the action simulation result of the m-th dialogue round, that is, the device status of the m-th dialogue round (i.e., the device operating parameters).

[0175] Preferably, in addition to receiving device control commands, the task scenario simulator can also receive the current task set, so as to quickly narrow down the device search range during the device positioning process and improve simulation efficiency.

[0176] As discussed above, the execution verification conditions in device control tasks provide verifiable evaluation criteria for the task scenario simulation process. Based on this, after obtaining the device state in the m-th dialogue round, the device state in the m-th dialogue round is fed back to the instruction generator. The instruction generator then determines the task simulation result for the m-th dialogue round based on the device state in the m-th dialogue round and the execution verification conditions of the task set. The task simulation result indicates whether each device control task in the task set was correctly executed during the simulation. The execution verification conditions of the task set include the execution verification conditions corresponding to each device control task in the task set.

[0177] Optionally, the task simulation results include success and failure.

[0178] Optionally, the process of determining the task simulation result of the m-th dialogue round is not limited to being completed by an instruction executor, or it can be completed by a task scenario simulator. That is, after the task scenario simulator obtains the device status of the m-th dialogue round, it further combines the execution verification conditions of the task set to determine the task simulation result of the m-th dialogue round and feeds it back to the instruction generator. This application embodiment does not limit this.

[0179] If the task simulation result for the m-th dialogue round is successful, it means that the device state for the m-th dialogue round matches the execution verification conditions. If the task simulation result for the m-th dialogue round is unsuccessful, it means that the device state for the m-th dialogue round does not match the execution verification conditions.

[0180] It should be understood that, as described above, the execution verification conditions are essentially the range of device parameters. Therefore, if the device status in the m-th dialogue round matches the execution verification conditions, it means that the device operating parameters in the m-th dialogue round are within the parameter range corresponding to the execution verification conditions; conversely, if the device status in the m-th dialogue round does not match the execution verification conditions, it means that the device operating parameters in the m-th dialogue round are not within the parameter range corresponding to the execution verification conditions.

[0181] For example, if the execution verification condition of the task set is that the temperature range of the main bedroom air conditioner is 20~22℃, and after the task scenario simulator performs the simulation, the device status in the m-th dialogue round is that the main bedroom air conditioner temperature is 30℃, which is not within the above temperature range, then the smart terminal determines that the task simulation result in the m-th dialogue round is a failure. Conversely, if the execution verification condition of the task set is that the temperature range of the main bedroom air conditioner is 20~22℃, and after the task scenario simulator performs the simulation, the device status in the m-th dialogue round is that the main bedroom air conditioner temperature is 21℃, which is within the above temperature range, then the smart terminal determines that the task simulation result in the m-th dialogue round is a success.

[0182] It should be understood that when there are multiple virtual devices corresponding to the device control command in the m-th dialogue round, the above-mentioned determination of the task simulation result in the m-th dialogue round specifically means that for each virtual device, it is necessary to determine whether its device operation parameters in the m-th dialogue round belong to the parameter range corresponding to the execution verification conditions.

[0183] For example, if the verification conditions are that the temperature range of the master bedroom air conditioner is 20~22℃ and the brightness range of the corridor light is 60%~75%, and after the task scenario simulation, the temperature of the master bedroom air conditioner is 30℃, which is not within the above temperature range; and the brightness range of the corridor light is 80%, which is not within the above brightness range, then the smart terminal determines that the task simulation result is a failure.

[0184] In one scenario, when the task simulation result of the m-th dialogue round is determined to be successful, the smart terminal determines that the execution result of the current task set is successful. Thus, the task scenario simulation process of the current task set ends, and the task scenario simulation process of the next task set begins.

[0185] In the aforementioned technical solution, when simulating task scenarios for the task set, both the question generation and instruction generation processes incorporate the virtual control scenario, the target user's attribute characteristics, the task set, and historical dialogue data. This ensures that both the question text and device control instructions are strictly constrained by the current virtual control scenario and user profile, guaranteeing that the question text matches the user's speaking style and behavioral habits, and that the device control instructions are generated around the devices in the current virtual control scenario while also taking into account the target user's usage habits. Furthermore, based on the execution verification conditions of the task set, the device state in the m-th dialogue round is verified after the task scenario simulation, ensuring that the training data is executable and conforms to the device adjustment logic, thus guaranteeing the reliability and effectiveness of the training data.

[0186] If the execution result of the current task set is successful, the smart terminal can obtain at least one round of dialogue training data for that task set through the dialogue generation module.

[0187] In one possible implementation, the dialogue data of the task set during the task scenario simulation process is determined as the dialogue training data for at least one round corresponding to the task set, including: Based on historical dialogue data, the question text of the m-th dialogue round, and the device control command of the m-th dialogue round, generate at least one round of dialogue training data corresponding to the task set.

[0188] Specifically, the dialogue generation module can encapsulate historical dialogue data, the question text of the m-th dialogue round, and the device control command of the m-th dialogue round into at least one round of dialogue training data corresponding to the task set, according to the required format for dialogue training data. This at least one round of dialogue training data can be used as standard data for training the device control model. The device control model is specifically the large model in the Agent.

[0189] In another scenario, when the task simulation result for the m-th dialogue turn is determined to be a failure, the smart terminal needs to consider the preset maximum number of dialogue turns allowed during the task scenario simulation to determine whether further dialogue is possible, and then determine the execution result of the task set.

[0190] In one possible implementation, the generation method also includes: If the task simulation result in the m-th dialogue round fails, determine whether the m-th dialogue round is the preset maximum dialogue round. If the m-th dialogue round is not the preset maximum dialogue round, let m+1=m, and repeatedly execute the following steps: generate the question text for the m-th dialogue round based on the preset questioning style, virtual control scenario, target user attribute characteristics, task set, and historical dialogue data from the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; determine the device control command for the m-th dialogue round based on the question text, virtual control scenario, target user attribute characteristics, task set, and historical dialogue data; simulate device actions for the device adjustment command in the m-th dialogue round to obtain the device state in the m-th dialogue round; and determine the task simulation result for the m-th dialogue round based on the device state in the m-th dialogue round and the execution verification conditions contained in the task set. If the m-th dialogue round is the preset maximum number of dialogue rounds, the execution result of the task set is determined to be a failure.

[0191] The preset maximum number of dialogue rounds refers to the maximum number of rounds of questioning and instruction interaction that a given task set can perform during task scenario simulation, given fixed questioning rules. Specifically, the preset maximum number of dialogue rounds can be the same for all task sets, for example, 8 rounds.

[0192] When the m-th dialogue round is the preset maximum dialogue round, i.e. the last dialogue round, it means that the task set cannot continue to simulate the task scenario, and the smart terminal determines that the execution result of the task set is a failure.

[0193] Conversely, if the m-th dialogue round is not the preset maximum dialogue round, it indicates that the task set can continue to simulate the task scenario. Therefore, the dialogue generation module can repeat the aforementioned task scenario simulation steps, that is, update the (m+1)-th dialogue round to the m-th dialogue round, and repeatedly execute the steps of generating the question text for the m-th dialogue round based on the preset questioning style, virtual control scenario, target user attribute characteristics, task set, and historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; determining the device control command for the m-th dialogue round based on the question text for the m-th dialogue round, virtual control scenario, target user attribute characteristics, task set, and historical dialogue data; simulating device actions for the device adjustment command for the m-th dialogue round to obtain the device state for the m-th dialogue round; and determining the task simulation result for the m-th dialogue round based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set.

[0194] Through the above process, the smart terminal can obtain the execution results of each task set.

[0195] As described above, if the execution result of a certain task set is successful, the intelligent terminal can obtain at least one round of dialogue training data for that task set through the dialogue generation module. Therefore, when at least one task set has completed the task scenario simulation, each of the at least one target task set with a successful execution result corresponds to at least one round of dialogue training data. The at least one round of dialogue training data corresponding to at least one target task set constitutes the complete dialogue training dataset under the current virtual control scenario and the target user profile.

[0196] The above describes the process of generating training data in the entire data generation process.

[0197] In some embodiments, training data can be generated by combining three stages: data generation, data augmentation, and reinforcement learning, and then used to train the device control model.

[0198] Figure 5 This is a schematic diagram of a layered system architecture provided in an embodiment of this application.

[0199] For example, such as Figure 5 As shown, the layered system architecture 500 includes an application layer and a foundation layer; the application layer includes a data generation module 501, a data augmentation module 502, and a reinforcement learning module 503; the foundation layer includes a device simulation engine 504 and a whole-house data generator 505.

[0200] The functions of each module in the layered system architecture will be further explained below.

[0201] The data generation module 501 specifically corresponds to... Figure 2The overall functional module, consisting of the task generation module and the dialogue generation module, is used to generate training data. When generating training data, to ensure that the generated device control tasks include household and user information, a scene blueprint for the virtual control scenario is generated using a scene generator (i.e., the whole-house data generator 505), and training data for this stage is obtained through task scenario simulation.

[0202] The training data generated by the data generation module 501 can be sent to the data augmentation module 502 to explore different response content of the device control model for a single user question, improve the diversity of response statements, and explore different user questions to increase the difficulty of device control tasks, thereby exploring and generating diverse dialogue trajectories.

[0203] The training data generated by the data augmentation module 502 can be sent to the reinforcement learning (RL) module 503 to calculate the reward value of each dialogue link including multi-turn interactive statements based on the training data, and combine the reward value with each dialogue link to train the device control model, thereby obtaining the trained device control model.

[0204] The device simulation engine (i.e., MHA) 504 is used to receive tool call instructions from the device control model, simulate virtual home devices to perform operations, and return the device status change results to the device control model in the form of tool responses.

[0205] It should be noted that the reinforcement learning module 503 can directly obtain training data from the data generation module 501, or obtain data-augmented training data from the data augmentation module 502. In addition, a trusted sample generation module can be configured in the reinforcement learning module to generate training data through the trusted sample generation module in the reinforcement learning module 503; and the device control model is trained based on the training data. That is, the data generation module 501, the data augmentation module 502 and the reinforcement learning module 503 can be decoupled independent functional modules.

[0206] The following is through Figure 6 The overall implementation flow of the training data generation process in the data generation stage of this application embodiment is described.

[0207] Figure 6 This is a schematic flowchart illustrating another method for generating training data provided in an embodiment of this application.

[0208] For example, such as Figure 6 As shown, the generation method 600 includes the following steps 601 to 612.

[0209] Step 601: Based on preset training requirements, construct the attribute characteristics of the virtual control scenario and the target user.

[0210] Step 602: Based on the attribute characteristics of the virtual control scenario and the target user, generate descriptive text for N scene fragments of the target user in the virtual control scenario.

[0211] Step 603: For any scene segment among the N scene segments, perform intent parsing on the description text of the scene segment to obtain the control intent corresponding to the scene segment.

[0212] Step 604: Based on the control intent corresponding to the scene segment, determine the target smart device from the M virtual smart devices, and determine the adjustment action of the target smart device.

[0213] Step 605: Determine the execution verification conditions for the target smart device based on the attribute characteristics of the target smart device and the target user.

[0214] Step 606: Based on the control intent, the target intelligent device, the adjustment action of the target intelligent device, and the execution verification conditions, generate the device control task corresponding to the scene segment. The device control task is the device control task that has been successfully verified.

[0215] Step 607: Determine the device control tasks corresponding to N scene segments as multiple device control tasks.

[0216] Step 608: Simulate the task scenario for the task set and determine the execution result of the task set.

[0217] Step 609: Determine whether the execution result of the task set is successful.

[0218] If the task set is executed successfully, proceed to step 611; If the execution of the task set fails, proceed to step 610.

[0219] Step 610: Determine that the task scenario simulation has ended.

[0220] Step 611: Determine all dialogue data of the task set during the task scenario simulation process as the dialogue training data for at least one round corresponding to the task set.

[0221] Step 612: Generate dialogue training data based on at least one round of dialogue training data corresponding to at least one target task set whose execution result is successful in at least one task set.

[0222] Steps 601 to 612 in the above-mentioned generation method 600 have the same inventive concept as steps 401 to 403 in the generation method 400. For details, please refer to the above-mentioned description of generation method 400, which will not be repeated here.

[0223] In summary, this application proposes a method for generating dialogue training data for intelligent device control models. First, by constructing a virtual control scenario and the attribute features of the target user, this method provides spatial environment constraints and user behavior style constraints for the generation of training samples, preventing training data from deviating from real interaction scenarios and user habits. Furthermore, based on the target user's control intentions, device adjustment actions, and device execution verification conditions in the virtual scenario, the device control task is determined and task scenario simulation is performed to obtain dialogue training data. Here, the control intention reflects the user's adjustment needs for the device, the device adjustment actions reflect the specific content that needs to be adjusted, and the execution verification conditions provide verifiable quantitative evaluation criteria for the task scenario simulation. This ensures that the generated dialogue training data not only meets the interaction needs of real users but also possesses rationality and executability, providing reliable training samples for subsequent device control model training and significantly reducing the proportion of invalid training data generated.

[0224] Figure 7 This is a schematic diagram of the structure of a training data generation device provided in an embodiment of this application.

[0225] For example, such as Figure 7 As shown, the generating apparatus 700 includes: Module 701 is used to construct the attribute features of the virtual control scene and the target user based on preset training requirements; The determination module 702 is used to determine multiple device control tasks of the target user in the virtual control scenario based on the virtual control scenario and the attribute characteristics of the target user. The device control tasks include the control intention of the target user in the virtual control scenario, the adjustment action of the device corresponding to the control intention, and the execution verification conditions. The generation module 703 is used to simulate the task scenarios of the multiple device control tasks and generate a dialogue training dataset. This dialogue training dataset is used to train the device control model, which enables the user to control the smart device through control commands.

[0226] In one possible implementation, the determining module 702 is specifically used to: generate descriptive text for N scene segments of the target user in the virtual control scenario based on the virtual control scenario and the attribute characteristics of the target user, where N is a positive integer greater than or equal to 1, and the descriptive text of the scene segments is used to represent the spatiotemporal information, device status and device adjustment requirements associated with the behavioral trajectory of the target user in the virtual control scenario; and determine the multiple device control tasks based on the descriptive text of the N scene segments and the attribute characteristics of the target user.

[0227] In one possible implementation, the virtual control scenario includes M virtual smart devices, where M is a positive integer greater than or equal to 1. The determining module 702 is specifically used for: for any scene segment among the N scene segments, performing intent parsing on the description text of the scene segment to obtain the control intent corresponding to the scene segment; determining the target smart device from the M virtual smart devices based on the control intent corresponding to the scene segment, and determining the adjustment action of the target smart device; determining the execution verification conditions of the target smart device based on the attribute characteristics of the target smart device and the target user; generating the device control task corresponding to the scene segment based on the control intent, the target smart device, the adjustment action of the target smart device, and the execution verification conditions of the target smart device; and, after the device control tasks corresponding to the N scene segments have been generated, determining the device control tasks corresponding to the N scene segments as the plurality of device control tasks.

[0228] In one possible implementation, the generation module 703 is specifically used to: combine the multiple device control tasks to generate at least one task set, the task set containing at least one device control task; and perform task scenario simulation on the at least one task set to obtain the dialogue training dataset.

[0229] In one possible implementation, the generation module 703 is specifically used to: for any task set in the at least one task set, perform task scenario simulation on the task set and determine the execution result of the task set; if the execution result of the task set is successful, determine all dialogue data of the task set during the task scenario simulation process as at least one round of dialogue training data corresponding to the task set; if the at least one task set has completed the task scenario simulation, generate the dialogue training dataset based on at least one round of dialogue training data corresponding to at least one target task set, wherein the target task set is the task set in the at least one task set whose execution result is successful.

[0230] In one possible implementation, the generation module 703 is specifically used to: acquire the virtual control scenario, the attribute characteristics of the target user, and the execution verification conditions contained in the task set; for the m-th dialogue round of the task set in the task scenario simulation process, generate the question text for the m-th dialogue round based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; determine the device control command for the m-th dialogue round based on the question text, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data; simulate the device action of the device adjustment command for the m-th dialogue round to obtain the device state for the m-th dialogue round, and determine the task simulation result for the m-th dialogue round based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set; if the task simulation result for the m-th dialogue round is successful, determine that the execution result of the task set is successful.

[0231] In one possible implementation, the generation module 703 is specifically used to: generate at least one round of dialogue training data corresponding to the task set based on the historical dialogue data, the question text of the m-th dialogue round, and the device control command of the m-th dialogue round.

[0232] In one possible implementation, the generation module 703 is further configured to: if the task simulation result of the m-th dialogue round is a failure, determine whether the m-th dialogue round is a preset maximum dialogue round; if the m-th dialogue round is not the preset maximum dialogue round, let m+1=m, and repeatedly execute the generation of the question text for the m-th dialogue round based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; based on the m-th... The steps include: determining the device control command for the m-th dialogue round based on the question text of the dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data; simulating device actions for the device adjustment command in the m-th dialogue round to obtain the device state in the m-th dialogue round; determining the task simulation result for the m-th dialogue round based on the device state in the m-th dialogue round and the execution verification conditions contained in the task set; and determining the execution result of the task set as a failure if the m-th dialogue round is the preset maximum dialogue round.

[0233] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0234] For example, such as Figure 8As shown, the electronic device 800 includes a memory 801 and a processor 802. The memory 801 stores executable program code 8011, and the processor 802 is used to call and execute the executable program code 8011 to perform a training data generation method.

[0235] This embodiment can divide the electronic device into functional modules according to the above method example. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0236] When functional modules are divided according to their respective functions, the electronic device may include: a construction module, a determination module, and a generation module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0237] The electronic device provided in this embodiment is used to execute the above-described method for generating training data, and thus can achieve the same effect as the above-described implementation method.

[0238] When using integrated units, the electronic device may include a processing module and a storage module. The processing module is used to control and manage the operation of the electronic device. The storage module is used to support the execution of relevant program code and data by the electronic device.

[0239] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.

[0240] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement a training data generation method in the above embodiment.

[0241] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement a training data generation method described in the above embodiment.

[0242] In addition, the electronic device provided in the embodiments of this application may specifically be a chip, component or module. The electronic device may include a connected processor and a memory. The memory is used to store instructions. When the electronic device is running, the processor may call and execute the instructions to make the chip execute a training data generation method in the above embodiments.

[0243] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0244] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0245] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0246] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating training data, characterized in that, The generation method includes: Based on preset training requirements, construct the attribute characteristics of virtual control scenarios and target users; Based on the virtual control scenario and the attribute characteristics of the target user, multiple device control tasks of the target user in the virtual control scenario are determined. The device control tasks include the control intention of the target user in the virtual control scenario, the adjustment action of the device corresponding to the control intention, and the execution verification conditions. The multiple device control tasks are simulated to generate a dialogue training dataset. The dialogue training dataset is used to train a device control model, which enables users to control smart devices through control commands.

2. The generation method according to claim 1, characterized in that, The step of determining multiple device control tasks for the target user in the virtual control scenario based on the virtual control scenario and the attribute characteristics of the target user includes: Based on the virtual control scenario and the attribute characteristics of the target user, N scene fragment description texts of the target user under the virtual control scenario are generated, where N is a positive integer greater than or equal to 1. The scene fragment description texts are used to represent the spatiotemporal information, device status and device adjustment requirements associated with the target user's behavior trajectory under the virtual control scenario. Based on the descriptive text of the N scene fragments and the attribute characteristics of the target user, the multiple device control tasks are determined.

3. The generation method according to claim 2, characterized in that, The virtual control scenario includes M virtual intelligent devices, where M is a positive integer greater than or equal to 1; the step of determining the control tasks for the multiple devices based on the descriptive text of the N scenario fragments and the attribute characteristics of the target user includes: For any scene segment among the N scene segments, perform intent parsing on the description text of the scene segment to obtain the control intent corresponding to the scene segment; Based on the control intent corresponding to the scene segment, the target intelligent device is determined from the M virtual intelligent devices, and the adjustment action of the target intelligent device is determined; Based on the attribute characteristics of the target smart device and the target user, determine the execution verification conditions of the target smart device; Based on the control intent, the target intelligent device, the adjustment action of the target intelligent device, and the execution verification conditions of the target intelligent device, generate the device control task corresponding to the scene segment; Once the device control tasks corresponding to the N scene segments have been generated, the device control tasks corresponding to the N scene segments are determined as the plurality of device control tasks.

4. The generation method according to claim 1, characterized in that, The step of simulating task scenarios for the multiple device control tasks and generating a dialogue training dataset includes: The plurality of device control tasks are combined to generate at least one task set, the task set containing at least one of the device control tasks; The dialogue training dataset is obtained by simulating task scenarios on the at least one task set.

5. The generation method according to claim 4, characterized in that, The step of simulating task scenarios on the at least one task set to obtain the dialogue training dataset includes: For any task set in the at least one task set, perform task scenario simulation on the task set and determine the execution result of the task set; If the execution result of the task set is successful, all dialogue data of the task set during the task scenario simulation process shall be determined as at least one round of dialogue training data corresponding to the task set. If all tasks in the at least one task set have completed the task scenario simulation, the dialogue training dataset is generated based on at least one round of dialogue training data corresponding to at least one target task set, wherein the target task set is the task set in the at least one task set whose execution result is successful.

6. The generation method according to claim 5, characterized in that, The step of simulating task scenarios for the task set and determining the execution result of the task set includes: Obtain the virtual control scenario, the attribute characteristics of the target user, and the execution verification conditions contained in the task set; For the m-th dialogue round in the task scenario simulation process, the question text for the m-th dialogue round is generated based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1. Based on the question text of the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data, determine the device control command for the m-th dialogue round; The device adjustment command in the m-th dialogue round is simulated to obtain the device state in the m-th dialogue round. Based on the device state in the m-th dialogue round and the execution verification conditions contained in the task set, the task simulation result of the m-th dialogue round is determined. If the task simulation result of the m-th dialogue round is successful, the execution result of the task set is determined to be successful.

7. The generation method according to claim 6, characterized in that, The step of determining all dialogue data of the task set during the task scenario simulation process as at least one round of dialogue training data corresponding to the task set includes: Based on the historical dialogue data, the question text of the m-th dialogue round, and the device control command of the m-th dialogue round, at least one round of dialogue training data corresponding to the task set is generated.

8. The generation method according to claim 6 or 7, characterized in that, The generation method further includes: If the task simulation result of the m-th dialogue round is a failure, determine whether the m-th dialogue round is the preset maximum dialogue round; If the m-th dialogue round is not the preset maximum dialogue round, let m+1=m, and repeatedly execute the following steps: generate the question text for the m-th dialogue round based on the preset questioning style, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data of the previous m-1 dialogue rounds, where m is a positive integer greater than or equal to 1; determine the device control command for the m-th dialogue round based on the question text for the m-th dialogue round, the virtual control scenario, the attribute characteristics of the target user, the task set, and the historical dialogue data; simulate device actions for the device adjustment command for the m-th dialogue round to obtain the device state for the m-th dialogue round; and determine the task simulation result for the m-th dialogue round based on the device state for the m-th dialogue round and the execution verification conditions contained in the task set. If the m-th dialogue round is the preset maximum dialogue round, the execution result of the task set is determined to be a failure.

9. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the electronic device to perform the generation method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the generation method as described in any one of claims 1 to 8.

11. A computer program product, characterized in that, The computer program product includes: computer program code, which, when run on a computer, implements the generation method as described in any one of claims 1 to 8.