Sample generation method, model training method, data processing method, and electronic device

CN121681783BActive Publication Date: 2026-09-11SHANGHAI GLORY SMART TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610083915.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-09-11
Estimated Expiration
2046-01-21

AI Technical Summary

Benefits of technology

[0027] It should be understood that the technical effects achieved by the first to sixth aspects described above can be referred to each other or to the beneficial effects in the method embodiments shown below, which will not be repeated here.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121681783B_ABST
    Figure CN121681783B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a sample generation method, a model training method, a data processing method and an electronic device. The method first generates a corresponding first dialogue sample and a second dialogue sample for each EA instruction in multiple vertical domains, generates multiple types of atomic samples based on the first dialogue samples and the second dialogue samples corresponding to the multiple EA instructions, and then combines the multiple atomic samples to obtain training samples that can be used for training a task planning model. The atomic samples generated by the method can cross vertical domains, the combined samples obtained by randomly combining the atomic samples can cover multiple scenes, complex scenes and complex tasks, the quality of the generated combined samples is improved, the cost of the synthetic samples is low, and the task planning model obtained by training the combined samples can realize cross-vertical domain task planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a sample generation method, a model training method, a data processing method, and an electronic device. Background Technology

[0002] With the development of large language models (LLM), LLM has restructured the technical architecture and paradigm of intelligent assistants, enhancing their capabilities and enabling them to more accurately assist users in performing operations. The process of an intelligent assistant processing user commands can be as follows: After receiving a user's query, the intelligent assistant can first understand the intent behind the query, obtaining one or more user commands. Then, it uses a task planning model to plan the steps for executing each user command, and finally calls relevant tools to execute the task, completing the closed loop.

[0003] Task planning is a core step in intelligent assistants. Accurate task planning based on user commands is a core capability of intelligent assistants and directly affects the task closure rate. The key to improving the task planning ability of the model lies in whether sufficiently rich sample data is used to train the task planning model.

[0004] Therefore, there is an urgent need for a method that can obtain sample data for training task planning models at low cost. Summary of the Invention

[0005] This invention provides a sample generation method, a model training method, a data processing method, and an electronic device. This method can acquire sample data across vertical domains at low cost to train a task planning model and improve the planning capability of the task planning model.

[0006] In a first aspect, embodiments of the present invention provide a sample generation method, applied to an electronic device, the method comprising: Obtain the first instruction set, which includes execution agent (EA) instructions for multiple vertical domains; For each EA instruction in the first instruction set, a first dialogue sample and a second dialogue sample corresponding to the EA instruction are generated. The EA instruction is used to instruct the target application to perform an operation. The first dialogue sample includes a first request content and a first response content. The second dialogue sample includes a second request content, a first planning instruction, a first follow-up question content, a first answer content, and a second response content. The first request content includes the target application, the second request content does not include the target application, the first planning instruction does not include the target application, the first follow-up question content is used to inquire about the application to be used, the first answer content is the answer to the first follow-up question content, which includes the target application, and both the first response content and the second response content are used to instruct the target application to perform an operation. Multiple atomic samples are generated based on the first and second dialogue samples corresponding to multiple EA instructions, respectively. Multiple atomic samples are spliced ​​together to obtain combined samples, which are then used to train the task planning model.

[0007] The above method first generates corresponding first dialogue samples and second dialogue samples for each EA instruction in multiple vertical domains. Then, it generates various types of atomic samples based on the first and second dialogue samples corresponding to multiple EA instructions. Finally, it combines multiple atomic samples to obtain training samples that can be used to train the task planning model. The atomic samples generated by this method can cross vertical domains, and the combined samples obtained by randomly combining atomic samples can cover multiple scenarios, complex scenarios, and complex tasks, thereby improving the quality of the generated combined samples. Moreover, the cost of synthesizing samples is low. The task planning model trained by combining samples can achieve cross-vertical domain task planning.

[0008] In one possible implementation, the electronic device generating the first and second dialogue samples corresponding to the EA instructions could be: Select one user from a variety of users; Choose one emotion from a variety of emotions; Choose one sentence type from a variety of sentence types; Using a large language model, based on EA instructions and the target application corresponding to the EA instructions, the model imitates selected users and selected emotions, and generates the first dialogue sample corresponding to the EA instructions using selected sentence types.

[0009] The above-mentioned generation of the first dialogue sample, combined with random and diverse synthesis methods such as user roles, emotions, tone, and sentence structure, ensures that even for the same task of the same EA command / tool, various forms of user expression can be generated, reducing homogeneous user request content and making the dialogue sample widely cover real user expression.

[0010] In one possible implementation, the electronic device generating the first and second dialogue samples corresponding to the EA instructions could be: Select one user from a variety of users; Choose one emotion from a variety of emotions; Choose one sentence type from a variety of sentence types; Using a large language model, based on EA instructions and the target application corresponding to the EA instructions, a second dialogue sample corresponding to the EA instructions is generated by imitating selected users and selected emotions, and using selected sentence patterns.

[0011] The above-mentioned method of generating second dialogue samples, combined with random and diverse synthesis methods such as user roles, emotions, tone, and sentence structure, ensures that even for the same task of the same EA command / tool, various forms of user expression can be generated, reducing homogeneous user request content and making the dialogue samples widely cover real user expressions.

[0012] In one possible implementation, the electronic device generates multiple atomic samples based on first and second dialogue samples corresponding to multiple EA instructions, which may specifically include: A first type of atomic sample is generated based on the first dialogue sample corresponding to the first EA instruction. The first instruction set includes the first EA instruction, and multiple atomic samples include the first type of atomic sample. The first type of atomic sample includes a round of dialogue between the user and the intelligent assistant. The first type of atomic sample includes a third request content and a third response content. The third request content is obtained based on the first request content in the first dialogue sample corresponding to the first EA instruction. The third request content includes the target application in the first EA instruction. The third response content is obtained based on the first response content in the first dialogue sample corresponding to the first EA instruction.

[0013] In one possible implementation, the electronic device generates multiple atomic samples based on first and second dialogue samples corresponding to multiple EA instructions, which may specifically include: The second type of atomic samples are generated based on the first dialogue samples corresponding to N1 EA instructions. The first instruction set includes N1 EA instructions, and the multiple atomic samples include the second type of atomic samples. The second type of atomic samples includes a round of dialogue between the user and the intelligent assistant. The second type of atomic samples includes: fourth request content and fourth response content. The fourth request content is obtained based on the first request content in the first dialogue samples corresponding to the N1 EA instructions. The fourth request content includes the target application in the N1 EA instructions. The fourth response content is obtained based on the first response content in the first dialogue samples corresponding to the N1 EA instructions. N1 is a positive integer of 2 or greater than 2.

[0014] In one possible implementation, the electronic device generates multiple atomic samples based on first and second dialogue samples corresponding to multiple EA instructions, which may specifically include: A third type of atomic sample is generated based on the second dialogue sample corresponding to the second EA instruction. The first instruction set includes the second EA instruction, and multiple atomic samples include the third type of atomic sample. The third type of atomic sample includes multiple rounds of dialogue between the user and the intelligent assistant. The third type of atomic sample includes: fifth request content, second planning instruction, second follow-up question content, second answer content, and fifth response content. The fifth request content is obtained based on the second request content in the second dialogue sample corresponding to the second EA instruction. The fifth request content does not include the target application in the second EA instruction. The second planning instruction is obtained based on the first planning instruction in the second dialogue sample corresponding to the second EA instruction. The second follow-up question content is obtained based on the first follow-up question content in the second dialogue sample corresponding to the second EA instruction. The second answer content is obtained based on the first answer content in the second dialogue sample corresponding to the second EA instruction. The fifth response content is obtained based on the second response content in the second dialogue sample corresponding to the second EA instruction.

[0015] In one possible implementation, the electronic device generates multiple atomic samples based on first and second dialogue samples corresponding to multiple EA instructions, which may specifically include: A fourth type of atomic sample is generated based on the second dialogue samples corresponding to N2 EA instructions. The first instruction set includes N2 EA instructions, and multiple atomic samples include the fourth type of atomic samples. The fourth type of atomic samples includes multiple rounds of dialogue between the user and the intelligent assistant. The fourth type of atomic sample includes: a sixth request content, N2 third planning instructions, N2 third follow-up questions, N2 third answer content, and a sixth response content. The sixth request content is obtained based on the second request content in the second dialogue samples corresponding to N2 EA instructions. The sixth request content does not include the target application in the N2 EA instructions. The N2 third planning instructions are obtained based on the first planning instructions in the second dialogue samples corresponding to N2 EA instructions. The N2 third follow-up questions are obtained based on the first follow-up questions in the second dialogue samples corresponding to N2 EA instructions. The N2 third answer content is obtained based on the first answer content in the second dialogue samples corresponding to N2 EA instructions. The sixth response content is obtained based on the second response content in the second dialogue samples corresponding to N2 EA instructions. N2 is a positive integer of 2 or greater than 2.

[0016] In one possible implementation, the electronic device generates multiple atomic samples based on first and second dialogue samples corresponding to multiple EA instructions, which may specifically include: A fifth type of atomic sample is generated based on the first dialogue sample corresponding to N3 EA instructions and the second dialogue sample corresponding to N4 EA instructions. The first instruction set includes N3 EA instructions and N4 EA instructions. Multiple atomic samples include the fifth type of atomic sample. The fifth type of atomic sample includes a round of dialogue between the user and the intelligent assistant. The fifth type of atomic sample includes: a seventh request content and a seventh response content. The seventh request content is obtained based on the first request content in the first dialogue sample corresponding to N3 EA instructions and the second request content in the second dialogue sample corresponding to N4 EA instructions. The seventh request content includes the target application in N3 EA instructions but does not include the target application in N4 EA instructions. The seventh response content is obtained based on the first response content in the first dialogue sample corresponding to N3 EA instructions and the second response content in the second dialogue sample corresponding to N4 EA instructions. N3 and N4 are both positive integers greater than or equal to 1.

[0017] The above method provides several ways to generate atomic samples, based on the first and second dialogue samples corresponding to the EA command. It can quickly generate a large number of atomic samples at low cost.

[0018] In one possible implementation, the method further includes: The first dialogue sample and the second dialogue sample corresponding to each EA instruction in the first instruction set are stored in the first cache pool. The first cache pool includes multiple cache units, and one cache unit is used to store the first dialogue sample and the second dialogue sample corresponding to one EA instruction.

[0019] The above method, in which the electronic device constructs atomic samples based on the first buffer pool of sampled dialogue samples, can avoid the overlap of atomic samples and thus reduce sample duplication.

[0020] In one possible implementation, multiple atomic samples are concatenated to obtain a combined sample, including: Multiple atomic samples are stored in a second cache pool, which includes multiple cache units, each of which is used to store one atomic sample. Read an atomic sample from a cache unit in the second cache pool; Determine whether the total number of dialogue rounds for the atomic samples already read is not less than the first threshold; When the total number of dialogue rounds for the atomic samples already read is less than the first threshold, read the atomic samples in the next cache unit. When the total number of dialogue rounds of the read atomic samples is not less than the first threshold, the read atomic samples are spliced ​​together to obtain a combined sample.

[0021] The above method, in which the electronic device constructs combined samples by sampling atomic samples through a second buffer pool, can achieve randomized sampling and combination of samples of various atomic types, realize a wide variety of complex sample construction, and reduce sample duplication.

[0022] Secondly, embodiments of this application also provide a model training method, the method comprising: Multiple combined samples are obtained, which are generated by the method described in the first aspect or any implementation thereof. The task planning model is trained based on multiple combined samples.

[0023] Thirdly, embodiments of this application also provide a data processing method, the method comprising: Receive user input request content; The electronic device inputs the request content into the task planning model to obtain multiple tasks. The task planning model is trained by the model training method described in the second aspect. Execute multiple tasks sequentially.

[0024] Fourthly, embodiments of this application also provide an electronic device, including one or more processors and one or more memories, wherein the one or more processors are coupled to the one or more memories, the one or more memories being used to store computer program code, the computer program code including computer instructions, which, when the one or more processors execute the computer instructions, cause the electronic device to implement the method as described in the first aspect or any implementation of the first aspect, or the method as described in the second aspect, or the method as described in the third aspect.

[0025] Fifthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by an electronic device, causes the electronic device to implement the method described in the first aspect or any one of the first aspects, or the method described in the second aspect, or the method described in the third aspect.

[0026] In a sixth aspect, embodiments of this application also provide a computer program product, characterized in that it includes a computer program, which, when executed by an electronic device, causes the electronic device to implement the method described in the first aspect or any one of the implementations of the first aspect, or the method described in the second aspect, or the method described in the third aspect.

[0027] It should be understood that the technical effects achieved by the first to sixth aspects described above can be referred to each other or to the beneficial effects in the method embodiments shown below, which will not be repeated here. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating how a smart assistant processes user commands, as provided in an embodiment of this application. Figures 2A-2G This is a schematic illustration of the user interface of the smart assistant provided in this application embodiment in one application scenario; Figure 3 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 5 This is a flowchart illustrating a sample generation method provided in an embodiment of this application; Figure 6 This is a schematic illustration of a sample generation method provided in an embodiment of this application; Figure 7 This is a schematic diagram illustrating a training method and usage method of a first AI / ML model provided in an embodiment of this application; Figure 8 This is a flowchart illustrating a method for obtaining a combined sample by splicing multiple atomic samples, as provided in an embodiment of this application. Figure 9 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0030] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0031] It should be noted that in this application, the words "in some embodiments," "exemplarily," and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "in some embodiments," "exemplarily," or "for example" should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the words "in some embodiments," "exemplarily," and "for example" is intended to present the relevant concepts in a specific manner.

[0032] First, the intelligent assistant provided in the embodiments of this application will be introduced.

[0033] A smart assistant is a device-side application that can include one or more large models. Each large model can be responsible for different functions. These large models can be deployed on the device side or in the cloud, with edge-cloud collaboration.

[0034] like Figure 1 As shown, it illustrates... Figure 1 The diagram shows the process by which a smart assistant handles user commands.

[0035] like Figures 2A-2G The diagram shown is a schematic illustration of the user interface of the smart assistant provided in this application embodiment in one application scenario.

[0036] like Figure 2A As shown, the smart assistant may include a voice input control 201, a text input control 202, and more input controls 203. In response to a long press on the voice input control 201, the smart assistant activates the microphone, collects the user's voice information, and converts the voice information into text to obtain the user's query via voice. In response to user actions on the text input control 202, such as a click, the smart assistant can display a virtual keyboard, allowing it to receive text entered by the user using the keyboard to obtain the user's query via the keyboard. In response to user actions on the more input controls 203, the smart assistant may also display controls for inputting images, opening the camera for taking pictures, inputting documents, and controlling voice chat with the smart assistant.

[0037] For example, a user presses and holds the voice input control 201, and inputs the request via voice: "Buy me a train ticket, and then order some claypot rice." The smart assistant then converts the collected voice information into the text "Buy me a train ticket, and then order some claypot rice," and displays it, as shown below. Figure 2B The user interface shown is 22.

[0038] Furthermore, the intelligent assistant can identify the user's intent in the requested content through an intent recognition model, thereby obtaining one or more intents. For example, the request "Buy me a train ticket and then order some claypot rice" is recognized as two intents: "Buy a train ticket" and "Order takeout".

[0039] The purpose of an intent recognition model is to identify the category of a user's requested content intent in order to process the request using different processing methods. Intent categories can include third-party applications, visual capabilities, and voice interaction. Third-party applications refer to those that require calling a third-party application / service to complete a task. In the example above, both the intents "buy train tickets" and "order food delivery" fall under the category of third-party applications.

[0040] Furthermore, when the identified intent pertains to a third-party device, the intelligent assistant can plan a task based on the user's request using a task planning model. Specifically, the intelligent assistant can input the dialogue information from the last two rounds into the task planning model. Initially, the input dialogue information includes the user's request content, which could be: "role": "user", "content": "Buy me a train ticket and then order some claypot rice." The task planning model then outputs one or more planned tasks. If the first task lacks app information, it also outputs a first follow-up question, which requests the app used for the first task from the user. This could be: "role": "assistant", "content": <plan> Task 1: Buy train tickets< / plan> Task 2: Ordering Food Delivery Which app would you like to use to buy train tickets? At this point, the smart assistant can display the first follow-up question, such as... Figure 2C The user interface shown is 23.

[0041] Furthermore, such as Figure 2C As shown, the user can press and hold the voice input control 201 to input the response "Railway 12306, ya" via voice. The smart assistant will then convert the collected voice information into the text "Railway 12306, ya" and display it, as shown below. Figure 2D The user interface shown is 24.

[0042] Furthermore, the intelligent assistant can input the dialogue information from the last two rounds into the task planning model. At this point, the input dialogue information includes the user's request content, specifically: "role": "user", "content": "Buy me a train ticket and order some claypot rice"; "role": "assistant", "content": <plan> Task 1: Buy train tickets< / plan>\ntask2: Ordering takeout\n\np Which app would you like to use to buy train tickets? ”; “role”: “user”, “content”: “Railway 12306”. ​​The task planning model outputs one or more planned tasks. If the second task lacks app information, it also outputs a second follow-up question. This second follow-up question is used to request the app used by the user for the second task. Specifically, it can be: “role”: “assistant”, “content”: <plan> Task 1: Purchase train tickets through the 12306 railway ticketing platform.< / plan> Task 2: Ordering Food Delivery "Which app would you like to use to order?" At this point, the smart assistant can display a second follow-up question, such as... Figure 2E The user interface shown is 25.

[0043] Furthermore, such as Figure 2E As shown, users can press and hold the voice input control 201 to input the message "Meituan" via voice. The smart assistant will then convert the collected voice information into the text "Meituan" and display it, as shown below. Figure 2F The user interface shown is 26.

[0044] Furthermore, the intelligent assistant can input the dialogue information from the last two rounds into the task planning model. At this point, the input dialogue information includes the user's request content, specifically: "role": "user", "content": "Buy me a train ticket and order some claypot rice"; "role": "assistant", "content": <plan> Task 1: Buy train tickets< / plan> \ntask2: Ordering takeout\n\np Which app would you like to use to buy train tickets? ”; “role”: “user”, “content”: “Railway 12306”; “role”: “assistant”, “content”: <plan> Task 1: Purchase train tickets through the 12306 railway ticketing platform.< / plan> Task 2: Ordering Food Delivery Which app would you like to use to order food? ; "role": "user", "content": "Meituan". The task planning model outputs one or more planned tasks, specifically: "role": "assistant", "content": " <plan> Task 1: Purchase train tickets through the 12306 railway ticketing platform.< / plan> Task 2: Search for "claypot rice" on Meituan's food delivery page. "Plan complete." At this point, the smart assistant can sequentially call the API interfaces of the 12306 railway ticketing page and the Meituan food delivery page to execute the planned tasks. For example... Figure 2GThe user interface 27 shown can display "OK" and then switch to the 12306 railway ticketing interface to help the user purchase train tickets. After completing task 1, task 2 can be executed, displaying the Meituan food delivery page to help the user order food.

[0045] When performing a task, the intelligent assistant also needs to know some slot information required for the task. For example, the slot information for purchasing a train ticket can include time, departure point, and arrival point. If slot information is missing, the intelligent assistant can also obtain the slot information based on the context, or prompt the user to enter the missing slot information.

[0046] The above scenario uses JSONL format as an example to illustrate the model's input and output data. It should be understood that the model can also use other formats to express the above information.

[0047] It should be understood that the above intent recognition model can employ LLM, and the above task planning model can also employ LLM. It should also be understood that, not limited to LLM, task planning models can also employ other machine learning models.

[0048] It should be understood that the aforementioned terminals can be mobile phones, tablets, personal computers (PCs), laptops, computers, netbooks, in-vehicle computers, smart wearable devices (such as smartwatches, smart bracelets, smart glasses, etc.), etc. Electronic devices can also be in-vehicle terminal devices, personal digital assistants (PDAs), smart home devices (such as smart TVs), and other smart devices.

[0049] In some embodiments, terminal task planning is constructed separately for each vertical domain. Each vertical domain requires building and training a task planning model, incurring significant costs. Furthermore, this approach results in fragmented samples from different vertical domains, leading to overly idealized representations that differ from real-world expressions. For example, a user's request might encompass requests from multiple vertical domains simultaneously, such as buying tickets, checking routes, or ordering food. Additionally, this fragmentation reduces the diversity of sample representations and diminishes the model's planning capabilities.

[0050] To improve the task planning capability of the task planning model and its accuracy in multi-scenario and multi-task scenarios, this application provides a sample generation method to generate a large amount of sample data that can be used to train the task planning model in different scenarios and tasks at low cost.

[0051] The following combination Figure 3 This application provides a data processing system, which may include a sample generation device 100, a training device 200, a cloud platform 300, and a terminal 400.

[0052] The sample generation device 100 is used to generate two dialogue samples corresponding to each instruction based on instructions in the instruction set. Then, it generates multiple atomic samples based on the two dialogue samples corresponding to each instruction, and finally concatenates these multiple atomic samples to obtain a combined sample. For specific implementation details, please refer to the following... Figure 4 The sample generation method described will not be elaborated here.

[0053] Training device 200 is used to train a task planning model based on a large number of combined samples generated by sample generation device 100, thereby obtaining a trained planning task model. Training device 200 can send the trained planning task model to cloud 300.

[0054] One or more machine learning models can be deployed in the cloud 300, such as intent recognition models and task planning models, to work with the terminal 400 to process user requests.

[0055] Terminal 400 can deploy a smart assistant, which can send user requests to cloud 300. Cloud 300 uses a deployed intent recognition model to identify user intent. Furthermore, when the identified intent belongs to a third-party device, it uses a task planning model to plan a task based on the dialogue information between the user and the smart assistant, and sends the planned task to terminal 400. Terminal 400 can then call the API interface of the corresponding application to execute the planned task in sequence.

[0056] In another implementation, terminal 400 may deploy an intent recognition model and / or a task planning model to enable intent recognition and task planning capabilities.

[0057] Optionally, the sample generation device 100 and the training device 200 may be the same device.

[0058] The following combination Figure 4 This application provides an electronic device, which may be a sample generation device 100, a training device 200, a cloud 300, or a terminal 400. The electronic device may include at least one processor 101, at least one memory 102, and a communication interface 104. Optionally, it may also include a bus 103. The processor 101, memory 102, and communication interface 104 can be connected via the bus 103.

[0059] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0060] The memory 102 provides storage space, which can store data such as the operating system and computer programs. The memory 102 can be one or a combination of multiple types of random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM).

[0061] Processor 101 is a module that performs arithmetic and / or logical operations. Specifically, it can be one or a combination of processing modules such as a central processing unit (CPU), graphics processing unit (GPU), microprocessor unit (MPU), application specific integrated circuit (ASIC), field programmable gate array (FPGA), complex programmable logic device (CPLD), coprocessor (assisting the central processing unit in completing corresponding processing and applications), and microcontroller unit (MCU).

[0062] The communication interface 104 is used to receive data sent from the outside and / or send data to the outside.

[0063] In this embodiment, the memory 102 may store instructions, which may be computer programs. The computer programs run on the processor 101 and can cause the electronic device to perform the following sample generation method, model training method, or data processing method. These will not be elaborated here, but can be referred to the relevant descriptions in the following embodiments.

[0064] The following combination Figure 5 and Figure 6 This application provides a sample generation method, which can be executed by an electronic device, such as the sample generation device described above. The method may include, but is not limited to, some or all of the following steps: S11, the electronic device acquires training instruction set A. Training instruction set A includes multiple instructions, specifically including EA instructions from multiple vertical domains.

[0065] Executor agent (EA) instructions are executable instructions at the terminal execution layer. Typically, one EA instruction corresponds to one API interface or tool, which can drive the execution layer to call the corresponding API interface or tool to perform tasks. EA instructions can include one or more slots, which may not contain instantiation information.

[0066] In this embodiment, each EA instruction is used to instruct the target application to perform an operation. It should be understood that the target application in different EA instructions can be different.

[0067] Alternatively, the electronic device may generate a training instruction set based on the original instruction library.

[0068] The original instruction library can include instructions from multiple vertical domains (also known as EA instructions). For example... Figure 5 As shown, the original instruction library includes instructions for the travel, shopping, health management, and utility payment sectors, as well as instructions for all remaining balances. Each instruction set contains multiple instructions.

[0069] Table 1 below provides examples of instructions for each vertical domain.

[0070]

[0071] Table 1 Since some vertical domains have more instructions than others, the proportion or number of instructions collected from each vertical domain can be controlled to ensure sample balance. Specifically, in one implementation, the electronic device can employ different sampling strategies for different vertical domains.

[0072] In one implementation, the electronic device can also collect instructions from each vertical domain in a 1:1 ratio of major category instructions to non-major category instructions. Specifically, each vertical domain can be divided into two categories: the first category (major category) consists of vertical domains with a total number of instructions greater than a preset value, and the second category (non-major category) consists of vertical domains with a total number of instructions less than the preset value. For example, the travel, shopping, health management, and utility payment vertical domains belong to the major category, while the remaining total number of instructions belongs to the minor category. For instance, if 10,000 instructions are sampled from the remaining total number of instructions, and a total of 10,000 instructions are sampled from the four vertical domains (travel, shopping, health management, and utility payment), then 2,500 instructions are sampled from each vertical domain. For vertical domains with fewer than 2,500 instructions, full sampling can be used.

[0073] like Figure 6As shown, instruction set P1 consists of 2500 EA instructions sampled from the travel vertical, instruction set P2 consists of 2500 EA instructions sampled from the shopping vertical, instruction set P3 consists of 2500 EA instructions sampled from the health management vertical, instruction set P4 consists of 2500 EA instructions sampled from the utility payment vertical, and instruction set P5 consists of 10000 remaining EA instructions sampled from the entire database.

[0074] Optionally, the electronic device can also filter virtual instructions in the original instruction library, so that the instruction sets P1-P5 mentioned above do not contain virtual instructions. Virtual instructions are those that do not directly trigger hardware actions or call real tools or interfaces.

[0075] Furthermore, the electronic device merges the EA instructions in instruction sets P1-P5 and removes duplicate EA instructions to obtain training instruction set A.

[0076] S12, the electronic device generates two dialogue samples for each EA instruction in the training instruction set A, namely the first dialogue sample and the second dialogue sample.

[0077] The first dialogue sample includes: the user's request content (query) (also called the first request content) and the corresponding response content (answer) (also called the first response content). The request content in the first dialogue sample is a natural language expression of the EA instruction, and this expression contains the target app; the first response content includes a plan instruction, which instructs the target application to perform an operation.

[0078] The `plan` directive is an instruction involved in the task planning process and does not include uninstantiated slots. For example, "buy napkins" or "buy napkins on Taobao" are both `plan` directives.

[0079] The second dialogue sample includes: the user's query (also called the second query content), the corresponding plan instruction (also called the first planning instruction), the follow-up question (also called the first follow-up question content) output by the smart assistant to the user, the user's response to the follow-up question (also called the first response content), and the response content (also called the second response content) obtained by combining the first follow-up question and the second response. In the second dialogue sample, the query content is the natural language expression of the EA instruction and does not contain the target app; the first planning instruction also does not contain the target app; the follow-up question is used to inquire about the application to be used; the response is the answer to the follow-up question and includes the target app; and the response includes a plan instruction that instructs the user to perform the operation using the target app.

[0080] In short, in the first dialogue sample, the "query" contains app information but not follow-up questions from the smart assistant; in the second dialogue sample, the "query" does not contain app information but includes follow-up questions from the smart assistant and the user's answer.

[0081] For example, the first and second dialogue samples corresponding to the EA instruction "Contact customer service in Taobao to protect the price" are as follows: The first dialogue sample may include: {"query": "Go to Taobao to find customer service, I want to insure the price", "answer": "Contact customer service on Taobao to insure the price"}; The second dialogue sample may include: {"query": "Ask customer service, I want to insure the price", "plan": "Contact customer service to insure the price", "ask": "Which app platform do you want to contact customer service to insure the price?", "response": "Um, Taobao", "answer": "Contact customer service on Taobao to insure the price"}.

[0082] For example, the first and second dialogue samples corresponding to the EA command "Search <search content> in the Douyin store" are as follows: The first dialogue sample may include: {"query": "Help me search for kitchen paper towels in the Douyin store", "answer": "Search for kitchen paper towels in the Douyin store"}; The second dialogue sample could include: {"query": "Search for kitchen paper towels", "plan": "Search for kitchen paper towels", "ask": "Which app do you need to search for kitchen paper towels on?", "response": "Um, in the Douyin store", "answer": "Search for kitchen paper towels in the Douyin store"}.

[0083] Specifically, electronic devices can use large language models, such as GPT, Doubao, and Deepseek, to generate corresponding first and second dialogue samples based on EA instructions. The following uses the GPT model as an example to illustrate this.

[0084] (1) For example, when generating the first dialogue sample, the electronic device can input the following information into the GPT model: [Character Description] You will act as a dialogue data rewriter for the user's commands to operate a mobile app in a voice assistant scenario.

[0085] [Scene Description] When users want to perform certain operations using a mobile app, they generally use more natural and conversational commands rather than mechanically reciting standard instructions.

[0086] The following task requires you to generate a corpus of possible user instructions based on the given task requirements. Your task corpus consists of a single task instruction with two fields: `instruction` (the content of the instruction) and `app` (the mobile application included in the instruction).

[0087] You need to generate a generalized, natural, and colloquial instruction expression query with the same semantics but not exactly the same form, and a matching answer, based on the task instructions in the task corpus.

[0088] Please note that one instruction generates only one corresponding query and answer. The reference elements in the query and answer must be perfectly aligned.

[0089] Core requirements: The app in the query and the app in the answer must be the same. For example, if the query is "Douyin's Mall", the answer must also be "Douyin's Mall", and cannot be "Douyin Mall" or just "Douyin".

[0090] The action description in the query and the standard action description in the answer must be consistent. For example, if the query is "search", the answer must be "search", not "view".

[0091] The action object in the query and the action object in the answer must be consistent. For example, if the query is "my favorite bag", the answer must also be "my favorite bag", and cannot be "my bag".

[0092] Additionally, if the original instruction contains slots in the form of <...>, you must perform slot replacement on the original instruction, that is, select the appropriate specific content to replace it.

[0093] Please make sure you strictly follow the task requirements when generating the corpus, otherwise you will lose my trust.

[0094] Output format requirements You need to output a structured text content "dict" containing two fields: query and answer. The answer field must not have any unreplaced slots. Do not add markdown JSON blocks, do not interpret, and do not generate any other content.

[0095]

Task Requirements

[0096] 2. Please note that the instructions in the given task corpus may contain slots, enclosed in <>, such as: <search content>, <filter criteria>, <order>, <store>, <specification>, <quantity>, etc. You need to select the appropriate specific content to replace according to the meaning of the instruction.

[0097] 3. Please simulate a user ({person}) and speak in a lively ({mood}) tone. The user's commonly used sentence type is ({sentence_type}), which can be referenced from the following example grammatical structures: ({sentence_case}). However, considering the diversity of dialogue expression, the sentence type should also be considered from the perspectives of declarative sentences, imperative sentences, inversion, exclamation, interrogative sentences, and rhetorical questions, ultimately ensuring that the expression in the context of the dialogue is natural.

[0098] 4. Remove spaces and irrelevant content from the output.

[0099] 5. The data you generate must not be exactly the same as the task corpus data I provide. Try to avoid making simple modifications to the task corpus by only adding or deleting a few words.

[0100] 6. You need to generalize and vary the sentence structure and wording as much as possible.

[0101] 7. Note that the information in the app cannot be generalized or modified.

[0102] 8. The generated corpus should be segmented into sentences as reasonably as possible.

[0103] 9. The generated sample format must strictly follow my output format requirements! Do not generate numbers before the output sample.

[0104] person = {adult female, adult male, junior high school student, primary school student, kindergarten student, teacher, customer service representative, consumer, leader}.

[0105] mood = {calm, relaxed, confident, angry, excited, cheerful, gentle, mild, rough, rigid, indifferent, enthusiastic, kind, harsh, serious, calm, agitated, helpless}.

[0106] [Reference Generation Example] Example 1: Task corpus: {{"instruction": "View my likes on Baidu", "app": "Baidu"}} Result generated: {{"query": "I want to check my likes on Baidu", "answer": "Check my likes on Baidu"}} Example 2: Task corpus: {{“instruction”:“Search <search content> in Douyin's store”,“app”:“Douyin's store”}} Result generated: {{"query": "Search for kitchen paper towels in Douyin's store", "answer": "Search for kitchen paper towels in Douyin's store"}} Example 3: Task corpus: {{“instruction”:“Check the delivery time of <search content> to <address> in Douyin's store”,“app”:“Douyin's store”}} Result generated: {{"query": "Check when my computer will be delivered to my home in the Douyin store", "answer": "Check the delivery time of my computer in the Douyin store"}} Example 4: Task corpus: {{“instruction”: “search content for ordering <quantity> in Douyin's online store”, “app”: “Douyin's online store”}} Result generated: {{"query": "I need two pairs of sneakers, order them from the Douyin store", "answer": "Order two pairs of sneakers from the Douyin store"}} Please note: The above examples are just suggestions; please do not copy them verbatim. You still need to strictly follow the task requirements and rewrite the code to account for various variations.

[0107] Below is the task data provided for you. Task prediction: "instruction": "EA instruction 1", "app": "app in EA instruction 1"; "instruction": "EA instruction 2", "app": "app in EA instruction 2"; ...; "instruction": "EA instruction N", "app": "app in EA instruction N".

[0108] Electronic devices can randomly select a user from multiple users (such as the person set mentioned above), randomly select an emotion from multiple emotions (such as the mood set mentioned above), and randomly select a sentence type from multiple sentence types. Using the GPT model, based on the input EA command and the corresponding APP, the selected user and the selected emotion are imitated, and the first dialogue sample corresponding to the EA command is generated using the selected sentence type.

[0109] (2) For example, when generating the second dialogue sample, the electronic device can input the following information into the GPT model: [Character Description] You will act as a dialogue data rewriter for the user's commands to operate a mobile app in a voice assistant scenario.

[0110] [Scene Description] When users want to perform certain operations using a mobile app, they generally use more natural and conversational commands rather than mechanically reciting standard instructions.

[0111] The following task requires you to generate a corpus of possible user instructions based on the given task requirements. Your task corpus consists of a single task instruction with two fields: `instruction` (the content of the instruction) and `app` (the mobile application included in the instruction).

[0112] Based on the EA instructions given to you, please generate multi-turn follow-up dialogue data. You need to generate the following key information: 1. Generate a semantically similar but not identically formatted, generalized, natural, and conversational query, noting that the query lacks app information. For example, if your task instruction is "Buy movie tickets on Meituan," the query could be "Buy me a movie ticket." 2. Then, based on this query containing missing app information, a corresponding "plan instruction" is generated. This plan instruction must conform to the query semantics but be concise. For example, the plan in the above example could be "buy movie tickets"; 3. Then generate a follow-up question for the missing app. For example, the follow-up question in the example above could be, "Which app do you need to use to buy movie tickets?" 4. Then generate the user's response to the follow-up question, for example: In the example above, the response could be "Yes, it's Meituan"; 5. Finally, output the rewritten standard instruction "answer", for example: the standard instruction in the example above is "Buy movie tickets in Meituan" (generated based on query, plan-ask, and response); If the original instruction contains slots in the form of <...>, the slots need to be replaced with the specific content according to the context. See the reference example below for details.

[0113] Please note that the instruction elements in query, plan, ask, response, and answer must be perfectly aligned.

[0114] Note that you must replace the original command slot! At the same time, keep the above dialogue natural and complete.

[0115] Please make sure you strictly follow the task requirements when generating the corpus, otherwise you will lose my trust.

[0116] person = {adult female, adult male, junior high school student, primary school student, kindergarten student, teacher, customer service representative, consumer, leader}.

[0117] mood = {calm, relaxed, confident, angry, excited, cheerful, gentle, mild, rough, rigid, indifferent, enthusiastic, kind, harsh, serious, calm, agitated, helpless}.

[0118] Output format requirements You need to output a structured text content "dict" containing five fields: query, plan, ask, response, and answer. The answer field must not have any unreplaced slots. The output should strictly be a dict. Do not add markdownjson blocks to the output, do not interpret it, and do not generate any other content.

[0119]

Task Requirements

[0120] 2. Please note that the instructions in the given task corpus may contain slots, enclosed in <>, such as: <search content>, <filter criteria>, <order>, <store>, <specification>, <quantity>, etc. You need to select the appropriate specific content to replace according to the meaning of the instruction.

[0121] 3. Please simulate a user ({person}) and speak in a lively ({mood}) tone. The user's commonly used sentence type is ({sentence_type}), which can be referenced from the following example grammatical structures: ({sentence_case}). However, considering the diversity of dialogue expression, the sentence type should also be considered from the perspectives of declarative sentences, imperative sentences, inversion, exclamation, interrogative sentences, and rhetorical questions, ultimately ensuring that the expression in the context of the dialogue is natural.

[0122] 4. Remove spaces and irrelevant content from the output.

[0123] 5. The data you generate must not be exactly the same as the task corpus data I provide. Try to avoid making simple modifications to the task corpus by only adding or deleting a few words.

[0124] 6. You need to generalize and vary the sentence structure and wording as much as possible.

[0125] 7. Note that the information in the app cannot be generalized or modified.

[0126] 8. The generated corpus should be segmented into sentences as reasonably as possible.

[0127] 9. The generated sample format must strictly follow my output format requirements! Do not generate numbers before the output sample.

[0128] [Reference Generation Example] Example 1: Task corpus: {{“instruction”:“Manage shopping cart in Douyin Mall”,“app”:“Douyin Mall”}} Result generated: {{"query":"yoyo I want to manage my shopping cart","plan":"Manage shopping cart","ask":"Which app do you want to use to manage your shopping cart?"","response":"Of course, Douyin Mall","answer":"Manage shopping cart in Douyin Mall"}} Example 2: Task corpus: {{“instruction”:“Complete musician certification in Soda Music”,“app”:“Soda Music”}} Result generated: {{"query": "Quickly complete my musician verification", "plan": "Complete musician verification", "ask": "What app do you need to use to complete musician verification?", "response": "Soda Music", "answer": "Complete musician verification in Soda Music"}} Example 3: Task corpus: {{“instruction”:“Search <search content> in Douyin's store”,“app”:“Douyin's store”}} Result generated: {{"query": "Search for kitchen paper towels", "plan": "Search for kitchen paper towels", "ask": "Which app do you want to search for kitchen paper towels on?", "response": "Um, in Douyin's store", "answer": "Search for kitchen paper towels in Douyin's store"}} Example 4: Task corpus: {{“instruction”:“Check the delivery time of <search content> to <address> in Douyin's store”,“app”:“Douyin's store”}} Result generated: {{"query":"Check when my computer will be delivered","plan":"Check the delivery time of my computer","ask":"Which app do you need to check the delivery time of my computer","response":"Douyin Mall","answer":"Check the delivery time of my computer in Douyin Mall"}} Please note: The above examples are just suggestions; please do not copy them verbatim. You still need to strictly follow the task requirements and rewrite the code to account for various variations.

[0129] Below is the task data provided for you. Task prediction: "instruction": "EA instruction 1", "app": "app in EA instruction 1"; "instruction": "EA instruction 2", "app": "app in EA instruction 2"; ...; "instruction": "EA instruction N", "app": "app in EA instruction N".

[0130] Electronic devices can randomly select a user from multiple users (such as the person set mentioned above), randomly select an emotion from multiple emotions (such as the mood set mentioned above), and randomly select a sentence type from multiple sentence types. Using the GPT model, based on the input EA command and the corresponding APP, they can generate a second dialogue sample that imitates the selected user and the selected emotion, and use the selected sentence type to generate the second dialogue sample corresponding to the EA command.

[0131] The above-mentioned methods for generating the first and second dialogue samples, combined with random and diverse synthesis methods such as user roles, emotions, tone, and sentence structure, ensure that even for the same task of the same EA command / tool, various forms of user expression can be generated, reducing homogeneous user request content and enabling dialogue samples to widely cover real user expressions.

[0132] Optionally, the first and second dialogue samples corresponding to the EA instruction may also include app, class, and tag fields. The app field indicates the app used to execute the instruction, the class field indicates the type or domain of the instruction, and the tag field indicates additional characteristics of the instruction for more refined classification.

[0133] S13, the electronic device merges the first dialogue sample and the second dialogue sample according to the EA instruction to obtain the single instruction sample corresponding to each EA instruction, and the single instruction samples corresponding to each EA instruction form a single instruction sample set B.

[0134] like Figure 6As shown, the first dialogue samples corresponding to each EA instruction in training instruction set A form the first sample set B1, and the second dialogue samples corresponding to each EA instruction in training instruction set A form the second sample set B2. Further, the first and second dialogue samples can be merged according to the EA instruction to obtain the single instruction sample corresponding to that EA instruction. The single instruction samples corresponding to each EA instruction in training instruction set A form the single instruction sample set B. In other words, the single instruction sample set B includes the single instruction sample corresponding to each EA instruction in training instruction set A, which also includes the first and second dialogue samples corresponding to each EA instruction in training instruction set A.

[0135] S14, the electronic device generates multiple atomic samples based on the single instruction sample set B. These multiple atomic samples can include various types. These various types of atomic samples can include single-round single instruction (also known as type 1 atomic samples), single-round multiple instruction (also known as type 2 atomic samples), multiple-round single instruction (also known as type 3 atomic samples), multiple-round multiple instruction (also known as type 4 atomic samples), multiple-round mixed instruction (also known as type 5 atomic samples), etc., which are explained one by one below.

[0136] (1) Single-round single-instruction (also known as the first type of atomic sample) The first type of atomic sample is a single-turn, single-instruction atomic sample, containing only one round of dialogue between the user and the intelligent assistant, and a single plan instruction. An electronic device can generate a first-type atomic sample based on the first dialogue sample corresponding to an EA instruction. Here, a single round of dialogue refers to a process containing only one user request and intelligent assistant response.

[0137] Taking the first EA instruction as an example, the first EA instruction is any EA instruction in the training instruction set A. The first type of atomic sample generated based on the first dialogue sample corresponding to the first EA instruction can include the third request content and the third response content. The third request content is obtained based on the request content (i.e. the first request content) in the first dialogue sample corresponding to the first EA instruction. The third request content includes the target app in the first EA instruction. The third response content is obtained based on the response content (i.e. the first response content) in the first dialogue sample corresponding to the first EA instruction. The third response content includes a plan instruction.

[0138] Specifically, the third request content includes the "query" in the first dialogue sample corresponding to the first EA instruction, and the third response content includes the "answer" in the first dialogue sample corresponding to the first EA instruction.

[0139] For example, the first EA instruction is "Buy a train ticket on 12306". The first dialogue sample corresponding to the first EA instruction is: {"query": "Buy a train ticket on 12306", "answer": "Buy a train ticket on 12306"}.

[0140] The first type of atomic sample generated from the first dialogue sample corresponding to the first EA instruction may include the following (described using JSON format): [{"role":"system","content":"..."} {"role":"user","content":"Help me buy a train ticket on 12306"} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> \n\nplan completed"}] In this context, "role": "system" represents the system role, and {"role": "system", "content": "..."} is used to set the background, rules, or restrictions for the model; it can also be omitted. "role": "user" represents the user role, and {"role": "user", "content": "Buy me a train ticket on 12306"} is the user's request content. "role": "assistant" represents the intelligent assistant role, and {"role": "assistant", "content": "..."} is the user's request content. <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> "\nplan complete" is the response content of the model / intelligent assistant. <plan>The label defines the task planning ("Buy train tickets on 12306 railway") and marks "plan completed".

[0141] (2) Single-round multiple instructions (also known as type II atomic sample) The second type of atomic sample is a single-turn, multi-instruction atomic sample, which contains only one round of dialogue between the user and the intelligent assistant and multiple plan instructions. An electronic device can generate a second type of atomic sample based on the first dialogue sample corresponding to multiple EA instructions.

[0142] Taking N1 EA instructions as an example, N1 EA instructions are any N1 EA instructions in the training instruction set A, where N1 is a positive integer of 2 or greater. The second type of atomic samples generated based on the first dialogue samples corresponding to the N1 EA instructions include: fourth request content and fourth response content. The fourth request content is obtained based on the request content (i.e., the first request content) in the first dialogue samples corresponding to the N1 EA instructions, and includes the target app in the N1 EA instructions. The fourth response content is obtained based on the response content (i.e., the first response content) in the first dialogue samples corresponding to the N1 EA instructions, where N1 is a positive integer of 2 or greater.

[0143] Specifically, the third request content includes the "query" from the first dialogue sample corresponding to the merged N1 EA instructions, and the third response content includes a combination of the "answer" from the first dialogue sample corresponding to the N1 EA instructions. Optionally, the electronic device can also use models such as GPT, Beanbag, and Deepseek to refine the merged "query" pairs to ensure that the expression of the user's request content is natural and fluent.

[0144] For example, taking N1=2, the first dialogue samples corresponding to these two EA instructions are as follows: EA instruction 1 is "Buy a train ticket on 12306". The first dialogue sample corresponding to EA instruction 1 is: {"query": "Buy a train ticket on 12306", "answer": "Buy a train ticket on 12306"}.

[0145] EA instruction 2 is "Open an external page in Meituan". The first dialogue sample corresponding to EA instruction 2 is: {"query": "Order a meal with Meituan", "answer": "Open an external page in Meituan"}.

[0146] The second type of atomic samples generated based on the first dialogue samples corresponding to EA instruction 1 and EA instruction 2 respectively can include the following (described in JSON format): [{"role": "system", "content": "..."} {"role": "user", "content": "Buy me a train ticket on 12306 and order food from Meituan"} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on 12306 (China Railway's official website). Task 2: Open the external page in Meituan (a Chinese online marketplace).< / plan> \n\nplan completed"}] In this context, "role": "system" represents the system role, and {"role": "system", "content": "..."} is used to set background, rules, or restrictions for the model; it can also be omitted. "role": "user" represents the user role, and {"role": "user", "content": "Buy me a train ticket on 12306 and order food from Meituan"} is the user request content, a combination of the "queries" from the first dialogue sample corresponding to the two EA commands. "role": "assistant" represents the intelligent assistant role, and {"role": "assistant", "content": "..."} is used to set background, rules, or restrictions for the model; it can also be omitted. <plan> Task 1: Purchase train tickets on 12306 (China Railway's official website). Task 2: Purchase train tickets on 12306 (China Railway's official website).< / plan> "\nplan complete" is the response content of the model / intelligent assistant. <plan>The tags define Task 1 ("Buy train tickets on 12306" and Task 2 (Open an external page in Meituan) and mark "plan completed".

[0147] (3) Multi-round single instruction (also known as third type of atomic sample) The third type of atomic sample is a multi-turn single-instruction atomic sample, which includes a multi-turn dialogue between the user and the intelligent assistant and a plan instruction. The electronic device can generate a third type of atomic sample based on the second dialogue sample corresponding to an EA instruction.

[0148] Taking the second EA instruction as an example, the second EA instruction is any EA instruction in the training instruction set A. The third type of atomic sample generated based on the second dialogue sample corresponding to the second EA instruction can include: the fifth request content, the second planning instruction, the second follow-up question content, the second answer content, and the fifth response content. Among them, the fifth request content is obtained based on the request content (i.e., the second request content) in the second dialogue sample corresponding to the second EA instruction; the second planning instruction is obtained based on the plan instruction (i.e., the first planning instruction) in the second dialogue sample corresponding to the second EA instruction; the fifth request content does not include the target app in the second EA instruction; the second follow-up question content is obtained based on the follow-up question content (i.e., the first follow-up question content) in the second dialogue sample corresponding to the second EA instruction; the second answer content is obtained based on the answer content (i.e., the first answer content) in the second dialogue sample corresponding to the second EA instruction; and the fifth response content is obtained based on the response content (i.e., the second response content) in the second dialogue sample corresponding to the second EA instruction.

[0149] Specifically, the fifth request content includes "query" in the second dialogue sample corresponding to the second EA instruction, the second planning instruction includes "plan" in the second dialogue sample corresponding to the second EA instruction, the second follow-up question content includes "ask" in the second dialogue sample corresponding to the second EA instruction, the second answer content includes "response" in the second dialogue sample corresponding to the second EA instruction, and the fifth response content includes "answer" in the second dialogue sample corresponding to the second EA instruction.

[0150] For example, the second EA instruction is "Buy a train ticket on 12306". The second dialogue sample corresponding to the second EA instruction is: {"query": "Buy me a train ticket", "plan": "Buy a train ticket", "ask": "Which app would you like to use to buy a train ticket?", "response": "Of course, it's 12306", "answer": "Buy a train ticket on 12306"}.

[0151] The third type of atomic sample generated from the second dialogue sample corresponding to the first EA instruction may include the following (described using JSON format): [{"role": "system", "content": "..."} {"role": "user", "content": "Buy me a train ticket"} {"role":"assistant","content":" <plan> Task 1: Buy train tickets< / plan> Which app would you like to use to buy train tickets? {"role": "user", "content": "Of course it's Railway 12306!"} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> \n\nplan completed"}] In this context, "role": "system" represents the system role, and {"role": "system", "content": "..."} is used to set background, rules, or restrictions for the model; it can also be omitted. "role": "user" represents the user role, and {"role": "user", "content": "Buy me a train ticket on 12306 and order food from Meituan"} is the user request content, a combination of the "query" from the first sample corresponding to the two EA instructions. "role": "assistant" represents the intelligent assistant role, and {"role": "assistant", "content": "..."} is used to set background, rules, or restrictions for the model; it can also be omitted. <plan> Task 1: Purchase train tickets on 12306 (China Railway's official website). Task 2: Purchase train tickets on 12306 (China Railway's official website).< / plan> "\nplan complete" is the response content of the model / intelligent assistant. <plan>The tags define Task 1 ("Buy train tickets on 12306" and Task 2 (Open an external page in Meituan) and mark "plan completed".

[0152] (4) Multi-round multi-instruction (also known as the fourth type of atomic sample) The fourth type of atomic sample is a multi-turn, multi-instruction atomic sample, which includes multiple rounds of dialogue between the user and the intelligent assistant and multiple plan instructions. Electronic devices can generate a fourth type of atomic sample based on the second dialogue samples corresponding to multiple EA instructions.

[0153] Taking N² EA instructions as an example, N² EA instructions are any N² EA instructions in the training instruction set A, where N² is a positive integer of 2 or greater than 2. The fourth type of atomic samples generated based on the second dialogue samples corresponding to the N² EA instructions includes: a sixth request content, N² third planning instructions, N² third follow-up questions, N² third answers, and a sixth response content. The sixth request content is obtained based on the second request content in the second dialogue samples corresponding to the N² EA instructions, and does not include the target app in the N² EA instructions. The N² third follow-up questions are each obtained based on the first follow-up questions in the second dialogue samples corresponding to the N² EA instructions. The N² second answers are each obtained based on the first answers in the second dialogue samples corresponding to the N² EA instructions. The sixth response content is obtained based on the second response content in the second dialogue samples corresponding to the N² EA instructions.

[0154] Specifically, in the first round of dialogue, the sixth request is derived by merging the "query" from the second dialogue samples corresponding to N2 EA commands. The intelligent assistant's response includes "plan" from the second dialogue samples corresponding to N2 EA commands and "ask" from the second dialogue sample corresponding to the first EA command. In the second round of dialogue, the user's response includes "response" from the second dialogue sample corresponding to the first EA command, and the intelligent assistant's response includes "answer" from the second dialogue sample corresponding to the first EA command, "plan" from the second to Nth second dialogue samples, and "ask" from the dialogue sample corresponding to the second EA command. In the third round of dialogue, the user's response includes "response" from the second dialogue sample corresponding to the second EA command, and the intelligent assistant's response includes "answer" from the second dialogue sample corresponding to the first EA command, "answer" from the second dialogue sample corresponding to the second EA command, "plan" from the second dialogue samples corresponding to the third to N2 EA commands, and "ask" from the dialogue sample corresponding to the third EA command. Similarly, in the i-th round of dialogue, the user's response includes the "response" from the second dialogue sample corresponding to the (i-1)-th EA command, and the intelligent assistant's response includes the "answer" from the second dialogue samples corresponding to the 1st to (i-1)-th EA commands, the "plan" from the second dialogue samples corresponding to the i-th to N2-th EA commands, and the "ask" from the i-th dialogue sample. In the final round of dialogue, the user's response includes the "response" from the second dialogue sample corresponding to the N2-th EA command, and the intelligent assistant's response includes the "answer" from the second dialogue samples corresponding to the 1st to N2-th EA commands.

[0155] For example, taking N2=2 as an example, the second dialogue samples corresponding to these two EA instructions are as follows: EA instruction 1 is "Buy a train ticket on 12306". The second dialogue sample corresponding to EA instruction 1 is: {"query": "Buy me a train ticket", "plan": "Buy a train ticket", "ask": "Which app would you like to use to buy a train ticket?", "response": "12306, of course", "answer": "Buy a train ticket on 12306"}; EA instruction 2 is "Open an external page in Meituan". The second dialogue sample corresponding to EA instruction 2 is: {"query": "Let's order some food", "plan": "Order food", "ask": "Which app would you like to use to order food?", "response": "Um, Meituan", "answer": "Open an external page in Meituan"}; The fourth type of atomic samples generated based on the second dialogue samples corresponding to EA instruction 1 and EA instruction 2 respectively can include the following (described in JSON format): [{"role": "system", "content": "..."} {"role": "user", "content": "Buy me a train ticket and order some food."} {"role":"assistant","content":" <plan> Task 1: Buy train tickets< / plan> Task 2: Ordering takeout< / plan> Which app would you like to use to buy train tickets? {"role": "user", "content": "Of course it's 12306 (China Railway)!"} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets through the 12306 railway ticketing platform.< / plan> Task 2: Ordering takeout< / plan> Which app would you like to use to order food? {"role": "user", "content": "Yes, Meituan"} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing platform. Task 2: Open the food delivery page on Meituan.< / plan> \n\nplan completed"}] (5) Single-round mixed instructions (also known as fifth type of atomic sample) The fifth type of atomic sample is a single-round mixed instruction atomic sample, which includes a round of dialogue between the user and the intelligent assistant and multiple plan instructions. The electronic device can generate a fifth type of atomic sample based on a first dialogue sample corresponding to at least one EA instruction and a second dialogue sample corresponding to at least one EA instruction.

[0156] Taking the example of an electronic device generating a fifth-type atomic sample based on a first dialogue sample corresponding to N3 EA instructions and a second dialogue sample corresponding to N4 EA instructions, where N3 EA instructions are any N3 EA instructions in training instruction set A, and N4 EA instructions are any N4 EA instructions in training instruction set A, and N3 and N4 are positive integers not less than 1. Generating a fifth-type atomic sample based on the first dialogue sample corresponding to N3 EA instructions and the second dialogue sample corresponding to N4 EA instructions can include: a seventh request content and a seventh response content. The seventh request content is obtained based on the first request content in the first dialogue sample corresponding to N3 EA instructions and the second request content in the second dialogue sample corresponding to N4 EA instructions. The seventh request content includes the target application in the N3 EA instructions but does not include the target application in the N4 EA instructions. The seventh response content is obtained based on the first response content in the first dialogue sample corresponding to N3 EA instructions and the second response content in the second dialogue sample corresponding to N4 EA instructions.

[0157] Specifically, the content of the seventh request is obtained by merging the "query" in the first dialogue sample corresponding to N3 EA commands and the "query" in the second dialogue sample corresponding to N4 EA commands. The response content of the intelligent assistant includes the "answer" in the first dialogue sample corresponding to N3 EA commands and the "answer" in the second dialogue sample corresponding to N4 EA commands. N3 is a positive integer of 1 or greater than 1, and N4 is a positive integer of 1 or greater than 1.

[0158] For example, taking N3=1 and N4=1 as an example, the first dialogue sample corresponding to EA instruction 1 is: EA instruction 1 is "Buy a train ticket on 12306". The first dialogue sample corresponding to EA instruction 1 is: {"query": "Buy a train ticket on 12306", "answer": "Buy a train ticket on 12306"}.

[0159] The second dialogue sample corresponding to EA instruction 2 is: EA instruction 2 is "Open an external page in Meituan". The second dialogue sample corresponding to EA instruction 2 is: {"query": "Let's order some food", "plan": "Order food", "ask": "Which app would you like to use to order food?", "response": "Um, Meituan", "answer": "Open an external page in Meituan"}; The fifth type of atomic sample generated based on the first dialogue sample corresponding to EA instruction 1 and the second dialogue sample corresponding to EA instruction 2 may include the following (described in JSON format): [{"role": "system", "content": "..."} {"role": "user", "content": "Please buy me a train ticket on 12306 and then order some food."} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing platform. Task 2: Open the food delivery page on Meituan.< / plan> \n\nplan completed"}] Missing app information in the fifth type of atomic sample is handled through inheritance.

[0160] In some embodiments, such as Figure 7 As shown, one implementation of step S14 above may include, but is not limited to, some or all of the following steps: S141, the electronic device stores the sample set B corresponding to each EA instruction into the first cache pool.

[0161] The first cache pool comprises multiple first cache units, each storing a single instruction sample (including a first dialogue sample and a second dialogue sample corresponding to the EA instruction) for a given EA instruction. The electronic device can randomly store the single instruction sample corresponding to each EA instruction into the first cache pool. It should be understood that the sizes of the cache units can differ. The size of a cache unit can be determined by the size of the single instruction sample it stores.

[0162] S142, the electronic device uses a sliding window of length 1 to slide on the cache unit of the first cache pool, reads the first dialogue sample in the sliding window each time it slides, and generates a first type of atomic sample based on the first dialogue sample.

[0163] Electronic devices can be generated according to the data format shown in the first type of atom sample above. They can also be generated according to other data formats.

[0164] Each type 1 atom sample forms a set C1 of single-round single-instruction atom samples (that is, the set of type 1 atom samples).

[0165] S143, the electronic device slides a sliding window of length N1 across the cache units of the first cache pool. Each slide reads N1 first dialogue samples within the sliding window and generates a second type of atomic sample based on the N1 first dialogue samples. N1 is a positive integer of 2 or greater than 2.

[0166] Electronic devices can be generated according to the data format shown in the second type of atom samples above. They can also be generated according to other data formats.

[0167] Optionally, when generating user request content for the second type of atomic samples, the electronic device can merge the "queries" of the N1 first dialogue samples. It can also use models such as GPT, Doubao, and Deepseek to refine the merged "query" pairs, ensuring the expression of the user request content is natural and fluent. It should be understood that when using models for refinement, the models need to be constrained to ensure that the app in the query remains consistent before and after refinement.

[0168] Each second-class atom sample forms a set C2 of single-round single-instruction atom samples (that is, a set of second-class atom samples).

[0169] It should be understood that the length N1 can be a variable, meaning that the size of the sliding window used for different second-type atom samples can be different.

[0170] S144, the electronic device slides a sliding window of length 1 on the cache unit of the first cache pool, reads the second dialogue sample in the sliding window each time it slides, and generates a third type of atomic sample based on the read second dialogue sample.

[0171] Electronic devices can be generated according to the data format shown in the third type of atom sample above. They can also be generated according to other data formats.

[0172] Each third-class atom sample forms a set C3 of multi-round single-instruction atom samples (that is, a set of third-class atom samples).

[0173] S145, the electronic device slides a sliding window of length N2 across the cache cells of the first cache pool. Each slide reads N2 second dialogue samples within the sliding window and generates a fourth type of atomic sample based on these N2 second dialogue samples. N2 is a positive integer of 2 or greater than 2.

[0174] Optionally, the electronic device can be generated according to the data format shown in the fourth type of atom sample above. It can also be generated according to other data formats.

[0175] Optionally, when generating the user request content in the first round of dialogue in the fourth type of atomic samples, the electronic device can merge the "queries" of the N2 second dialogue samples. It can also use models such as GPT, Doubao, and Deepseek to refine the merged "query" pairs to ensure the expression of the user request content is natural and fluent. It should be understood that when using models for refinement, the models need to be constrained so that if the query does not contain "app," the refined version will also not contain "app."

[0176] Each fourth-class atom sample forms a set C4 of multi-round multi-instruction atom samples (that is, a set of fourth-class atom samples).

[0177] It should be understood that the length N2 can be a variable, meaning that the size of the sliding window used for different fourth-type atom samples can be different.

[0178] S146, the electronic device uses a sliding window of length N3+N4 to slide across the cache unit of the first cache pool. Each slide reads N3 first dialogue samples and N4 second dialogue samples within the sliding window, and generates a fifth type of atomic sample based on the N3 first dialogue samples and N4 second dialogue samples. N3 is a positive integer of 1 or greater than 1, and N4 is a positive integer of 1 or greater than 1.

[0179] Optionally, the electronic device can be generated according to the data format shown in the fifth type of atom sample above. It can also be generated according to other data formats.

[0180] Optionally, when generating the user request content in the first round of dialogue in the fifth type of atomic samples, the electronic device can merge the "queries" of the N3 first dialogue samples and the N4 second dialogue samples. It can also use models such as GPT, Doubao, and Deepseek to refine the merged "query" pairs to ensure the expression of the user request content is natural and fluent. It should be understood that when using models for refinement, the model needs to be constrained so that the "app" in the query remains consistent before and after refinement; if the query does not contain "app," then it will not contain "app" after refinement.

[0181] Each fifth-class atom sample forms a set C5 of multi-round mixed instruction atom samples (that is, a set of fifth-class atom samples).

[0182] It should be understood that the lengths N3 and N4 can be variables, meaning that the size of the sliding window used for different fifth-class atomic samples can be different, and the number of first and second dialogue samples used can be different.

[0183] It should be understood that the execution order of S142-S145 above is not important.

[0184] S15, the electronic device is based on splicing multiple atomic samples to obtain a combined sample.

[0185] In one possible implementation of step S15, such as Figure 8 As shown, it may include some or all of the following steps: S151, the electronic device may randomly sample atomic samples from the aforementioned sets C1-C5 and store them sequentially in a second buffer pool, wherein the second buffer pool comprises multiple buffer units, each buffer unit being used to store one atomic sample. The electronic device may also randomly store the sampled atomic sample set into the second buffer pool. It should be understood that the sizes of the various buffer units may differ. The size of a buffer unit may be determined by the size of the atomic samples it stores.

[0186] S152, the electronic device reads the atomic sample of the i-th cache unit from the second cache pool.

[0187] Initially, i = 1. I is no greater than the total number of cache units in the second cache.

[0188] S153, the electronic device determines whether the total number of dialogue rounds of the read atomic samples is not less than a first threshold (such as 4, 5 or other values). If not, then set i=i+1 and execute S152-S153; otherwise, if yes, then execute S154.

[0189] S154, the electronic device splices the read atomic samples to obtain a combined sample.

[0190] After S154, the electronic device can set i=i+1 to construct the next combined sample.

[0191] The above method involves the electronic device selecting and combining atomic samples stored in one or more consecutive cache units from the second cache pool, ensuring that the total number of dialogue rounds in the combined samples is not less than the first threshold, thereby guaranteeing the complexity of the combined samples and improving the quality of the samples.

[0192] For example, atomic sample 1 is as follows (described using JSON format): [{"role": "system", "content": "..."}, {"role": "user", "content": "Please buy me a train ticket on 12306"} {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> \n\nplan completed"}] Atomic sample 2 is as follows: [{"role":"user","content":"Order some claypot rice and check out hotels for outbound travel"}] {"role":"assistant","content":" <plan> Task 1: Ordering takeout< / plan> Task 2: Check outbound hotels< / plan> Which app would you like to use to order food? {"role":"user","content":"Meituan Bar"}, {"role":"assistant","content":" <plan> Task 1: View claypot rice on Meituan< / plan> Task 2: Check out outbound hotels\n\np On which platform do you plan to check out outbound hotels? {"role":"user","content":"Let's stick with Meituan"} {"role":"assistant","content":" <plan> Task 1: View claypot rice on Meituan Task 2: View outbound hotels on Meituan< / plan> \n\nplan completed"}] The combined sample obtained by concatenating atomic sample 1 and atomic sample 2 can include the following content (using JSON format as an example): {"role":"system","content":"..."}, {"role":"user","content":"Help me buy a train ticket on 12306"}, {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> "\nplan completed" {"role":"user","content":"Order some claypot rice and check out hotels for outbound travel"} {"role":"assistant","content":" <plan> Task 1: Ordering takeout< / plan> Task 2: Check out hotels for outbound travel. Which app would you like to use to order food? {"role":"user","content":"Meituan Bar"}, {"role":"assistant","content":" <plan> Task 1: View claypot rice on Meituan< / plan> Task 2: Check out outbound hotels\n\np On which platform do you plan to check out outbound hotels? {"role":"user","content":"Let's stick with Meituan"} {"role":"assistant","content":" <plan> Task 1: View claypot rice on Meituan Task 2: View outbound hotels on Meituan< / plan> \n\nplan completed"}] The above method first generates corresponding first dialogue samples and second dialogue samples for each EA instruction in multiple vertical domains. Then, it generates various types of atomic samples based on the first and second dialogue samples corresponding to multiple EA instructions. Finally, it combines multiple atomic samples to obtain training samples that can be used to train the task planning model. The atomic samples generated by this method can cross vertical domains, and the combined samples obtained by randomly combining atomic samples can cover multiple scenarios, complex scenarios, and complex tasks, thereby improving the quality of the generated combined samples and reducing the cost of synthesizing samples.

[0193] In addition, the above method, which constructs atomic samples based on sampling dialogue samples in the first buffer pool, can avoid the overlap of atomic samples. The method of constructing combined samples by sampling atomic samples in the second buffer pool can realize randomized sampling and combination of various atomic types of samples, realize the construction of a wide variety of complex samples, and reduce sample duplication.

[0194] After obtaining multiple combined samples, the training device can use these combined samples to train the task planning model.

[0195] Among them, task planning models can employ LLM, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), etc.

[0196] The method for training a task planning model can be as follows: the training device can use the dialogue content of the intelligent assistant "assistant" in the combined samples as the ground truth (i.e., the label), use the dialogue information before the label as the input of the task planning model, obtain the prediction information, calculate the loss between the prediction information and the ground truth, and adjust the model parameters of the task planning model based on the loss to minimize the loss.

[0197] For example, inputting the user's dialogue information {"role": "user", "content": "Buy me a train ticket on 12306"} from the combined sample in the above example into the task planning model, we obtain the first predicted information. Then, we calculate the first predicted information and the first truth value {"role": "assistant", "content": " <plan> Task 1: View claypot rice on Meituan Task 2: View outbound hotels on Meituan< / plan> The first loss after "\n\nplan completed"}.

[0198] The dialogue information is: [{"role": "user", "content": "Help me buy a train ticket on 12306"}, {"role": "assistant", "content": " <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> The second prediction information is obtained by inputting the following into the task planning model: {"role": "user", "content": "Order a claypot rice and check outbound hotels"}. The second prediction information is then compared with the second true value {"role": "assistant", "content":}. <plan> Task 1: View claypot rice on Meituan< / plan> \ntask2: Check outbound hotels\n\np On which platform do you plan to check outbound hotels? "} The second loss.

[0199] The dialogue information is: [{"role":"user","content":"Help me buy a train ticket on 12306"}, {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> plan completed”}, {“role”:“user”,“content”:“Order a claypot rice and check out departure hotels”}, {“role”:“assistant”,“content”:<plan> Task 1: Ordering takeout< / plan> \ntask2: View outbound hotels\n\np Which app would you like to order food from? ”}, {“role”:“user”, “content”:“Meituan Bar”}] are input into the task planning model to obtain the third prediction information, and the third prediction information and the third truth value {“role”:“assistant”, “content”: <plan> Task 1: View claypot rice on Meituan< / plan> \ntask2: Check outbound hotels\n\np On which platform do you plan to check outbound hotels? "} is the third loss.

[0200] The dialogue information is: [{"role":"user","content":"Help me buy a train ticket on 12306"}, {"role":"assistant","content":" <plan> Task 1: Purchase train tickets on the 12306 railway ticketing website.< / plan> plan completed”}, {“role”:“user”,“content”:“Order a claypot rice and check out departure hotels”}, {“role”:“assistant”,“content”: <plan> Task 1: Ordering takeout< / plan> \ntask2: View outbound hotels\n\np Which app would you like to order food from? ”}, {“role”:“user”,“content”:“Meituan Bar”}, {“role”:“assistant”,“content”: <plan> Task 1: View claypot rice on Meituan< / plan> \ntask2: View outbound hotels\n\np Which platform do you plan to use to view outbound hotels? ”}, {“role”:“user”,“content”:“Meituan”}] are input into the task planning model to obtain the fourth prediction information. The fourth prediction information and the fourth truth value {“role”:“assistant”,“content”: <plan> Task 1: View claypot rice on Meituan Task 2: View outbound hotels on Meituan< / plan> The fourth loss is "\n\nplan completed"}.

[0201] Furthermore, the training device can calculate the total loss based on the first loss, second loss, third loss, and fourth loss, and adjust the model parameters of the task planning model to minimize the total loss.

[0202] The above example uses a single combined sample. In actual training, multiple samples can be used for multiple rounds of training.

[0203] After training the task planning model, the terminal can use the trained model to plan tasks based on user requests. Specifically, see the following... Figure 9 The data processing method shown may include, but is not limited to, some or all of the following steps: S21, the terminal receives the request content input by the user.

[0204] S22, The terminal identifies the intent behind the requested content.

[0205] S23, when the terminal recognizes that the intent belongs to a third-party user, it inputs the request content into the task planning model and obtains multiple tasks.

[0206] S24, the terminal executes these multiple tasks sequentially.

[0207] In another implementation, the terminal can also send the request content to the cloud, where the cloud identifies the intent of the request. If the identified intent pertains to a third-party user, the cloud inputs the request content into a task planning model, generating multiple tasks, which are then sent to the terminal. The terminal then executes these tasks sequentially.

[0208] This data processing method can also be found in the above-mentioned... Figures 2A-2G The description of the scenario shown will not be repeated here.

[0209] This application also provides a computer program product, which includes a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the method executed by the electronic device or terminal in any of the above embodiments.

[0210] This application also provides a computer-readable storage medium storing a computer program (also referred to as code or instructions). When the computer program is run, it causes the computer to perform the method executed by the electronic device or terminal in any of the above embodiments.

[0211] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.

[0212] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to this application are generated, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive).

[0213] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A sample generation method, characterized in that, The method includes: Obtain a first instruction set, which includes execution agent (EA) instructions for multiple vertical domains; For each EA instruction in the first instruction set, a first dialogue sample and a second dialogue sample corresponding to the EA instruction are generated. The EA instruction is used to instruct the use of the target application to perform an operation. The first dialogue sample includes a first request content and a first response content. The second dialogue sample includes a second request content, a first planning instruction, a first follow-up question content, a first answer content, and a second response content. The first request content includes the target application, the second request content does not include the target application, the first planning instruction does not include the target application, the first follow-up question content is used to inquire about the application to be used, the first answer content is the answer to the first follow-up question content and includes the target application, and both the first response content and the second response content are used to instruct the use of the target application to perform an operation. Multiple atomic samples are generated based on the first and second dialogue samples corresponding to the multiple EA instructions, respectively. These multiple atomic samples include various types of atomic samples: first type, second type, third type, fourth type, and fifth type. The first type of atomic sample is a single-round, single-instruction atomic sample generated from the first dialogue sample corresponding to one EA instruction. The second type of atomic sample is a single-round, multi-instruction atomic sample generated from the first dialogue samples corresponding to multiple EA instructions. The third type of atomic sample is a multi-round, single-instruction atomic sample generated from the second dialogue sample corresponding to one EA instruction. The fourth type of atomic sample is a multi-round, multi-instruction atomic sample generated from the second dialogue samples corresponding to multiple EA instructions. The fifth type of atomic sample is a single-round, mixed-instruction atomic sample generated from at least one first dialogue sample corresponding to at least one EA instruction and at least one second dialogue sample corresponding to at least one EA instruction. The multiple atomic samples are concatenated to obtain combined samples, which are used to train a task planning model. The task planning model, when a user's input request is identified as belonging to a task completed by calling a third-party application, plans one or more tasks based on the user's input request, including a first task. When the first task lacks application information, the task planning model outputs follow-up questions to request the application used by the first task from the user. When the first task contains application information, the model outputs a task containing information about the application used to execute the output task.

2. The method as described in claim 1, characterized in that, The generation of the first dialogue sample and the second dialogue sample corresponding to the EA instruction includes: Select one user from a variety of users; Choose one emotion from a variety of emotions; Choose one sentence type from a variety of sentence types; Using a large language model, based on the EA instruction and the target application corresponding to the EA instruction, a selected user and a selected emotion are simulated, and the first dialogue sample corresponding to the EA instruction is generated using a selected sentence type.

3. The method as described in claim 1, characterized in that, The generation of the first dialogue sample and the second dialogue sample corresponding to the EA instruction includes: Select one user from a variety of users; Choose one emotion from a variety of emotions; Choose one sentence type from a variety of sentence types; Using a large language model, based on the EA instruction and the target application corresponding to the EA instruction, a selected user and a selected emotion are simulated, and a second dialogue sample corresponding to the EA instruction is generated using a selected sentence type.

4. The method according to any one of claims 1-3, characterized in that, Multiple atomic samples are generated based on the first dialogue sample and the second dialogue sample corresponding to the multiple EA instructions, including: A first type of atomic sample is generated based on a first dialogue sample corresponding to a first EA instruction. The first instruction set includes the first EA instruction, and the plurality of atomic samples include the first type of atomic samples. The first type of atomic sample includes a round of dialogue between the user and the intelligent assistant. The first type of atomic sample includes a third request content and a third response content. The third request content is obtained based on the first request content in the first dialogue sample corresponding to the first EA instruction. The third request content includes the target application in the first EA instruction. The third response content is obtained based on the first response content in the first dialogue sample corresponding to the first EA instruction.

5. The method according to any one of claims 1-3, characterized in that, Multiple atomic samples are generated based on the first dialogue sample and the second dialogue sample corresponding to the multiple EA instructions, including: A second type of atomic sample is generated based on the first dialogue sample corresponding to N1 EA instructions. The first instruction set includes the N1 EA instructions, and the multiple atomic samples include the second type of atomic sample. The second type of atomic sample includes a round of dialogue between the user and the intelligent assistant. The second type of atomic sample includes: a fourth request content and a fourth response content. The fourth request content is obtained based on the first request content in the first dialogue sample corresponding to each of the N1 EA instructions. The fourth request content includes the target application in the N1 EA instructions. The fourth response content is obtained based on the first response content in the first dialogue sample corresponding to each of the N1 EA instructions. N1 is a positive integer of 2 or greater than 2.

6. The method according to any one of claims 1-3, characterized in that, Multiple atomic samples are generated based on the first dialogue sample and the second dialogue sample corresponding to the multiple EA instructions, including: A third type of atomic sample is generated based on the second dialogue sample corresponding to the second EA instruction. The first instruction set includes the second EA instruction, and the multiple atomic samples include the third type of atomic sample. The third type of atomic sample includes multiple rounds of dialogue between the user and the intelligent assistant. The third type of atomic sample includes: a fifth request content, a second planning instruction, a second follow-up question content, a second answer content, and a fifth response content. The fifth request content is obtained based on the second request content in the second dialogue sample corresponding to the second EA instruction. The fifth request content does not include the target application in the second EA instruction. The second planning instruction is obtained based on the first planning instruction in the second dialogue sample corresponding to the second EA instruction. The second follow-up question content is obtained based on the first follow-up question content in the second dialogue sample corresponding to the second EA instruction. The second answer content is obtained based on the first answer content in the second dialogue sample corresponding to the second EA instruction. The fifth response content is obtained based on the second response content in the second dialogue sample corresponding to the second EA instruction.

7. The method according to any one of claims 1-3, characterized in that, Multiple atomic samples are generated based on the first dialogue sample and the second dialogue sample corresponding to the multiple EA instructions, including: A fourth type of atomic sample is generated based on the second dialogue samples corresponding to N2 EA instructions. The first instruction set includes the N2 EA instructions, and the multiple atomic samples include the fourth type of atomic samples. The fourth type of atomic samples includes multiple rounds of dialogue between the user and the intelligent assistant. The fourth type of atomic samples includes: a sixth request content, N2 third planning instructions, N2 third follow-up questions, N2 third answer content, and a sixth response content. The sixth request content is obtained based on the second request content in the second dialogue samples corresponding to the N2 EA instructions. The sixth request content does not include the target application in the N2 EA instructions. The N2 third planning instructions are obtained based on the first planning instructions in the second dialogue samples corresponding to the N2 EA instructions. The N2 third follow-up questions are obtained based on the first follow-up questions in the second dialogue samples corresponding to the N2 EA instructions. The N2 third answer content is obtained based on the first answer content in the second dialogue samples corresponding to the N2 EA instructions. The sixth response content is obtained based on the second response content in the second dialogue samples corresponding to the N2 EA instructions. N2 is a positive integer of 2 or greater than 2.

8. The method according to any one of claims 1-3, characterized in that, Multiple atomic samples are generated based on the first dialogue sample and the second dialogue sample corresponding to the multiple EA instructions, including: A fifth type of atomic sample is generated based on the first dialogue sample corresponding to N3 EA instructions and the second dialogue sample corresponding to N4 EA instructions. The first instruction set includes the N3 EA instructions and the N4 EA instructions. The multiple atomic samples include the fifth type of atomic sample. The fifth type of atomic sample includes a round of dialogue between the user and the intelligent assistant. The fifth type of atomic sample includes: a seventh request content and a seventh response content. The seventh request content is obtained based on the first request content in the first dialogue sample corresponding to the N3 EA instructions and the second request content in the second dialogue sample corresponding to the N4 EA instructions. The seventh request content includes the target application in the N3 EA instructions but does not include the target application in the N4 EA instructions. The seventh response content is obtained based on the first response content in the first dialogue sample corresponding to the N3 EA instructions and the second response content in the second dialogue sample corresponding to the N4 EA instructions. N3 and N4 are both positive integers greater than or equal to 1.

9. The method according to any one of claims 1-3, characterized in that, The method further includes: The first dialogue sample and the second dialogue sample corresponding to each EA instruction in the first instruction set are stored in the first cache pool. The first cache pool includes multiple cache units, and one cache unit is used to store the first dialogue sample and the second dialogue sample corresponding to one EA instruction.

10. The method according to any one of claims 1-3, characterized in that, The process of splicing together the multiple atomic samples to obtain a combined sample includes: The plurality of atomic samples are stored in a second cache pool, the second cache pool comprising a plurality of cache units, each cache unit being used to store one atomic sample; Read an atomic sample from a cache unit in the second cache pool; Determine whether the total number of dialogue rounds for the atomic samples already read is not less than the first threshold; When the total number of dialogue rounds for the atomic samples already read is less than the first threshold, read the atomic samples in the next cache unit; When the total number of dialogue rounds of the read atomic samples is not less than the first threshold, the read atomic samples are spliced ​​together to obtain a combined sample.

11. A model training method, characterized in that, The method includes: Obtain multiple combined samples, wherein the combined samples are generated by the method described in any one of claims 1-10; The task planning model is trained based on the multiple combined samples.

12. A data processing method, characterized in that, include: Receive user input request content; The request content is input into the task planning model to obtain multiple tasks. The task planning model is trained using the model training method described in claim 11. The multiple tasks are executed sequentially.

13. An electronic device, characterized in that, The device includes one or more processors and one or more memories, the one or more processors being coupled to the one or more memories, the one or more memories being used to store computer program code, the computer program code including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by an electronic device, implements the method as described in any one of claims 1-12.

15. A computer program product, characterized in that, Includes a computer program that, when executed by an electronic device, implements the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Dialogue model training method and device, electronic equipment and storage medium

    CN117033582A

  • Large language model training method, training data acquisition method and intention recognition method

    CN120429382A

  • Multi-modal sample data generation method and device, electronic equipment and storage medium

    CN121365246A