Training data generation method, electronic equipment, storage medium and program product

By generating system-level prompt information and conducting multiple rounds of dialogue expansion, the problems of low efficiency in training dataset construction and poor domain adaptability in existing technologies are solved, and efficient and automated training data generation is achieved, which is suitable for business needs in different fields.

CN120632467AActive Publication Date: 2025-09-12SHANGHAI COOPERS TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511101859.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-09-12
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing technologies for constructing high-quality command-response training datasets suffer from low efficiency, poor domain adaptability, and a lack of consistency in conversational semantics, making it difficult to meet the needs of large-scale data construction.

Method used

By obtaining the domain identification information of the target domain, system-level prompt information is generated. The input instruction set and response set are automatically generated using the aligned and trained large language model, multi-round dialogue expansion is performed, and a multi-round dialogue training dataset with semantic coherence is constructed.

Benefits of technology

It significantly improves the automation level of the command-response construction process, enhances the domain adaptability of data, and has excellent versatility, portability, and cross-domain scalability. The generated data can quickly adapt to business needs in different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632467A_ABST
    Figure CN120632467A_ABST
Patent Text Reader

Abstract

The invention provides a training data generation method, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining domain identification information of a target domain, and generating system-level prompt information related to the target domain; inputting the system-level prompt information into the large language model after alignment training is completed, and driving the model to generate an input instruction set related to the target domain; based on the input instruction set, generating a response set semantically related to the input instruction set, thereby forming a first training data set; constructing multi-round dialogue training data with semantic coherence and rich context for each instruction-response pair in the first training data set through a multi-round dialogue extension generation mode; and summarizing the multi-round dialogue training data to construct a second training data set. The training data generated by the method does not depend on manual prompting engineering, expert writing or seed instruction presetting, can quickly adapt to business requirements in different fields, and has excellent universality, mobility and cross-field expansibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training data generation method, electronic device, storage medium, and program product. Background Art

[0002] In recent years, large language models (LLMs) have made significant progress in natural language processing, demonstrating powerful language understanding and generation capabilities in tasks such as general conversation, question-answering, and code generation. To further enhance the model's adaptability to specific tasks, researchers commonly use instruction tuning techniques. This technique aligns the model's behavior with high-quality instruction-response data, enhancing its generalization capabilities and enabling it to handle new tasks not encountered during training.

[0003] At present, although some open source models such as GPT, Qwen, and LLaMA have opened their parameter weights to provide developers with a model foundation, most of the corresponding high-quality aligned training datasets are still private, limiting the open research progress of artificial intelligence technology and the expansion of its practical application scope.

[0004] In the existing technology, the instruction data construction method still has the following problems: 1. One type of method relies on experts to manually write instructions and annotate responses. Although the data quality is relatively high, it has problems such as high cost, low efficiency, and poor scalability, making it difficult to meet the needs of large-scale data construction; Another type of automated method, such as Self-Instruct, attempts to use the model itself to generate training samples to reduce labor costs. Although this improves data generation efficiency, in practice it still relies heavily on the design of initial seed instructions or prompt templates, and its generation quality still requires a lot of manual intervention. 3. As the scale of automatically synthesized data continues to expand, existing generation mechanisms tend to be homogenized in terms of language expression and task types, making it difficult to maintain language diversity and the breadth of task coverage.

[0005] Therefore, existing instruction dataset construction methods still have problems such as strong manual dependence, weak professional adaptation, insufficient multi-round construction capabilities and decreased data diversity. How to efficiently construct an instruction dataset with wide coverage and stable quality has become an urgent problem that needs to be solved in this field. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present application provides a training data generation method, electronic device, storage medium and program product, which are at least used to solve the problems in the existing technology of low efficiency, poor domain adaptability and lack of semantic coherence in constructing high-quality command-response training data.

[0007] In order to achieve the above objectives and other advantages, some embodiments of the present application provide the following aspects: In a first aspect, some embodiments of the present application provide a method for generating training data, comprising: Acquire domain identification information of a target domain, and generate system-level prompt information related to the target domain based on the domain identification information; Inputting the system-level prompt information into a large language model that has completed alignment training to drive the model to generate an input instruction set related to the target domain; The input instruction set is used as an input prompt and is further input into the large language model to generate a response set semantically related to the input instruction set, thereby forming a first training data set including the input instruction set and the response set; For each command-response pair in the first training dataset, construct multi-turn dialogue training data with semantic coherence and rich context through the multi-turn dialogue extension generation method of the large language model; Aggregate the plurality of multi-round dialogue training data to construct a second training dataset for instruction tuning training or supervised alignment tasks of a large language model.

[0008] In a second aspect, some embodiments of the present application further provide an electronic device, comprising: One or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processors perform any one of the training data generation methods described above.

[0009] In a third aspect, some embodiments of the present application further provide a computer-readable storage medium having stored thereon a computer program and / or instructions, which, when executed by a processor, implements any of the training data generation methods described above.

[0010] In a fourth aspect, some embodiments of the present application further provide a computer program product, comprising a computer program and / or instructions, which, when executed by a processor, implements any of the training data generation methods described above.

[0011] Compared with the related art, in the solution provided by the embodiment of the present application, by introducing system-level prompt information constructed based on domain identification information of different target fields, the large language model that has completed alignment training is guided to automatically generate a user instruction set with semantic diversity and expression consistency, and further generate response content that is highly consistent with the instruction semantics, thereby constructing a first training data set with standardized structure and accurate semantics. This method does not rely on manual prompt engineering, expert writing or preset seed instructions, and significantly improves the degree of automation and task generalization ability of the instruction-response construction process. Furthermore, by performing multiple rounds of context expansion on the instruction-response pairs in the first training data set, multi-round dialogue training data with semantic coherence and logical progression is constructed, which enhances the data's adaptability to complex interactive scenarios and context understanding ability. Therefore, the training data generated by this method has good domain adaptability, so that the generated data content can quickly adapt to the business needs of different fields, and has excellent versatility, portability and cross-domain scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other implementation methods can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 This is a flow chart of a training data generation method provided in an embodiment of the present application; Figure 2 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0014] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0015] First embodiment

[0016] The first embodiment of the present application relates to a method for generating training data, referring to Figure 1 As shown, the method may include the following steps: Step S1: Acquire domain identification information of a target domain, and generate system-level prompt information related to the target domain based on the domain identification information.

[0017] For step S1, specifically, the target domain refers to the specific task or application scenario to which the training data generation method of this application is applicable, which represents a specific topic or professional scope in the data generation process. For example, medical, legal, financial, educational, automotive and other industries. In the medical field, tasks may involve disease diagnosis, treatment methods, medical image analysis, etc. In the legal field, tasks may include contract analysis, legal consulting, case judgment analysis, etc.; they can also be professional knowledge areas of specific disciplines such as computer science, physics, chemistry, such as computer science programming code generation, technical document writing and other tasks; or they can be specific applications or task requirements, such as technical support, customer service, educational counseling, etc., involving FAQs, technical support and other tasks.

[0018] Domain identification information refers to key information used to describe and define the characteristics of the target domain in a specific application scenario. Core attributes of a domain can be identified through specific labels, parameters, or metadata, helping the system understand and generate data tailored to the domain's requirements. Domain identification information is a core input in the automated data generation process, directly impacting the model's generative capabilities and the relevance of the generated data. For example, in the medical domain, domain identification information could include "medicine," "disease type," "treatment methods," and "preventive measures." This information can be used to generate system-level prompts that are more focused on the medical task, such as "Please describe common medical diseases and their treatments." For the educational domain, the generated system-level prompt might be "Please describe key teaching methods and learning assessment strategies in this field." Compared to other prompts, which typically require the user to directly input instructions to the model to complete a simple task, system-level prompts are designed to guide the large language model, providing comprehensive task guidance and context, ensuring the structure and consistency of the generated content.

[0019] Step S2: Input the system-level prompt information into the large language model that has completed alignment training, and drive the model to generate an input instruction set related to the target domain.

[0020] Specifically, in step S2, alignment training involves fine-tuning model parameters using data from the target domain to better understand that domain's terminology and generate domain-specific content. Therefore, in many applications, especially domain-specific tasks such as healthcare, law, and education, manually designing and generating task instructions and responses is both time-consuming and complex. By utilizing large language models (such as GPT and BERT) that have undergone alignment training, the system can automatically generate high-quality domain-specific data, reducing the need for manual design and annotation and significantly improving the efficiency and scale of data generation.

[0021] After receiving a system-level prompt, the model generates a set of input instructions related to the target domain based on the prompt. These instructions typically consist of a series of task requirements or questions that effectively guide the model in generating responses. For example, in the medical field, the generated input instructions might be: "Please describe the symptoms and treatment of diabetes" or "Please list the common types of cancer and early warning signs."

[0022] To ensure the diversity and coverage of the generated instruction set, the instruction content can be iteratively optimized through multiple rounds of generation or parameter adjustment. For example, by adjusting the model's generation parameters, such as sampling temperature, maximum output length, and repetition penalty coefficient, the model can generate a more diverse instruction set, ensuring that each instruction is semantically unique and practical. The instruction set generated in this way should not be limited to common tasks or problems, but should cover different types of tasks and difficulty levels within the domain.

[0023] Step S3: The input instruction set is used as an input prompt and is further input into the large language model to generate a response set semantically related to the input instruction set, thereby forming a first training data set including the input instruction set and the response set.

[0024] Regarding step S3, specifically, in step S2, an input instruction set related to the target domain has been generated. These instruction sets contain descriptions of specific tasks or problems, which can guide the model to generate responses in specific domains. These instruction sets are input as input prompts and are input one by one into the aligned and trained large language model. Each input instruction contains a clear task goal, domain background and semantic requirements. After receiving the input instructions, the large language model first parses the semantic information of each instruction, including the task goal, context requirements and answer format requirements in the instructions. Based on the task information conveyed in the instructions and the system-level prompt information, the model generates response content related to the semantics of the instructions. These responses should match the instructions semantically and conform to the language style, format and task requirements specified in the system-level prompts. For example, if the input instruction is "Please describe the common symptoms and treatment methods of the disease", the generated response content can include a description of the symptoms of the disease and corresponding treatment recommendations.

[0025] The generated response set undergoes a quality check to ensure semantic integrity and linguistic fluency, avoiding duplicate, irrelevant, or substandard responses. If a generated response does not meet the expected quality standards or semantic alignment requirements, the system will make appropriate adjustments or generate a new response until it meets the standards. Each input command and its corresponding generated response are bound together to form a command-response pair. All generated command-response pairs are aggregated to form the first training dataset.

[0026] Step S4: For each command-response pair in the first training dataset, construct multi-round dialogue training data with semantic coherence and rich context through the multi-round dialogue expansion generation method of the large language model.

[0027] Specifically, in step S4, a first training dataset containing a certain number of command-response pairs has been generated in step S3. For each command-response pair in the first training dataset, step S4 uses a multi-turn dialogue expansion generation method to enhance the semantic coherence and contextual depth of the training data.

[0028] Multi-turn dialogue expansion refers to the system continuously generating subsequent rounds of commands and responses based on a single command-response pair, expanding the breadth and depth of the conversation. When generating multi-turn dialogues, the system retains the content of the previous round as historical context, allowing the semantic continuity and development of the command-response pairs in the next round. This allows the model to not only understand the current task but also semantically expand upon the previously generated content, ensuring logical coherence in the conversation.

[0029] During the multi-round generation process, the system evaluates the quality of each command-response pair to ensure consistency between the response and the command, clarity of meaning, and compliance with the task requirements. If a response is irrelevant, semantically ambiguous, or of low quality, the system will adjust or regenerate it to ensure data quality. If a model-generated response deviates from the task topic, the system will screen it based on semantic similarity and regenerate a more relevant response.

[0030] Step S5: Aggregate multiple multi-round dialogue training data to construct a second training dataset for instruction tuning training or supervised alignment tasks of a large language model.

[0031] Specifically, in step S5, before aggregating the multi-round conversation data, a quality check is performed to ensure that the command-response pairs in each conversation dataset are semantically consistent, task-specific, and coherent. If any semantic errors, task deviations, or formatting issues occur during the conversation generation process, the system will screen and correct them through a quality control mechanism.

[0032] Each multi-turn conversation data sample is converted into a standardized format, typically including information such as the input command, response, task label, and conversation turn. All samples are unified into structured data units, for example: {"Input Command": "Describe diabetes treatment", "Response": "Diabetes treatment includes medication and lifestyle adjustments", "Turn": 1, "Task Type": "Disease Treatment"}. This structured data ensures that the system clearly understands the context and task objectives of each training sample during subsequent model training.

[0033] Through multi-round dialogue expansion, each generated command-response pair is constructed into a complete dialogue sequence. Ultimately, these multi-round dialogue training data are aggregated to form a complete multi-round dialogue dataset, which will serve as training data for different target domains.

[0034] Compared with the related art, in the solution provided by the embodiment of the present application, by introducing system-level prompt information constructed based on domain identification information of different target fields, the large language model that has completed alignment training is guided to automatically generate a user instruction set with semantic diversity and expression consistency, and further generate response content that is highly consistent with the instruction semantics, thereby constructing a first training data set with standardized structure and accurate semantics. This method does not rely on manual prompt engineering, expert writing or preset seed instructions, and significantly improves the degree of automation and task generalization ability of the instruction-response construction process. Furthermore, by performing multiple rounds of context expansion on the instruction-response pairs in the first training data set, multi-round dialogue training data with semantic coherence and logical progression is constructed, which enhances the data's adaptability to complex interactive scenarios and context understanding ability. Therefore, the training data generated by this method has good domain adaptability, so that the generated data content can quickly adapt to the business needs of different fields, and has excellent versatility, portability and cross-domain scalability.

[0035] Second embodiment

[0036] The second embodiment of the present application relates to a method for generating training data. The second embodiment is an improvement on the first embodiment. The specific improvement is that: in the second embodiment of the present application, a specific implementation method for automatically generating system-level prompt information based on target domain identification information is provided, that is, step S1 can further include the following steps: Step S101: Based on the domain classification library or knowledge graph defined by different target fields, domain identification information corresponding to different application scenarios is extracted therefrom. The domain identification information includes: industry name, business type, task category and knowledge subcategory label.

[0037] Specifically, domain identification information can be key information extracted from the domain classification library or knowledge graph of the target domain to identify and describe a specific domain or application scenario. The domain classification library or knowledge graph includes the core knowledge of the domain, task types, common application scenarios, and industry-specific knowledge subcategories.

[0038] Domain identification information includes: industry name, business type, task category, and knowledge subcategory label. For example, in the medical field, domain identification information may include "medicine" as the industry name, "disease diagnosis" as the business type, "symptom identification" as the task category, and "diabetes" as the knowledge subcategory label. In the legal field, domain identification information may include "law" as the industry name, "contract review" as the business type, "intellectual property" as the task category, and "patent law" as the knowledge subcategory label.

[0039] Step S102: Based on at least one content in the domain identification information, a prompt structure entry matching the content is retrieved from a preset prompt template library, and the prompt structure entry is combined with the domain identification information through template filling and parameter binding to dynamically generate system-level prompt information optimized for different target domains.

[0040] A prompt template library is a pre-designed and stored collection of standardized templates that provide a structured framework for different tasks or domains. A prompt structure entry is a specific template instance retrieved from the prompt template library. For example, for a diabetes task in the medical field, relevant templates would be selected from the prompt template library and populated to form a prompt structure entry suitable for the diabetes task.

[0041] In this embodiment, in step S102, the prompt structure entries and the domain identification information are combined through template filling and parameter binding to dynamically generate system-level prompt information optimized for different target domains, specifically including: Step S1021: identifying multiple placeholder fields set in the prompt structure entry, the placeholder fields including a model role field, a task target field, an answer format field, and a style constraint field; Step S1022: semantically map and parameterize each semantic element in the domain identification information with one or more items in the placeholder field, and semantically enhance and modify the answer format and style constraint fields according to different task contexts; Step S1023: Through content insertion and context constraint rules, the filled prompt structure entries are converted into system-level prompt information with complete structure and clear semantics.

[0042] Specifically, based on the prompt structure entries retrieved from the prompt template library, a series of placeholder fields are identified. These placeholder fields are predefined template tags and can include the following categories: Model role field: used to specify the role that the model should assume, helping the model understand its role in generating tasks, such as "Please play the role of a [doctor]" or "As a [legal expert], please answer the following questions."

[0043] Task objective field: This field specifies the specific task that the model is instructed to perform, such as “Please describe the common symptoms of [disease type]” or “Please explain the scope of application of [legal provision]”.

[0044] Answer format field: Set the format requirements when generating answers, such as "Answer should include [symptoms], [treatment methods], [preventive measures]" or "Answer should be concise and up to 150 words."

[0045] Style Constraints: Specifies style constraints to be followed when generating responses, such as "use clear and concise language" or "avoid using technical terms."

[0046] The system performs semantic mapping and parameter filling on each semantic element in the domain identification information and the placeholder fields in the prompt structure entries. Specifically, based on the domain identification information (such as "medicine", "disease type", etc.), the system will match this information with the relevant fields in the placeholder fields (such as model roles, task objectives, etc.) to determine which placeholder field should be filled with which information. For example, "diabetes" contained in the domain identification information is mapped to the "[disease type]" placeholder field in the task objective field. Parameter filling is to fill the specific values ​​or parameters in the domain identification information into the placeholder fields in the template, replace the placeholders determined by semantic mapping with actual content, and generate the final system-level prompt information.

[0047] At the same time, these placeholder fields will be semantically enhanced and modified according to different task contexts. The specific type and goal of the task determine the structural requirements of the generated content. For example, for a disease diagnosis task, it may be necessary to describe it based on dimensions such as symptoms, treatment plans, and preventive measures, while for a legal task, it may be necessary to list the key points of the legal provisions one by one. If the task requires describing the treatment of diabetes, the original prompt structure entry of "Please summarize the treatment of [disease type]" can be enhanced through the task context to "Please summarize the treatment of diabetes in the order of drug treatment, lifestyle adjustment, and surgical treatment." Through this enhancement and modification, the answer format becomes more organized and in line with the task objectives.

[0048] The context and style of different tasks require different linguistic styles. For example, legal tasks may require formal and rigorous language, while tasks in the educational field may prefer a concise and friendly style. If the task requires the model to generate an explanation of a legal provision, the style constraints may require the generated content to "use formal legal terminology" and "avoid overly colloquial expressions." Therefore, the original prompt structure of "Please explain the core content of [legal provision]" can be modified through style enhancement to "Please explain the core content of [legal provision] rigorously, ensuring the use of accurate legal terminology."

[0049] During the content insertion phase, the system further supplements the task with more details based on domain knowledge and task requirements. For example, in a medical task, the system might not only insert the disease name but also, depending on the complexity of the task, further insert information such as treatment plans and symptom descriptions to ensure that the generated instructions are comprehensive and detailed.

[0050] Contextual constraints are rules that ensure logical consistency and semantic coherence when generating system-level prompts. This ensures that the generated prompts are not only semantically clear but also properly guide the model in executing the task. For example, for a medical problem, the system ensures that the generated task instructions are presented in the correct order (e.g., describing the symptoms first, followed by treatment, and then explaining preventive measures). Contextual constraints also ensure the professionalism and applicability of the content. For example, in legal tasks, the generated prompts must adhere to the legal text, and in education, the generated instructions must meet the teaching objectives. Combining content insertion and contextual constraints, the system ultimately generates system-level prompts with well-structured and semantically clear task instructions. These instructions clearly guide the model in performing the corresponding task and ensure the semantic accuracy and logical consistency of the generated content.

[0051] It's easy to see that in this application's examples, the combination of domain identification information and a pre-set template library automatically generates system-level prompt information related to the target domain, avoiding the reliance on manual design and annotation in traditional methods. This allows for rapid generation of task instructions, significantly improving their efficiency and scalability, and meeting the needs of large-scale data generation tasks.

[0052] Third embodiment

[0053] The third embodiment of the present application relates to a method for generating training data. The third embodiment is an improvement on the first embodiment. The specific improvement is that: in the third embodiment of the present application, a specific implementation method for automatically generating batches of input instructions is provided, that is, step S2 can further include the following steps: Step S201: inputting the system-level prompt information as guidance content into the input interface of the large language model, so that the large language model parses and obtains a comprehensive semantic expression for the target domain; Step S202: Based on the comprehensive semantic expression and the configured generation parameters, the model is guided to perform multiple rounds of batch generation to automatically construct an input instruction set of an order of magnitude that matches the set task requirements. The generation parameters include: sampling temperature, maximum output length, repetition penalty coefficient, and diversity sampling strategy; Step S203: The input instruction set can cover multiple sub-class tasks and diverse application scenarios in the target domain, and has diversity in task semantic structure and maintains consistency in expression style.

[0054] Specifically, we first designed an input interface for the large language model to receive formatted system-level prompt information. This interface not only needs to process text information but also be able to parse and understand placeholder fields and template fields within the system-level prompt information. The large language model uses natural language processing (NLP) technology to parse the input system-level prompt information, extracting domain information, task type, format requirements, and other content. Using pre-trained parameters and language knowledge, the model generates task instructions related to the target domain, obtaining a comprehensive semantic representation suitable for that target domain and ensuring that the comprehensive semantic representation is correctly transmitted to the large model.

[0055] Generation parameters are tools for regulating the model's generation of task instructions, helping the system control the diversity and consistency of generated content. By adjusting parameters such as sampling temperature, output length, repetition penalty coefficient, and diverse sampling strategy, the system can generate input instruction sets of a certain scale to match task requirements, ensuring that the generated content is both rich and well-structured.

[0056] The sampling temperature is used to control the randomness and diversity of the generated content. To generate diverse responses, the system sets a higher temperature; to generate more concise, standardized task instructions, the system sets a lower temperature.

[0057] Output length controls the length of each generated instruction, preventing overly long or short output. The system automatically adjusts the maximum output length based on the complexity of the task and the amount of detail required. For example, for tasks requiring detailed descriptions, the system increases the maximum output length; for brief instructions, the system limits the maximum output length.

[0058] The repetition penalty coefficient is used to control the model's ability to avoid duplicate content during the generation process. When the similarity of generated content exceeds a set threshold, the system will penalize duplicate content to avoid redundancy. For example, for tasks requiring high diversity, the system will set a higher repetition penalty coefficient to increase the diversity of generated content.

[0059] A diversity sampling strategy is used to control the diversity in task instruction generation to ensure that the content in the instruction set is semantically and formally diverse. This strategy is adjusted according to the complexity of the task and the required breadth. The system automatically adjusts the appropriate diversity sampling strategy based on the diversity requirements of the target task. For example, when generating multiple instructions related to diabetes, the system will generate different instructions based on the requirements of the task scenario, such as "Description of Diabetes Symptoms", "Treatment Methods of Diabetes", and "Preventive Measures for Diabetes", while avoiding generating instructions with semantic repetition or overly similar content.

[0060] Based on the comprehensive semantic representation and configured generation parameters, the system repeatedly guides the large language model to generate instructions in multiple batches. During each generation round, the system adjusts the generation strategy based on previously generated content or task context to ensure that each instruction is highly relevant to the target task and meets the generation parameter requirements. The system controls the size of the output task instruction set by setting the generation batch size. Specifically, based on the task scale, the system dynamically adjusts the generation batch size based on the required number of task instructions. Furthermore, to avoid excessive resource consumption caused by generating too much content in a single batch, the system generates task instructions in batches. To handle large-scale task instruction generation (e.g., automatic batch generation of 10,000, 100,000, or 1 million instructions), the system can leverage distributed computing frameworks (such as Apache Spark or Hadoop) to distribute the task instruction generation process across multiple compute nodes for parallel processing. Based on the current generation load, generation tasks are automatically assigned to different compute nodes for processing. Large-scale data generation consumes significant computing resources, particularly memory and processing power. The system needs to monitor resource usage in real time and dynamically allocate computing resources. Priority scheduling should be implemented when resources are insufficient to ensure that the generation process is not interrupted due to resource constraints.

[0061] The generation of large-scale task instructions also requires ensuring the quality of the generated content. Quality control and screening should be performed on each batch of generated data. For example, automated quality assessment tools can be used to evaluate the generated task instructions based on criteria such as semantic consistency, task coverage, and generation diversity. Post-processing screening mechanisms can also be used to filter out instructions that do not meet quality requirements. For example, task instructions with duplicate content, incorrect grammar, or unclear semantics should be removed.

[0062] Use professional knowledge graphs or domain classification libraries to analyze the core structure of the target task, understand its key components, and then break it down into different sub-tasks. For example, the task of "health data management" can be broken down into multiple sub-tasks, such as "health data collection," "health data storage," and "health data analysis." For larger and more complex tasks, a recursive approach can be used to split the task into sub-categories. This approach ensures that the task sub-categories meet multiple levels of requirements and provide a more refined task scope.

[0063] Application scenarios are often closely related to the task's execution environment, user needs, or usage context. Identifying specific application scenarios within the target domain allows you to determine the specific requirements for tasks within each scenario. For example, a "health management" scenario might emphasize long-term health tracking and personalized recommendations, while a "disease diagnosis" scenario might focus more on accurate symptom identification and clinical support. Tasks in different scenarios may require different knowledge, technologies, or solutions, so different application scenarios often lead to different task breakdown requirements.

[0064] During the generation process, the system utilizes a diverse sampling strategy for semantic mapping and configuration parameters to ensure that the generated task instructions possess diverse semantic structures, thereby generating multiple instructions with varying semantic expressions. Specifically, during each generation round, by modeling the different levels and details of the task, each instruction will describe the target task from a different perspective or level, ensuring that the generated content is broad enough to cover all sub-tasks and application scenarios in the target domain. At the same time, by adjusting the sampling temperature and repetition penalty coefficient in the generation parameters, the generated content is ensured to have appropriate randomness and diversity.

[0065] It is not difficult to see that in the embodiment of the present application, the comprehensive semantic expression is combined with the generation parameters to effectively improve the efficiency and quality of instruction generation. Through multiple rounds of batch generation, the system can flexibly generate task instructions covering different subtasks and application scenarios according to the specific needs of the target field, avoiding the high cost and low efficiency of manually writing task instructions, while ensuring the quality and consistency of the instruction content. This makes the generation of input instructions more automated, especially suitable for large-scale data generation tasks, and can efficiently support the training data generation needs of multi-field tasks, and improve the adaptability and diversity of tasks while ensuring data quality.

[0066] It should be noted that the third embodiment of the present application may also be an improvement based on any one or more of the first to second embodiments.

[0067] Fourth embodiment

[0068] The fourth embodiment of the present application relates to a method for generating training data. The fourth embodiment is an improvement on the first embodiment. The specific improvement is that: in the fourth embodiment of the present application, a specific implementation method for generating and storing command-response data is provided, that is, step S3 can further include the following steps: Step S301: Each input instruction in the input instruction set is input into the large language model one by one as input content, so that the large language model analyzes the semantic information in the input instruction and generates a response content with structural specifications and semantic consistency based on the answer format and language style requirements set in the system-level prompt information; Step S302: Data binding is performed on each input instruction and its generated response content to establish a unique key-value mapping relationship, and the unique key-value mapping relationship is stored as a structured sample unit to form a command-response pair. Step S303: Serialize and organize the multiple structured sample units according to a preset format, and aggregate them to construct a first training data set.

[0069] Specifically, the system inputs the instructions one by one into the large language model, and relies on the model's natural language understanding ability for the input instructions to parse out the semantic information and generate a structured response to ensure that the response content meets the requirements of the instructions and is consistent with the task objectives in the instructions. At the same time, according to the answer format requirements and language style set in the system-level prompt information, for example, the answer may be required to be concise, clearly structured, and use formal language, to provide format guidance for the large language model. When generating the response content, the system adjusts the output of the model according to preset generation parameters (such as sampling temperature, maximum output length, repetition penalty coefficient, etc.). For example, through a higher sampling temperature, the generated content will show more diversity, avoiding the task instruction generation content being too single and repetitive, and ensuring that the generated instruction set covers a wider range of task requirements.

[0070] Each generated task instruction and its corresponding response are bound together using a key-value pair. For example, a unique identifier (ID) can be used to identify each instruction-response pair. Through this mapping, the system ensures that each input instruction and its response can be uniquely referenced. The system then stores the input instruction, response, and task objective information as a structured sample unit (e.g., in JSON format, CSV format, etc.), allowing the generated data to be efficiently managed and processed.

[0071] Each structured sample unit is processed according to a pre-set format. For example, in JSON format, the system ensures that each field conforms to a predefined standard type (string, integer, array, etc.) and is not missing. During the collation process, the system further sorts, filters, and group the data based on the requirements of different sub-tasks. For example, sorting and grouping can be performed by factors such as task category, generation time, and task difficulty.

[0072] Data filtering includes screening and filtering the command-response pairs in the first training dataset based on a set of quality assessment indicators, wherein the quality assessment indicators include: Set the minimum and maximum length thresholds for input commands and responses, and trim or remove commands and responses that exceed the threshold.

[0073] For example, if the input command length is set to 10-100 characters, samples below this threshold may lack sufficient information and will be eliminated. Samples exceeding this threshold will be trimmed or eliminated to avoid generating overly long and redundant content. Similarly, the response content length is limited to 20-200 characters, and samples exceeding the threshold will be trimmed or eliminated.

[0074] The input instructions and response content are semantically embedded based on the pre-trained language model, and the semantic similarity between the two is calculated to evaluate whether the response content is overall aligned with the semantic intent of the input instructions.

[0075] By converting input commands and responses into high-dimensional vectors, semantic information is captured. Methods such as cosine similarity or Euclidean distance are used to calculate the semantic similarity between the command and response. If the similarity falls below a set threshold, it indicates that the response is semantically inconsistent with the input command, and the system will reject these unqualified samples.

[0076] Based on the semantic points extracted from the input instructions, the coverage of the response content in the context information is evaluated to determine the semantic completeness of the response content.

[0077] Semantic key points are extracted from the input instruction and used as criteria for evaluating the response content. The following criteria can be set to determine the semantic completeness of the response content, such as whether the response covers all key elements mentioned in the input instruction, whether the response fully and comprehensively answers the core question of the instruction, and whether the response content is consistent with the nature and objectives of the task. If the coverage meets the predetermined criteria and the information is sufficient, the response is considered semantically complete.

[0078] Perform discretized statistical analysis on the task objectives or semantic expressions in the input instruction set to measure their coverage breadth in task categories and expression structures.

[0079] Discrete analysis is performed on the task objectives. Task objectives usually refer to the core goals or requirements of the input instructions. For example, in the medical field, the task objectives may be "describe symptoms", "explain treatment methods", "answer health questions", etc. By counting the frequency of occurrence of each task objective, the system can identify the bias of certain task objectives in the data and ensure a balanced distribution of various task objectives. The semantic expression of the task also needs to be discretized to ensure the diversity of expression structure and grammar. For example, analyze whether the generated input instructions use different sentence structures (such as "please describe", "please explain", "give examples", etc.), so as to ensure that the generated data has a diversity of grammar and expression.

[0080] By calculating the minimum semantic distance or word-level difference between adjacent samples, the similarity between adjacent instruction-response pairs is evaluated and highly repetitive or redundant samples are eliminated.

[0081] Semantic distance measures the semantic similarity between two texts, specifically the degree of difference in the actual meaning expressed in the texts. Common methods for calculating semantic distance include cosine similarity, Euclidean distance, and Manhattan distance. Models such as word embeddings or sentence embeddings can be used to convert texts into vector representations, from which similarity can be calculated. Word-level dissimilarity measures similarity by calculating the lexical differences between texts. Methods such as Jaccard similarity and Levenshtein distance can be used to calculate lexical differences between texts. A similarity threshold (such as 0.85) is set. If the similarity between two command-response pairs exceeds the threshold, they are considered redundant and should be eliminated. To more efficiently eliminate redundant examples, the system can use batch processing to calculate the semantic dissimilarity between command-response pairs at a large data scale. For example, parallel processing or distributed computing frameworks can be used to accelerate similarity calculation and data processing, improving system processing efficiency.

[0082] The trained reward model is used to evaluate the quality of the response content, and the instruction-response pairs with scores higher than the set scoring threshold are retained as valid samples.

[0083] Reward models are typically fine-tuned using large-scale pre-trained language models (such as GPT and BERT). During training, a large number of high-quality samples are used to annotate the generated responses and learn how to evaluate their quality. The annotated data used in training may include manual ratings or automatically generated ratings based on specific criteria. After generating a command-response pair, the system uses the trained reward model to evaluate the quality of the response content. The model generates a score for each response based on multiple dimensions, such as grammar, content relevance, and semantic consistency between the command and response. As the amount of data generated increases, the reward model's scoring accuracy can be iteratively optimized by continuously collecting new generated data and manual evaluation feedback, thereby further improving data quality.

[0084] It is not difficult to see that in the embodiment of the present application, by inputting instructions one by one and parsing their semantic information, combined with the format and language style requirements in the system-level prompt information, it is ensured that the response content meets high standards in terms of semantic consistency and structural norms, thereby optimizing the quality of the data. By binding instructions and responses and uniformly storing them as structured sample units, the operability and flexibility of the data set are enhanced, providing a high-quality, easy-to-manage data foundation for subsequent large language model training. The batch-generated and serialized training data sets greatly improve the efficiency of data construction, reduce manual intervention, and promote the rapid generation of large-scale, high-quality training data.

[0085] It should be noted that the fourth embodiment of the present application may also be an improvement based on any one or more of the first to third embodiments.

[0086] Fifth embodiment

[0087] The fifth embodiment of the present application relates to a method for generating training data. The fifth embodiment is an improvement on the first embodiment. Specifically, the improvement is as follows: In the fourth embodiment of the present application, a specific implementation method for expanding the multi-round dialogue of command-response pair data is provided, that is, step S4 can further include the following steps: Step S401: Input each command-response pair in the first training dataset as the context of the dialogue initiation turn into the large language model; Step S402: Based on retaining the previous round of conversation content as historical context, the model is guided to generate a new round of instructions and corresponding response content that are semantically related to the previous round. The new round of user instructions semantically include question asking, task deepening, content reflection, or knowledge generalization; Step S403: Repeat the above generation process to continuously construct multiple rounds of command-response sequences until a preset number of dialogue rounds is reached or a dialogue generation interruption condition is reached; Step S404: Multiple rounds of command-response pairs constitute a multi-round dialogue training data as a whole, which has contextual semantic coherence and logical progressive structure, and is used to enhance the understanding and response capabilities of the large language model in complex dialogue tasks.

[0088] Specifically, the system extracts the initial context from each instruction-response pair in the first training data set, and uses it as the input content for the initial round of dialogue, and inputs it into the large language model one by one. Before inputting the data, the instructions and responses are first preprocessed, such as removing redundant punctuation, irrelevant characters and formatting issues to ensure that the text contains only meaningful content, and the input text is broken down into words or subwords through the vocabulary. After the preprocessing step, the semantic vector representation of the instructions and responses is generated by the pre-trained language model. The language model converts each word or phrase into a corresponding embedding vector, which captures the contextual information and semantic features of the word. The system passes the generated semantic embedding as input to the input interface of the large language model.

[0089] The system converts the previous round's commands and responses into vectors or embedded representations, concatenates them with the newly generated commands and responses, and then feeds them into the model. This allows the model to understand the context of the current task and the previous round of conversation. By retaining the previous round of conversation as historical context, the large language model is guided to generate new commands and responses that are semantically related to the previous round.

[0090] To achieve semantic progression across multiple rounds of conversation, the system dynamically adjusts generation parameters, such as temperature, maximum output length, content diversity, and repetition penalty coefficient, based on task complexity and conversation progress. If the generated commands and responses in the previous round were already relatively basic, the system can appropriately increase the temperature (to improve diversity) and use Top-p sampling to ensure that the generated content is diverse without deviating from the task objective.

[0091] Based on historical context, the large language model generates commands relevant to the current conversational goal and generates corresponding responses based on the semantics of the generated commands. The system ensures that new commands and responses meet the requirements of semantic progression and avoids generating overly simplistic or irrelevant content. Specifically, based on the content of the previous round of conversation, the system generates new commands of various types, such as follow-up questions, task deepening, reflection, or knowledge generalization. Based on the generated commands, the large language model generates responses that are semantically consistent with the commands, ensuring that the responses thoroughly address the task objectives or provide relevant information. In each round of conversation, the generation of commands and responses is accomplished using natural language generation technology. The system analyzes the context of the conversation content and generates language expressions that meet the task objectives. By continuously accumulating historical conversation content, the commands and responses generated in each round are stored in a structured manner. With each round of conversation, the historical context is updated and passed to the next round to ensure conversational coherence and logical progression.

[0092] To ensure that the generated multi-round dialogues meet the needs of actual applications, the system sets conditions for the number of rounds generated or the progression of dialogue content. During the multi-round dialogue generation process, the system will set a maximum number of rounds to ensure that the generated dialogue does not exceed the needs of the actual application. For example, the system can be set to generate no more than 5 rounds of dialogue content. If the generated instructions and responses have covered the required task objectives, the system can also automatically stop generation based on the degree of progression of the dialogue content. When the content reaches a certain depth or complexity, the model will stop generation to avoid meaningless expansion. In the process of each round of generation, the system does not just copy the content of the previous round, but instead semantically progresses based on the historical dialogue content and context to increase the depth of the dialogue.

[0093] When compiling data, the system verifies the semantic coherence of each command-response pair generated in each round. For example, semantic embedding technology is used to calculate the similarity and correlation between the current response and the previous round's response. If a response's semantics are inconsistent with the previous round, or if there are logical gaps, the system will regenerate or adjust the response. The system also analyzes whether each round of instructions effectively expands or deepens the content of the previous round, ensuring that multiple rounds of conversation are not only semantically consistent but also logically progressive, avoiding stagnation or repetition.

[0094] The collated and verified data will be integrated into the final multi-turn dialogue training dataset and serialized and stored in a pre-set format. This data will provide rich training samples for subsequent large-scale language model training, enabling the model to understand and generate coherent and logical responses when handling complex dialogue tasks.

[0095] It is not difficult to find that in the embodiment of the present application, by introducing a multi-round dialogue generation mechanism, the system can efficiently construct a complex dialogue data set. The generation of each round of command-response pairs not only ensures semantic consistency and logical progression, but also effectively enhances the model's ability to respond to complex tasks. By expanding round by round based on historical context, the system can generate dialogue content that contains progressiveness and depth, avoiding simple copying or a single answer method, thereby improving the richness and diversity of the dialogue. This method greatly improves the efficiency and quality of training data generation, provides more high-quality and diverse task data for the training of large language models, and significantly enhances the model's understanding and response capabilities.

[0096] It should be noted that the fifth embodiment of the present application may also be an improvement based on any one or more of the first to fourth embodiments.

[0097] The step division of the above various methods is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0098] In addition, some embodiments of the present application further provide an electronic device. The electronic device may be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device may also be various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0099] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes a training data generation method provided by any one or more of the above embodiments. Figure 2An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing some of the necessary operations. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0100] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means. Figure 2 The bus connection is taken as an example.

[0101] Input device 1103 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. Examples include a touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, and other input devices. Output device 1104 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). The display device may include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0102] To provide user interaction, the electronic device may be a computer. The computer includes a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse) through which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback), and input from the user may be received in any form (e.g., voice input or tactile input).

[0103] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium. When executed by a processor, the computer program / instruction implements a training data generation method provided in any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the device. The computer-readable medium carries one or more computer-readable instructions.

[0104] The memory 1102 can be used as a non-transitory computer-readable storage medium to store non-transitory software programs, non-transitory computer executable programs, and modules. The processor 1101 executes the non-transitory software programs, instructions, and modules stored in the memory 1102 to execute various functional applications and data processing of the server, thereby implementing the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.

[0105] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely located relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0106] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. Computer-readable media may be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.

[0107] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, compact discs, digital versatile discs or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0108] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] In the above embodiments, all or part of the steps or functions of the present invention may be implemented using software, hardware, firmware, or any combination thereof. For example, implementation may be achieved using a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present application may be implemented using hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.

[0110] The computer program product provided in the embodiments of the present application includes one or more computer programs / instructions that, when executed by a processor, fully or partially produce the processes or functions described in accordance with the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).

[0111] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-specific system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0112] The scope of this application is defined by the appended claims rather than the foregoing description and is therefore intended to encompass within this application all changes that come within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to which they relate. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim may also be implemented by one unit or device through software or hardware. Words such as "first" and "second" are only used to distinguish the description and do not indicate any particular order, nor should they be understood as indicating or implying relative importance.

[0113] The above descriptions are merely specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art may easily propose variations or substitutions within the technical scope disclosed in the present application, and such variations or substitutions shall be encompassed within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims, and the above descriptions shall be regarded as exemplary and non-limiting.

Claims

1. A training data generation method, characterized in that: include: Acquire domain identification information of a target domain, and generate system-level prompt information related to the target domain based on the domain identification information; Inputting the system-level prompt information into a large language model that has completed alignment training to drive the model to generate an input instruction set related to the target domain; The input instruction set is used as an input prompt and is further input into the large language model to generate a response set semantically related to the input instruction set, thereby forming a first training data set including the input instruction set and the response set; For each command-response pair in the first training dataset, construct multi-turn dialogue training data with semantic coherence and rich context through the multi-turn dialogue extension generation method of the large language model; Aggregate the plurality of multi-round dialogue training data to construct a second training dataset for instruction tuning training or supervised alignment tasks of a large language model.

2. The training data generation method according to claim 1, characterized in that The step of obtaining domain identification information of the target domain and generating system-level prompt information related to the target domain based on the domain identification information includes: Based on the domain classification library or knowledge graph defined for different target fields, domain identification information corresponding to different application scenarios is extracted from it. The domain identification information includes: industry name, business type, task category and knowledge subcategory label; Based on at least one content in the domain identification information, a prompt structure entry that matches it is retrieved from a preset prompt template library, and the prompt structure entry is combined with the domain identification information through template filling and parameter binding to dynamically generate system-level prompt information optimized for different target domains.

3. The training data generation method according to claim 2, characterized in that The step of combining the prompt structure entries with the domain identification information through template filling and parameter binding to dynamically generate system-level prompt information optimized for different target domains includes: Identifying a plurality of placeholder fields set in the prompt structure entry, the placeholder fields comprising a model role field, a task goal field, an answer format field, and a style constraint field; Perform semantic mapping and parameter filling on each semantic element in the domain identification information and one or more items in the placeholder field, and perform semantic enhancement and modification on the answer format and style constraint fields according to different task scenarios; Through content insertion and context constraint rules, the filled prompt structure entries are converted into system-level prompt information with complete structure and clear semantics.

4. The training data generation method according to claim 1, wherein: The step of inputting the system-level prompt information into the large language model that has completed alignment training, and driving the model to generate an input instruction set related to the target domain, includes: Inputting the system-level prompt information as guidance content into the input interface of the large language model, so that the large language model parses and obtains a comprehensive semantic expression oriented to the target domain; Based on the comprehensive semantic expression and configured generation parameters, guiding the model to perform multiple rounds of batch generation to automatically construct an input instruction set of an order of magnitude that matches the set task requirements, wherein the generation parameters include: sampling temperature, maximum output length, repetition penalty coefficient, and diversity sampling strategy; The input instruction set can cover multiple sub-class tasks and diverse application scenarios in the target domain, and has diversity in task semantic structure while maintaining consistency in expression style.

5. The training data generation method according to claim 1, wherein: The step of inputting the input instruction set as an input prompt into the large language model and generating a response set semantically related to the input instruction set, thereby forming a first training data set including the input instruction set and the response set, includes: Inputting each input instruction in the input instruction set as input content into the large language model one by one, causing the large language model to parse semantic information in the input instruction and generate response content with structural specifications and semantic consistency based on the answer format and language style requirements set in the system-level prompt information; Bind each input command to its generated response content, establish a unique key-value mapping relationship, and store them uniformly as structured sample units to form a command-response pair; The plurality of structured sample units are serialized and arranged according to a preset format, and the first training data set is constructed by aggregation.

6. The training data generation method according to claim 1, characterized in that Before constructing multi-turn dialogue training data with semantic coherence and rich context using the multi-turn dialogue expansion generation method of the large language model for each command-response pair in the first training dataset, the method further includes screening and filtering the command-response pairs in the first training dataset based on a set of quality evaluation indicators, wherein the quality evaluation indicators include: Set minimum and maximum length thresholds for input commands and responses, and trim or remove commands and responses that exceed the thresholds; The input command and response content are semantically embedded based on the pre-trained language model, and the semantic similarity between the two is calculated to evaluate whether the response content is aligned with the semantic intent of the input command as a whole; Based on the semantic points extracted from the input command, the coverage of the response content in the context information is evaluated to determine the semantic completeness of the response content; Conduct discretized statistical analysis on the task objectives or semantic expressions in the input instruction set to measure their coverage breadth in terms of task categories and expression structures; By calculating the minimum semantic distance or word-level difference between adjacent samples, the similarity between adjacent instruction-response pairs is evaluated and highly repetitive or redundant samples are eliminated; The trained reward model is used to evaluate the quality of the response content, and the instruction-response pairs with scores higher than the set scoring threshold are retained as valid samples.

7. The training data generation method according to claim 1, characterized in that The step of constructing multi-turn dialogue training data with semantic coherence and rich context by using the multi-turn dialogue extension generation method of the large language model for each command-response pair in the first training dataset includes: Input each command-response pair in the first training dataset into the large language model as the context of a dialogue initiation turn; While retaining the previous round of conversation as historical context, the model is guided to generate a new round of instructions and corresponding responses that are semantically related to the previous round. The new round of instructions semantically include question-asking, task deepening, content reflection, or knowledge generalization. Repeat the process of guiding the model to generate a new round of commands and corresponding response content related to the previous round of semantics, and continue to build multiple rounds of command-response pairs until the preset number of dialogue rounds is reached or the dialogue generation interruption condition is met; The multiple rounds of command-response pairs constitute a multi-round dialogue training data as a whole, which has contextual semantic coherence and logical progressive structure, and is used to enhance the understanding and response capabilities of the large language model in complex dialogue tasks.

8. An electronic device, characterized in that: The electronic device comprises: One or more processors; and a memory storing computer program instructions, which, when executed, cause the processor to perform the training data generation method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that: When the computer program and / or instructions are executed by a processor, the training data generation method according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or instruction is executed by a processor, the training data generating method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Speech processing method and server

    CN108563633A

  • Table information extraction method, system and equipment based on compressed pre-training language model

    CN116701619A

  • Medical big language model training method and device, electronic equipment and storage medium

    CN118098562A

  • Machine translation quality improving method and device based on large language model

    CN118839702A

  • Intelligent question answering system construction method and system based on LLM large language model

    CN119691140A

Cited By

  • Sample generation method, model training method, data processing method and electronic equipment

    CN121681783A