Training data generation method and device, electronic equipment and storage medium

By generating initial questions associated with seed information using a large model and supplementing them with user feature information, the problems of low efficiency and high cost in training data generation are solved, achieving efficient and personalized training data generation and improving the model's adaptability to user features.

CN122114223APending Publication Date: 2026-05-29BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and high cost in generating training data, making it difficult to generate data in real time based on individual user characteristics. This results in low model personalization and an inability to adapt to different user characteristics.

Method used

The initial question is generated by a large model and associated with multiple seed information. User feature information is supplemented to generate the target question. The answer is then generated by the large model to form a question-answer pair to determine the training data.

Benefits of technology

It enables automatic generation of training data, improves generation efficiency, reduces costs, and enhances the adaptability of large models to user features and the effect of personalized generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114223A_ABST
    Figure CN122114223A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a training data generation method and device, electronic equipment and storage medium. The method comprises: generating initial questions associated with a plurality of seed information respectively by a large model, wherein the seed information is pre-configured example data; supplementing user feature information in the initial questions to obtain target questions; generating answers corresponding to the target questions by the large model; and determining training data from question and answer pairs composed of a plurality of target questions and the answers. The embodiments of the present application automatically generate training data, do not require manual annotation, can improve the generation efficiency of training data, reduce the cost of model training, and can supplement user feature information in the initial questions, and then generate answers based on the target questions after supplementing the user feature information, thereby generating corresponding training data based on different user features, improving the adaptability of the large model to user features, and improving the personalized generation effect of the large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a training data generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] In existing technologies, the training data used to train large models usually relies on manual annotation and data collection. This approach not only makes the generation of training data inefficient but also increases the cost of model training. Furthermore, existing training data is difficult to generate in real time based on individual user characteristics, resulting in low model personalization and an inability to adapt to different user characteristics. Summary of the Invention

[0003] This application provides a training data generation method, apparatus, electronic device, and storage medium, which helps to improve the efficiency of training data generation, reduce model training costs, and enhance the adaptability of large models to user features.

[0004] To address the aforementioned problems, in a first aspect, embodiments of this application provide a training data generation method, including:

[0005] Initial questions are generated by a large model and associated with multiple seed information, which are pre-configured sample data;

[0006] By supplementing the initial question with user characteristic information, the target question is obtained;

[0007] The large model generates an answer corresponding to the target question.

[0008] Training data is determined from question-answer pairs consisting of multiple target questions and their answers.

[0009] Secondly, embodiments of this application provide a training data generation apparatus, comprising:

[0010] An initial question generation module is used to generate initial questions associated with multiple seed information items from a large model. The seed information items are pre-configured sample data.

[0011] The target question generation module is used to supplement user feature information into the initial question to obtain the target question;

[0012] The answer generation module is used to generate an answer corresponding to the target question using the large model.

[0013] The training data determination module is used to determine training data from question-answer pairs consisting of multiple target questions and answers.

[0014] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the training data generation method described in embodiments of this application.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the training data generation method disclosed in embodiments of this application.

[0016] The training data generation method, apparatus, electronic device, and storage medium provided in this application generate initial questions associated with multiple seed information through a large model, supplement user feature information into the initial questions to obtain target questions, generate answers corresponding to the target questions through the large model, and determine training data from question-answer pairs composed of multiple target questions and answers. This achieves automatic generation of training data without manual annotation, which can improve the efficiency of training data generation and reduce model training costs. Moreover, by supplementing user feature information into the initial questions and then generating answers based on the target questions with supplemented user feature information, it can generate corresponding training data based on different user features, which can improve the adaptability of the large model to user features and enhance the personalized generation effect of the large model. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a training data generation method provided in this application and this embodiment;

[0019] Figure 2 This is a schematic diagram of the training data generation process in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of the structure of a training data generation device provided in an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] Figure 1 This is a flowchart of a training data generation method provided in this embodiment of the application, such as... Figure 1 As shown, the method includes steps 110 to 140.

[0024] Step 110: Generate initial questions associated with multiple seed information items through a large model, wherein the seed information items are pre-configured sample data.

[0025] In one exemplary embodiment, seed information is sample text data used to generate the initial question. Seed information can be a question, a pre-set topic, or user information. This user information is not associated with a specific user but merely represents information that a user might possess. For example, sample data could be a sample question, a sample topic, or a sample user profile, etc.

[0026] In one exemplary embodiment, multiple seed information can be input into a large model, which then generates initial questions associated with the seed information based on prompts. One or more initial questions can be generated based on each seed information. The large model, also known as a large language model (LLM), is a pre-trained, large-scale language model capable of generating corresponding natural language responses based on input.

[0027] In one exemplary embodiment, the plurality of seed information may include at least one of: seed questions, seed topics, and seed user profiles. The seed question is a pre-configured example question; the seed topic is a pre-configured example topic, such as food or travel; and the seed user profile is a pre-configured example user profile. The user profile contains basic user information, including age and interests, used for personalized content generation.

[0028] Step 120: Supplement the initial question with user feature information to obtain the target question.

[0029] In one exemplary embodiment, the user characteristic information is not associated with a specific user, but merely describes certain... user Possible characteristic information. For example, the user characteristic information may include at least one of: user memory point information, historical time information, and in-vehicle user information, wherein the user memory point information represents historical data information and is memory data information that the user may have.

[0030] In one exemplary embodiment, the initial problem can be analyzed, and user feature information can be added to the initial problem. The initial problem with the added user feature information can then be used as the target problem. For example, a large model can be used to add user feature information to the initial problem.

[0031] Step 130: Generate the answer corresponding to the target question using the large model.

[0032] In one exemplary embodiment, the target question is input into a large model, which then generates an answer corresponding to the target question based on the initial question and user feature information within the target question. For example, after inputting the target question into the model, the large model can integrate complete contextual information (including the initial question, user memory points, historical time information, and in-vehicle user information, etc.) to generate a personalized and context-adaptive answer.

[0033] Step 140: Determine training data from question-answer pairs consisting of multiple target questions and answers.

[0034] In one exemplary embodiment, each target question and its corresponding answer constitute a question-answer pair. Multiple question-answer pairs are deduplicated, and those meeting quality requirements are selected as the final training data.

[0035] The training data generation method provided in this application generates initial questions associated with multiple seed information through a large model, supplements user feature information into the initial questions to obtain target questions, generates answers corresponding to the target questions through the large model, and determines training data from question-answer pairs composed of multiple target questions and answers. This achieves automatic generation of training data without manual annotation, which can improve the efficiency of training data generation and reduce model training costs. Moreover, by supplementing user feature information into the initial questions and generating answers based on the target questions with supplemented user feature information, the method can generate corresponding training data based on different user features, which can improve the adaptability of the large model to user features and enhance the personalized generation effect of the large model.

[0036] Based on the above technical solution, the multiple seed information may include at least one of: seed questions, seed topics, and seed user profiles;

[0037] The initial questions generated by the large model and associated with multiple seed information can include at least one of the following:

[0038] Multiple seed problems are input into the large model, and the large model expands the seed problems according to the problem expansion prompts to generate the initial problem;

[0039] Multiple seed topics are input into the large model, which expands the seed topics based on topic expansion prompts to generate expanded topics, and generates initial questions associated with the expanded topics through the large model.

[0040] Multiple seed user profiles are input into the large model. The large model expands the seed user profiles based on profile expansion prompts to generate expanded user profiles. The large model then generates initial questions associated with the expanded user profiles.

[0041] In one exemplary embodiment, the seed questions are a small subset of manually designed seed queries that can cover typical problems or tasks from different domains. All seed questions are input into a large model, which expands them based on question expansion prompts to obtain initial questions. For example, the question expansion prompts can include expansion requirements and expansion directions. Expansion requirements may include ensuring that the questions do not repeat each other, expansion based on semantic similarity and logical relevance, etc. Expansion directions may include changing sentence structure, introducing synonyms, adjusting language style, etc., or they may also include focusing on specific age groups or professions. For example, by inputting multiple seed questions into the large model, the model can associate and rewrite the seed questions based on semantic similarity, logical relevance, etc., generating diverse questions as initial questions by changing sentence structure, introducing synonyms, adjusting language style, etc., thereby expanding the coverage of the seed questions. The generated initial questions not only possess content diversity but also reflect different language styles and expressions, further improving the diversity of the training data.

[0042] In one exemplary embodiment, seed topics are a small number of representative topics, such as those related to age, family, and food, which can represent the user's interests or characteristics. These seed topics are input into a large model, which receives them and expands upon them, generating sub-topics related to the seed topics through association. For example, if the seed topic is "food," the large model can associate it with sub-topics such as "baking," "international cuisine," and "healthy eating." These generated sub-topics are used to construct more detailed and personalized user question scenarios. Furthermore, the large model can simulate reasonable user inquiries based on the sub-topics, serving as initial questions. These initial questions reflect the questions users might ask under different topics, enhancing the diversity and semantic coverage of data generation.

[0043] In one exemplary embodiment, the seed user profile is a small number of manually generated seed user profiles, including basic information (such as age, occupation, interests, etc.), serving as the initial setting for user characteristics. These seed user profiles are input into a large model, which receives and expands them, generating more diverse extended user profiles through association. These generated extended user profiles reflect different combinations of user characteristics, such as a young housewife or a retired food enthusiast. Based on the expanded user profiles, the large model simulates user queries that match the characteristics of different profiles, obtaining initial questions associated with the extended user profiles. The initial questions generated by the large model can reflect personalized needs under specific user characteristics; for example, a young tech enthusiast might ask questions about the latest electronic products. Through expansion and simulation, the large model not only generates richer user profiles but also enhances the personalization of the training data, making subsequent data synthesis more aligned with user needs.

[0044] By progressively increasing the diversity and quantity of queries, topics, and user profiles, a large number of diverse initial questions are generated, broadly covering different topics, language styles, and user characteristics, thereby significantly improving the richness and adaptability of the model training data. This method effectively reduces the steps of manual intervention and rapidly generates highly diverse and personalized initial questions through the associative and rewriting capabilities of the large model, providing high-quality data support for model training.

[0045] Based on the above technical solution, the user feature information may include at least one of user memory point information, historical time information, and in-vehicle user information, wherein the user memory point information represents historical data information;

[0046] The step of supplementing the initial question with user characteristic information to obtain the target question may include at least one of the following:

[0047] The target problem is obtained by supplementing the initial problem with user memory points using the large model.

[0048] The large model determines historical time information based on the initial question and user memory point information, and supplements the initial question with the historical time information to obtain the target question;

[0049] The target problem is obtained by supplementing the in-vehicle user information into the initial problem using the large model.

[0050] In one exemplary embodiment, before responding to an answer, to ensure that the large model can generate accurate answers that conform to the memory question, it is necessary to generate and supplement user characteristic information for each initial question. User characteristic information may include supporting evidence for the answer, reasonable time information, and user context information, etc.

[0051] In one exemplary embodiment, the initial question can be generated based on a seed question, seed topic, or seed user profile. An initial question generated based on a seed question or seed topic does not contain user profile information. In this case, the large model can analyze the content of the initial question and, combined with current relevant data, supplement the initial question with relevant data as user memory points, and use the initial question with supplemented user memory points as the target question. An initial question generated based on a seed user profile contains user profile information. In this case, the large model can combine the user profile information and the initial question to supplement user memory points, and use the initial question with supplemented user memory points as the target question.

[0052] In an exemplary embodiment, historical time information refers to time information corresponding to user memory point information. A large model can be used to analyze the initial question and user memory point information to determine the historical time information corresponding to the user memory point information. This historical time information is then added to the initial question to obtain the target question. For example, if no time information appears in the initial question, a random time is used as the historical time information; if the initial question includes time information, time information preceding that time information is constructed as historical data information. For example, the initial question and user memory points can be input into the large model. The large model generates reasonable historical time information based on the background and content of the initial question. This historical time information is an important contextual part of the data generation process, especially for questions related to memory (user memory points). The large model will infer the time period based on the query content, such as "yesterday", "last month" or "last year", or provide more specific time information at specific time points. For example, if the query asks for the time when an event occurred, the large model will generate a time label that matches the time of the event. For time-sensitive queries, such as "What was the name of the restaurant I went to last time?", the large model will generate a general time range as a reference.

[0053] In one exemplary implementation, the application scenario for training a large model using the training data generated by the training data generation method provided in this application is in-vehicle dialogue. In addition, in-vehicle user information can be supplemented in the initial question. In-vehicle user information may include, for example, the relationship between multiple users in the vehicle, the seat position of the users in the vehicle, etc.

[0054] By generating user memory points, historical time information, and in-vehicle user information, the large-scale model ensures that it has sufficient contextual basis and supporting evidence when generating answers. This step not only increases the accuracy of the large-scale model's answers but also ensures personalization and contextual adaptability. The large-scale model can generate answers with specific time contexts and user environments based on different user characteristics and needs, thereby improving user satisfaction and trust in the answers.

[0055] Based on the above technical solution, the step of supplementing the initial problem with user memory point information through the large model to obtain the target problem may include:

[0056] If the initial question includes a user profile, then the large model is used to supplement the initial question with user memory point information associated with the user profile, thereby obtaining the target question; or

[0057] If the initial question does not include a user profile, the target question is obtained by supplementing the initial question with user memory point information associated with the initial question through the large model.

[0058] In one exemplary embodiment, if the initial question is generated based on a seed user profile, this initial question includes the user profile (such as age, interests, historical behavior, etc.). In this case, the large model can generate relevant user memory point information based on the user profile. The large model combines the user profile with the initial question, simulating the user's past memories, experiences, or habits to generate specific and personalized user memory point information. For example, if the user profile includes information such as "likes outdoor activities," and the initial question asks for recommendations for holiday activities, the large model will generate user memory point information related to outdoor activities as the basis for the answer.

[0059] In another exemplary embodiment, if the initial question is generated based on a seed question or seed topic, then the initial question does not include a user profile. In this case, the large model can directly generate user memory point information based on the content of the initial question. The large model infers the user's possible memories or experiences based on the semantics and logic of the initial question and generates answer evidence that fits the context of the question. For example, for an initial question like "What movies are worth watching recently?", the large model would generate information on currently popular movies as user memory point information.

[0060] By supplementing the initial question with information that includes the user's memory points as the basis for the answer, support can be provided for subsequent answer generation.

[0061] Based on the above technical solution, the step of supplementing the initial problem with in-vehicle user information through the large model to obtain the target problem includes:

[0062] If the initial question includes a single user profile, then the large model is used to supplement the initial question with the single user's identity information and seat location, as the in-vehicle user information, to obtain the target question; or

[0063] If the initial problem involves multiple user scenarios, then the large model is used to supplement the initial problem with the relationships between multiple users, which serve as the in-vehicle user information, thus obtaining the target problem.

[0064] When a large model trained on generated training data is used for in-vehicle dialogue, the model can generate in-vehicle user information as contextual background to simulate user needs in a real-world environment. This in-vehicle user information can include the identities of the users, the number of users, and their interests and preferences.

[0065] In one exemplary embodiment, if the initial question includes a single user profile, the generated in-vehicle user information will include the single user's identity background, seat position, etc. When supplementing the seat position, if the initial question involves a seat position, that seat position is used as the supplementary seat position; if the initial question does not involve a seat position, a seat position is randomly determined as the supplementary seat position.

[0066] In another exemplary embodiment, if the initial question involves multiple user scenarios, the large model simulates relationships between the multiple users based on information about them in the initial question, such as family members or colleagues, as in-vehicle user information to enhance the understanding and accuracy of the initial question. For example, if the initial question mentions family information, the large model can simulate family member relationships as in-vehicle user information; or, if the initial question mentions friends or colleagues, the large model can simulate friend or colleague relationships as in-vehicle user information. For instance, if the initial question is "Where is a suitable place for a family trip?", the generated in-vehicle user information will include "family" scenarios, such as "parents + children," to provide a reference for the recommendation results.

[0067] By supplementing the in-vehicle user information with contextual information applicable to the current query, it supports the generation of personalized, context-sensitive answers for large models.

[0068] Based on the above technical solution, the step of determining training data from multiple question-answer pairs consisting of the target questions and the answers includes at least one of the following:

[0069] According to preset rules, multiple question-answer pairs are filtered, and question-answer pairs that conform to the preset rules are retained as training data;

[0070] The large model is used to evaluate the quality of multiple question-answer pairs, and the question-answer pairs that meet the quality requirements are selected as the training data.

[0071] The multiple question-answer pairs are deduplicated, and the deduplicated question-answer pairs are used as the training data.

[0072] After generating answers using a large model, the generated target questions and answers need to undergo a series of post-processing steps to ensure data quality, accuracy, and diversity. The post-processing workflow includes rule filtering, large model quality assessment, and similarity deduplication.

[0073] In one exemplary embodiment, the generated question-and-answer pairs can be filtered according to preset rules to remove those that do not conform to the preset rules and retain those that do. The retained question-and-answer pairs can then be used as training data. Alternatively, the retained question-and-answer pairs can be subjected to other post-processing. The purpose of using rule-based filtering is to eliminate data that does not conform to preset specifications and standards, ensuring that the generated data meets basic quality requirements.

[0074] In one exemplary embodiment, a large model can be used to perform quality assessment on multiple question-answer pairs to filter out those that meet quality requirements. These quality-compliant question-answer pairs can then be used as training data. Alternatively, further post-processing can be applied to the quality-compliant question-answer pairs. Utilizing a large model for automatic data quality assessment can filter out high-quality training data.

[0075] In one exemplary embodiment, question-answer pairs generated by a large model may contain duplicates. In this case, the generated question-answer pairs can be deduplicated to remove duplicate or overly similar question-answer pairs and ensure data diversity.

[0076] By performing post-processing such as rule filtering, large model quality assessment, and deduplication, the quality, relevance, and diversity of the generated training data can be effectively improved. The datasets processed through these steps have higher accuracy and diversity, providing a more reliable foundation for the training and application of large models, and improving the model's adaptability to different scenarios and user needs.

[0077] Based on the above technical solution, the step of filtering multiple question-answer pairs according to preset rules and retaining question-answer pairs that conform to the preset rules as training data may include: filtering multiple question-answer pairs according to preset content rules and retaining question-answer pairs that conform to the preset content rules; selecting question-answer pairs that conform to preset format rules from the question-answer pairs that conform to the preset content rules; and selecting question-answer pairs that have a logical relationship between the target question and the answer from the question-answer pairs that conform to the preset format rules as training data.

[0078] In an exemplary embodiment, the preset rules may include preset content rules and preset format rules. First, based on the preset content rules (such as sensitive word filtering, violation content detection, etc.), the generated question-and-answer pairs are checked, eliminating those that do not conform to the preset content rules and retaining those that do. For example, data entries involving sensitive topics, content unsuitable for dissemination, or containing incorrect information are excluded. Next, the question-and-answer pairs conforming to the preset content rules are checked against the preset format rules to ensure that the format of the question-and-answer pairs meets requirements, such as word count limits and sentence structure, to ensure uniform format and avoid inconsistencies affecting data processing and use. A logical consistency check is also performed on the question-and-answer pairs conforming to both the content rules and the preset format rules to ensure a logical relationship between the answer and the target question, preventing irrelevant or inconsistent answers. A large model can be used to filter question-and-answer pairs that have a logical relationship between the target question and the answer from those conforming to the preset format rules; that is, the large model determines whether the answer correctly answers the target question.

[0079] By filtering the generated question-and-answer pairs according to rules, we can ensure that the data conforms to basic content specifications and logical consistency, laying the foundation for subsequent processing steps.

[0080] Based on the above technical solution, the step of using the large model to perform quality assessment on multiple question-answer pairs and selecting question-answer pairs that meet the quality requirements as training data includes: using the large model to assess the relevance of the target question and answer in each question-answer pair, obtaining a relevance score, and retaining question-answer pairs whose relevance score is greater than or equal to a first score threshold; using the large model to assess the language fluency of the question-answer pairs whose relevance score is greater than or equal to the first score threshold, and selecting question-answer pairs without grammatical errors from the question-answer pairs whose relevance score is greater than or equal to the first score threshold as training data; and / or, for multi-turn question-answer pairs in multiple question-answer pairs, using the large model to assess the coherence of the multi-turn question-answer pairs, obtaining a coherence score, and retaining multi-turn question-answer pairs whose coherence score is greater than a second score threshold as training data.

[0081] In one exemplary embodiment, a large model is used to evaluate the content relevance between the target question and the answer in a question-answer pair, resulting in a relevance score for the question-answer pair. The scoring criteria may include the accuracy, reasonableness, and semantic consistency of the answer. Question-answer pairs with a relevance score lower than a first scoring threshold are filtered out, while those with a relevance score greater than or equal to the first scoring threshold are retained.

[0082] In one exemplary embodiment, question-answer pairs with a relevance score greater than or equal to a first scoring threshold also need to undergo a language fluency assessment. By analyzing the language fluency of the target document and the answer through a large model, it is determined whether there are any grammatical errors, logical inconsistencies, or inappropriate word choices, ensuring that the retained question-answer pairs can meet the user's understanding and reading experience through fluent expression.

[0083] In one exemplary embodiment, multiple question-answer pairs can be concatenated as multi-turn question-answering. For the multi-turn question-answering in the generated question-answer pairs, a large model is used to evaluate the coherence of the multi-turn question-answering, in order to assess whether the generated answers conform to the contextual logic and whether the answers are consistent with the preceding text. The coherence of the generated multi-turn question-answering is scored using the multi-turn question-answering capabilities of the large model, resulting in a coherence score for the multi-turn question-answering. Multi-turn question-answering with a coherence score lower than a second scoring threshold is eliminated, while multi-turn question-answering with a coherence score greater than or equal to the second scoring threshold is retained.

[0084] By using a large model to evaluate the quality of the generated question-and-answer pairs, high-quality answer pairs can be retained while low-quality or unacceptable question-and-answer pairs can be removed.

[0085] Based on the above technical solution, the step of deduplicating multiple question-answer pairs and using the deduplicated question-answer pairs as the training data includes:

[0086] Determine the similarity between any two question-answer pairs. If the similarity is greater than a similarity threshold, retain one of the two question-answer pairs.

[0087] The retained question-answer pairs are filtered for semantic similarity using the large model, and the filtered question-answer pairs are used as the training data.

[0088] And / or, for multiple question-answer pairs, the large model filters two multiple question-answer pairs based on the similarity between each pair, and uses the filtered multiple question-answer pairs as the training data.

[0089] In one exemplary embodiment, question-answer pairs can be converted into vectors, and based on the vectors corresponding to the question-answer pairs, a similarity algorithm (such as cosine similarity or Jaccard similarity) can be used to calculate the similarity of the generated question-answer pairs. For question-answer pairs with a similarity greater than a similarity threshold, one of them is retained; if the similarity of two question-answer pairs is less than the similarity threshold, both are retained. For example, higher-quality or preferred question-answer pairs can be retained, while other similar pairs can be deleted. Higher-quality question-answer pairs can be those with the highest relevance scores. For example, when calculating the similarity between question-answer pairs, the similarity between the two target questions in the question-answer pair can be calculated as the similarity between the question-answer pairs, because if two target questions are dissimilar, the answers must also be dissimilar. This reduces the amount of data processed and improves processing efficiency.

[0090] In an exemplary embodiment, in addition to detecting textual similarity, a large model can also be used to perform semantic similarity detection on the question-answer pairs retained from the similarity detection, identifying question-answer pairs with different expressions but similar content. For example, the target questions of two question-answer pairs can be input into the large model, which performs semantic similarity detection on these two target questions to determine whether they are semantically similar, and retains one of the question-answer pairs if they are semantically similar. Semantic similarity filtering can effectively improve the accuracy of deduplication.

[0091] In one exemplary embodiment, for multi-turn question answering, in addition to checking the similarity of each question-answer pair, it is also necessary to detect the similarity of the entire dialogue stream through a large model to ensure that different dialogue streams have sufficient content differences in order to provide diverse data in subsequent model training. For example, when performing similarity filtering on multi-turn question answering using a large model, the target questions from two multi-turn question answering sessions can be input into the large model. The large model determines whether the target questions in these two sessions are similar, and if the target questions in these two sessions are similar, one of the multi-turn question answering sessions is retained.

[0092] By deduplicating the generated question-answer pairs, we can ensure the diversity and uniqueness of the final training data, providing richer materials for training large models.

[0093] Based on the above technical solution, the step of determining training data from multiple question-answer pairs consisting of multiple target questions and answers may further include: determining question-answer pairs that meet a preset category ratio from multiple question-answer pairs according to the category to which each question-answer pair belongs, and using these as the training data.

[0094] In one exemplary embodiment, post-processing of the generated question-answer pairs may further include data category balancing to maintain a balance between data of different categories, thereby preventing the model training data from being biased towards certain specific categories, which could affect the fairness and generalization ability of the model.

[0095] In one exemplary embodiment, the data volume distribution of different categories can be statistically analyzed for all generated question-answer pairs, including topic categories (such as travel, education, entertainment, etc.), user profile categories (such as different age groups, interest types, etc.), and seed topic types (such as question-based, recommendation-based, reminiscing-based, etc.). Based on a preset category ratio, the data of different categories can be balanced. For categories with less data, data augmentation (such as generating more question-answer pairs of this category) can be used to fill the gaps. For categories with excessive data, downsampling can be used to reduce redundant data, preventing large models from over-relying on certain categories. Depending on specific needs, the proportion of data in each category can be pre-adjusted. For example, on datasets with broad user interests, different proportion strategies can be set to ensure that the model can cover different domains and user characteristics.

[0096] By balancing the generated question-answer pairs according to a preset category ratio, data from different categories can be reasonably distributed, thus providing balanced input for large model training and improving the model's generalization ability.

[0097] Figure 2 This is a schematic diagram illustrating the training data generation process in an embodiment of this application. For example... Figure 2 As shown, the generation of initial questions includes seed question generation, seed topic generation, and seed user profile generation. For seed question-based initial topic generation, the seed question is input into a large model, which expands the seed question to generate the initial question. For seed topic-based initial question generation, the seed topic is input into a large model, which expands the seed topic to obtain an extended topic, and an initial question associated with the extended topic is generated. For seed user profile-based initial question generation, the seed user profile is input into a large model, which expands the seed user profile to generate an extended user profile, and an initial question associated with the extended user profile is generated. After generating the initial questions, user feature information, including user memory point information, historical time information, and in-vehicle user information, is added to the generated initial questions to obtain the target question. For target questions with contextual information, the answer to the target question is generated using the large model. Post-processing is performed on the question-answer pair consisting of the target question and answer, including rule filtering, large model quality assessment, similarity deduplication, and data category balancing. The specific process of each processing step can be referred to the above embodiment, and will not be repeated here.

[0098] This application's embodiments can automatically synthesize user memory data, generating training data that fits user characteristics without manual intervention, thus improving data generation efficiency. By synthesizing time and user features, it achieves dynamic data adaptability, enhancing the personalized generation effect of the model. Through rule filtering, large model quality assessment, similarity screening, and category balancing, it ensures the quality of the generated data. By automatically generating user feature data, it reduces the cost of manual data collection. It can improve the personalized adaptability of the model; by combining user memory information, the content generated by the model is closer to the user's actual needs. It accelerates data updates and responses; through automated construction, it can quickly generate new training data to meet the changing needs of different users and environments.

[0099] Figure 3 This is a schematic diagram of the structure of a training data generation device provided in an embodiment of this application, as shown below. Figure 3 As shown, the device includes:

[0100] The initial question generation module 310 is used to generate initial questions associated with multiple seed information items through a large model, wherein the seed information items are pre-configured sample data;

[0101] The target question generation module 320 is used to supplement user feature information into the initial question to obtain the target question;

[0102] Answer generation module 330 is used to generate an answer corresponding to the target question through the large model;

[0103] The training data determination module 340 is used to determine training data from question-answer pairs consisting of multiple target questions and answers.

[0104] Optionally, the plurality of seed information includes at least one of: seed questions, seed topics, and seed user profiles;

[0105] The initial question generation module is used to perform at least one of the following:

[0106] Multiple seed problems are input into the large model, and the large model expands the seed problems according to the problem expansion prompts to generate the initial problem;

[0107] Multiple seed topics are input into the large model, which expands the seed topics based on topic expansion prompts to generate expanded topics, and generates initial questions associated with the expanded topics through the large model.

[0108] Multiple seed user profiles are input into the large model. The large model expands the seed user profiles based on profile expansion prompts to generate expanded user profiles. The large model then generates initial questions associated with the expanded user profiles.

[0109] Optionally, the user feature information includes at least one of user memory point information, historical time information, and in-vehicle user information, wherein the user memory point information represents historical data information;

[0110] The target question generation module includes at least one of the following:

[0111] A memory point supplementation unit is used to supplement user memory point information in the initial problem through the large model to obtain the target problem;

[0112] The time supplementation unit is used to determine historical time information based on the initial question and user memory point information through the large model, and supplement the historical time information into the initial question to obtain the target question;

[0113] The in-vehicle user information supplementation unit is used to supplement the in-vehicle user information in the initial problem through the large model to obtain the target problem.

[0114] Optionally, the memory point supplementation unit is specifically used for:

[0115] If the initial question includes a user profile, then the large model is used to supplement the initial question with user memory point information associated with the user profile, thereby obtaining the target question; or

[0116] If the initial question does not include a user profile, the target question is obtained by supplementing the initial question with user memory point information associated with the initial question through the large model.

[0117] Optionally, the in-vehicle user information supplementation unit is specifically used for:

[0118] If the initial question includes a single user profile, then the large model is used to supplement the initial question with the single user's identity information and seat location, as the in-vehicle user information, to obtain the target question; or

[0119] If the initial problem involves multiple user scenarios, then the large model is used to supplement the initial problem with the relationships between multiple users, which serve as the in-vehicle user information, thus obtaining the target problem.

[0120] Optionally, the training data determination module includes at least one of the following:

[0121] The rule filtering unit is used to filter multiple question-answer pairs according to preset rules, and retain the question-answer pairs that conform to the preset rules as the training data;

[0122] The large model quality assessment unit is used to assess the quality of multiple question-answer pairs through the large model and select question-answer pairs that meet the quality requirements as the training data.

[0123] The deduplication unit is used to deduplicatize multiple question-answer pairs and use the deduplicated question-answer pairs as the training data.

[0124] Optionally, the rule filtering unit is specifically used for:

[0125] According to preset content rules, multiple question-answer pairs are filtered, and question-answer pairs that conform to the preset content rules are retained;

[0126] Filter out question-and-answer pairs that conform to the preset format rules from the question-and-answer pairs that conform to the preset content rules;

[0127] The training data is obtained by selecting question-answer pairs that have a logical relationship between the target question and the answer from the question-answer pairs that conform to the preset format rules.

[0128] Optionally, the large model quality assessment unit is specifically used for:

[0129] The large model is used to evaluate the relevance of the target question and answer in each question-answer pair to obtain a relevance score, and question-answer pairs with a relevance score greater than or equal to a first score threshold are retained.

[0130] The large model is used to evaluate the fluency of question-answer pairs whose relevance scores are greater than or equal to the first score threshold. Question-answer pairs without grammatical errors are selected from the question-answer pairs whose relevance scores are greater than or equal to the first score threshold and used as the training data.

[0131] And / or, for multiple rounds of question-and-answer pairs, the coherence of the multiple rounds of question-and-answer is evaluated by the large model to obtain a coherence score, and the multiple rounds of question-and-answer with a coherence score greater than or equal to a second score threshold are retained as the training data.

[0132] Optionally, the deduplication unit is specifically used for:

[0133] Determine the similarity between any two question-answer pairs. If the similarity is greater than a similarity threshold, retain one of the two question-answer pairs.

[0134] The retained question-answer pairs are filtered for semantic similarity using the large model, and the filtered question-answer pairs are used as the training data.

[0135] And / or, for multiple question-answer pairs, the large model filters two multiple question-answer pairs based on the similarity between each pair, and uses the filtered multiple question-answer pairs as the training data.

[0136] Optionally, the training data determination module includes:

[0137] A category balancing unit is used to determine, based on the category to which each question-answer pair belongs, question-answer pairs that satisfy a preset category ratio from multiple question-answer pairs, and use them as training data.

[0138] The training data generation apparatus provided in this application embodiment is used to implement the steps of the training data generation method described in this application embodiment. The specific implementation of each module of the apparatus is described in the corresponding steps, and will not be repeated here.

[0139] The training data generation device provided in this application generates initial questions associated with multiple seed information through a large model, supplements user feature information into the initial questions to obtain target questions, generates answers corresponding to the target questions through the large model, and determines training data from question-answer pairs composed of multiple target questions and answers. This achieves automatic generation of training data without manual annotation, which can improve the efficiency of training data generation and reduce model training costs. Moreover, by supplementing user feature information into the initial questions and then generating answers based on the target questions with supplemented user feature information, the device can generate corresponding training data based on different user features, which can improve the adaptability of the large model to user features and enhance the personalized generation effect of the large model.

[0140] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 4 As shown, the electronic device 400 may include one or more processors 410 and one or more memories 420 connected to the processors 410. The electronic device 400 may also include an input interface 430 and an output interface 440 for communicating with another device or system. Program code executed by the processor 410 may be stored in the memory 420.

[0141] The processor 410 in the electronic device 400 calls the program code stored in the memory 420 to execute the training data generation method in the above embodiment.

[0142] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the training data generation method as described in this application.

[0143] This application also provides a computer program product that, when executed by a processor, implements the steps of the training data generation method described in this application.

[0144] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0145] The above provides a detailed description of a training data generation method, apparatus, electronic device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

Claims

1. A method for generating training data, characterized in that, include: Initial questions are generated by a large model and associated with multiple seed information, which are pre-configured sample data; By supplementing the initial question with user characteristic information, the target question is obtained; The large model generates an answer corresponding to the target question. Training data is determined from question-answer pairs consisting of multiple target questions and their answers.

2. The method according to claim 1, characterized in that, The multiple seed information includes at least one of: seed questions, seed topics, and seed user profiles; The initial problem generated by the large model and associated with multiple seed information respectively includes at least one of the following: Multiple seed problems are input into the large model, and the large model expands the seed problems according to the problem expansion prompts to generate the initial problem; Multiple seed topics are input into the large model, which expands the seed topics based on topic expansion prompts to generate expanded topics, and generates initial questions associated with the expanded topics through the large model. Multiple seed user profiles are input into the large model. The large model expands the seed user profiles based on profile expansion prompts to generate expanded user profiles. The large model then generates initial questions associated with the expanded user profiles.

3. The method according to claim 1, characterized in that, The user feature information includes at least one of user memory point information, historical time information, and in-vehicle user information, wherein the user memory point information represents historical data information. The step of supplementing the initial question with user characteristic information to obtain the target question includes at least one of the following: The target problem is obtained by supplementing the initial problem with user memory points using the large model. The large model determines historical time information based on the initial question and user memory point information, and supplements the initial question with the historical time information to obtain the target question; The target problem is obtained by supplementing the in-vehicle user information into the initial problem using the large model.

4. The method according to claim 3, characterized in that, The process of supplementing the initial problem with user memory point information using the large model to obtain the target problem includes: If the initial question includes a user profile, then the large model is used to supplement the initial question with user memory point information associated with the user profile, thereby obtaining the target question; or If the initial question does not include a user profile, the target question is obtained by supplementing the initial question with user memory point information associated with the initial question through the large model.

5. The method according to claim 3, characterized in that, The process of supplementing the initial problem with in-vehicle user information using the large model to obtain the target problem includes: If the initial question includes a single user profile, then the large model is used to supplement the initial question with the single user's identity information and seat location, as the in-vehicle user information, to obtain the target question; or If the initial problem involves multiple user scenarios, then the large model is used to supplement the initial problem with the relationships between multiple users, which serve as the in-vehicle user information, thus obtaining the target problem.

6. The method according to any one of claims 1-5, characterized in that, The process of determining training data from question-answer pairs consisting of multiple target questions and answers includes at least one of the following: According to preset rules, multiple question-answer pairs are filtered, and question-answer pairs that conform to the preset rules are retained as training data; The large model is used to evaluate the quality of multiple question-answer pairs, and the question-answer pairs that meet the quality requirements are selected as the training data. The multiple question-answer pairs are deduplicated, and the deduplicated question-answer pairs are used as the training data.

7. The method according to claim 6, characterized in that, The step of filtering multiple question-answer pairs according to preset rules and retaining question-answer pairs that conform to the preset rules as training data includes: According to preset content rules, multiple question-answer pairs are filtered, and question-answer pairs that conform to the preset content rules are retained; Filter out question-and-answer pairs that conform to the preset format rules from the question-and-answer pairs that conform to the preset content rules; The training data is obtained by selecting question-answer pairs that have a logical relationship between the target question and the answer from the question-answer pairs that conform to the preset format rules.

8. The method according to claim 6, characterized in that, The step of evaluating the quality of multiple question-answer pairs using the large model and selecting those that meet the quality requirements as training data includes: The large model is used to evaluate the relevance of the target question and answer in each question-answer pair to obtain a relevance score, and question-answer pairs with a relevance score greater than or equal to a first score threshold are retained. The large model is used to evaluate the fluency of question-answer pairs whose relevance scores are greater than or equal to the first score threshold. Question-answer pairs without grammatical errors are selected from the question-answer pairs whose relevance scores are greater than or equal to the first score threshold and used as the training data. And / or, for multiple rounds of question-and-answer pairs, the coherence of the multiple rounds of question-and-answer is evaluated by the large model to obtain a coherence score, and the multiple rounds of question-and-answer with a coherence score greater than or equal to a second score threshold are retained as the training data.

9. The method according to claim 6, characterized in that, The step of deduplicating multiple question-answer pairs and using the deduplicated question-answer pairs as the training data includes: Determine the similarity between any two question-answer pairs. If the similarity is greater than a similarity threshold, retain one of the two question-answer pairs. The retained question-answer pairs are filtered for semantic similarity using the large model, and the filtered question-answer pairs are used as the training data. And / or, for multiple question-answer pairs, the large model filters two multiple question-answer pairs based on the similarity between each pair, and uses the filtered multiple question-answer pairs as the training data.

10. The method according to any one of claims 1-5, characterized in that, The step of determining training data from question-answer pairs consisting of multiple target questions and answers includes: Based on the category to which each question-answer pair belongs, question-answer pairs that meet the preset category ratio are determined from multiple question-answer pairs and used as the training data.

11. A training data generation device, characterized in that, include: An initial question generation module is used to generate initial questions associated with multiple seed information items from a large model. The seed information items are pre-configured sample data. The target question generation module is used to supplement user feature information into the initial question to obtain the target question; The answer generation module is used to generate an answer corresponding to the target question using the large model. The training data determination module is used to determine training data from question-answer pairs consisting of multiple target questions and answers.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the training data generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the training data generation method according to any one of claims 1 to 10.