Multi-language supervision data set generation method and device based on data enhancement

By obtaining data from the multilingual knowledge base and using large language models to generate Q&A data, the problem of high labeling of multilingual supervision data sets is solved, and stronger language expression coverage and model robustness is achieved, and it is suitable for multilingual dialogue, translation and Q&A systems.

CN120336502APending Publication Date: 2025-07-18启元实验室
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510401313.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the labeling of multilingual supervision data sets is expensive and time-consuming, and the automated labeling method is less exploring multilingual capabilities, resulting in a lack of knowledge about specific topics in non-English areas, direct translation may lead to distortion, and scarcity of multilingual SFT resources.

Method used

By obtaining dialogue generation requirements and multilingual knowledge bases, splitting language data, generating question-and-answer data using large language models, and streamlining the model through adjustments to generate multilingual supervision data sets.

Benefits of technology

It improves the diversity and accuracy of the data set, enhances the robustness and generalization capabilities of the model in different language environments, and supports the efficient application of multilingual dialogue systems, translation systems and question-and-answer systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336502A_ABST
    Figure CN120336502A_ABST
Patent Text Reader

Abstract

The invention provides a multi-language supervision data set generation method and device based on data enhancement, and relates to the technical field of data processing. The multi-language supervision data set generation method based on data enhancement comprises the steps of obtaining a dialogue generation demand and a preset multi-language knowledge base, and determining expected language data from the multi-language knowledge base according to the dialogue generation demand; splitting the expected language data according to a dialogue generation demand to determine segmented language texts; generating multiple groups of question and answer data based on a preset large language model, the dialogue generation demand and the segmented language text; and according to a preset adjustment mode and the multiple groups of question and answer data, adjusting a preset simplified model to determine target question and answer data from the multiple groups of question and answer data, and determining the target question and answer data as a multi-language supervision data set. By generating multiple groups of question and answer data, fine adjustment is performed on the simplified model in combination with an adjustment mode, low-quality or irrelevant data is filtered, and the overall quality of the data set is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, for example, to a method and device for generating a multilingual supervised data set based on data augmentation. Background Art

[0002] In recent years, large language models (LLMs) have made remarkable progress in the field of natural language processing (NLP). These models are pre-trained on large-scale data sets and further optimized for specific tasks using supervised fine-tuning (SFT). Supervised fine-tuning has become an important part of building powerful LLMs, and its main purpose is to guide the learning process of the model by providing labeled data, thereby improving its accuracy and robustness on specific tasks. Currently, there are mainly three categories of SFT data annotation methods. The first category is to manually annotate large-scale data to generate training data. The second category is to use pre-trained models to automatically generate question-and-answer pairs or other task-related data. The third category is to translate existing English data into other languages to generate multilingual data sets. However, manual annotation in these several categories of SFT data annotation methods is costly and time-consuming, so it is currently rarely used for data sets. Automated annotation methods are widely used, but mainly in English SFT data sets, with relatively little exploration of multilingual capabilities, resulting in relatively limited multilingual SFT resources and causing the problem of data scarcity.

[0003] In related technologies, to enhance the global multilingual usability of LLMs, many multilingual SFT data sets have been created. Some tasks use ChatGPT (Chat Generative Pre-trained Transformer) to translate Alpaca (Alpaca Model) into various languages. There are also some tasks that combine ShareGPT (Shared Generative Pre-trained Transformer) with Alpaca and then translate these two data sets. However, the translation of English data may not fully cover topics specific to non-English regions, resulting in LLMs lacking language-specific knowledge. In addition, for certain instructions, the answers vary in different cultural backgrounds, so directly translating all English conversations may lead to a large number of distorted translations. Summary of the Invention

[0004] This application aims to provide a method and device for generating a multilingual supervised data set based on data augmentation.

[0005] According to one aspect of the present application, a method for generating a multilingual supervised data set based on data augmentation is proposed, including:

[0006] Obtain the dialogue generation requirements and a preset multilingual knowledge base, and determine the expected language data from the multilingual knowledge base according to the dialogue generation requirements;

[0007] Split the expected language data according to the dialogue generation requirements to determine segmented language texts;

[0008] Generate multiple sets of Q&A data based on a preset large language model, dialogue generation requirements, and segmented language texts;

[0009] Adjust a preset reduction model according to the preset adjustment method and multiple sets of Q&A data to determine target Q&A data from the multiple sets of Q&A data, and determine the target Q&A data as the multilingual supervised data set.

[0010] According to one aspect of the present application, a device for generating a multilingual supervised data set based on data augmentation is proposed, including:

[0011] A data acquisition module, configured to obtain the dialogue generation requirements and a preset multilingual knowledge base, and determine the expected language data from the multilingual knowledge base according to the dialogue generation requirements;

[0012] A segmentation module, configured to split the expected language data according to the dialogue generation requirements to determine segmented language texts;

[0013] A Q&A data generation module, configured to generate multiple sets of Q&A data based on a preset large language model, dialogue generation requirements, and segmented language texts;

[0014] A data reduction module, configured to adjust a preset reduction model according to the preset adjustment method and multiple sets of Q&A data to determine target Q&A data from the multiple sets of Q&A data, and determine the target Q&A data as the multilingual supervised data set.

[0015] According to one aspect of the present application, an electronic device is proposed, which includes: a processor; a memory storing a computer program, and when the computer program is executed by the processor, the processor is caused to execute the method as described above.

[0016] According to one aspect of the present application, a non-transitory computer-readable medium is proposed, on which readable instructions are stored, and when the instructions are executed by the processor, the processor is caused to execute the method as described above.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application.

[0018] Beneficial effects:

[0019] Through the above embodiments provided by the present application, by obtaining expected language data from a multilingual knowledge base and splitting it according to the requirements of dialogue generation, multiple sets of Q&A data are generated, effectively increasing the diversity of the dataset. This method can cover more language phenomena and dialogue scenarios, enabling the trained model to have stronger generalization ability. Using a preset large language model to generate Q&A data can introduce rich language patterns and relationships, enabling the model to come into contact with more diverse language expressions during the learning process. This not only helps the model better capture language features but also improves its performance in different language environments, thereby enhancing the robustness and generalization ability of the model. By adjusting the streamlined model with preset adjustment methods and the generated Q&A data, the accuracy and relevance of the target Q&A data can be ensured. This method can filter out low-quality or irrelevant data, thereby improving the overall quality of the dataset and making the trained model more accurate and reliable. The generated multilingual supervised dataset can be directly applied to scenarios such as multilingual dialogue systems, translation systems, and Q&A systems, and can support the model to switch and adapt between different languages. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without exceeding the scope protected by the present application.

[0021] Figure 1 It is a flowchart of a method for generating a multilingual supervised dataset based on data augmentation provided by an embodiment of the present application;

[0022] Figure 2 It is a block diagram of a device for generating a multilingual supervised dataset based on data augmentation provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0025] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0026] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0027] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps may be decomposed, while some operations / steps may be combined or partially combined, so the actual execution order may change according to the actual situation.

[0028] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below may be referred to as the second component without departing from the teachings of the concept of the present application. As used herein, the term "and / or" includes any one and all combinations of one or more of the associated listed items.

[0029] Specific implementation manners may refer to the following embodiments.

[0030] Figure 1 The flowchart of the method for generating a multilingual supervised data set based on data augmentation provided for the embodiments of the present application. The method of this embodiment can be applied to a data generation server. As Figure 1 shown, the method includes: step S10, step S11, step S12, and step S13.

[0031] In step S10, obtain the dialogue generation requirement and the preset multilingual knowledge base, and determine the expected language data from the multilingual knowledge base according to the dialogue generation requirement.

[0032] In this application, the dialogue generation requirement can be used to characterize the dialogue topic, the specific types of multiple languages, and the selection method of the expected language data, which can be pre-set or input by the user. Among them, the expected language data can be used to characterize the relevant language data that meets the user's expectations. In some implementation manners, a public database such as the Wikipedia database can be used as the multi-language knowledge base in this application. The expected language data is selected from the multi-language knowledge base according to the dialogue generation requirement.

[0033] In step S11, the expected language data is split according to the dialogue generation requirement to determine segmented language texts.

[0034] In this application, to achieve the best effect of the model processing data, it is necessary to limit the data length. Therefore, the dialogue generation requirement can include the method of splitting the expected language data to split the relevant language data into segmented language texts of appropriate lengths.

[0035] In step S12, based on a pre-set large language model, the dialogue generation requirement, and the segmented language texts, multiple sets of Q&A data are generated.

[0036] In some implementation manners, an existing language model, such as the GPT3.5 (Generative Pre-trained Transformer 3.5) model, can be used as the large language model in this application. The question generation ability and answer generation ability of the large language model are utilized. The dialogue generation requirement can include prompts for the large language data, such as system prompts, principle prompts, and dialogue history prompts, etc. Among them, the system prompt can be used to describe the task of generating the initial question, such as reading what type of papers, asking what topic questions and answering. The rule prompt can provide some detailed principles to improve the quality of the generated data.

[0037] In some implementation manners, the dialogue generation requirement and the segmented language texts can be used as the input of the large language model. Based on the dialogue generation requirement, the processes of reading, understanding, and data generation of the large language model can be restricted so that it learns the segmented language texts according to the requirement and outputs multiple sets of Q&A data.

[0038] In step S13, according to the pre-set adjustment method and multiple sets of Q&A data, the pre-set refined model is adjusted to determine the target Q&A data from the multiple sets of Q&A data, and the target Q&A data is determined as the multi-language supervised data set.

[0039] In this application, the adjustment method can be preset. In some implementation manners, the adjustment method can include two types. The first type is: first, fine-tune the base model using English code and mathematical data, and then further fine-tune the model using different amounts of Chinese data; the second type is: directly fine-tune the base model using Chinese code and mathematical data. The fine-tuned models can all become streamlined models.

[0040] By fine-tuning the model, the ability to obtain corresponding answer-related data with the least amount of input data can be achieved, and the redundant data among them can be removed. Finally, the least amount of input data and the corresponding answer-related data are used as target Q&A data, that is, a multilingual supervised dataset.

[0041] In some implementation manners, the performance of the target Q&A data can be evaluated according to a preset multilingual benchmark model to output a performance score. Some multilingual SFT datasets can be obtained and stored in advance. This application can use, for example, the Okapi dataset, the Guanaco dataset, the Multialpaca dataset, and the Phoenix SFT dataset. The multilingual benchmark model can be a currently publicly available model. Using this model to verify the preset multilingual SFT dataset and the multilingual supervised dataset obtained in this application, and output an evaluation score for subsequent public use.

[0042] This application obtains expected language data from a multilingual knowledge base and splits it according to the dialogue generation requirements to generate multiple groups of Q&A data, effectively increasing the diversity of the dataset. This method can cover more language phenomena and dialogue scenarios, making the trained model have stronger generalization ability. Using a preset large language model to generate Q&A data can introduce rich language patterns and relationships, enabling the model to come into contact with more diverse language expressions during the learning process. This not only helps the model better capture language features but also improves its performance in different language environments, thereby enhancing the robustness and generalization ability of the model. By adjusting the streamlined model through the preset adjustment method and the generated Q&A data, the accuracy and relevance of the target Q&A data can be ensured. This method can filter out low-quality or irrelevant data, thereby improving the overall quality of the dataset and making the trained model more accurate and reliable. The generated multilingual supervised dataset can be directly applied to scenarios such as multilingual dialogue systems, translation systems, and Q&A systems, and can support the model to switch and adapt between different languages.

[0043] According to some embodiments, the dialogue generation requirements and the multilingual knowledge base can be obtained; the specified topic, expected language features, and target translation language can be determined from the dialogue generation requirements; according to the specified topic, the corresponding topic text data can be extracted from the multilingual knowledge base; and the expected language data can be determined according to the topic text data, expected language features, and target translation language.

[0044] In this application, the specified topic can be used to characterize the topic related to the question-and-answer data in the multilingual supervised dataset. The expected language features can be used to characterize the features of the data that meet the user's expectations. The target translation language can be used to characterize the display language of the final dataset, which can include Chinese, English, Russian, Spanish, etc.

[0045] According to an exemplary embodiment, the pre-stored multilingual knowledge base and the dialogue generation requirements sent by the user can be obtained first, and the specified topic, expected language features, and target translation language can be extracted from the dialogue generation requirements. Match the topic text data corresponding to the specified topic in the multilingual knowledge base. Then, according to the expected language features, non-expected languages can be removed from the topic text data, and text translation can also be performed according to the target translation language to obtain the expected language data.

[0046] In some implementation manners, non-expected language features can be preset, such as full names of people, countries, regions, address-related, customs, politics, religions, poems, foods, clothes, furniture, and product brands, etc., to remove non-expected languages.

[0047] By extracting the text data related to the specified topic from the preset multilingual knowledge base, this application can significantly increase the diversity of the dataset, and can guide the model to better capture the complexity and diversity of languages. At the same time, covering more language phenomena and dialogue scenarios enables the trained model to have stronger generalization ability and be able to handle various unknown language inputs. After extracting the topic text data, screening and converting the text data according to the expected language features helps the model learn a wider range of language patterns and relationships. By combining the requirements of the target translation language, further screening and conversion of the topic text data are carried out to ensure the accuracy and relevance of the expected language data, and low-quality or irrelevant data can be filtered out, thereby improving the overall quality of the dataset. At the same time, the data closely related to the target translation language can help the model better understand and generate the target language, and improve the performance of the model in a specific language environment.

[0048] According to some embodiments, when the target translation language includes Chinese, font conversion can be performed on the topic text data; when the target translation language does not include Chinese or the font conversion is completed, according to the expected language features, non-expected language data in the topic text data is removed; the filtering requirements are determined from the dialogue generation requirements, and the topic text data after removal is filtered according to the filtering requirements to determine the expected language data.

[0049] In some implementation manners, the initial task of this application can be the processing of English texts. If the target translation language is Chinese, traditional Chinese can be converted into simplified Chinese.

[0050] In some other implementations, if the target translation language does not include Chinese or the font conversion has been completed, the non-expected language data can be excluded based on the expected language features in the above-mentioned manner. Filtering requirements can be extracted from the dialogue generation requirements. For the expected language data, documents with less than 1K tokens or more than 10K tokens can be excluded. Here, the number of tokens is calculated by a tool that splits the text into the smallest semantic units. The filtered data is used as the expected language data.

[0051] In the case where the target translation language includes Chinese, font conversion is performed on the subject text data in this application, which can ensure the accurate representation of Chinese text in the target translation language. In the case where the target translation language does not include Chinese or the font conversion is completed, non-expected language data in the subject text data is excluded according to the expected language features. This step can remove data that does not meet the requirements, such as colloquial expressions and informal terms, thereby enhancing the purity of the dataset and making model training more efficient and accurate. Filtering requirements are determined from the dialogue generation requirements, and the subject text data after exclusion is filtered according to the filtering requirements. This step can select data that meets specific requirements, such as data in a specific domain and of a specific length, thereby enhancing the diversity and coverage of the dataset.

[0052] According to some embodiments, the default text length can be determined from the dialogue generation requirements; according to the default text length, the expected language data is split into text segments to determine multiple segments of expected language text and multiple split nodes; for the multiple segments of expected language text, the split positions of the multiple split nodes in the expected language data are detected; according to the split positions, the multiple split nodes are adjusted to determine the segmented language text.

[0053] In some implementations, since the context length of most LLMs has priority, the default text length can be preset, and this length can be a value between 1K tokens and 2K tokens. First, according to the default text length, each complete text in the expected language data is split in the order of arrangement to obtain multiple segments of expected language text, and these segments of expected language text are connected by split nodes.

[0054] The split positions of the multiple split nodes in the expected language data can be detected. Requirements for the split positions can be preset, and it is determined whether the split nodes need to be adjusted according to the requirements. If adjustment is needed, new segmented language text can be obtained.

[0055] This application can make full use of long text data by determining the default text length from the dialogue-generated requirements and splitting the expected language data according to this length. Even if the original text is long, multiple training samples can be generated through splitting, thus significantly improving the utilization rate of the dataset. By detecting the splitting positions of the splitting nodes in the expected language data and adjusting the splitting nodes according to the splitting positions, segmented language texts with different lengths and structures can be generated. This flexible segmentation method can meet diverse dialogue generation requirements, such as generating responses of different lengths, dialogues of different structures, etc., thereby improving the flexibility and practicality of the model.

[0056] According to some embodiments, if the splitting position is in a complete sentence in the expected language data, the corresponding splitting node can be deleted, and the multi-segment expected language texts corresponding to the remaining splitting nodes among the multiple splitting nodes are determined as the segmented language text; if the splitting position is not in a complete sentence in the expected language data, the multi-segment expected language texts are determined as the segmented language text.

[0057] In this application, the expected language data may contain multiple texts, and each text contains multiple complete sentences.

[0058] In some implementation manners, if the splitting position is in a complete sentence, that is, the splitting node at this position breaks a complete sentence, at this time, the previous splitting node can be obtained in the order of splitting. Delete this splitting node, and use the multi-segment expected language texts obtained by splitting the remaining splitting nodes as the segmented language text.

[0059] In other implementation manners, if the splitting position is not in a complete sentence, at this time, the multi-segment expected language texts can be directly determined as the segmented language text.

[0060] This application can ensure the semantic integrity and accuracy of the segmented language text by detecting the splitting positions of the splitting nodes in the expected language data and determining whether this position is in a complete sentence. This adjustment method avoids splitting in the middle of a sentence, thus ensuring the semantic coherence and readability of the segmented language text. The semantic integrity and accuracy of the segmented language text help the model learn more accurate semantic representations. When the model encounters semantically coherent segmented language texts during the training process, it can better understand the semantic structure and relationships of the language, thereby enhancing the model's semantic understanding ability.

[0061] According to some embodiments, historical dialogue data can be determined based on the dialogue generation requirements; in the case where the historical dialogue data is empty, initial Q&A prompts are extracted from the dialogue generation requirements; the initial Q&A prompts and segmented language text are input into a large language model so that the large language model outputs the first round of dialogue, and the first round of dialogue is added to the dialogue generation requirements as historical dialogue data; in the case where the historical dialogue data is not empty, associated Q&A prompts and the number of iterations are extracted from the dialogue generation requirements; the associated Q&A prompts, historical dialogue data, segmented language text, and a preset random Q&A direction are input into the large language model so that the large language model outputs multiple rounds of dialogue according to the number of iterations, and the multiple rounds of dialogue are determined as multiple sets of Q&A data.

[0062] This application can detect whether the dialogue generation requirements contain historical dialogue data. If the historical dialogue data is not included, that is, the historical dialogue data is empty, it can indicate that the currently generated is the first round of dialogue; if the historical dialogue data is included, that is, the historical dialogue data is not empty, it can indicate that the currently generated is not the first round of dialogue.

[0063] In some implementation manners, in the case where the historical dialogue data is empty, the prompt extracted from the dialogue generation requirements is the initial Q&A prompt. The initial Q&A prompt and segmented language text are input into the large language model, and the large language model learns and outputs the first round of dialogue, and then the first round of dialogue can be added to the dialogue generation requirements as historical dialogue data.

[0064] In the case where the historical dialogue data is not empty, the prompt extracted from the dialogue generation requirements is the associated Q&A prompt, and at this time, the number of iterations can also be obtained; the associated Q&A prompts, historical dialogue data, segmented language text, and a preset random Q&A direction are input into the large language model. Among them, the random Q&A direction can include two types. One is to conduct a more detailed exploration of the same topic based on the previous round of Q&A; the other is to explore other topics based on the current round of Q&A for expansion. The large language model conducts iterative learning and stops at the preset number of iterations, outputting multiple rounds of dialogue, which are used as multiple sets of Q&A data.

[0065] This application can generate diverse Q&A data by combining segmented language texts and large language models. By using large language models to generate multi-turn conversations, it can simulate real conversation scenarios and enhance the model's conversation generation ability. During the training process, the model learns how to generate coherent and reasonable responses based on historical conversation data and segmented language texts, enabling it to generate more natural and fluent conversations. By combining conversation generation requirements and segmented language texts, the generated Q&A data is highly relevant to actual application scenarios. The conversation generation requirements specify requirements such as the theme, language features, and target translation language of the Q&A data, while the segmented language texts provide language materials that meet these requirements. This relevance ensures the quality of the Q&A data, enabling it to better serve actual application scenarios.

[0066] According to some embodiments, a streamlined model can be selected in an adjustment manner; according to a preset data set division method, multiple groups of Q&A data can be divided into a training set and a validation set; based on the training set and the validation set, the streamlined model can be trained to determine target streamlined data that matches the validation set from the training set; the target streamlined data can be determined as the target Q&A data, and the target Q&A data can be determined as a multilingual supervised data set.

[0067] In this application, different adjustment methods target different streamlined models, and there may be overlaps. According to the data set division method, multiple groups of Q&A data are divided into a training set and a validation set. The training set and the validation set are used to train the streamlined model so that the obtained Q&A data is consistent with the Q&A data in the validation set. In the case of consistency, the least amount of data used during training is obtained as the target streamlined data. This target streamlined data is determined as the target Q&A data, that is, the multilingual supervised data set.

[0068] In some implementation manners, the minimum amount of data obtained by different adjustment methods is different, and the least amount of data can be used as the target streamlined data.

[0069] By selecting a suitable simplified model according to a preset adjustment method, the present application can significantly improve the training efficiency and inference speed of the model. Simplified models usually have fewer parameters and a simpler structure, and can process a large amount of data faster while ensuring certain performance. Training the simplified model based on the training set and the validation set can ensure the accuracy and generalization ability of the model. The training set is used to train the model to learn language patterns and relationships; the validation set is used to evaluate the performance of the model to ensure that the model can also perform well on unseen data. By determining the target simplified data in the training set that matches the validation set, the quality and relevance of the target Q&A data can be ensured. This supervision method helps the model learn more accurate language patterns and relationships because the target simplified data has been verified by the validation set and has high credibility and representativeness. Determining the target simplified data as the target Q&A data and determining it as a multilingual supervision data set can provide high-quality training data for subsequent multilingual applications and services.

[0070] The device embodiments of the present application are described below, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, reference can be made to the method embodiments of the present application.

[0071] Figure 2 It is a block diagram of a device for generating a multilingual supervision data set based on data augmentation provided by an embodiment of the present application. As Figure 2 shown, the device 200 for generating a multilingual supervision data set based on data augmentation includes a data acquisition module 201, a segmentation module 202, a Q&A data generation module 203, and a data simplification module 204.

[0072] The data acquisition module 201 is configured to acquire the dialogue generation requirement and a preset multilingual knowledge base, and determine the expected language data from the multilingual knowledge base according to the dialogue generation requirement;

[0073] The segmentation module 202 is configured to split the expected language data according to the dialogue generation requirement to determine segmented language texts;

[0074] The Q&A data generation module 203 is configured to generate multiple groups of Q&A data based on a preset large language model, the dialogue generation requirement, and the segmented language texts;

[0075] The data simplification module 204 is configured to adjust a preset simplified model according to a preset adjustment method and multiple groups of Q&A data, so as to determine target Q&A data from the multiple groups of Q&A data, and determine the target Q&A data as a multilingual supervision data set.

[0076] Optionally, the data acquisition module 201 is specifically configured to:

[0077] Acquire the dialogue generation requirement and the multilingual knowledge base;

[0078] Determine a specified topic, expected language features, and target translation language from the dialogue-generated requirements;

[0079] Extract corresponding topic text data from the multilingual knowledge base according to the specified topic;

[0080] Determine the expected language data based on the topic text data, expected language features, and target translation language.

[0081] Optionally, when the data acquisition module 201 determines the expected language data based on the topic text data, expected language features, and target translation language, it is specifically used for:

[0082] When the target translation language includes Chinese, perform font conversion on the topic text data;

[0083] When the target translation language does not include Chinese or the font conversion is completed, eliminate the unexpected language data in the topic text data according to the expected language features;

[0084] Determine the filtering requirements from the dialogue-generated requirements, and perform data filtering on the eliminated topic text data according to the filtering requirements to determine the expected language data.

[0085] Optionally, the segmentation module 202 is specifically used for:

[0086] Determine the default text length from the dialogue-generated requirements;

[0087] According to the default text length, split the expected language data to determine multiple segments of expected language text and multiple split nodes;

[0088] For multiple segments of expected language text, detect the split positions of the multiple split nodes in the expected language data;

[0089] Adjust the multiple split nodes according to the split positions to determine the segmented language text.

[0090] Optionally, when the segmentation module 202 adjusts the multiple split nodes according to the split positions to determine the segmented language text, it is used for:

[0091] If the split position is in a complete sentence in the expected language data, delete the corresponding split node, and determine the multiple segments of expected language text corresponding to the remaining multiple split nodes among the multiple split nodes as the segmented language text;

[0092] If the split position is not in a complete sentence in the expected language data, determine the multiple segments of expected language text as the segmented language text.

[0093] Optionally, the Q&A data generation module 203 is specifically used for:

[0094] Generate requirements based on the conversation and determine the historical conversation data;

[0095] In the case where the historical conversation data is empty, extract the initial Q&A prompts from the conversation-generated requirements;

[0096] Input the initial Q&A prompts and segmented language text into the large language model so that the large language model outputs the first round of conversation, and add the first round of conversation as historical conversation data to the conversation-generated requirements;

[0097] In the case where the historical conversation data is not empty, extract the associated Q&A prompts and the iteration count from the conversation-generated requirements;

[0098] Input the associated Q&A prompts, historical conversation data, segmented language text, and a preset random Q&A direction into the large language model so that the large language model outputs multiple rounds of conversation according to the iteration count, and determine the multiple rounds of conversation as multiple sets of Q&A data.

[0099] Optionally, the data reduction module 204 is specifically configured to:

[0100] Select a reduction model according to the adjustment method;

[0101] Divide the multiple sets of Q&A data into a training set and a validation set according to a preset data set division method;

[0102] Train the reduction model based on the training set and the validation set to determine target reduced data that matches the validation set from the training set;

[0103] Determine the target reduced data as target Q&A data and determine the target Q&A data as a multilingual supervised data set.

[0104] The device performs functions similar to those of the method provided above. For other functions, refer to the previous description and will not be elaborated here.

[0105] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 300 in this embodiment may include: a memory 301 and a processor 302.

[0106] A computer program is stored on the memory 301. When the computer program is executed by the processor 302, the aforementioned processor 302 executes the method in the above embodiment.

[0107] Among them, the processor 302 and the memory 301 are connected, such as through a bus.

[0108] Optionally, the electronic device 300 may further include a transceiver. It should be noted that in practical applications, the number of transceivers is not limited to one, and the structure of the electronic device 300 does not constitute a limitation to the embodiments of the present application.

[0109] The processor 302 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 302 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0110] The bus may include a path for transmitting information between the above components. The bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0111] The memory 301 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0112] The memory 301 is used to store the application program code for executing the solution of this application, and is controlled by the processor 302 to execute. The processor 302 is used to execute the application program code stored in the memory 301 to implement the content shown in the foregoing method embodiments.

[0113] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0114] The electronic device of this embodiment can be used to execute the method of any of the foregoing embodiments, and its implementation principle and technical effects are similar, and will not be elaborated here.

[0115] This application also provides a non-transitory computer-readable storage medium, on which computer-readable instructions are stored. When the foregoing instructions are executed by a processor, the processor is caused to execute the method in the above embodiments.

[0116] Those of ordinary skill in the art can understand that all or part of the steps for implementing the foregoing method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a non-transitory computer-readable storage medium. When the program is executed, it executes the steps including the foregoing method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0117] The embodiments of this application have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, changes or deformations made by those skilled in the art based on the idea of this application in terms of the specific implementation manner and application scope of this application all belong to the protection scope of this application. In summary, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating a multilingual supervised dataset based on data augmentation, characterized in that, include: Acquire a dialogue generation requirement and a preset multilingual knowledge base, and determine expected language data from the multilingual knowledge base according to the dialogue generation requirement; Splitting the expected language data according to the dialogue generation requirements to determine segmented language texts; Generate multiple sets of question-answering data based on a preset large language model, the dialogue generation requirements, and the segmented language texts; According to a preset adjustment method and the multiple sets of question and answer data, the preset streamlined model is adjusted to determine target question and answer data from the multiple sets of question and answer data, and the target question and answer data is determined as a multilingual supervision data set.

2. The method according to claim 1, wherein The acquiring of the dialogue generation requirement and a preset multilingual knowledge base, and determining the expected language data from the multilingual knowledge base according to the dialogue generation requirement, includes: Acquiring the dialogue generation requirement and the multilingual knowledge base; Determining a designated topic, expected language features, and a target translation language from the dialogue generation requirements; According to the specified topic, extracting corresponding topic text data from the multilingual knowledge base; The expected language data is determined according to the subject text data, the expected language features and the target translation language.

3. The method according to claim 2, characterized in that, The step of determining the expected language data according to the subject text data, the expected language features and the target translation language includes: When the target translation language includes Chinese, performing font conversion on the subject text data; When the target translation language does not include Chinese or the font conversion is completed, according to the expected language features, the non-expected language data in the subject text data is eliminated; A filtering requirement is determined from the dialogue generation requirement, and the eliminated subject text data is filtered according to the filtering requirement to determine the expected language data.

4. The method according to claim 1, wherein The step of splitting the expected language data according to the dialogue generation requirement to determine segmented language texts includes: Determining a default text length from the dialogue generation requirements; According to the default text length, the expected language data is subjected to text splitting to determine a plurality of expected language text segments and a plurality of splitting nodes; For the multiple sections of expected language text, detecting the split positions of the multiple split nodes in the expected language data; The multiple split nodes are adjusted according to the split positions to determine the segmented language text.

5. The method according to claim 4, wherein The step of adjusting the plurality of split nodes according to the split positions to determine the segmented language text includes: If the split position is in a complete sentence in the expected language data, the corresponding split node is deleted, and the multiple segments of expected language text corresponding to the remaining split nodes in the multiple split nodes are determined as the segmented language text; If the split position is not in a complete sentence in the expected language data, the multiple sections of expected language text are determined as the segmented language text.

6. The method according to claim 1, wherein The generating of multiple sets of question-answering data based on the preset large language model, the dialogue generation requirement and the segmented language text includes: Determining historical conversation data according to the conversation generation requirement; In the case where the historical dialogue data is empty, extract an initial Q&A prompt from the dialogue generation requirements; Input the initial Q&A prompt and the segmented language text into the large language model so that the large language model outputs the first round of dialogue, and add the first round of dialogue as historical dialogue data to the dialogue generation requirements; In the case where the historical dialogue data is not empty, extract the associated Q&A prompt and the iteration count from the dialogue generation requirements; Input the associated Q&A prompt, the historical dialogue data, the segmented language text, and a preset random Q&A direction into the large language model so that the large language model outputs multiple rounds of dialogue according to the iteration count, and determine the multiple rounds of dialogue as the multiple sets of Q&A data; 7. The method according to claim 1, characterized in that, Adjust the preset refined model according to the preset adjustment method and the multiple sets of Q&A data to determine target Q&A data from the multiple sets of Q&A data, and determine the target Q&A data as a multilingual supervised dataset, including: Select the refined model according to the adjustment method; Divide the multiple sets of Q&A data into a training set and a validation set according to a preset dataset division method; Train the refined model based on the training set and the validation set to determine target refined data from the training set that matches the validation set; Determine the target refined data as the target Q&A data, and determine the target Q&A data as the multilingual supervised dataset.

8. A generating device for a multi - language supervised data set based on data augmentation, characterized in that, Include: A data acquisition module for acquiring dialogue generation requirements and a preset multilingual knowledge base, and determining expected language data from the multilingual knowledge base according to the dialogue generation requirements; A segmentation module for splitting the expected language data according to the dialogue generation requirements to determine segmented language text; A Q&A data generation module for generating multiple sets of Q&A data based on a preset large language model, the dialogue generation requirements, and the segmented language text; A data refinement module for adjusting a preset refined model according to a preset adjustment method and the multiple sets of Q&A data to determine target Q&A data from the multiple sets of Q&A data, and determining the target Q&A data as a multilingual supervised dataset.

9. An electronic device, characterized in that, Include: A processor; A memory storing a computer program, which when executed by the processor causes the processor to execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium, characterized in that, A computer-readable instruction is stored thereon, which when executed by the processor causes the processor to execute the method according to any one of claims 1-7.

Citation Information

Cited By

  • Training data generation method, electronic equipment, storage medium and product

    CN120994799A