Synthetic data set construction method and electronic equipment
By selecting highly representative and important word segmentation units during the construction of synthetic datasets to generate question-answer pairs, the problem of low coverage and domain relevance of synthetic datasets is solved, and efficient fine-tuning of LLM is achieved in data-scarce scenarios.
Patent Information
- Application Number
- CN202511494562.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-18
AI Technical Summary
Existing synthetic dataset construction methods generate synthetic datasets with low data coverage and domain relevance, resulting in poor fine-tuning performance of large language models (LLMs) in target technology scenarios with limited publicly available data.
By acquiring original multi-source documents in the target domain, using a word segmenter to divide the documents into multiple word units, filtering candidate keywords based on representativeness score and importance score, calling a pre-trained language model to generate question-answer pairs, and constructing a synthetic dataset with high coverage and high relevance.
It improves the data coverage and domain relevance of synthetic datasets, enhances the fine-tuning effect of LLM in the target domain, and reduces manual maintenance costs.
Smart Images

Figure CN120975247A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a synthetic dataset construction method and an electronic device. BACKGROUND
[0002] Large language models (LLM) not only perform outstandingly in general fields, but also achieve significant breakthroughs in some professional fields, becoming the core tool in the field of natural language processing. However, in target technical scenarios with less public data, the direct application of LLM is severely limited.
[0003] In order to solve the problem that the direct application of LLM is severely limited due to less public data, synthetic dataset construction has become the mainstream response method. The core logic is to generate sufficient training resources from limited available data for the supervised fine-tuning of LLM, thereby improving the adaptation ability of LLM in the target technical scenario. The synthetic dataset construction method in the related art obtains cut segments by fixed length cutting or fixed natural segment cutting of the limited available data of the target technical scenario, inputs the cut segments into the LLM as prompt words, obtains the corresponding question and answer pairs, and further obtains the synthetic dataset. The synthetic dataset generated in this way has low data coverage and domain relevance, which further leads to poor fine-tuning effect of LLM. SUMMARY
[0004] The present application provides a synthetic dataset construction method and an electronic device to at least solve the problem of low data coverage and domain relevance of the synthetic dataset generated by the synthetic dataset construction method in the related art.
[0005] The present application provides a synthetic dataset construction method, comprising: obtaining original multi-source documents of a target domain of a synthetic dataset to be constructed; dividing the original multi-source documents into a plurality of token units using a tokenizer; obtaining representative scores of the plurality of token units on the original multi-source documents; determining, based on the representative scores, that a token unit with a representative score higher than a first score threshold is a candidate keyword; determining, based on the representative score of the candidate keyword, an importance score of the candidate keyword; determining, based on the importance score, that a candidate keyword with an importance score higher than a second score threshold is a target keyword; calling a pre-trained language model to generate a question and answer pair corresponding to the target keyword based on the target keyword, to obtain a synthetic dataset of the target domain.
[0006] The application also provides an electronic device, comprising a memory for storing a computer program; and a processor for implementing the steps of any of the above synthetic data set construction methods when executing the computer program.
[0007] The application also provides a computer-readable storage medium having a computer program stored therein, wherein the computer program, when executed by a processor, implements the steps of any of the above synthetic data set construction methods.
[0008] The application also provides a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above synthetic data set construction methods.
[0009] According to the application, the original multi-source document of the target field to be constructed into a synthetic data set is obtained; the original multi-source document is divided into a plurality of token units by using a tokenizer; the representative score of the plurality of token units on the original multi-source document is obtained; based on the representative score, the token unit with a representative score higher than a first score threshold is determined as a candidate keyword; based on the representative score of the candidate keyword, the importance score of the candidate keyword is determined; based on the importance score, the candidate keyword with an importance score higher than a second score threshold is determined as a target keyword; a pre-trained language model is called to generate a question and answer pair corresponding to the target keyword based on the target keyword to obtain a synthetic data set of the target field. Therefore, the technical problem that the synthetic data set generated by the synthetic data set construction method in the related art has low data coverage and field relevance, resulting in poor fine-tuning effect of the LLM, can be solved, and the technical effects of improving the data coverage and field relevance of the generated synthetic data set and further improving the fine-tuning effect of the LLM are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0011] Figure 1 A structural schematic diagram of a synthetic data set construction system provided for the embodiments of the application; Figure 2 A flowchart of a synthetic data set construction method provided for the embodiments of the application; Figure 3 A flowchart of another synthetic data set construction method provided for the embodiments of the application; Figure 4 A flowchart of another synthetic data set construction method provided for the embodiments of the application; Figure 5 A structural schematic diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0013] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0014] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0015] LLM faces the bottleneck of insufficient data in technical scenarios with extremely small amount of public data, such as high-precision industrial simulation, analysis of scarce professional literature, extraction of specialized domain knowledge, etc. Among them, technical scenarios with extremely small amount of public data can also be technical scenarios in which original corpus is protected by privacy, commercial secrecy or has extremely high collection cost. Exemplarily, high-precision industrial simulation can be semiconductor simulation.
[0016] In order to solve the problem that the direct application of LLM is severely limited due to the small amount of public data, synthetic dataset construction has become the mainstream response method. The core logic is to take limited public text data in the target technical scenario with small amount of public data as the basis, generate instruction-answer pairs (i.e. question-answer pairs) in a specific task format (such as Alpaca format) through a large language model, and use them for supervised fine-tuning of LLM to improve the adaptation ability of LLM in the target technical scenario with small amount of public data.
[0017] The synthetic data set construction method in the related art specifically includes: 1, data preprocessing: the original corpus in the target technical scene is divided into multiple segments, so as to generate a synthetic data set, that is, a synthetic question and answer pair. One segmentation method can be a fixed-length segmentation method, such as dividing the original corpus into multiple segments according to the number of characters, times or Token. This segmentation method is simple and efficient, and is commonly used to construct synthetic data sets in standard formats such as Alpaca, Self-Instruct, etc. In order to enhance semantic coherence, a fixed overlap window is introduced, that is, the last several Token of the previous segment will appear at the beginning of the next segment, so as to alleviate the problem of semantic fragmentation. Another method can be to divide the original corpus in the target technical scene into multiple segments through structural information such as natural paragraphs, periods or line breaks. Even the topic boundary detection algorithm such as TextTiling is used to identify topic changes, and the original corpus in the target technical scene is divided into multiple segments, however, this method is usually based on heuristic rules such as word frequency distribution similarity, boundary value sliding window between paragraphs, etc., and lacks the ability to quantify the "information amount" of the paragraph, and it is also difficult to adapt to structured data or optical character recognition (Optical Character Recognition, abbreviated as: OCR) text and other non-standard inputs. 2, input the segmented multiple segments as prompts into the LLM in turn, generate the corresponding question and answer pair, and obtain the synthetic data set.
[0018] Among them, the fixed-length segmentation method treats all texts equally, cannot identify situations where some paragraph information is highly concentrated and some paragraph information is sparse, and is prone to truncate important information or waste generation resources on invalid paragraphs.
[0019] The above two segmentation methods are prone to split key sentences or miss topic sentences, resulting in synthetic data sets generated by the method lacking necessary context support and poor quality.
[0020] The above two segmentation methods have poor adaptability to structured texts such as formatted tables, nested long sentences, and list descriptions, and are prone to mis-cut sentences and mis-break blocks, affecting the effectiveness of downstream tasks.
[0021] The fixed overlap window method fails to dynamically determine the necessity of the overlap part, often resulting in excessive redundant context input, leading to poor quality of the synthetic data set generated.
[0022] The above two segmentation methods lack a unified evaluation index for quantifying the richness of segment information, that is, it is impossible to determine whether a segment is worth generating a question and answer pair before segmentation, resulting in poor quality of the synthetic data set generated.
[0023] In summary, the above-mentioned segmentation method causes the data coverage and domain relevance of the generated synthetic data set to be low, and cannot effectively mine potential multi-dimensional knowledge points in the target technical scenario, thereby causing the fine-tuning effect of the LLM to be poor. The data coverage is the complete coverage of the core knowledge points of the target technical scenario in the question and answer pairs in the synthetic data set. The domain relevance is the precise matching degree of the question and answer pairs in the synthetic data set and the target technical scenario.
[0024] To solve the above technical problems, the embodiment of the present application provides a synthetic data set construction method and an electronic device. The synthetic data set construction method comprises: obtaining original multi-source documents of a target domain of a synthetic data set to be constructed; dividing the original multi-source documents into a plurality of token units by using a tokenizer; obtaining representative scores of the plurality of token units on the original multi-source documents; determining, based on the representative scores, that a token unit with a representative score higher than a first score threshold is a candidate keyword; determining, based on the representative scores of the candidate keywords, an importance score of the candidate keywords; determining, based on the importance scores, that a candidate keyword with an importance score higher than a second score threshold is a target keyword; and calling a pre-trained language model to generate a question and answer pair corresponding to the target keyword based on the target keyword, to obtain a synthetic data set of the target domain. The method provided by the above scheme determines the candidate keywords in the original multi-source documents according to the representative scores, determines the target keywords based on the importance scores of the candidate keywords, generates the corresponding question and answer pairs based on the target keywords, and then obtains the synthetic data set of the target domain, thereby improving the data coverage and domain relevance of the synthetic data set, and further improving the fine-tuning effect of the large language model.
[0025] In combination with the specific application environment architecture or specific hardware architecture on which the synthetic data set construction method is dependent, the specific application environment architecture or specific hardware architecture is described herein.
[0026] The synthetic data set construction method and the electronic device provided by the embodiment of the present application are suitable for constructing a synthetic data set of a target domain, and are used for supervised fine-tuning of an LLM to improve the adaptation capability of the LLM in the target domain. For example, Figure 1As shown, a structural schematic diagram of a synthetic dataset construction system based on the synthetic dataset construction system based on which the present application is built is shown, which includes a client and a synthetic dataset construction apparatus, wherein the client is configured to send original multi-source documents of a target field of a synthetic dataset to be constructed to the synthetic dataset construction apparatus. The synthetic dataset construction apparatus is configured to obtain original multi-source documents of a target field of a synthetic dataset to be constructed; divide the original multi-source documents into a plurality of tokenized units using a tokenizer; obtain representative scores of the plurality of tokenized units on the original multi-source documents; determine, based on the representative scores, that a tokenized unit with a representative score higher than a first score threshold is a candidate keyword; determine, based on the representative scores of the candidate keywords, an importance score of the candidate keywords; determine, based on the importance scores, that a candidate keyword with an importance score higher than a second score threshold is a target keyword; invoke a pre-trained language model to generate a question and answer pair corresponding to the target keyword based on the target keyword to obtain a synthetic dataset of the target field.
[0027] Embodiments of the present application provide a synthetic dataset construction method, Figure 2 A flowchart of the synthetic dataset construction method provided by the embodiments of the present application is shown as Figure 2 The synthetic dataset construction method includes the following steps: Step S201, obtaining original multi-source documents of a target field of a synthetic dataset to be constructed.
[0028] The target field is a field with less public data, such as high-precision industrial simulation, rare professional literature analysis, and proprietary domain knowledge extraction.
[0029] It should be noted that the original corpus of the target field is the data that has been disclosed, and the original corpus can be a technical manual in PDF format, such as an industrial equipment operation manual, a professional standard specification, a scientific research paper, etc. The original corpus can also be structured data, such as tables, database export files (files in CSV, XLSX, SQL Dump, etc. formats). The original corpus often exists in the form of scanned copies, pictures, or handwritten manuscripts.
[0030] The text in the original corpus is extracted using OCR technology to obtain the original multi-source documents. It can be understood that the OCR technology can convert structured data into natural language description text (such as "the table lists the effects of temperature changes on the electrical resistivity of materials").
[0031] Step S202, dividing the original multi-source documents into a plurality of tokenized units using a tokenizer.
[0032] The language of the original multi-source documents is recognized using a fast text classification model fastText or a language recognition tool langid to determine the tokenizer used by the original multi-source documents.
[0033] Specifically, the Chinese part in the original multi-source document is divided into multiple token units using a BERT-based tokenizer, and part-of-speech tagging information of each token unit is generated. The BERT-based tokenizer can be BERT-Tokenizer.
[0034] The English part in the original multi-source document is divided into multiple token units using a natural language processing tool spaCy or NLTK (Natural Language Toolkit), and part-of-speech tagging information of each token unit is generated.
[0035] The part-of-speech tagging information is used to represent the part of speech of the token unit, such as noun, adjective, verb, etc.
[0036] Step S203, obtaining representative scores of the multiple token units on the original multi-source document.
[0037] The representative scores of the multiple token units on the original multi-source document are determined using a TF-IDF algorithm. It can be understood that the higher the representative score, the more representative the corresponding token unit is of the original multi-source document, i.e., the higher the value and the more critical the corresponding token unit is in the original multi-source document.
[0038] Step S204, determining a token unit with a representative score higher than a first score threshold as a candidate keyword based on the representative scores.
[0039] The first score threshold is set by a technician and is not specifically limited here.
[0040] It should be noted that the TextRank algorithm can also be used to determine the candidate keywords in the multiple token units. The TextRank algorithm is a text keyword extraction and abstract generation algorithm based on graph theory.
[0041] The candidate keywords can also be determined from the multiple token units through pre-set special rules of the target field (such as regular matching command format, variable name format).
[0042] Step S205, determining an importance score of the candidate keyword based on the representative score of the candidate keyword. The higher the importance score, the more important the candidate keyword is in the original multi-source document.
[0043] In step S206, based on the importance score, the candidate keywords with an importance score higher than a second score threshold are determined as target keywords.
[0044] The second score threshold is set by a technician and is not specifically limited herein.
[0045] In step S207, a pre-trained language model is called to generate a question and answer pair corresponding to the target keyword based on the target keyword to obtain a synthetic data set of the target domain.
[0046] The pre-trained language model is a large language model. It can be understood that the synthetic data set includes all question and answer pairs corresponding to the target keywords. The large language model is a deep learning model trained based on massive text data, and the core goal is to understand and generate human language. It learns from large-scale text corpus (such as books, web pages, conversations, etc.) to capture grammar rules, semantic associations, logical structures, and even cultural background knowledge in language, thereby having multiple natural language processing capabilities. Pre-training is a strategy for training deep learning models, which involves using large data sets to initially train the model so that the model learns general feature representations. This process is similar to the basic learning stage before humans learn new knowledge, through extensive reading, observation, and experience accumulation.
[0047] After obtaining the synthetic data set, the synthetic data set needs to be automatically detected. Specifically, it is detected whether the synthetic data set conforms to the specified format, whether the variable naming in the synthetic data set is uniform, whether the units in the synthetic data set are standardized, etc.
[0048] A to-be-verified data set is extracted from the obtained synthetic data set according to a preset proportion (such as 1%), and the domain correctness of the to-be-verified data set is verified manually.
[0049] If the manual verification is passed, the obtained synthetic data set is stored in a database supporting version management (such as MongoDB or Parquet file), and meta information such as source document, target keyword, generation time, and generation model version is recorded to facilitate traceability and rollback.
[0050] The synthetic data set construction method provided by the embodiment of the application determines the candidate keywords in the original multi-source document according to the representative scores of the plurality of segmentation units, determines the target keywords based on the importance scores of the candidate keywords, generates corresponding question and answer pairs based on the target keywords, and then obtains the synthetic data set of the target domain, thereby improving the data coverage and domain relevance of the synthetic data set in the scene where the domain data is scarce, and further improving the fine-tuning effect of the large language model. The target keywords are automatically determined, and the manual maintenance cost is reduced.
[0051] The embodiment of the application provides a synthetic data set construction method,Figure 3 This is a flowchart illustrating the synthetic dataset construction method provided in the embodiments of this application, as shown below. Figure 3 As shown, the method for constructing this synthetic dataset includes the following steps: Step S301: Obtain the original multi-source documents from the target domain of the synthetic dataset to be constructed. See details below. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0052] Step S302: Using a word segmenter, the original multi-source document is divided into multiple word units. For details, please refer to [link to details]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0053] Step S303: Obtain representative scores of multiple word segmentation units for the original multi-source document.
[0054] Specifically, step S303 includes: Step S3031: For any word segmentation unit, obtain the total number of word segmentation units in the target document containing the word segmentation unit and the number of times the word segmentation unit appears in the target document.
[0055] It is understandable that the original multi-source document includes multiple documents.
[0056] Step S3032: Determine the frequency of occurrence of the segmentation unit in the target document based on the total number of segmentation units and the number of occurrences.
[0057] Step S3033: Obtain the number of first documents in the original multi-source documents that include the word segmentation unit.
[0058] Step S3034: Determine the rarity of the word segmentation unit based on the total number of documents included in the original multi-source document and the number of first documents.
[0059] Step S3035: Based on the frequency of occurrence of the segmentation unit in the target document and the rarity of the segmentation unit, determine the representative score of the segmentation unit for the original multi-source document.
[0060] Step S304: Based on the representativeness score, identify word segmentation units with a representativeness score higher than the first score threshold as candidate keywords. For details, please refer to [link to relevant documentation]. Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0061] Step S305: Determine the importance score of the candidate keywords based on their representativeness score.
[0062] Specifically, step S305 includes: Step S3051, obtaining the position of the candidate keyword in the original multi-source document, the information amount of the target segment where the candidate keyword is located, and the salience of the candidate keyword in the original multi-source document.
[0063] Wherein, the length of the target segment does not exceed the maximum Token length.
[0064] Step S3052, determining the contribution degree of the candidate keyword based on the position of the candidate keyword in the original multi-source document, the information amount of the target segment where the candidate keyword is located, and the salience of the candidate keyword in the original multi-source document, wherein the more important the position of the candidate keyword in the original multi-source document, the more dense the information amount of the target segment where the candidate keyword is located, and the more salient the candidate keyword in the original multi-source document, the higher the contribution degree of the candidate keyword.
[0065] Wherein, the title position and the table header position in the original multi-source document are more important than the text position. The more formulas, parameters, and professional terms contained in the target segment, the more dense the information amount of the target segment. The more salient the candidate keyword in the original multi-source document in the case of the candidate keyword being at the beginning of a sentence or being bolded.
[0066] Determining the contribution degree of the candidate keyword based on the position of the candidate keyword in the original multi-source document, the information amount of the target segment where the candidate keyword is located, and the salience of the candidate keyword in the original multi-source document, comprises: Pre-setting the position importance scores of data in different positions in the original multi-source document. Pre-setting the corresponding relationship between different information amounts and dense scores. Pre-setting the corresponding relationship between different salient modes and salient scores.
[0067] Determining the target position importance score of the candidate keyword based on the pre-set position importance scores of data in different positions in the original multi-source document and the position of the candidate keyword in the original multi-source document.
[0068] Determining the target dense score of the target segment where the candidate keyword is located based on the pre-set corresponding relationship between different information amounts and dense scores and the information amount of the target segment where the candidate keyword is located.
[0069] Determining the target salient score of the candidate keyword based on the pre-set corresponding relationship between different salient modes and salient scores and the salience of the candidate keyword in the original multi-source document. Wherein, the corresponding salient mode is determined based on the salience, and then the target salient score is determined.
[0070] Based on the target position importance score of the candidate keyword, the target density score of the target segment where the candidate keyword is located, and the target saliency score of the candidate keyword, the contribution degree of the candidate keyword is determined. Specifically, the weights of the target position importance score, the target density score, and the target saliency score are pre-set, and the target position importance score of the candidate keyword, the target density score of the target segment where the candidate keyword is located, and the target saliency score of the candidate keyword are weighted and summed based on the pre-set weights of the target position importance score, the target density score, and the target saliency score to obtain the contribution degree of the candidate keyword. The contribution degree can be a value between 0 and 1.
[0071] It should be noted that the contribution degree of the candidate keyword is the contribution degree of the candidate keyword in the context information density analysis.
[0072] By means of "pre-set standard + quantitative calculation + weighted integration", the position importance of the candidate keyword, the information amount of the segment where the candidate keyword is located, and the saliency of the candidate keyword are converted into scores that can be objectively measured, avoiding subjective bias in determining the contribution degree. At the same time, the weights of the scores can be flexibly adjusted according to the characteristics of the target field, so that the contribution degree calculation is more suitable for professional scene requirements, and finally the candidate keywords with higher value for the field knowledge are accurately identified, providing a reliable basis for subsequent target keyword screening and high-quality synthetic dataset construction.
[0073] In step S3053, the part-of-speech tagging weighting coefficient of the candidate keyword is determined based on the part-of-speech tagging information of the candidate keyword. The part-of-speech tagging weighting coefficient is a global parameter, and the corresponding relationship table of the part-of-speech tagging information and the part-of-speech tagging weighting coefficient is pre-set by experts in the target field. Based on the pre-set corresponding relationship table of the part-of-speech tagging information and the part-of-speech tagging weighting coefficient and the part-of-speech tagging information of the candidate keyword, the part-of-speech tagging weighting coefficient of the candidate keyword is determined.
[0074] In step S3054, the importance score of the candidate keyword is determined based on the representative score of the candidate keyword, the contribution degree of the candidate keyword, and the part-of-speech tagging weighting coefficient of the candidate keyword.
[0075] In step S306, based on the importance score, the candidate keyword with an importance score higher than a second score threshold is determined as the target keyword. For details, please refer to Figure 2 The step S206 of the embodiment shown in FIG. 2 will not be repeated here.
[0076] In step S307, the pre-trained language model is called, and based on the target keyword, the question and answer pair corresponding to the target keyword is generated to obtain the synthetic dataset of the target field. For details, please refer to Figure 2 The step S207 of the embodiment shown in FIG. 2 will not be repeated here.
[0077] The synthetic data set construction method provided by the embodiment of the application determines the representative score of the word segmentation unit on the original multi-source document through the appearance frequency of the word segmentation unit in the target document and the rarity of the word segmentation unit, and then determines the candidate keywords according to the representative score, so that the word segmentation unit that appears frequently in a single target document and is rare in the whole document set is accurately screened out. These word segmentation units are often core terms in the target field, which can effectively improve the accuracy of subsequent target keyword determination and reduce the interference of general words or low-frequency irrelevant words.
[0078] By comprehensively evaluating the importance score of the candidate keywords from multiple dimensions of position importance, fragment information density and self saliency, the screening deviation of the target keywords caused by a single dimension is avoided, the accuracy of the target keywords screened based on the importance score is ensured, the core keywords with high value and high relevance in the target field are accurately identified, and the data coverage and field relevance of the generated synthetic data set are improved.
[0079] In some optional embodiments, the synthetic data set construction method further includes: Step a1, generating part-of-speech tagging information of a plurality of word segmentation units by using a word segmenter.
[0080] For details, refer to the description of the foregoing step S202, which will not be repeated here.
[0081] Step a2, inputting the plurality of word segmentation units and the part-of-speech tagging information of the plurality of word segmentation units into a target named entity recognition model to obtain professional word entities in the target field, the target named entity recognition model being a named entity recognition model fine-tuned based on the target field.
[0082] The target named entity recognition (NER) model identifies professional word entities in the target field, such as physical model names, chemical elements, device models, professional abbreviations, etc.
[0083] Step a3, determining the professional word entities in the target field as candidate keywords.
[0084] It can be understood that there are two sources of candidate keywords, one is the word segmentation unit with a representative score higher than a first score threshold based on the representative score of the plurality of word segmentation units on the original multi-source document, and the other is the professional word entities in the target field output by the target named entity recognition model.
[0085] The synthetic data set construction method provided by the embodiment of the application determines the professional word entity of the target field through a target named entity recognition model, determines the professional word entity of the target field as a candidate keyword, and supplements the candidate keyword determined based on the representative score, so as to ensure the accuracy of the candidate keyword.
[0086] In some optional embodiments, the step S3054 comprises: Step b1, determining the importance score of the candidate keyword based on a first formula, the first formula being:
[0087] wherein, is the importance score of the i th candidate keyword, is the representative score of the i th candidate keyword, is the part-of-speech tagging weighting coefficient of the i th candidate keyword, is the contribution degree of the i th candidate keyword, , and is a preset weight parameter.
[0088] It should be noted that, , and is a weight parameter set based on experience, which can be automatically adjusted through grid search or Bayesian optimization, The initial default value of can be set to 0.5, The initial default value of can be set to 0.3, The initial default value of can be set to 0.2.
[0089] In some optional embodiments, the step S307 comprises: Step c1, determining the type of the target keyword by using a pre-trained classifier.
[0090] The pre-trained classifier can be a Robustly Optimized BERT Pretraining Approach (RoBERTa) model.
[0091] It should be noted that the type of the target keyword can also be determined according to a preset type matching rule.
[0092] The type of the target keyword can be: a concept type, such as a field core definition and a principle noun; a formula type, characterized by mathematical symbols and physical units; an instruction type, such as set_param and pdbSet; a configuration type, such as a parameter name and a configuration option value; and other types, such as specific experimental conditions and material specifications.
[0093] The pre-trained classifier outputs the type of the target keyword based on the target keyword, and the specific output format is as follows: Output JSON format: { "keywords": [ {"term":"Hydrodynamic Model","type": "Concept", "score":0.92}, {"term": "pdbSet","type":"Command","score": 0.87} ]}.
[0094] Step c2, based on the type of the target keyword, determine the prompt word template of the target keyword.
[0095] Different types correspond to different prompt word templates to ensure the accuracy of the output question and answer pair in the field and the executable type.
[0096] Specifically, the prompt word template of the concept class: requires to generate definition, background principle, application scenario, and an extension question. The prompt word template of the formula class: generate the formula itself, the physical meaning of each variable, the application condition of the formula, and if necessary, attach an example calculation. The prompt word template of the instruction class: generate the function description of the command, the complete parameter list, and the example code block. The prompt word template of the configuration class: generate the meaning of the configuration item, the optional value, and the effect comparison of different configurations.
[0097] Step c3, based on the target document segment where the target keyword is located, the target keyword and the prompt word template, determine the target prompt word.
[0098] Wherein, the length of the target document segment does not exceed the maximum Token length. Embed the target document segment where the target keyword is located, the target keyword into the corresponding prompt word template to determine the target prompt word.
[0099] Step c4, input the target prompt word into the pre-trained language model to obtain the question and answer pair corresponding to the target keyword.
[0100] Wherein, the pre-trained language model can be LLaMA, GPT, DeepSeek.
[0101] It should be noted that the pre-trained language model automatically formats the output question and answer pair corresponding to the target keyword into Alpaca JSON format: {"instruction": "Please explain the meaning and application of the Hydrodynamic Model.", "input": "", "output": "Hydrodynamic Model is..."}.
[0102] It should be further explained that after obtaining the question and answer pair corresponding to the target keyword, the question and answer pair needs to be checked to ensure the quality of the question and answer pair.
[0103] Specifically, in the case where the type of the target keyword is a formula type, a regular check is performed on the question and answer pair corresponding to the target keyword to ensure that the formula symbol is closed without error.
[0104] In the case where the type of the target keyword is an instruction type, a domain grammar parser is used to verify the legality of the command in the question and answer pair corresponding to the target keyword.
[0105] In the case where the type of the target keyword is a concept type or a configuration type, keyword matching is performed on the question and answer pair corresponding to the target keyword to check whether the core information point is covered.
[0106] The synthetic data set construction method provided by the embodiments of the present application improves the accuracy and usability of the professional knowledge content of the question and answer pair by using a differentiated strategy to generate corresponding question and answer pairs based on the type of the target keyword.
[0107] In some optional embodiments, the above step S307 comprises: Step d1, for any target keyword, based on the question and answer pair corresponding to the target keyword, determining the cosine similarity between the question and answer pair corresponding to the target keyword and each question and answer pair corresponding to a target keyword in the question and answer pairs corresponding to other target keywords.
[0108] It can be understood that in order to maximize data coverage, all target keywords of the same document segment will trigger independent question and answer pair generation, but in order to avoid redundancy, the embodiments of the present application calculate the cosine similarity between the question and answer pair corresponding to the target keyword and each question and answer pair corresponding to a target keyword in the question and answer pairs corresponding to other target keywords based on a sentence-level bidirectional encoder representation converter (Sentence Bidirectional Encoder Representations from Transformers, referred to as: Sentence-BERT).
[0109] The question and answer pair corresponding to other target keywords is the question and answer pair corresponding to other target keywords other than the question and answer pair corresponding to the target keyword.
[0110] Step d2, if there is at least one question and answer pair corresponding to the other target keyword in the question and answer pair corresponding to the target keyword, and the cosine similarity between the question and answer pair corresponding to the target keyword and the question and answer pair corresponding to the other target keyword is greater than the preset similarity threshold, determining the target keyword to be deleted based on the importance score of the target keyword to be compared and the importance score of the target keyword.
[0111] The preset similarity threshold is set by a technician. For example, the preset similarity threshold is 0.9, so as to retain the question and answer pair with the highest quality.
[0112] It can be understood that the target keyword to be compared and the target keyword with the highest importance score are determined as the target keyword to be retained, and the other target keywords are determined as the target keyword to be deleted.
[0113] If there is no question and answer pair corresponding to the target keyword in the question and answer pair corresponding to the other target keyword, and the cosine similarity between the question and answer pair corresponding to the target keyword and the question and answer pair corresponding to the other target keyword is greater than the preset similarity threshold, it is determined that the synthetic data set includes the question and answer pair corresponding to the target keyword.
[0114] It should be noted that the target keyword to be deleted can also be determined according to the integrity, accuracy, clarity and format compliance of the question and answer pair corresponding to the target keyword to be compared and the question and answer pair corresponding to the target keyword.
[0115] Step d3, based on the target keyword to be deleted, the question and answer pair corresponding to the target keyword to be deleted is removed to obtain a synthetic data set of a target field.
[0116] It can be understood that the question and answer pair corresponding to the target keyword to be retained is retained.
[0117] The synthetic data set construction method provided by the embodiment of the application can trigger independent question and answer pair generation for all target keywords in the same document segment, and significantly improve the production efficiency of the synthetic data set.
[0118] By calculating the cosine similarity of the question and answer pairs corresponding to different target keywords, the semantic repetition or highly similar content can be accurately identified, the same or similar information can be avoided from being repeatedly present in the data set, the invalid learning cost during subsequent large language model supervision and fine-tuning can be reduced, and the data utilization efficiency can be improved. When similar question and answer pairs are identified, the target to be deleted is screened in combination with the keyword importance score, so as to ensure that the final synthetic data set focuses on the knowledge with “high importance and high representativeness” in the target field, avoid dilution of the model learning effect by low-value data, and help the large language model to more accurately master the field knowledge in the data-scarce professional field.
[0119] In some optional embodiments, the above step S307 comprises: Step e1, for any target keyword, obtaining a semantic consistency degree between the target keyword and each of other target keywords.
[0120] The other target keywords are other target keywords except the target keyword.
[0121] Step e2, if there is at least one target keyword to be merged in the other target keywords, which has a semantic consistency degree higher than a preset consistency degree threshold with the target keyword, merging the target keyword to be merged with the target keyword to obtain a merged target keyword.
[0122] The preset consistency degree threshold is set by a technician and is not specifically limited herein.
[0123] Step e3, calling a pre-trained language model to generate a question and answer pair corresponding to the merged target keyword based on the merged target keyword.
[0124] It can be understood that if there is no target keyword in the other target keywords, which has a semantic consistency degree higher than the preset consistency degree threshold with the target keyword, the pre-trained language model is called to generate a question and answer pair corresponding to the target keyword based on the target keyword.
[0125] The synthetic data set construction method provided by the embodiments of the present application avoids the dispersion of semantically repeated keywords by calculating the semantic consistency degree between the target keywords and merging the high-similarity keywords, so that the keywords participating in the question and answer generation are more focused on the core terms of the target field, the redundant cost of subsequent data processing is reduced, and the knowledge density of the synthetic data set is ensured.
[0126] In some optional embodiments, before the original multi-source document is divided into a plurality of token units by using a tokenizer, the above synthetic data set construction method further includes: Step f1, encoding the original multi-source document into a unified format.
[0127] The unified format can be a Unicode Transformation Format - 8-bit (abbreviated as: UTF-8) format, so as to eliminate the inconsistency of the encoding of documents from different sources.
[0128] Step f2, performing format cleaning on the original multi-source document to remove invalid information in the original multi-source document, the invalid information including at least one of a header and footer, a page number, a repeated title, and a table of contents page.
[0129] Specifically, the regular expression is used to batch remove the invalid information such as the header and footer, the page number, the repeated title, the table of contents page, and the like in the original multi-source document.
[0130] Step f3, based on the representative score of the original multi-source document corresponding to the plurality of segmented units, determine the segmented units with a representative score lower than the first score threshold, if the first paragraph in the original multi-source document includes segmented units with a representative score lower than the first score threshold, delete the first paragraph; if the second paragraph in the original multi-source document does not include the core words in the target domain dictionary, delete the second paragraph.
[0131] It can be understood that the original multi-source document usually contains format noise and recognition errors, and the domain noise filtering process eliminates the paragraphs irrelevant to the target domain in the original multi-source document, reducing the invalid generation proportion. The domain noise filtering process is the content of step f3. The target domain dictionary includes a plurality of core words of the target domain.
[0132] It should be noted that the original multi-source document also needs to be subjected to recognition noise removal processing, data formula format recovery processing, line merging processing, and error correction processing. It is also necessary to use a statistical language model-based sentence segmentation algorithm (such as a maximum entropy sentence segmenter) to recover the sentence boundaries of the original multi-source document, avoiding semantic fragmentation caused by line breaks.
[0133] It can be understood that the synthetic dataset construction method provided by the present application is suitable for technical scenarios with data scarcity and strong domain specialization, realizes the automatic generation of high-quality question and answer pairs from original multi-source documents, is compatible with structured, semi-structured and OCR noise text, and reduces the difficulty of early data cleaning.
[0134] The synthetic dataset construction method provided by the embodiment of the present application optimizes the quality of the original multi-source document from the source by performing format unification processing, format cleaning processing and domain noise filtering processing, providing a high-quality text input of "format unification, noise removal, domain focus" for the subsequent word segmentation and keyword extraction steps.
[0135] The embodiment of the present application provides a synthetic dataset construction method, Figure 4 The flowchart of the synthetic dataset construction method provided by the embodiment of the present application is shown in Figure 4 As shown in the figure, the synthetic dataset construction method includes the following steps: First, the original multi-source document input. For details, see the related description of the aforementioned step S301, which will not be repeated here.
[0136] Second, data preprocessing. For details, see the related description of the aforementioned steps f1 to f3, which will not be repeated here.
[0137] Third, keyword extraction. That is, to determine the target keywords, for details, see the related description of the aforementioned steps S302 to S306, which will not be repeated here.
[0138] Step 4: Keyword type classification. For details, refer to the description of step c1 above.
[0139] Step 5: Keyword-driven question and answer generation. For details, refer to the description of steps c2 to c4 above.
[0140] Step 6: Multi-keyword combination and deduplication. For details, refer to the description of steps d1 to d3 and e1 to e3 above.
[0141] Step 7: Quality control and format verification. For details, refer to the description of steps S207 and c4 above.
[0142] Step 8: Structured storage and version management. For details, refer to the description of step S207 above.
[0143] The synthetic dataset construction method provided by the embodiments of the present application determines candidate keywords in the original multi-source document according to the representative scores of the plurality of segmentation units for the original multi-source document, determines target keywords based on the importance scores of the candidate keywords, generates corresponding question and answer pairs based on the target keywords, and further obtains a synthetic dataset of a target field, thereby improving the data coverage and field relevance of the synthetic dataset in a scenario where field data is scarce, and further improving the fine-tuning effect of a large language model.
[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation.
[0145] The embodiments of the present application also provide an electronic device, as shown in the figure, comprising a processor 501 and a memory 502, the memory 502 storing a computer program, and the processor 501 is configured to run the computer program to perform the steps in any of the above synthetic dataset construction method embodiments. Figure 5
[0146] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above synthetic dataset construction method embodiments when running.
[0147] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0148] Embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program, when executed by a processor, implements the steps in any of the synthetic data set construction method embodiments described above.
[0149] Embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in any of the synthetic data set construction method embodiments described above.
[0150] The skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0151] The above describes in detail a synthetic data set construction method and an electronic device provided by the present application. The principles and implementation modes of the present application are described by applying specific examples in this paper, and the above description of the examples is only applicable to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A synthetic dataset construction method, characterized by, The method comprises: obtaining original multi-source documents of a target field to be constructed into a synthetic data set; dividing the original multi-source documents into a plurality of token units by using a tokenizer; obtaining representative scores of the plurality of token units for the original multi-source documents; determining, based on the representative scores, that a token unit with a representative score higher than a first score threshold is a candidate keyword; determining, based on the representative scores of the candidate keywords, an importance score of the candidate keyword; determining, based on the importance score, that a candidate keyword with an importance score higher than a second score threshold is a target keyword; calling a pre-trained language model to generate a question and answer pair corresponding to the target keyword based on the target keyword to obtain a synthetic data set of the target field.
2. The method of claim 1, wherein, The method further comprises: generating part-of-speech tagging information of the plurality of token units by using a tokenizer; inputting the plurality of token units and the part-of-speech tagging information of the plurality of token units into a target named entity recognition model to obtain professional word entities of the target field, the target named entity recognition model being a named entity recognition model fine-tuned based on the target field; determining the professional word entities of the target field as the candidate keywords. The method further comprises: obtaining a position of the candidate keyword in the original multi-source documents, an information amount of a target segment in which the candidate keyword is located, and a significance of the candidate keyword in the original multi-source documents; 3. The method of claim 1, wherein, determining a contribution degree of the candidate keyword based on the position of the candidate keyword in the original multi-source documents, the information amount of the target segment in which the candidate keyword is located, and the significance of the candidate keyword in the original multi-source documents, wherein the more important the position of the candidate keyword in the original multi-source documents, the more concentrated the information amount of the target segment in which the candidate keyword is located, and the more significant the candidate keyword in the original multi-source documents, the higher the contribution degree of the candidate keyword; determining a part-of-speech tagging weighting coefficient of the candidate keyword based on the part-of-speech tagging information of the candidate keyword; determining the importance score of the candidate keyword based on the representative score of the candidate keyword, the contribution degree of the candidate keyword, and the part-of-speech tagging weighting coefficient of the candidate keyword. 4. The method of claim 1, wherein, 5. The method of claim 4, wherein, The importance score of the candidate keyword is determined based on the representative score of the candidate keyword, the contribution degree of the candidate keyword, and the part-of-speech tagging weighting coefficient of the candidate keyword, and includes: The importance score of the candidate keyword is determined based on a first formula, and the first formula is: wherein, is an importance score of the i-th candidate keyword, is a representativeness score of the i-th candidate keyword, is a part-of-speech tagging weighting coefficient of the i-th candidate keyword, is a contribution degree of the i-th candidate keyword, , and is a preset weight parameter.
6. The method of claim 1, wherein, The pre-trained language model is called to generate the question and answer pair corresponding to the target keyword based on the target keyword, and includes: The type of the target keyword is determined by using a pre-trained classifier; The prompt word template of the target keyword is determined based on the type of the target keyword; The target prompt word is determined based on the target document segment where the target keyword is located, the target keyword, and the prompt word template; The target prompt word is input into the pre-trained language model to obtain the question and answer pair corresponding to the target keyword.
7. The method of claim 1, wherein, The question and answer pair corresponding to the target keyword is generated based on the target keyword to obtain the synthesized data set of the target domain, and includes: For any target keyword, the cosine similarity between the question and answer pair corresponding to the target keyword and the question and answer pair corresponding to each target keyword in the question and answer pairs corresponding to other target keywords is determined based on the question and answer pair corresponding to the target keyword; If the cosine similarity between at least one question and answer pair corresponding to a target keyword to be compared and the question and answer pair corresponding to the target keyword in the question and answer pairs corresponding to other target keywords is greater than a preset similarity threshold, a target keyword to be deleted is determined based on the importance score of the target keyword to be compared and the importance score of the target keyword. Based on the target keyword to be deleted, the question and answer pair corresponding to the target keyword to be deleted is removed to obtain the synthesized data set of the target domain.
8. The method of claim 1, wherein, The pre-trained language model is called to generate the question and answer pair corresponding to the target keyword based on the target keyword, and includes: For any target keyword, the semantic consistency between the target keyword and each target keyword in other target keywords is obtained; If the semantic consistency between at least one target keyword to be merged and the target keyword in the other target keywords is higher than a preset consistency threshold, the target keyword to be merged and the target keyword are merged to obtain a merged target keyword; The pre-trained language model is called to generate the question and answer pair corresponding to the merged target keyword based on the merged target keyword.
9. The method of claim 1, wherein, Before the original multi-source document is divided into a plurality of tokenization units by using a tokenizer, the method further includes: The original multi-source document is encoded into a unified format; The original multi-source document is format cleaned to remove invalid information in the original multi-source document, and the invalid information includes at least one of header and footer, page number, repeated title, and directory page. Determine, based on representative scores of a plurality of segmented units corresponding to the original multi-source document, segmented units whose representative scores are lower than a first score threshold, and delete a first paragraph in the original multi-source document if the first paragraph includes only segmented units whose representative scores are lower than the first score threshold; delete a second paragraph in the original multi-source document if the second paragraph does not include a core word in a target domain dictionary.
10. An electronic device, comprising: Comprise: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the synthetic data set construction method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Keyword extraction method and device, storage medium and equipment
CN112257424A
Text similarity calculation method fusing improved YAKE and neural network
CN115129815A
Home service support method and device based on large language model, and medium
CN119557401A
Intelligent prospecting question-answering system based on LLM and RAG
CN119938814A
RAG system evaluation data set automatic synthesis method and device based on reinforcement learning
CN120596663A
Cited By
Data set construction method and device, equipment, readable storage medium and program product
CN121614870A
Problem set generation method and device, storage medium and program product
CN121615623A
Question and answer data synthesis method and system based on plan driving
CN122019731A