Method, apparatus, electronic device and storage medium for constructing training data set
Through adaptive field evaluation functions and large-model assisted discrimination, the quality and efficiency problems of vertical large-model training data sets are solved, and the dynamic evaluation and screening of high-quality data is realized, which meets the needs of multiple scenarios and deep-level specialization.
Patent Information
- Application Number
- CN202510288650.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-12
AI Technical Summary
When building training data sets for vertical large models, the existing technology has problems such as uneven data quality, insufficient professionalism, inconsistent evaluation standards, limited data sources and high cleaning costs, which are difficult to meet the needs of multiple scenarios and deep-level specialized data.
Adaptive field evaluation function is adopted, and data and screening are dynamically evaluated through large models and combined with text classification models and keyword matching to build a high-quality training data set, including data acquisition, preprocessing, comprehensive classification scoring and target data set screening.
It broadens the data source, reduces cleaning costs, unifies quality standards, improves the purity of data in professional fields, and meets the needs of multiple scenarios and in-depth professional data.
Smart Images

Figure CN119782830B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for constructing a training data set. Background Art
[0002] In the training process of large models, the pre-training link of the model is particularly crucial, accounting for more than 95% of the training time. During this process, the model constructs a basic knowledge framework through large-scale data. The higher the quality of the data used for model training, the stronger the model's cognitive and generalization abilities.
[0003] However, currently, the training data used for pre-training mainly relies on web crawling. Although the quantity is large, there are generally problems such as uneven quality and insufficient professionalism. Taking a certain autoregressive language model as an example, the high-quality training data used for its training accounts for less than 20%, and it is scattered in different scenarios and fields. The proportion of high-quality training data in professional fields is even scarcer.
[0004] With the rapid development of the Internet, this current situation makes it face the dilemma of lacking high-quality training data when constructing vertical large models (high-quality models focusing on specific fields), thus restricting the improvement of model performance and the in-depth development in professional applications. Summary of the Invention
[0005] The present invention provides a method, device, electronic device, and storage medium for constructing a training data set, aiming to solve the defects in the prior art in constructing a model training set, such as limited data sources, inconsistent evaluation criteria, insufficient professional depth, and difficulty in fully meeting the professional data requirements of multiple scenarios and deep levels.
[0006] The present invention provides a method for constructing a training data set, including the following steps:
[0007] Collect a first data set;
[0008] Preprocess the first data set to obtain a second data set, where the preprocessing includes converting non-text type data in the first data set into text type data;
[0009] Obtain the comprehensive classification scores of each data in the second data set as data within the target field, where the target field is the field to which the vertical large model to be trained belongs;
[0010] Based on the comprehensive classification scores of each data in the second data set, screen out the target training data set from the second data set.
[0011] According to the method for constructing a training data set provided by the present invention, the collecting of the first data set includes:
[0012] For a target data source in any vertical domain related to the target domain, perform the following discrimination operations: Extract a part of the data samples to construct a vertical domain data set, and convert each of the data samples into a text data sample; Obtain the correlation score between each of the text data samples and the target domain; Based on the comparison result between the correlation score of each of the text data samples and a first preset threshold, determine whether to completely collect the remaining data in the target data source;
[0013] If it is determined, add all the data in the collected target data source to the initial data set, and set the data source in the next vertical domain related to the target domain as the target data source, and re-perform the discrimination operation;
[0014] If it is not determined, directly set the data source in the next vertical domain related to the target domain as the target data source, and re-perform the discrimination operation;
[0015] Iteratively execute until a preset stop condition is met, and use the obtained initial data set as the first data set.
[0016] According to the method for constructing a training data set provided by the present invention, the obtaining the correlation score between each of the text data samples and the target domain includes:
[0017] Generate domain guiding text related to the correlation score of the target domain;
[0018] Input each of the text data samples and the domain guiding text into a pre-trained correlation evaluation model respectively, and obtain the domain discrimination result and the correlation score for each of the text data samples output by the correlation evaluation model;
[0019] The domain discrimination result is used to represent whether the text data sample belongs to the target domain; The correlation score is used to represent the matching degree between the text data sample and the target domain.
[0020] According to the method for constructing a training data set provided by the present invention, the determining whether to completely collect the remaining data in the target data source based on the comparison result between the correlation score of each of the text data samples and a first preset threshold includes:
[0021] Obtain the first quantity of all text data samples whose domain discrimination result is not passed;
[0022] Obtain the second quantity of all text data samples whose correlation score is lower than a second preset threshold;
[0023] Calculate the proportion of unqualified data according to the first quantity, the second quantity, and the total amount of data in the vertical domain dataset;
[0024] Determine whether to completely collect the remaining data in the target data source according to the comparison result between the proportion of unqualified data and the first preset threshold.
[0025] According to the method for constructing a training dataset provided by the present invention, for any data in the second dataset, obtaining the comprehensive classification score for classifying each data in the second dataset as data in the target domain includes:
[0026] Input the any data into a text classification model, and obtain the classification probability that the any data is classified as data in the target domain output by the text classification model;
[0027] Construct a keyword set related to the target domain, and perform a word segmentation operation on the any data to obtain a vocabulary set, so as to determine the matching score between the vocabulary set and the keyword set;
[0028] Determine the comprehensive classification score of the any data according to the weighted combination of the classification score determined by the classification probability and the matching score.
[0029] According to the method for constructing a training dataset provided by the present invention, screening out a target training dataset from the second dataset based on the comprehensive classification scores of the data in the second dataset includes:
[0030] Exclude the data in the second dataset whose comprehensive classification score is less than the third preset threshold to obtain the target training dataset.
[0031] According to the method for constructing a training dataset provided by the present invention, the text classification model is obtained by pre-training with multiple data samples in a vertical domain labeled with classification probability labels, and further includes:
[0032] Obtain the model performance index of the text classification model after training each sample data, and the comprehensive classification score test value obtained based on each sample data;
[0033] Analyze the model performance index and the comprehensive classification score test value of all the sample data by using a heat map and a confusion matrix;
[0034] Determine the weight combination corresponding to each vertical domain and the value of the third preset threshold according to the analysis result, and the weight combination is the weight distribution relationship between the classification probability and the matching score.
[0035] According to the method for constructing a training data set provided by the present invention, the preprocessing of the first data set includes:
[0036] Obtain each HTML-formatted data in the first data set;
[0037] Locate and strip out the body part in each of the HTML-formatted data;
[0038] Extract the text data in each of the body parts, and remove the redundant formats and useless information in the text data.
[0039] According to the method for constructing a training data set provided by the present invention, the preprocessing of the first data set further includes:
[0040] Obtain each PDF-formatted data in the first data set;
[0041] Identify the text blocks in the PDF-formatted data based on optical character recognition technology;
[0042] According to the positions of the text blocks in the PDF-formatted data, perform layout restoration on all the text blocks and convert them into single-column mode text data;
[0043] Remove the useless information in the single-column mode text data.
[0044] The present invention also provides a device for constructing a training data set, including the following modules:
[0045] A data source collection module, configured to collect a first data set;
[0046] A data preprocessing module, configured to preprocess the first data set to obtain a second data set, and the preprocessing includes converting non-text type data in the first data set into text type data;
[0047] An adaptive domain evaluation module, configured to obtain a comprehensive classification score for classifying each data in the second data set as data within a target domain, where the target domain is the domain to which the vertical large model to be trained belongs;
[0048] A data set screening module, configured to screen out a target training data set from the second data set based on the comprehensive classification scores of the data in the second data set.
[0049] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method for constructing a training data set as described in any one of the above.
[0050] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for constructing a training data set as described in any one of the above is implemented.
[0051] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for constructing a training data set as described in any one of the above is implemented.
[0052] The method, device, electronic device and storage medium for constructing a training data set provided by the present invention calculate a comprehensive classification score for each data by introducing an adaptive domain evaluation function, and can dynamically evaluate and screen data according to the requirements of each scenario and domain, so as to have obvious technical improvement effects in aspects such as broadening the data source, reducing the cleaning cost, unifying the quality standard and improving the purity of data in the professional field. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0054] Figure 1 is a schematic flowchart of the method for constructing a training data set provided by the present invention.
[0055] Figure 2 is a schematic flowchart of collecting a first data set provided by the present invention.
[0056] Figure 3 is a first schematic structural diagram of the device for constructing a training data set provided by the present invention.
[0057] Figure 4 is a second schematic structural diagram of the device for constructing a training data set provided by the present invention.
[0058] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0060] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and thus cannot be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0061] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0062] A vertical large model refers to a high-quality network model focusing on a specific field. In the prior art, in terms of improving the quality of pre-training data for vertical large models, it mainly focuses on the following two points:
[0063] Optimizing data source selection: That is, try to obtain text data from authoritative, professional and screened channels, such as professional databases, scientific research paper databases and high-confidence cleaned data sets, so as to fundamentally reduce the mixing ratio of low-quality and irrelevant data.
[0064] Multi-round refined cleaning and filtering: First, use automated tools (such as regular expressions, keyword filtering) to perform basic cleaning on the initial data set, and eliminate advertisements, redundant and obviously irrelevant data. Subsequently, with the help of manual sampling inspection and small-scale quality classifiers for multiple filtering, continuously eliminate suspicious content and low-quality segments, and finally form a relatively pure and more professional training data set.
[0065] Although the above solutions improve the quality of training data to a certain extent, they are still insufficient when dealing with vertical large models and are difficult to fully meet the professional data needs of multi-scenarios and deep levels. Their defects are mainly manifested in:
[0066] (1) The evaluation criteria are not unified. Different scenarios and fields have different quality requirements for training data, making it difficult to establish a general evaluation standard, which leads to the complication of data integration and management.
[0067] (2) The lack of professional depth. Existing methods mostly focus on removing low-quality or redundant data, lacking refined optimization means for specific professional fields and high-difficulty scenarios, which restricts the professional improvement space of vertical large models.
[0068] (3) The limitation of data sources. The sources of high-quality training data are limited in themselves, with high acquisition costs and insufficient coverage, making it difficult to meet the requirements of large models in multiple scenarios and fields.
[0069] (4) The high cost of data cleaning. Multiple rounds of data filtering, manual quality inspection, and classifier optimization require a large amount of manpower and time, with low efficiency and being not conducive to rapid iteration.
[0070] In view of the above problems, the present invention proposes a method for constructing a training data set for a vertical large model. By introducing an adaptive domain evaluation function, it can dynamically evaluate and screen data according to the requirements of each scenario and field, thereby significantly improving the deficiencies of the existing technology in terms of broadening data sources, reducing cleaning costs, unifying quality standards, and improving the purity of data in professional fields. The following combines Figures 1-5 to describe the construction method, device, electronic device, and storage medium of the training data set provided by the present invention.
[0071] Figure 1 is a schematic flowchart of the construction method of the training data set provided by the present invention. As Figure 1 shown, it includes but is not limited to the following steps:
[0072] Step 101, collect the first data set.
[0073] The present invention can obtain potential data sources through various channels, including authoritative databases, industry reports, professional journals, academic paper databases, standard specification texts, and vertical domain data sets provided by industry partners, etc. To ensure the breadth and diversity of the collected data, various methods such as Application Programming Interface (API) calls, targeted web crawler scraping, and batch import can be used to achieve this.
[0074] For example, for a vertical large model in the medical field, medical-related text data can be collected from the PubMed database, medical journal websites, and the official websites of major medical institutions.
[0075] Step 102, preprocess the first data set to obtain a second data set. The preprocessing includes converting non-text type data in the first data set into text type data.
[0076] Preprocess the collected first data set. The preprocessing steps include, but are not limited to, data cleaning, format unification, type conversion, etc. All the data after preprocessing constitutes a new data set, called the second data set.
[0077] Among them, data cleaning is mainly used to remove irrelevant content, such as advertisements, redundant information, text data that is significantly irrelevant and obviously incorrect to the vertical large model to be trained. Format unification means unifying the data formats from different sources into a format convenient for subsequent processing, such as plain text format or JSON format in a specific format, etc. In particular, in this embodiment, it is necessary to convert non-text type data (such as pictures, tables, PDFs, etc.) in the collected first data set into text type data.
[0078] Step 103: Obtain the comprehensive classification scores of each data in the second data set as data within the target field.
[0079] Among them, the target field is the field to which the vertical large model to be trained belongs. Since the data quality requirements vary in different fields and it is difficult to establish a unified standard, in order to evaluate the matching degree of each data in the second data set with the target field, the present invention proposes an adaptive domain evaluation function. By selecting benchmark source data as a reference, for example, selecting a well-known data set within the target field as the benchmark source data, and performing matching scoring by processing the benchmark source data to make the classification standard consistent with the benchmark. At the same time, combining the classification probabilities obtained by the network classifier for domain classification of each data, calculate the comprehensive classification score of each data, so as to use the comprehensive classification score to achieve a unified evaluation of all the data in the second data set.
[0080] Step 104: Based on the comprehensive classification scores of each data in the second data set, screen out the target training data set from the second data set.
[0081] First, a preset threshold can be set according to experience or historical data, combined with the characteristics of the target field where the vertical large model to be trained is located. Retain the data in the second data set whose comprehensive classification score is greater than or equal to the preset threshold, and regard the remaining data whose comprehensive classification score is less than the preset threshold as data that does not meet the requirements and exclude them to construct the target training data set.
[0082] The method for constructing the training data set provided by the present invention can dynamically evaluate and screen data according to the requirements of each scenario and field by introducing an adaptive domain evaluation function to calculate the comprehensive classification score of each data, thereby having obvious technical improvement effects in aspects such as broadening the data source, reducing the cleaning cost, unifying the quality standard, and improving the purity of data in the professional field.
[0083] Figure 2 It is a schematic flowchart for collecting the first data set provided by the present invention. As Figure 2 shown, in this embodiment, the specific implementation manner of how to achieve the collection of the first data set will be described in detail.
[0084] The premise for the present invention to construct training data for a certain vertical large model to be trained is to adopt a large model assisted discrimination strategy to construct an initial data set (referred to as the first data set) with diversity and wide coverage. The main purpose is to quickly filter out unsuitable data sources and reduce the cost of obtaining data sources. The specific implementation manner is as follows:
[0085] First, obtain data sources. Potential data sources can be obtained through various channels, including authoritative databases, industry reports, professional journals, academic paper databases, standard specification texts, and vertical domain data sets provided by industry partners, etc. This module supports multiple methods such as API calls, targeted web crawler scraping, and batch import to ensure the diversity and wide coverage of data sources.
[0086] Among them, through the API call method, it can be directly docked with the databases of other platforms or other systems to obtain target data sources within any vertical domain related to the target domain. The advantage of this API call method is that the accuracy and integrity of the data source are relatively high, and at the same time, technical obstacles that may be encountered during the crawling process can be avoided.
[0087] The targeted web crawler scraping method refers to targeting specific websites or data sources (generally websites or data sources in vertical domains related to the target domain), and extracting the required data by writing a web crawler program. Through targeted web crawler scraping, data can be flexibly obtained from multiple different data sources, covering a wider range of fields and scenarios. At the same time, targeted web crawler scraping can also be customized according to needs to meet specific data requirements.
[0088] The batch import method means importing existing data sets or data files at one time, which is mainly applicable to the situation where there are already a large number of data sources. By batch import, a large-scale database can be quickly constructed, which can greatly save time and labor costs.
[0089] Since the amount of unsupervised data collected by the previous data collection method is relatively large, in order to avoid excessive bandwidth and storage consumption, the present invention will first determine multiple vertical domains related to the target domain, and select a data source within any vertical domain related to the target domain as the target data source, and further screen the target data source by the following steps.
[0090] First, randomly extract a very small portion of data from the target data source (e.g., 5%, with a data volume of m records), and then pause the extraction. Use the extracted data as data samples to construct a dataset (referred to as the vertical domain dataset). By detecting and analyzing this randomly extracted portion of data samples, determine whether to continue extracting data from the target data source, thereby effectively reducing unnecessary resource investment.
[0091] Subsequently, according to the characteristics and requirements of the target data source, select a suitable text extraction tool, such as Trafilatura. Configure the parameters of the tool according to specific extraction requirements to efficiently extract and process the text data in the data samples using these text extraction tools, and automatically identify and remove parts that are irrelevant to the required content, such as advertisements, navigation links, copyright information, etc., thereby ensuring that the extracted text is purer. For example, set the text extraction range, filtering conditions, etc. to ensure that the extracted text data meets expectations.
[0092] Input the sample data into the configured text extraction tool and perform the extraction operation. The tool will automatically analyze the text structure, extract the required content, and remove the irrelevant parts. The text data extracted from each data sample is called a text data sample, and these text data samples can be further verified to ensure that their structures are clear, the content is pure, and they meet the requirements of subsequent quality assessment and analysis.
[0093] Furthermore, for these extracted text data samples, domain guiding texts (which can be called Prompts) related to the determination of the target domain can be designed. For example, in the case where the target domain is the medical field, the Prompt can be designed as a question "Does this text data involve medical terms or medical cases?".
[0094] As a core improvement point, the present invention uses a large model to perform the relevance discrimination and scoring of each text data sample with the target domain. Based on the preset domain knowledge base and semantic understanding ability of the large model, analyze whether each text data sample belongs to the target domain or a specific application scenario closely related to the target domain, and obtain the relevance score between each text data sample and the target domain.
[0095] Then, based on the comparison result of the relevance score with a preset threshold, determine whether the association between each text data sample and the target domain is close. Furthermore, based on the relevance scores in all text data samples, determine whether to completely collect the remaining data in the target data source.
[0096] If it is determined that the data in the target data source has been completely collected, the collected data is added to the initial data set, and the data source in the next vertical domain related to the target domain is set as the new target data source, and the above discrimination operation is performed again.
[0097] If it is not determined, the current target data source is directly skipped, the data source in the next vertical domain related to the target domain is set as the new target data source, and the discrimination operation is performed again.
[0098] All the above processes will be iteratively executed until a preset stop condition is met, and the obtained initial data set is used as the first data set.
[0099] The preset stop condition may include but is not limited to the following situations: reaching a predetermined number of iterations, the initial data set reaching a predetermined data volume, and all data sources in the vertical domains related to the target domain having been traversed and processed.
[0100] The method for constructing the training data set provided by the present invention can effectively collect data from multiple vertical domains related to the target domain, screen according to the relevance score between the data and the target domain, and finally obtain a high-quality training data set, providing strong support for the training of the vertical large model.
[0101] Based on the content of the above embodiments, as an alternative embodiment, the obtaining of the relevance score between each text data sample and the target domain includes:
[0102] Generating domain guiding text related to the relevance score of the target domain;
[0103] Respectively inputting each text data sample and the domain guiding text into a pre-trained relevance evaluation model, and obtaining the domain discrimination result and the relevance score of the relevance evaluation model output for each text data sample;
[0104] The domain discrimination result is used to represent whether the text data sample belongs to the target domain; the relevance score is used to represent the matching degree between the text data sample and the target domain.
[0105] The following combines Figure 2 as shown, to detail how to perform the relevance score between the text data sample and the target domain.
[0106] First, according to the characteristics of the target domain, a series of representative questions or instructions are designed and converted into domain guidance text Prompt, which is mainly used to provide specific guidance on the target domain or specific application scenarios related to the target domain to the vertical large model during auxiliary discrimination, so as to help the model more accurately evaluate the relevance of text content. For example, for the medical field, the Prompt can be designed as "Does this text involve medical terms or medical cases?".
[0107] Optionally, the Prompt is used as a part of the text and spliced with the text data sample to be discriminated to form a complete input sequence, which is input into a pre-trained relevance evaluation model. When processing this input sequence, the relevance evaluation model will consider both the Prompt and the text data sample to be discriminated, and perform relevance discrimination according to the requirements or instructions in the Prompt and give a specific relevance score.
[0108] As another alternative embodiment, assuming that the relevance evaluation model is a conditional generation model (such as a conditional variational autoencoder, a conditional generative adversarial network, etc.), the Prompt can be used as an additional input condition to guide the discrimination process of the relevance evaluation model.
[0109] The relevance evaluation model outputs a relevance score for each text data sample that passes the relevance discrimination. The score range can be set from 1 to 100, and the higher the score, the higher the degree of matching with the target domain.
[0110] As an alternative embodiment, determining whether to completely collect the remaining data in the target data source based on the comparison result between the relevance score of each text data sample and a first preset threshold includes:
[0111] Obtaining a first quantity of all text data samples whose domain discrimination result is not passed;
[0112] Obtaining a second quantity of all text data samples whose relevance score is lower than a second preset threshold;
[0113] Calculating the proportion of unqualified data according to the first quantity, the second quantity, and the total amount of data in the vertical domain dataset;
[0114] Determining whether to completely collect the remaining data in the target data source according to the comparison result between the proportion of unqualified data and the first preset threshold.
[0115] Specifically, the present invention calculates the proportion of text data samples that meet the domain relevance standard in the previously sampled vertical domain dataset according to the discrimination result of the relevance evaluation model. The specific steps include:
[0116] Perform relevance screening, count the number of text data samples in the vertical domain dataset where the domain discrimination result is not passed (i.e., does not meet the relevance criteria), and record it as the first quantity 。
[0117] Perform relevance scoring evaluation, count the number of all text data samples in the vertical domain dataset where the domain discrimination result is passed but the relevance score is lower than the second preset threshold (such as 60), and record it as the second quantity 。
[0118] Finally, according to the first quantity, the second quantity, and the total amount of data in the vertical domain dataset, calculate the proportion of unqualified data. The calculation formula is: 。
[0119] Set the first preset threshold , if the proportion of unqualified data is greater than or equal to the first preset threshold In this case, terminate the crawling of the target data source and switch to a new data source of other potential high-quality data sources, and continue to repeat the above process.
[0120] Conversely, if the proportion of unqualified data is less than the first preset threshold It indicates that the quality of the text data in the target data source is relatively high, and then it is possible to continue to obtain all the data for the target source data until all the data in the entire target source data is collected and then switch to a new data source of other potential high-quality data sources, or meet the above preset stop conditions.
[0121] The method for constructing the training dataset provided by the present invention uses a large model to assist in domain relevance discrimination and scoring, performs partial sampling on each data source related to the target domain, and determines whether to continue to complete the full-scale collection of the data source through a small amount of data discrimination, so as to achieve rapid filtering of unsuitable data sources and reduce the data source acquisition cost.
[0122] Based on the content of the above embodiments, as an optional embodiment, for any data in the second dataset, the comprehensive classification score for obtaining each data in the second dataset as data within the target domain includes:
[0123] Input the any data into the text classification model, and obtain the classification probability that the any data is classified as data within the target domain output by the text classification model;
[0124] Construct a keyword set related to the target domain, and perform word segmentation on the any data to obtain a vocabulary set, so as to determine the matching score between the vocabulary set and the keyword set;
[0125] Determine the comprehensive classification score of any one of the data according to the weighted combination of the classification score determined by the classification probability and the matching score.
[0126] As an innovation point of the present invention, an adaptive domain evaluation function is proposed. . Since different fields have different requirements for data quality and it is difficult to establish a unified standard, the present invention selects benchmark source data (such as well-known data sets in the field) as a benchmark. By processing the benchmark data and training a classifier using a text classification model (such as the FastText model), the classification standard is made consistent with the benchmark to achieve unified evaluation.
[0127] It should be noted that compared with deep learning models with large parameters and slow processing (such as BERT), FastText has higher classification efficiency. To make up for its slightly insufficient language understanding ability, the present invention proposes to combine FastText with a keyword matching mechanism, which not only maintains high-efficiency classification but also improves the accuracy of screening, ensuring that the data is close to the benchmark scenario. The specific execution process is as follows:
[0128] Any one of the data in the second data set obtained after data preprocessing can be converted into a vector or an embedding representation and then input into the FastText model to obtain the classification probability that the data output by the model is classified as data within the target domain. , this classification probability is converted into a classification score within the range of 0 - 100 :
[0129] .
[0130] Optionally, before using the FastText model, it can be trained by means of supervised learning or unsupervised learning.
[0131] Furthermore, the specific implementation steps for performing a word segmentation operation on any one of the data to obtain a vocabulary set and determining the matching score between the vocabulary set and the keyword set are as follows:
[0132] Construct a keyword set K = { } related to the target domain, and perform a word segmentation operation on any one of the data to be judged to obtain the corresponding vocabulary set.
[0133] Then, by matching the words in the vocabulary set with the keywords in the keyword set, calculate the matching score between the vocabulary set and the keyword set according to the number or weight of the associated words matched:
[0134] ;
[0135] Among them, is the number of associated words appearing in the vocabulary set, is the total number of words in the vocabulary set.
[0136] Finally, the comprehensive classification score of any one of the data can be determined according to the weighted combination of the classification score determined by the classification probability and the matching score:
[0137] ;
[0138] where: is the comprehensive classification score (value range: 0 - 100), is the classification score determined by FastText (value range: 0 - 100), is the matching score of the keyword (range: 0 - 100%). and are the weight coefficients, satisfying , and can be dynamically adjusted according to different domain characteristics.
[0139] As an optional embodiment, screening out the target training data set from the second data set based on the comprehensive classification score of each data in the second data set includes:
[0140] Eliminating the data with a comprehensive classification score less than the third preset threshold from the second data set to obtain the target training data set.
[0141] Specifically, if , then retain the data, otherwise eliminate the data from the second data set. Finally, the required target training data set can be obtained. Among them, is the third preset threshold, and the value for different domains can be set by itself.
[0142] Based on the content of the above embodiment, as an optional embodiment, the text classification model is obtained by pre-training with multiple data samples in multiple vertical domains labeled with classification probability labels, and further includes:
[0143] Obtaining the model performance index of the text classification model after training each sample data, and the comprehensive classification score test value obtained based on each sample data;
[0144] Analyzing the model performance index and the comprehensive classification score test value of all the sample data by using a heat map and a confusion matrix;
[0145] Determining the weight combination corresponding to each vertical domain and the value of the third preset threshold according to the analysis result, where the weight combination is the weight distribution relationship between the classification probability and the matching score.
[0146] As another innovative point of the present invention, a dynamic optimization strategy based on feature importance is proposed, covering dynamic weight adjustment and methods such as automatically optimizing the first preset threshold, the second preset threshold, the third preset threshold, integrating active learning and data augmentation, etc. Through continuous iterative optimization, the optimization of the data of the vertical large model is achieved. The specific process is as follows:
[0147] (1)Model performance monitoring.
[0148] The present invention visualizes the model performance metrics (such as accuracy, recall, F1 score) monitored in real time through TensorBoard or MLflow, that is, during the process of training the text classification model, the key model performance metrics (such as accuracy, recall, F1 score, etc.) can be displayed in the graphical interface in real time, so that users can intuitively understand the performance of the model and make adjustments and optimizations accordingly.
[0149] In addition, during the process of model training, various abnormal situations may occur, such as overfitting. In order to detect this situation in a timely manner, an automatic alarm system can be set. When any model performance metric (such as the accuracy on the validation set) is lower than a certain threshold after a certain training round, the alarm will be automatically triggered to remind the user to intervene and adjust.
[0150] In addition, during the process of model training, it is also very important to regularly save the current state (i.e., checkpoint) of the text classification model. In this way, even if an abnormal interruption occurs during the training process (such as power failure, hardware failure, etc.), the training can be resumed from the nearest checkpoint, avoiding starting from scratch. At the same time, saving multiple checkpoints can also be used for subsequent performance analysis and model selection.
[0151] (2)Feedback collection and analysis.
[0152] The present invention also conducts error analysis based on the key model performance metrics and log data of the collected text classification model to identify common misclassifications and performance bottlenecks.
[0153] Use heatmaps and confusion matrices to show the screening effects of data in different fields, providing guidance for subsequent adjustments to ensure the pertinence and effectiveness of optimization.
[0154] Among them, heatmaps can be used to show the classification effects of data in different fields, and the differences in colors can intuitively reflect which fields of data are better classified by the model and which fields of data have greater classification difficulties. By observing the heatmap, it is possible to quickly identify which fields of data are crucial for improving the model performance, so as to pay more attention to them in future data collection and optimization.
[0155] In addition, the confusion matrix can display metrics such as accuracy, recall, and precision of the model for each category, thereby comprehensively evaluating the classification performance of the text classification model. Through the confusion matrix, it is possible to intuitively see which domain's data is prone to misclassification and the specific situation of the misclassification (such as misclassifying the data of a certain domain as another domain). Based on the results of the confusion matrix, it is possible to identify the performance bottlenecks of the text classification model in which domains, thus guiding subsequent optimization work. For example, if the recall rate of the data source within a certain domain is low, it is possible to consider increasing the training data of this data source or adjusting the above-mentioned preset threshold.
[0156] (3) Parameter adjustment and strategy optimization.
[0157] Based on the above feedback and analysis results, the present invention evaluates the sensitivity of the text classification model to different evaluation parameters (such as and ). Through feature importance analysis (such as SHAP values), determine the and of different domains.
[0158] Furthermore, grid search or Bayesian optimization can be used to adjust the threshold , balancing data quality and quantity.
[0159] It is also possible to introduce active learning to preferentially label data samples that improve model performance; apply data augmentation techniques to generate diverse training samples.
[0160] The method for constructing the training data set provided by the present invention can effectively improve the accuracy and efficiency of data screening, ensure the high quality of the data set, and promote the continuous improvement of the model performance through the above-mentioned measures.
[0161] Based on the content of the above embodiments, as an alternative embodiment, the preprocessing of the first data set includes:
[0162] Obtain each HTML-formatted data in the first data set;
[0163] Locate and strip out the body part in each of the HTML-formatted data;
[0164] Extract the text data in each of the body parts and remove the redundant formats and useless information from the text data.
[0165] As an innovation point of the present invention, an efficient data processing tool is combined with a large model assistance strategy to perform refined processing on the first data set obtained in the above embodiments, ensuring the structuring of the data and the purity of the content.
[0166] Due to the complex web page data structure in HTML format, which contains a large amount of irrelevant information such as advertisements, navigation bars, footnotes, etc., the present invention adopts the following methods for efficient processing:
[0167] (1) Adopt HTML parsing technology and natural language processing (NLP) algorithms to identify the structure and content of web pages, which can effectively analyze HTML tags to determine which parts contain the main content and which parts are irrelevant information such as advertisements, navigation bars, footnotes, etc.
[0168] (2) By identifying specific HTML tags (such as 、 、 <article>etc.) or a combination of tags, as well as analyzing the attributes of tags (such as class, id, etc.) to accurately locate the body of the web page.
[0169] (3) After identifying the main text, you can also use specialized text extraction tools or algorithms to separate the main text content from HTML tags and remove redundant formatting and useless information, such as removing irrelevant content such as HTML tags, CSS styles, JavaScript codes, and removing redundant spaces, line breaks, special characters, etc. in the text data.
[0170] (4) In order to ensure that the structure of the extracted text data is clear, techniques such as paragraph segmentation and sentence division can be used to organize the text data into a format that is easy to understand and process. At the same time, grammar checking and spelling correction can be performed on the text data to improve its quality and readability.
[0171] (5) The content of the text data can be further purified. During the process, a quality assessment algorithm or model can be used to assess the quality of the extracted text data. These algorithms or models can analyze the language characteristics, content quality and other aspects of the text data, thereby screening out high-quality text data and ensuring that the final output text data is accurate, clear and of high quality.
[0172] The training data set construction method provided by the present invention, when realizing HTML format data processing, can accurately identify the main content area of the web page, eliminate irrelevant information, and purify the extracted text by adopting advanced HTML parsing technology, NLP algorithm, machine learning model and other intelligent processing strategies, thereby ensuring that the final output text data is of high quality and easy to understand and process.
[0173] Based on the content of the above embodiment, as an optional embodiment, the preprocessing of the first data set further includes:
[0174] Obtain each PDF format data in the first data set;
[0175] Recognize text blocks in the PDF format data based on optical character recognition technology;
[0176] According to the positions of the text blocks in the PDF format data, all the text blocks are converted into single-column mode text data after the layout is restored;
[0177] Remove useless information from the single-column mode text data.
[0178] PDF files often contain complex tables, images or embedded text, and special methods are required to extract the text content. The invention's innovation in PDF data processing includes the following steps:
[0179] (1) An Optical Character Recognition (OCR) engine is adopted to efficiently recognize the text in the PDF and obtain the position information of each text block. Even if the PDF format data contains complex tables, images, or embedded text, etc., the text content therein can be accurately extracted.
[0180] (2) The recognized text blocks are subjected to layout restoration processing and converted into a single-column mode. This step simplifies the text structure, making subsequent processing and analysis more efficient and convenient.
[0181] Among them, the single-column mode refers to converting the text content that may originally contain multiple columns and complex layouts into a single, linear text stream. For example, a document in a PDF file originally contains a two-column or three-column layout, including various headings, paragraphs, pictures, and tables, etc. When this document is converted into the single-column mode, all the content will be rearranged into a continuous, single-column text stream.
[0182] (3) Additionally, for the irregular data commonly found in PDF format data, such as useless information like headers, footers, and blank lines, etc., structural processing and removal are performed to keep the logical structure of the single-column mode text data clear.
[0183] In addition, the present invention can also perform some appropriate optimization processing on other format data in the first data set, and specifically, scripts can be written according to the data source.
[0184] Figure 3 is one of the structural schematic diagrams of the training data set construction device provided by the present invention, as Figure 3 shown, mainly including but not limited to the following components:
[0185] The data source collection module 31 is mainly used to collect the first data set;
[0186] The data preprocessing module 32 is mainly used to preprocess the first data set to obtain a second data set, and the preprocessing includes converting non-text type data in the first data set into text type data;
[0187] The adaptive domain evaluation module 33 is mainly used to obtain the comprehensive classification scores of each data in the second data set being classified as data within the target domain, and the target domain is the domain to which the vertical large model to be trained belongs;
[0188] The data set screening module 34 is mainly used to screen out the target training data set from the second data set based on the comprehensive classification scores of each data in the second data set.
[0189] It should be noted that the construction device of the training data set provided by the present invention, when specifically running, can execute the construction method of the training data set described in any of the above embodiments, which will not be elaborated in this embodiment.
[0190] The construction device of the training data set provided by the present invention calculates the comprehensive classification score of each data by introducing an adaptive domain evaluation function, and can dynamically evaluate and screen data according to the requirements of each scenario and domain, so as to have obvious technical improvement effects in aspects such as broadening the data source, reducing the cleaning cost, unifying the quality standard, and improving the purity of data in the professional field.
[0191] Figure 4 is the second structural schematic diagram of the construction device of the training data set provided by the present invention, as Figure 4 shown, on the basis of the data source collection module, the data preprocessing module, and the adaptive domain evaluation and screening module, it further includes a data iterative optimization module. Among them, the adaptive domain evaluation and screening module can be regarded as obtained by combining the adaptive domain evaluation module and the data set screening module.
[0192] Among them, the functions performed by the data source collection module, the data preprocessing module, and the adaptive domain evaluation and screening module will not be elaborated here, and their corresponding functions recorded in the above embodiments are executed during actual operation.
[0193] The data iterative optimization module is mainly used to perform the following functions:
[0194] (1) Model performance monitoring, including obtaining the model performance indicators of the text classification model after training each sample data, and the comprehensive classification score test values obtained based on each sample data.
[0195] (2) Feedback collection and analysis, including analyzing the model performance indicators and comprehensive classification score test values of all the sample data by using a heat map and a confusion matrix;
[0196] (3) Parameter adjustment and strategy optimization, including determining the weight combination corresponding to each vertical domain and the value of the third preset threshold according to the analysis results, where the weight combination is the weight distribution relationship between the classification probability and the matching score.
[0197] Figure 5 is the structural schematic diagram of the electronic device provided by the present invention, as Figure 5 As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete their mutual communication through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute the method for constructing a training data set. The method includes: collecting a first data set; preprocessing the first data set to obtain a second data set, where the preprocessing includes converting non-text type data in the first data set into text type data; obtaining a comprehensive classification score for classifying each data in the second data set as data within a target field, where the target field is the field to which the to-be-trained vertical large model belongs; and screening out a target training data set from the second data set based on the comprehensive classification scores of the respective data in the second data set.
[0198] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0199] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method for constructing a training data set provided in the above-mentioned various embodiments. The method includes: collecting a first data set; preprocessing the first data set to obtain a second data set, where the preprocessing includes converting non-text type data in the first data set into text type data; obtaining a comprehensive classification score for classifying each data in the second data set as data within a target field, where the target field is the field to which the to-be-trained vertical large model belongs; and screening out a target training data set from the second data set based on the comprehensive classification scores of the respective data in the second data set.
[0200] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the method for constructing a training data set provided in the above embodiments. The method includes: collecting a first data set; preprocessing the first data set to obtain a second data set, where the preprocessing includes converting non-text type data in the first data set into text type data; obtaining a comprehensive classification score for classifying each data in the second data set as data within a target field, where the target field is the field to which the vertical large model to be trained belongs; and screening out a target training data set from the second data set based on the comprehensive classification scores of the data in the second data set.
[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0202] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / article>
Claims
1. A method for constructing a training data set, characterized in that: include: collecting a first data set; Preprocessing the first data set to obtain a second data set, wherein the preprocessing includes converting non-text data in the first data set into text data; Obtaining a comprehensive classification score of data in a target field for each data classification in the second data set, wherein the target field is the field to which the vertical large model to be trained belongs; Based on the comprehensive classification score of each data in the second data set, screening out a target training data set from the second data set; For any data in the second data set, obtaining a comprehensive classification score of each data classification in the second data set as data in the target field includes: Inputting any of the data into a text classification model, and obtaining a classification probability of the data output by the text classification model being classified as data within the target domain; Constructing a keyword set related to the target field, and performing a word segmentation operation on any of the data to obtain a vocabulary set, so as to determine a matching score between the vocabulary set and the keyword set; Determine the comprehensive classification score of any one of the data according to a weighted combination of the classification score determined by the classification probability and the matching score; The step of selecting a target training data set from the second data set based on the comprehensive classification score of each data in the second data set includes: Eliminate data whose comprehensive classification score is less than a third preset threshold from the second data set to obtain the target training data set; The text classification model is obtained by pre-training with data samples in multiple vertical fields annotated with classification probability labels, and also includes: Obtaining the model performance index of the text classification model after each sample data training, and the comprehensive classification score test value obtained based on each sample data training; Analyze the model performance indicators and comprehensive classification score test values of all the sample data using heat maps and confusion matrices; The weight combination corresponding to each of the vertical fields and the value of the third preset threshold are determined according to the analysis result, and the weight combination is the weight distribution relationship between the classification probability and the matching score.
2. The method for constructing a training data set according to claim 1, characterized in that: The collecting of the first data set includes: For a target data source in any vertical field related to the target field, the following discrimination operations are performed: extracting part of the data samples to construct a vertical field data set, and converting each of the data samples into a text data sample; obtaining a correlation score between each of the text data samples and the target field; based on a comparison result of the correlation score of each of the text data samples with a first preset threshold, determining whether to completely collect the remaining data in the target data source; If confirmed, all the collected data in the target data source are added to the initial data set, and the data source in the next vertical field related to the target field is set as the target data source, and the discrimination operation is re-executed; If uncertain, directly set the next vertical domain data source related to the target domain as the target data source, and re-execute the determination operation; The process is iteratively executed until a preset stop condition is met, and the obtained initial data set is used as the first data set.
3. The method for constructing a training data set according to claim 2, characterized in that: The obtaining of the correlation score between each of the text data samples and the target domain comprises: generating a domain-guided text associated with a relevance score of the target domain; Inputting each of the text data samples and the domain-guided text into a pre-trained relevance evaluation model, respectively, to obtain the domain discrimination result and the relevance score for each of the text data samples output by the relevance evaluation model; The domain discrimination result is used to indicate whether the text data sample belongs to the target domain; the relevance score is used to indicate the matching degree between the text data sample and the target domain.
4. The method for constructing a training data set according to claim 3, characterized in that: The determining whether to completely collect the remaining data in the target data source based on the comparison result of the relevance score of each of the text data samples with a first preset threshold value includes: Obtaining a first number of all text data samples whose domain discrimination results are failed; Obtain a second number of all text data samples whose relevance scores are lower than a second preset threshold; Calculate the proportion of unqualified data according to the first number, the second number, and the total amount of data in the vertical field data set; According to the comparison result of the unqualified data proportion and the first preset threshold, it is determined whether to completely collect the remaining data in the target data source.
5. The method for constructing a training data set according to claim 1, characterized in that: The preprocessing of the first data set includes: Obtain each HTML format data in the first data set; Locate and strip out the text part of each HTML format data; The text data in each of the body parts is extracted, and redundant formats and useless information in the text data are removed.
6. The method for constructing a training data set according to claim 1, characterized in that: The preprocessing of the first data set further includes: Obtain each PDF format data in the first data set; Recognize text blocks in the PDF format data based on optical character recognition technology; According to the positions of the text blocks in the PDF format data, all the text blocks are converted into single-column mode text data after the layout is restored; Remove useless information from the single-column mode text data.
7. A device for constructing a training data set, characterized in that: include: A data source collection module, used for collecting a first data set; A data preprocessing module, used for preprocessing the first data set to obtain a second data set, wherein the preprocessing includes converting non-text data in the first data set into text data; An adaptive domain evaluation module, used to obtain a comprehensive classification score of each data classification in the second data set as data in a target domain, wherein the target domain is the domain to which the vertical large model to be trained belongs; For any data in the second data set, obtaining a comprehensive classification score of each data classification in the second data set as data in the target field includes: Inputting any of the data into a text classification model, and obtaining a classification probability of the data output by the text classification model being classified as data within the target domain; Constructing a keyword set related to the target field, and performing a word segmentation operation on any of the data to obtain a vocabulary set, so as to determine a matching score between the vocabulary set and the keyword set; Determine the comprehensive classification score of any one of the data according to a weighted combination of the classification score determined by the classification probability and the matching score; A data set screening module, used for screening a target training data set from the second data set based on the comprehensive classification score of each data in the second data set, comprises: Eliminate data whose comprehensive classification score is less than a third preset threshold from the second data set to obtain the target training data set; The text classification model is obtained by pre-training with data samples in multiple vertical fields annotated with classification probability labels, and also includes: Obtaining the model performance index of the text classification model after each sample data training, and the comprehensive classification score test value obtained based on each sample data training; Analyze the model performance indicators and comprehensive classification score test values of all the sample data using heat maps and confusion matrices; The weight combination corresponding to each of the vertical fields and the value of the third preset threshold are determined according to the analysis result, and the weight combination is the weight distribution relationship between the classification probability and the matching score.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for constructing a training data set as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a training data set as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing a training data set as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Fine adjustment data screening method and device, computer equipment and readable storage medium
CN118504663A