Chinese pre-training sample quality evaluation method and device based on multiple classifiers

By evaluating the quality of Chinese pre-training samples through a combination of multiple classifiers, the problem of the scale of Chinese corpus resources and the lag in evaluation methods is solved, achieving efficient and accurate corpus selection and supporting the training of high-performance Chinese language models.

CN121093178APending Publication Date: 2025-12-09BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511035857.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Limited Chinese corpus resources and outdated quality assessment methods have restricted the development of high-performance Chinese language models. Existing single-model discrimination mechanisms suffer from low recall rate, high false positive rate, and weak generalization ability for high-quality samples.

Method used

Multiple heterogeneous classifiers are used to evaluate the quality of Chinese pre-trained samples. By combining models such as BERT and FastText, multiple quality evaluation results are integrated to improve the stability and reliability of the evaluation and avoid bias from a single model.

Benefits of technology

It significantly improves the ability to identify high-quality samples, supports the training needs of large-scale language models, provides a new technical path for the quality assessment of Chinese pre-trained data, and enhances the stability and reliability of the overall quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093178A_ABST
    Figure CN121093178A_ABST
Patent Text Reader

Abstract

The invention provides a Chinese pre-training sample quality evaluation method and device based on multiple classifiers. The method comprises the following steps: determining a Chinese pre-training sample to be subjected to quality evaluation; obtaining a plurality of quality evaluation results of the Chinese pre-training sample based on a plurality of pre-trained classifiers; wherein the plurality of classifiers are heterogeneous classifiers; and fusing the plurality of quality evaluation results to obtain a target quality evaluation result of the Chinese pre-training sample. According to the method, the quality of the Chinese pre-training sample is evaluated through the plurality of pre-trained heterogeneous classifiers, so that the prejudice of a single model can be avoided, the stability and the reliability of the overall quality evaluation are improved, the training requirements of a large-scale language model are effectively supported, and a new technical path is provided for the quality evaluation of the Chinese pre-training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for evaluating the quality of Chinese pre-trained samples based on a multi-classifier. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence technology, large language models (LLMs) have achieved significant breakthroughs in the field of natural language processing. The performance improvement of LLMs largely depends on the construction and continuous optimization of large-scale, high-quality pre-training corpora. These corpora not only provide the models with a broad knowledge base, but also significantly enhance their reasoning ability in complex tasks, enabling them to handle diverse language tasks, including creative writing, logical reasoning, and question answering.

[0003] In the process of promoting the development of LLM technology, open-source corpus datasets (such as The Pile and Common Crawl) have played an important role. They have not only provided rich training resources for academic research, but also established a standardized benchmark system for horizontal comparison and evaluation of model performance, thereby promoting continuous innovation in model training methods and architecture design.

[0004] As research deepens, the academic community's focus on pre-training corpora has gradually shifted from simply pursuing data scale to optimizing and improving data quality. Currently, the amount of data required for mainstream LLM training has reached trillions of labeled data, posing higher technical requirements for the acquisition, cleaning, and management of corpus resources.

[0005] In terms of English corpus resources, the scale of open-source corpora has continued to expand, from early medium-sized corpora such as The Pile (approximately 825GB) to ultra-large-scale corpora such as FineWeb (15TB) built on the Common Crawl platform. Simultaneously, data processing methods have evolved from traditional rule-based filtering strategies to more advanced model-driven approaches. For example, the FineWeb-Edu dataset, by introducing a language model-assisted quality assessment mechanism, significantly improved the usability and training efficiency of the corpus, becoming a representative achievement in the current field of English corpus processing.

[0006] However, the development of Chinese corpus resources still faces many challenges. First, due to the relative scarcity of Chinese internet data sources, existing open-source Chinese corpora (such as WuDao, SkyPile150B, and WanjuanV1) have significant bottlenecks in terms of data scale expansion, making it difficult to meet the rapidly increasing demands of current LLM training. Second, current research on the quality assessment of Chinese corpora is insufficient, resulting in inconsistent overall corpus quality, which is insufficient to support the training needs of high-performance Chinese language models.

[0007] Furthermore, although some high-quality Chinese corpora (such as CCI3-HQ and Chinese-FineWeb) have attempted to use classifiers for corpus quality screening, they generally employ a single-model discrimination mechanism, which suffers from low recall of high-quality samples, high false positive rate, and weak generalization ability, thus limiting their application effectiveness in large-scale corpus screening. These technical bottlenecks severely restrict the development of high-performance Chinese language models and also affect the performance of Chinese LLMs in practical applications.

[0008] In summary, current Chinese corpus resources suffer from significant problems in terms of both limited data scale and outdated quality assessment methods. There is an urgent need to propose a more efficient, accurate, and scalable corpus quality assessment and screening mechanism to support the high-quality development of large-scale Chinese language models and meet the growing demand for multi-task language understanding and generation. Summary of the Invention

[0009] This invention provides a method and apparatus for evaluating the quality of Chinese pre-trained samples based on a multi-classifier, which overcomes the shortcomings of existing Chinese pre-trained data quality evaluation methods that use a single classifier for quality discrimination, resulting in low recall of high-quality samples. It significantly improves the recognition ability of high-quality samples, effectively supports the training needs of large-scale language models, and provides a new technical path for the quality evaluation of Chinese pre-trained data.

[0010] On one hand, the present invention provides a method for quality assessment of Chinese pre-trained samples based on multiple classifiers, comprising: determining Chinese pre-trained samples to be quality assessed; obtaining multiple quality assessment results of the Chinese pre-trained samples based on multiple pre-trained classifiers; wherein the multiple classifiers are heterogeneous classifiers; and fusing the multiple quality assessment results to obtain a target quality assessment result of the Chinese pre-trained samples.

[0011] Furthermore, the plurality of classifiers includes a first classifier, a second classifier, and a third classifier; wherein, the first classifier is constructed based on a BERT pre-trained model and fine-tuned using a first quality-labeled sample set; the second classifier is constructed based on a BERT pre-trained model and fine-tuned using a second quality-labeled sample set; the third classifier includes a FastText classifier, which is trained using pre-selected seed positive samples and random negative samples; wherein, the first quality-labeled sample set and the second quality-labeled sample set are obtained by labeling different original sample sets using different preset models.

[0012] Further, the multiple quality assessment results include a first quality assessment result, a second quality assessment result, and a third quality assessment result; correspondingly, obtaining multiple quality assessment results of the Chinese pre-trained samples based on multiple pre-trained classifiers includes: obtaining a first quality assessment result based on the first classifier and the Chinese pre-trained samples; obtaining a second quality assessment result based on the second classifier and the Chinese pre-trained samples; and obtaining a third quality assessment result based on the third classifier and the Chinese pre-trained samples; wherein the first quality assessment result and the second quality assessment result are both multi-class classifications, and the third quality assessment result is a binary classification.

[0013] Furthermore, the step of fusing the multiple quality assessment results to obtain the target quality assessment result of the Chinese pre-training sample includes: determining the optimal quality assessment result among the multiple quality assessment results; and determining the optimal quality assessment result as the target quality assessment result of the Chinese pre-training sample.

[0014] Furthermore, the step of fusing the multiple quality assessment results to obtain the target quality assessment result of the Chinese pre-training sample includes: summing the multiple quality assessment results to obtain a comprehensive quality assessment result; and determining the comprehensive quality assessment result as the target quality assessment result of the Chinese pre-training sample.

[0015] Further, the process of fusing the multiple quality assessment results to obtain the target quality assessment result of the Chinese pre-training samples includes: classifying all Chinese pre-training samples into levels according to the target quality assessment result to obtain multiple quality levels of Chinese pre-training sample sets; wherein, the Chinese pre-training sample sets are used for prediction of the target model in downstream tasks.

[0016] Secondly, the present invention also provides a Chinese pre-training sample quality assessment device based on multiple classifiers, comprising: a Chinese pre-training sample determination module for determining Chinese pre-training samples to be quality assessed; a multiple quality assessment result acquisition module for acquiring multiple quality assessment results of the Chinese pre-training samples based on multiple pre-trained classifiers; wherein the multiple classifiers are heterogeneous classifiers; and a multiple quality assessment result fusion module for fusing the multiple quality assessment results to obtain a target quality assessment result of the Chinese pre-training samples.

[0017] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the Chinese pre-trained sample quality assessment method based on any of the above-described multi-classifier methods.

[0018] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the Chinese pre-trained sample quality assessment method based on a multi-classifier as described above.

[0019] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the Chinese pre-trained sample quality assessment method based on a multi-classifier as described above.

[0020] This invention provides a method for quality assessment of Chinese pre-trained samples based on multiple classifiers. The method identifies the Chinese pre-trained samples to be assessed and obtains multiple quality assessment results based on pre-trained classifiers. These classifiers are heterogeneous, and the multiple quality assessment results are then fused to obtain the target quality assessment result for the Chinese pre-trained samples. This method uses multiple pre-trained heterogeneous classifiers to assess the quality of Chinese pre-trained samples, avoiding the bias of a single model, improving the stability and reliability of the overall quality assessment, effectively supporting the training needs of large-scale language models, and providing a new technical path for quality assessment of Chinese pre-trained data. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1This is a flowchart illustrating the Chinese pre-trained sample quality assessment method based on a multi-classifier provided in this embodiment of the invention.

[0023] Figure 2 This is a schematic diagram of the structural parameters of the target model provided in the embodiment of the present invention.

[0024] Figure 3 This is a schematic diagram illustrating the correlation between the performance of the target model and its quality level, as provided in this embodiment of the invention.

[0025] Figure 4 This is a schematic diagram comparing the effectiveness of the Chinese pre-trained sample set provided in this embodiment of the invention with other existing training sample sets.

[0026] Figure 5 This is a schematic diagram of the structure of the Chinese pre-trained sample quality assessment device based on a multi-classifier provided in an embodiment of the present invention.

[0027] Figure 6 This is a schematic diagram of the physical structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0029] To address the current technical bottlenecks in Chinese pre-training data quality assessment, this invention proposes a method for assessing the quality of Chinese pre-training samples based on a multi-classifier. Specifically, Figure 1 The diagram illustrates a flowchart of the Chinese pre-trained sample quality assessment method based on a multi-classifier provided in an embodiment of the present invention.

[0030] like Figure 1 As shown, the method includes: S110, determining the Chinese pre-training samples to be quality evaluated; S120, obtaining multiple quality evaluation results of the Chinese pre-training samples based on multiple pre-trained classifiers; wherein, the multiple classifiers are heterogeneous classifiers; S130, fusing the multiple quality evaluation results to obtain the target quality evaluation result of the Chinese pre-training samples.

[0031] The following will provide a detailed description of steps S110-S130 and related steps.

[0032] S110, Identify the Chinese pre-training samples to be evaluated for quality.

[0033] It is easy to understand that high-quality pre-training corpora are a key factor in improving model performance. Especially in Chinese language modeling, due to the relative scarcity of Chinese Internet data sources and the uneven quality of corpora, how to select high-quality Chinese pre-training samples from massive amounts of data is an important prerequisite for building high-performance models.

[0034] In the sample quality assessment process, the first step is to identify the objects to be assessed, that is, to determine which Chinese pre-training samples from the original corpus will enter the subsequent quality assessment process.

[0035] Specifically, the first step is to obtain raw Chinese text data from open-source or internal corpora, including but not limited to web crawling content, encyclopedic texts, news website content, academic papers, technical documents, book corpora, social media texts, and Q&A platform content.

[0036] Then, in order to evaluate efficiency and quality, the original Chinese text data needs to be preliminarily cleaned and deduplicated. The specific sample set to be evaluated for quality is determined from the original corpus, that is, the sample set to be evaluated for quality, which includes a large number of Chinese pre-trained samples.

[0037] Randomly or selectively select Chinese pre-trained samples from the sample set to be evaluated for subsequent quality assessment processes, and further, execute step S120.

[0038] It should be noted that the number of Chinese pre-training samples determined in step S110 can be one or more, and no specific limitation is made here.

[0039] S120, Based on multiple pre-trained classifiers, obtain multiple quality assessment results of the Chinese pre-trained samples; wherein, the multiple classifiers are heterogeneous classifiers.

[0040] Traditional quality assessment methods often rely on a single model or rule for judgment, which has problems such as low recall, high false positive rate and weak generalization ability, making it difficult to meet the needs of large-scale corpus screening and refined quality control.

[0041] In view of this, this embodiment introduces multiple heterogeneous classifiers to collaboratively evaluate the quality of Chinese pre-trained samples. Specifically, the same Chinese pre-trained sample is input into multiple pre-trained classifiers, and multiple quality evaluation results can be obtained.

[0042] The multiple classifiers in this embodiment are pre-trained heterogeneous classifiers. This means that, on the one hand, these multiple classifiers differ in terms of structure, training data, training methods, and task objectives; on the other hand, these multiple classifiers have been trained using labeled high-quality / low-quality corpus data and have undergone validation and optimization, thus possessing good generalization ability.

[0043] The number of classifiers must be at least two, or more, depending on actual needs; no specific limit is set here.

[0044] Multiple quality assessment results can be specific scores or binary classification results, and no specific limitations are made here.

[0045] It is worth mentioning that this embodiment uses multiple pre-trained heterogeneous classifiers to evaluate the quality of Chinese pre-trained samples, which can avoid the bias of a single model and improve the stability and reliability of the overall quality evaluation; even if individual classifiers make misjudgments, the overall evaluation results still have high credibility.

[0046] After obtaining multiple quality assessment structures of Chinese pre-trained samples based on multiple pre-trained classifiers in step S120, step S130 is further executed.

[0047] S130, the multiple quality assessment results are fused to obtain the target quality assessment result of the Chinese pre-training sample.

[0048] When multiple quality assessment results have the same format, they can be merged using a maximum value strategy or a summation strategy. When multiple quality assessment results have different formats, they can first be transformed into a unified scoring range, and then merged using a maximum value strategy or a summation strategy. This yields the final quality assessment result for the Chinese pre-trained samples, which is also the target quality assessment result.

[0049] In one embodiment, different weights can be set for the quality evaluation performance of multiple classifiers, thereby determining the target quality evaluation result of the Chinese pre-trained samples by weighted summation during fusion.

[0050] In this embodiment, by identifying the Chinese pre-training samples to be quality evaluated, and based on multiple pre-trained classifiers, multiple quality evaluation results for the Chinese pre-training samples are obtained; wherein the multiple classifiers are heterogeneous classifiers, and then the multiple quality evaluation results are fused to obtain the target quality evaluation result for the Chinese pre-training samples. This method uses multiple pre-trained heterogeneous classifiers to evaluate the quality of Chinese pre-training samples, which can avoid the bias of a single model, improve the stability and reliability of the overall quality evaluation, effectively support the training needs of large-scale language models, and provide a new technical path for the quality evaluation of Chinese pre-training data.

[0051] Based on the above embodiments, the structure, training process and inference process of multiple classifiers will be described in detail below.

[0052] Taking three classifiers as an example, the multiple classifiers in this embodiment include a first classifier, a second classifier, and a third classifier. The first and second classifiers are both built based on BERT pre-trained models. The first classifier is fine-tuned using a first-quality labeled sample set, and the second classifier is fine-tuned using a second-quality labeled sample set. The first and second quality labeled sample sets are obtained by labeling different original sample sets using different preset models. The third classifier includes a FastText classifier, which is trained using pre-selected seed positive samples and random negative samples.

[0053] Specifically, the Qwen2.5-72B-Instruct tool can be used to independently label the first set of 460,000 original samples, resulting in a first quality-labeled sample set. Simultaneously, DeepSeek-V3 can be used to independently label the second set of 460,000 original samples, resulting in a second quality-labeled sample set. Furthermore, through rigorous screening, a total of 200,000 high-quality Chinese seed samples are obtained as the positive sample benchmark for the classifier, i.e., seed positive samples. Random negative samples are then randomly sampled from the original sample set.

[0054] The rigorous screening process for obtaining high-quality Chinese seed samples includes: 1) Professionally translating existing high-quality English seed datasets to ensure accuracy and fluency, thus obtaining a reliable foundation for Chinese data. This step directly utilizes validated high-quality English data resources; 2) Rigorously screening existing Chinese instruction datasets: By setting clear screening criteria, samples with lower educational levels are primarily removed, including but not limited to meaningless casual conversations, some creatively generated samples of varying quality, and low-quality cases in some stylized generation tasks. This screening process aims to improve the diversity of high-quality Chinese seed samples.

[0055] It's worth noting that choosing the best-performing prompts is crucial when constructing training samples for the three classifiers. The language and scoring rules of the prompts directly impact model performance. Specifically, the language can be Chinese or English, and the scoring rules can be summative (directly providing the final score) or cumulative (judging each rule and accumulating the scores). Since there are four possible combinations of prompts (Chinese + summative, Chinese + cumulative, English + summative, English + cumulative) when requesting large models, it's necessary to evaluate which combination is optimal. To accurately assess the performance of these four methods, this embodiment constructs a Ground Truth sample quality dataset based on manual annotations to compare and evaluate the effects of different prompt combinations. This process helps determine the optimal prompt design strategy, thereby improving classifier performance.

[0056] After comparison and evaluation, the prompts for the first and second classifiers use Chinese quality scores, while the third classifier uses binary classification. Based on the prompts, the output of the first classifier will be a Chinese quality score value, such as any score between 1 and 5, and the output of the second classifier will be a binary classification of Chinese, such as 0 or 1, or positive or negative.

[0057] After determining the training samples for each classifier, fine-tuning training of the first, second, and third classifiers begins. Specifically, this embodiment extends the open-source BGE-M3 vector model by adding a classification head for regression output, freezing the embedding and encoding layers, training only the classification head, and finally converting the BGE-M3 vector model into a binary classifier, namely the first and second classifiers.

[0058] The first and second classifiers were trained for 20 rounds and optimized with a learning rate of 3e-4. During training, the first classifier took samples from the first original sample set as input, the estimated quality score as output, and the difference between the estimated quality score and the true label as training loss; the second classifier took samples from the second original sample set as input, the estimated quality score as output, and the difference between the estimated quality score and the true label as training loss.

[0059] For the FastText model (third classifier), during training, seed positive samples or random negative samples are used as inputs, and the predicted binary classification (positive or negative, or 0 or 1) is used as output. The training loss is the difference between the predicted binary classification and the input true label. Multiple word segmenters, the number of word combinations, and the learning rate are optimized, and finally the parameter configuration with the best effect is selected.

[0060] The trained first, second, and third classifiers can be directly used for quality assessment of Chinese pre-training samples. Specifically, based on the first classifier and the Chinese pre-training samples, a first quality assessment result is obtained; based on the second classifier and the Chinese pre-training samples, a second quality assessment result is obtained; and based on the third classifier and the Chinese pre-training samples, a third quality assessment result is obtained.

[0061] The first and second quality assessment results are quality scores ranging from 1 to 5, while the third quality assessment result is 0 or 1.

[0062] In this embodiment, multiple quality assessment results for Chinese pre-trained samples are obtained based on multiple pre-trained classifiers. These classifiers are heterogeneous, and the multiple quality assessment results are then fused to obtain the target quality assessment result for the Chinese pre-trained samples. This method uses multiple pre-trained heterogeneous classifiers to assess the quality of Chinese pre-trained samples, avoiding the bias of a single model, improving the stability and reliability of the overall quality assessment, effectively supporting the training needs of large-scale language models, and providing a new technical path for the quality assessment of Chinese pre-trained data.

[0063] Based on the above embodiments, the following will further describe in detail the process of fusing multiple quality assessment results.

[0064] In one specific embodiment, multiple quality assessment results are fused to obtain the target quality assessment result of the Chinese pre-training samples, including: determining the optimal quality assessment result among the multiple quality assessment results; and determining the optimal quality assessment result as the target quality assessment result of the Chinese pre-training samples.

[0065] Specifically, the multiple quality assessment results output by multiple classifiers are all quality scores with specific numerical values. The highest quality score is the optimal quality assessment result. Therefore, the highest quality score is also the target quality assessment result of the current Chinese pre-training sample.

[0066] For example, for the current Chinese pre-training sample, the quality assessment result output by the first classifier is 3, the quality assessment result output by the second classifier is 5, and the quality assessment result output by the third classifier is 1. Therefore, the target quality assessment result of the current Chinese pre-training sample is 5.

[0067] In another specific embodiment, multiple quality assessment results are fused to obtain the target quality assessment result of the Chinese pre-training samples, including: summing the multiple quality assessment results to obtain a comprehensive quality assessment result; and determining the comprehensive quality assessment result as the target quality assessment result of the Chinese pre-training samples.

[0068] Specifically, the multiple quality assessment results output by multiple classifiers are all quality scores with specific numerical values. By directly summing these multiple quality scores, a comprehensive quality score can be obtained, which is the target quality assessment result of the current Chinese pre-training sample.

[0069] For example, for the current Chinese pre-training sample, the quality assessment result output by the first classifier is 3, the quality assessment result output by the second classifier is 5, and the quality assessment result output by the third classifier is 1. Then, the target quality assessment result of the current Chinese pre-training sample is 9.

[0070] In another specific embodiment, multiple quality assessment results are fused to obtain the target quality assessment result of the Chinese pre-training samples, including: determining the weight values ​​corresponding to the multiple quality assessment results; performing a weighted summation of the multiple quality assessment results and their corresponding weight values ​​to obtain the fused quality assessment result; and using the fused quality assessment result as the target quality assessment result of the Chinese pre-training samples.

[0071] Specifically, weight values ​​can be set according to the performance parameters (such as post-training optimization parameters) of multiple heterogeneous classifiers; the better the performance, the higher the weight value should be. During fusion, the weighted sum of multiple quality assessment results, each with a specific numerical quality score, can be obtained to get the fused quality assessment result, which is the target quality assessment result for the current Chinese pre-trained samples.

[0072] For example, for the current Chinese pre-training sample, the first classifier outputs a quality assessment result of 3, with a corresponding weight of 0.5; the second classifier outputs a quality assessment result of 5, with a corresponding weight of 0.3; and the third classifier outputs a quality assessment result of 1, with a corresponding weight of 0.2. Therefore, the target quality assessment result for the current Chinese pre-training sample is 3.2.

[0073] In this embodiment, multiple quality assessment results are fused using different fusion strategies to obtain the target quality assessment result for the Chinese pre-training samples. This method uses multiple pre-trained heterogeneous classifiers to assess the quality of Chinese pre-training samples, avoiding the bias of a single model, improving the stability and reliability of the overall quality assessment, effectively supporting the training needs of large-scale language models, and providing a new technical path for the quality assessment of Chinese pre-training data.

[0074] Based on the above embodiments, the post-processing (quality grade classification process) of all Chinese pre-trained samples will be described in detail below.

[0075] By integrating multiple quality assessment results, a target quality assessment result for the Chinese pre-training samples is obtained. This includes classifying all Chinese pre-training samples into different quality levels according to the target quality assessment results, resulting in multiple quality-level sets of Chinese pre-training samples. These Chinese pre-training sample sets are then used for prediction of the target model in downstream tasks.

[0076] Specifically, all Chinese pre-training samples are sorted according to the target quality assessment results. Using an equal-frequency partitioning method, the Chinese pre-training samples are dynamically mapped to 20 integer levels (0-19), with each level strictly maintaining a 5% sample proportion. Level 19 is specifically used to identify the top 5% of high-quality Chinese pre-training samples. This evenly divides all Chinese pre-training samples into 20 quality levels, resulting in a set of Chinese pre-training samples with 20 quality levels.

[0077] To evaluate the effectiveness of the quality level classification, this embodiment adopts a controlled variable experimental design: Chinese pre-training samples are randomly selected from different quality levels to construct training sets, and target models with a scale of 500 million parameters are trained from scratch. Figure 2 A schematic diagram of the structural parameters of the target model provided in an embodiment of the present invention is shown.

[0078] Experimental results show that the performance of the target model in downstream tasks is significantly positively correlated with the quality level of the training set, verifying the effectiveness of the quality assessment in the embodiments of the present invention. Figure 3 This diagram illustrates the correlation between the performance of the target model and its quality level, as provided in an embodiment of the present invention.

[0079] exist Figure 3 In the diagram, the horizontal axis represents different quality levels (quality score range / quality bucket), and the vertical axis represents the performance of the target model (average Chinese score / average Chinese rating). According to... Figure 3 It can be seen that in the lower quality score range (e.g., 0 to 6), the average Chinese score is relatively low, roughly between 28 and 29. As the quality score range increases, the average Chinese score gradually rises, showing a slight decrease in the 7th range before continuing to rise. In the higher quality score range (e.g., 14 to 19), the average Chinese score improves significantly, reaching over 34, and shows a clear upward trend.

[0080] In some other embodiments, in order to further evaluate the effectiveness of the Chinese pre-training sample quality assessment method based on multi-classifier provided by the embodiments of the present invention and its beneficial effect on training LLM, this embodiment also compares the training effect of the Chinese pre-training sample set constructed by the Chinese pre-training sample quality assessment method based on multi-classifier provided by the embodiments of the present invention with the datasets obtained by other quality assessment methods.

[0081] Figure 4 This diagram illustrates a comparison of the effectiveness of the Chinese pre-trained sample set provided in this embodiment of the invention with other existing training sample sets.

[0082] exist Figure 4 In the figure, the horizontal axis represents the number of training rounds, and the vertical axis represents the average score of the model in Chinese metrics on the dataset. The CCI4.0-Zh-HQ model is the target model trained using the Chinese pre-training sample set provided in this embodiment of the invention.

[0083] according to Figure 4 It can be seen that in the early training phase (0 to 20 rounds), the score of the CCI4.0-Zh-HQ model rises rapidly, demonstrating a fast learning speed. As the number of training rounds increases, its score gradually stabilizes, reaching a relatively high level (approximately 33 points) after 50 rounds, and maintains a steady improvement in subsequent training.

[0084] according to Figure 4 It can also be seen that the Chinese pre-training sample set provided by the embodiments of the present invention is significantly better than other training sample sets constructed by a single quality classifier, such as CCI3.0-HQ and Chinese-FineWeb, in terms of Chinese metrics. This directly proves the effectiveness of the Chinese pre-training sample quality assessment method based on multi-classifiers provided by the embodiments of the present invention in the pre-training scenario.

[0085] Corresponding to the Chinese pre-trained sample quality assessment method based on multi-classifiers described in the above embodiments, the present invention also provides a Chinese pre-trained sample quality assessment device based on multi-classifiers.

[0086] Specifically, Figure 5 A schematic diagram of the structure of the Chinese pre-trained sample quality assessment device based on a multi-classifier provided in an embodiment of the present invention is shown.

[0087] like Figure 5 As shown, the device includes: a Chinese pre-training sample determination module 510, used to determine Chinese pre-training samples to be quality evaluated; a multiple quality evaluation result acquisition module 520, used to acquire multiple quality evaluation results of the Chinese pre-training samples based on multiple pre-trained classifiers; wherein the multiple classifiers are heterogeneous classifiers; and a multiple quality evaluation result fusion module 530, used to fuse the multiple quality evaluation results to obtain the target quality evaluation result of the Chinese pre-training samples.

[0088] In this embodiment, the Chinese pre-training sample determination module 510 determines the Chinese pre-training samples to be quality evaluated. The multiple quality evaluation result acquisition module 520 acquires multiple quality evaluation results of the Chinese pre-training samples based on multiple pre-trained classifiers. These multiple classifiers are heterogeneous classifiers. Then, the multiple quality evaluation result fusion module 530 fuses the multiple quality evaluation results to obtain the target quality evaluation result of the Chinese pre-training samples. This device uses multiple pre-trained heterogeneous classifiers to perform quality evaluation on the Chinese pre-training samples, which can avoid the bias of a single model, improve the stability and reliability of the overall quality evaluation, effectively support the training needs of large-scale language models, and provide a new technical path for the quality evaluation of Chinese pre-training data.

[0089] It should be noted that the Chinese pre-trained sample quality assessment device based on multi-classifier provided in this embodiment of the invention can be referred to in correspondence with the Chinese pre-trained sample quality assessment method based on multi-classifier described in the above embodiments, and will not be repeated here.

[0090] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a Chinese pre-trained sample quality assessment method based on a multi-classifier. This method includes: determining the Chinese pre-trained samples to be quality assessed; obtaining multiple quality assessment results of the Chinese pre-trained samples based on multiple pre-trained classifiers; wherein the multiple classifiers are heterogeneous classifiers; and fusing the multiple quality assessment results to obtain a target quality assessment result for the Chinese pre-trained samples.

[0091] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0092] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the Chinese pre-trained sample quality assessment method based on multiple classifiers provided by the above methods. The method includes: determining the Chinese pre-trained sample to be quality assessed; obtaining multiple quality assessment results of the Chinese pre-trained sample based on multiple pre-trained classifiers; wherein the multiple classifiers are heterogeneous classifiers; and fusing the multiple quality assessment results to obtain the target quality assessment result of the Chinese pre-trained sample.

[0093] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the Chinese pre-trained sample quality assessment method based on a multi-classifier provided by the above methods. The method includes: determining the Chinese pre-trained sample to be quality assessed; obtaining multiple quality assessment results of the Chinese pre-trained sample based on multiple pre-trained classifiers; wherein the multiple classifiers are heterogeneous classifiers; and fusing the multiple quality assessment results to obtain a target quality assessment result of the Chinese pre-trained sample.

[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for evaluating the quality of Chinese pre-trained samples based on a multi-classifier, characterized in that, include: Identify the Chinese pre-training samples to be evaluated for quality; Based on multiple pre-trained classifiers, multiple quality assessment results of the Chinese pre-trained samples are obtained; wherein, the multiple classifiers are heterogeneous classifiers; By integrating the multiple quality assessment results, the target quality assessment result of the Chinese pre-trained samples is obtained.

2. The method for evaluating the quality of Chinese pre-trained samples based on a multi-classifier according to claim 1, characterized in that, The multiple classifiers include a first classifier, a second classifier, and a third classifier; The first classifier is built based on the BERT pre-trained model and is obtained by fine-tuning training using the first quality labeled sample set. The second classifier is built based on the BERT pre-trained model and is fine-tuned using a second quality labeled sample set; The third classifier includes the FastText classifier, which is trained using pre-selected seed positive samples and random negative samples; The first quality-labeled sample set and the second quality-labeled sample set are obtained by labeling different original sample sets using different preset models.

3. The method for evaluating the quality of Chinese pre-trained samples based on a multi-classifier according to claim 2, characterized in that, The multiple quality assessment results include a first quality assessment result, a second quality assessment result, and a third quality assessment result; Accordingly, the method of obtaining multiple quality assessment results for the Chinese pre-trained samples based on multiple pre-trained classifiers includes: Based on the first classifier, and according to the Chinese pre-trained samples, a first quality assessment result is obtained; Based on the second classifier, and according to the Chinese pre-trained samples, a second quality assessment result is obtained; Based on the third classifier, and according to the Chinese pre-trained samples, a third quality assessment result is obtained; The first quality assessment result and the second quality assessment result are both multi-classification, and the third quality assessment result is binary classification.

4. The method for evaluating the quality of Chinese pre-trained samples based on a multi-classifier according to claim 1, characterized in that, The fusion of the multiple quality assessment results yields the target quality assessment result for the Chinese pre-trained samples, including: Determine the optimal quality assessment result among the multiple quality assessment results; The optimal quality assessment result is determined as the target quality assessment result for the Chinese pre-training samples.

5. The method for evaluating the quality of Chinese pre-trained samples based on a multi-classifier according to claim 1, characterized in that, The fusion of the multiple quality assessment results yields the target quality assessment result for the Chinese pre-trained samples, including: The summation of the multiple quality assessment results yields a comprehensive quality assessment result. The comprehensive quality assessment result is determined as the target quality assessment result of the Chinese pre-training samples.

6. The method for evaluating the quality of Chinese pre-trained samples based on a multi-classifier according to any one of claims 1-5, characterized in that, The process of fusing the multiple quality assessment results to obtain the target quality assessment result for the Chinese pre-trained samples then includes: All Chinese pre-training samples are classified into different levels according to the target quality assessment results, resulting in Chinese pre-training sample sets with multiple quality levels. The Chinese pre-trained sample set is used for prediction of the target model in downstream tasks.

7. A device for evaluating the quality of Chinese pre-trained samples based on a multi-classifier, characterized in that, include: The Chinese pre-training sample determination module is used to determine the Chinese pre-training samples to be evaluated for quality. A multi-quality assessment result acquisition module is used to acquire multiple quality assessment results of the Chinese pre-trained samples based on multiple pre-trained classifiers; wherein, the multiple classifiers are heterogeneous classifiers; The multi-quality assessment result fusion module is used to fuse the multiple quality assessment results to obtain the target quality assessment result of the Chinese pre-training sample.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the Chinese pre-trained sample quality assessment method based on a multi-classifier as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the Chinese pre-trained sample quality assessment method based on a multi-classifier as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the Chinese pre-trained sample quality assessment method based on a multi-classifier as described in any one of claims 1 to 6.