Training sample generation method and device of large language model, electronic equipment and storage medium
By using a multi-model collaborative evaluation orchestrator and multi-source data cleaning, high-quality training samples are generated, which solves the problem of insufficient training data for large language models and improves the model training effect and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2025-11-14
- Publication Date
- 2026-04-17
AI Technical Summary
In large language model training, existing technologies suffer from insufficient training data quality, resulting in poor model training performance and security risks. Existing methods are unable to effectively screen and generate high-quality training samples.
A multi-model collaborative evaluation orchestrator is used to evaluate candidate corpora from multiple dimensions. Training samples are generated through unified evaluation instructions and explicit scoring identifiers. Combined with multi-source data cleaning and corpus processing, data quality and security are ensured.
It improves the quality and security of training data, enhances the training effect of large language models, reduces the risk of model output errors and biases, and ensures the credibility and diversity of training samples.
Smart Images

Figure CN121882142A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to a method, apparatus, electronic device, and storage medium for generating training samples for a large language model. Background Technology
[0002] It should be noted that the above description of the technical background is only for the purpose of providing a clear and complete explanation of the technical solutions of the present invention and facilitating understanding by those skilled in the art. It should not be assumed that the above technical solutions are known to those skilled in the art simply because they have been described in the background section of this invention.
[0003] Large Language Models (LLMs) are large neural network models capable of learning semantic representations and knowledge structures to achieve language understanding, generation, and various language-related tasks. With the development of artificial intelligence technology, LLMs are widely used in a variety of tasks, including question-answering systems, machine translation, text generation, dialogue interaction, code writing, and multimodal understanding, playing an increasingly important role in academic research, industrial applications, and social life.
[0004] The training performance of large language models largely depends on the quality of the training data. Therefore, improving the quality of training data, and consequently, the training performance of large language models, is a pressing issue. Summary of the Invention
[0005] In view of the above, the purpose of one or more embodiments of this disclosure is to provide a method, apparatus, electronic device and storage medium for generating training samples for a large language model, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the first aspect of this disclosure provides a method for generating training samples for a large language model, comprising: Obtain candidate corpora and convert them into a task prompt format; The candidate corpus in the task prompt format and the unified evaluation instructions are input into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus; the multi-model collaborative evaluation orchestrator is a cluster of models composed of multiple different sources, different architectures and different sizes, with each model running in an independent environment; The quantified score is converted into an explicit score identifier and embedded into the candidate corpus to generate training samples; the explicit score identifier is used to track the source and judgment criteria of the quantified score.
[0007] Optionally, the unified evaluation instructions include evaluation dimensions, scale boundaries, and discretization rules; The candidate corpus in the task prompt format and the unified evaluation instructions are input into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus, including: The candidate corpus and the unified evaluation instruction are synchronously input into all models of the multi-model collaborative evaluation orchestrator to obtain an independent score for each model. All models obtain independent scores based on the same evaluation dimension and the same scale boundary. The independent scores are mapped to the same data format according to the discretization rule; The quantitative score of the candidate corpus is obtained based on the independent score.
[0008] Optionally, after mapping the independent scores to the same data format according to the discretization rule, the method further includes: Cross-validation is performed based on the independent scores of different models, and the independent scores are then filtered.
[0009] Optionally, the quantified scores are converted into explicit score identifiers and embedded into the candidate corpus to generate training samples, including: The independent scores, the discretization rules, and the quantitative scores are recorded as a ternary lookup table; The ternary lookup table is embedded into the candidate corpus to generate training samples.
[0010] Optionally, the quantitative scoring includes a security assessment dimension and a quality assessment dimension; The method further includes: In response to the determination that the security assessment dimension score does not meet the preset requirements, the candidate corpus is marked as a dangerous candidate corpus; In response to determining that the security assessment dimension score meets the preset requirements, the candidate corpus is marked as a security candidate corpus; Based on the quality assessment dimensions of the secure candidate corpus, the secure candidate corpus is divided into high-quality candidate corpus, candidate corpus to be optimized, and low-quality candidate corpus.
[0011] Optionally, it also includes: Extract the defect vectors from the candidate corpus to be optimized; Generate targeted optimization suggestions based on the defect type of the defect vector; The targeted optimization hints and the defect vectors are input into the multi-model collaborative evaluation orchestrator to obtain the optimized candidate corpus.
[0012] Optionally, it also includes: The parameters of the multi-model collaborative evaluation orchestrator are adjusted based on the quality distribution of the training samples.
[0013] A second aspect of this disclosure provides a training sample generation apparatus for a large language model, comprising: The acquisition module is configured to acquire candidate corpora and convert the candidate corpora into a task prompt format; The first generation module is configured to input the candidate corpus in the task prompt format and the unified evaluation instruction into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus; the multi-model collaborative evaluation orchestrator is a cluster of models composed of multiple different sources, different architectures and different sizes, with each model running in an independent environment; The second generation module is configured to convert the quantized score into an explicit score identifier and embed it into the candidate corpus to generate training samples; the explicit score identifier is used to track the source and judgment criteria of the quantized score.
[0014] A third aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first aspect.
[0015] In a fourth aspect, this disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method as described in the first aspect.
[0016] As can be seen from the above, this disclosure provides a method, apparatus, electronic device, and storage medium for generating training samples for a large language model. Attached Figure Description To more clearly illustrate the technical solutions in one or more embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only one or more embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a method for generating training samples for a large language model according to one or more embodiments of this disclosure. Figure 2 This is a schematic diagram of the structure of a training sample generation device for a large language model according to one or more embodiments of the present disclosure. Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to one or more embodiments of this disclosure. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar words used in one or more embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] As described in the background section, large language models require massive general-purpose corpora for general training scenarios, while fine-tuning for specific application scenarios also requires high-quality corpora tailored to those scenarios. Therefore, the training performance of large language models largely depends on the quality of the training data. However, as the parameter scale of large language models continues to expand, the training and deployment costs increase significantly. Low-quality training samples not only affect the training performance but may also lead to security risks. Therefore, improving the quality of training data, and thus the training performance of large language models, has become an urgent problem to be solved.
[0021] Unprocessed raw training data often suffers from text redundancy, non-standard language, and a lack of logical coherence. It may even contain low-quality or harmful information, such as fake data, offensive language, or sensitive private content. These problems not only reduce the efficiency of model training but may also lead to the model outputting incorrect, biased, or even dangerous results in real-world applications.
[0022] Two types of data processing methods are typically used in related technologies to improve the quality of training samples. The first type uses data cleaning algorithms to filter the original data, such as regularization, keyword filtering, and deduplication algorithms to remove low-quality data. The second type uses a unified sample generation model to produce training data, ensuring a balanced data instruction.
[0023] However, the first type of method is often based on a single screening rule, which relies on manual design. This results in a rigid screening mechanism that cannot cope with complex and ever-changing application environments; the screening rules are difficult to set, and lenient rules can easily miss low-quality raw data, while strict rules may lead to excessive filtering, reducing the number of usable training samples.
[0024] The second type of method also has limitations. Related techniques propose using large language models (e.g., constructing question-answer pairs or instruction-response pairs) to generate or expand training samples. However, due to the structure and parameter settings of the sample generation model, a single sample generation model exhibits biases in fluency, relevance, and logicality, easily producing monotonous training sample data that lacks diversity and comprehensiveness. Furthermore, a single model lacks reliable external validation, making it difficult to comprehensively and objectively evaluate its own output results. There is a tendency to label low-quality output results as high-quality samples, thus reducing the credibility of the generated samples.
[0025] Furthermore, the related technologies combine the first and second types of methods. Taking instruction tuning as an example, these technologies utilize large models to generate instruction-response pairs in batches, which are then directly used for fine-tuning after a small amount of manual review. While this method can rapidly improve the model's dialogue capabilities in the initial stages, the generated questions or instructions often lack specificity, are semantically ambiguous or logically unclear, and the reference explanations are frequently too brief to fully express the reasoning process. The answer section may contain factual errors or unreasonable inferences, and there is a lack of cross-task, multi-dimensional evaluation and screening, leading to the retention of some potentially low-quality samples. In terms of security, if the generated data contains offensive, inappropriate, or privacy-disclosure information, it will directly affect the controllability and reliability of the model's application.
[0026] Improving the quality of training data, and thus enhancing the training performance of large language models, has become an urgent problem to be solved.
[0027] refer to Figure 1 This disclosure discloses a method for generating training samples for a large language model according to one or more embodiments, comprising the following steps: Step S101: Obtain candidate corpus and convert the candidate corpus into a task prompt format; Step S102: Input the candidate corpus in the task prompt format and the unified evaluation instruction into the pre-trained multi-model collaborative evaluation orchestrator to obtain the quantitative score of the candidate corpus; the multi-model collaborative evaluation orchestrator is a cluster of models composed of multiple different sources, different architectures and different sizes, and each model runs in an independent environment; Step S103: Convert the quantified score into an explicit score identifier and embed it into the candidate corpus to generate training samples; the explicit score identifier is used to track the source and judgment criteria of the quantified score.
[0028] In this embodiment of the disclosure, raw text can be obtained by aggregating multiple sources of text, such as the Internet, open source libraries, and manual compilation, through a unified access pipeline, and then candidate text can be obtained through processing by a large language model.
[0029] In this embodiment of the disclosure, a metadata header containing source fingerprint, collection batch, timestamp and copyright status can be generated based on the acquired raw corpus for subsequent use.
[0030] In some embodiments, after obtaining the original corpus, the data can also be cleaned. For example, one or more cleaning rules can be used to perform encoding detection and character set unification on the candidate data, remove control characters and broken segments, and perform reversible normalization on punctuation, whitespace, full / half form and uppercase / lowercase.
[0031] In some embodiments, to ensure reversibility and accurate positioning, a "replacement-offset" mapping table can be generated in real time, and the original span and new position of each replacement can be written into the side vehicle metadata. The cleaning rules adopt a versioning strategy and are fixed with each batch. Checksums are calculated before and after cleaning to verify the stability of the process. The product of this step, along with the metadata, is stored on disk, so that the original text or the aligned position can be safely replayed in any subsequent stage.
[0032] In some embodiments, the cleaning rules are versioned and embedded into the processing configuration, and strong checksums and light fingerprints are calculated before and after cleaning to ensure the reproducibility, traceability and interpretability of differences in subsequent links.
[0033] In this embodiment of the disclosure, in order to ensure end-to-end traceability, the entry point can be fixed with the parsing strategy and tool version, and environment variables, abnormal alarms and rate limiting information can be written into an immutable event log to form a trusted root for subsequent auditing and playback.
[0034] In this embodiment, the original corpus can be segmented into a three-layer structure of sentences, paragraphs, and documents using language-aware boundary recognition, maintaining strict parent-child hierarchies and cross-layer positional alignment. The system calculates a stable hash for each slice as an accurate identifier and generates a lightweight semantic fingerprint for near-duplicate detection and rapid clustering. The indexing service writes the hash, hierarchical relationship, and positional information into a searchable inverted index, ensuring determinism and high throughput for subsequent deduplication, tracing, sampling, and reassembly. The segmentation strategy employs deterministic rules and a fixed random seed, guaranteeing that multiple processing runs under the same configuration yield completely consistent hierarchical and identifier layouts.
[0035] In this embodiment, structured or semi-structured content such as formulas, tables, code blocks, long references, and embedded multimedia can be extracted in a typological manner, and the original text can be replaced with semantic placeholders. Simultaneously, the one-to-one mapping relationship between structural semantics (such as table grids, code languages and fences, formula syntax trees, and media metadata) and placeholders is stored in the side-view log. The training side can use placeholders to identify content types to preserve context boundaries and reduce noise; when recovery or auditing is required, accurate backfilling is performed based on the mapping, ensuring that "trainability, fidelity, and recoverability" are achieved simultaneously. This step provides a fault-tolerant unpacking strategy for abnormal nesting and broken structures, preventing structural fragmentation from causing subsequent contamination.
[0036] In this embodiment, language and script detection can be performed on slices and documents to form language coverage profiles and script diversity statistics, and mixed scripts and uncommon characters can be standardized. Based on the language quota and tolerance range of the target task, the target proportion of each language and script is calculated, and a stratified sampling and weighted ordering strategy is generated to avoid high-resource languages crowding out training quotas, while implementing upper limit protection and stable injection for low-resource and long-tail languages to maintain the representativeness and robustness of the coverage. This step writes language tags, suggested sampling weights, and priorities into the sample metadata as direct constraint signals for subsequent candidate generation, collaborative evaluation, and training scheduling, and archives them together with the aforementioned mapping, hashing, and structural sidetracking when the data is written to disk, forming a traceable and auditable preprocessing result.
[0037] In this embodiment of the disclosure, the original corpus can be converted into task prompt format data using a prompt engineer. The task prompt format has the advantages of being templated, structured, and auditable.
[0038] In this embodiment, the target task constraints, context pruning strategy, sensitive placeholders, and style / security clauses are explicitly injected into the task prompt format data, bound to generation parameters (e.g., temperature, kernel sampling, maximum length, penalty coefficient, etc.), and a random seed is injected to eliminate data incomparability. The prompt and decoding parameters use a "prompt-parameter-version" ternary binding, including controlled settings such as deterministic random seed, temperature / kernel sampling upper limit, length, and penalty coefficient. Sensitive placeholders and a domain terminology list can be injected simultaneously, providing hard constraints on candidate generation in terms of semantic boundaries and terminology consistency; all configurations are written to metadata to support subsequent replay and auditing.
[0039] In this embodiment of the disclosure, the large language model described above is required to ensure target relevance while suppressing topic detachment and redundant diffusion, so as to provide an alignable set of original corpus candidates for subsequent evaluation.
[0040] In this embodiment, the output of the large language model can perform format validity checks, disable pattern filtering, and security scanning, and trigger limited resampling or parameter rollback when necessary to suppress topic drift and redundant diffusion. Candidate corpora are forcibly distinguished and encapsulated into a three-part object: "prompt body—response body—control parameter body." This encapsulated object records the model version fingerprint, generation trajectory, runtime latency, and cost metrics to support performance reconciliation and regression comparison. Limited resampling or parameter rollback is performed as needed to stabilize output. Thus, all outputs are written to an immutable event log and carry a lineage ID, ensuring that the same input can be replayed under the same configuration, facilitating regression comparison and bias localization.
[0041] That is, in this embodiment of the disclosure, the candidate corpus can be encapsulated as a three-part object consisting of a "prompt body - response body - control parameter body". The encapsulator performs consistency repair at the text structure and punctuation level, while preserving the original generation snapshot and position mapping, ensuring reversibility and traceability without compromising readability.
[0042] In this embodiment of the disclosure, candidate corpora and preset unified evaluation instructions are routed to heterogeneous evaluation nodes of the multi-model collaborative evaluation orchestrator. Each heterogeneous evaluation node independently completes multi-dimensional structured scoring in a mutually isolated sandbox to avoid interference from other models or inconsistent evaluation dimensions in the scoring results of each model.
[0043] The orchestration layer provides a set of models for parallelism shaping, timeout and failure retries, node circuit breaking and rollback, ensuring stable throughput under high concurrency and partial failure scenarios. The output follows a unified data pattern, carrying node identifiers (IDs), scale configurations and anomaly markers to form a machine-readable and auditable multi-source scoring evidence set.
[0044] The aforementioned heterogeneous evaluation nodes refer to language models from different sources and architectures. Evaluation dimensions may include fluency, relevance, accuracy, and security, and the evaluation results are expressed in the form of scores (e.g., 1 to 5 points).
[0045] The output data includes field names, value ranges, conflict priority, and missing data handling.
[0046] In some embodiments, to reduce scale drift between models, the evaluation text's built-in anchor evaluation entries and sparse references to exemplified scale descriptions (non-training content) are strictly limited in their free expression to ensure the machine readability and comparability of the outputs of each evaluation node. The instructions also include safety rejection conditions and exception return code definitions, enabling abnormal paths to be identified and isolated at the orchestration layer.
[0047] In this embodiment of the disclosure, the scoring results can be converted into a controlled closed set score identifier (token) via a discrete mapping table. Each heterogeneous evaluation node corresponds to an evaluation result, and each evaluation result corresponds to an evaluation token. This evaluation token can be injected into the prefix of the candidate corpus by an "annotator".
[0048] In this embodiment, the "original score - discrete file - token" (independent score - discretization rule - quantization score) correspondence is retained to support traceability. To ensure cross-model consistency, the online calibration calibrator uses the anchor point set and nearest window statistics to scale the node output, and the calibration coefficients and residuals are recorded together to avoid drift over time that erodes the judgment boundary.
[0049] In this embodiment of the disclosure, in order to eliminate cross-model scale drift, an online scale calibrator can also perform periodic calibration with an anchor point set, output a comprehensive score token and confidence index, and continuously monitor distribution drift and consistency.
[0050] In some embodiments, the evaluation token undergoes threshold and conflict checks. This cross-validation ensures that the token is unique, free from out-of-bounds errors, and conflict-free.
[0051] In some embodiments, cross-validation may include: threshold value, mutual exclusion, and uniqueness checks. Node re-evaluation or cluster degradation is triggered when a missing or conflicting value is detected.
[0052] In some embodiments, the method may further include: performing robust anomaly processing (such as truncated mean, Huber / quantile method) on the scores output by each model to remove the influence of the highest / lowest or outlier nodes; then performing weighted fusion based on the node's historical reliability, recent calibration residuals, and task domain matching degree to generate a comprehensive score and its uncertainty index. The fusion result is discretized into a comprehensive score token, and written to metadata synchronously with the dimension-level token; simultaneously, drift monitoring signals and node weight write-back updates are output to drive the adaptive optimization of the evaluation cluster. The comprehensive score token serves as the main control variable for subsequent gating screening and training scheduling, and is written to disk along with the candidate's lineage and audit logs to ensure that the quality signal is "read-and-use, traceable, and replayable" throughout the entire process.
[0053] In this embodiment of the disclosure, candidate corpora are judged based on the scoring token in order to filter qualified candidate corpora and generate training samples.
[0054] In some embodiments, security is a hard veto criterion in the decision criteria; if the security dimension score of a candidate corpus is below a threshold, it is directly eliminated. The key quality dimension employs a tiered threshold and buffer to reduce false positives. The decision criteria for the security dimension may include whether the candidate corpus is harmful or shows signs of privacy breaches.
[0055] In some embodiments, when the key dimension score of any candidate corpus is at the level to be optimized or in the buffer zone, a defect vector of the candidate corpus is generated, and the defect vector and rewriting suggestions are input into a large language model for rewriting. The rewriting suggestions may include dimensions to be improved, semantic boundaries to be preserved, and safety / style constraints to be inherited. The large language model used for rewriting and the large language model used to generate the candidate corpus can be the same large language model or different large language models.
[0056] In some embodiments, to improve the effectiveness of optimization, the system generates measurable expected improvement magnitude and verification criteria for each dimension, and binds deterministic random seeds and decoding parameters to ensure that the optimization process is replayable, comparable, and regressible; all synthesis details are written to the optimization log, supporting subsequent offline evaluation and automatic tuning of the prompting strategy.
[0057] In some embodiments, the rewritten data can undergo format validity checks, disabled pattern scanning, and security reviews. Non-compliant products can be resampled a limited number of times and subjected to parameter annealing (lowering temperature and tightening kernel sampling). If necessary, conservative templates can be enabled to stabilize the output. The rewritten sample maintains a three-part encapsulation of "prompt body - response body - control parameter body" and retains the content alignment mapping and difference summary with the original version to facilitate accurate comparison and location later. All execution traces, latency, and resource bills are included in the immutable event log to support auditing.
[0058] The rewritten candidate corpus is re-evaluated until it meets the quality requirements.
[0059] In some embodiments, differential statistics and significance tests can be performed on the candidate corpus and the rewritten data. Only when the key dimensions reach a set improvement threshold and the safety dimensions are fully qualified can the data advance. If score fluctuations or insufficient improvement occur, the data will proceed to the next round of optimization or be transferred to the manual review process. During the re-evaluation phase, the calibration residuals and node weights are updated simultaneously to continuously correct the scale consistency of the evaluation cluster. The differential results, improvement trajectories, and reasons for non-compliance are written into the quality timeline, forming a closed-loop feedback loop for data and the model.
[0060] In some embodiments, when a sample meets any of the following conditions—quality threshold, reaches iteration limit, or stagnates in improvement—the system performs a convergence determination: qualified samples are added to a high-confidence corpus and their scoring tokens, suggested sampling weights, and loss weighting factors are frozen; samples that fail to meet the criteria in multiple rounds but are safe and qualified are marked as low-confidence and added to a manual review pool or isolation area for further processing after subsequent policy updates, preventing potentially valuable data from being discarded prematurely. During the archiving process, the system completes cross-layer deduplication, label consistency verification, and the implementation of training / validation / testing segmentation strategies. Simultaneously, course level, long-tail identifiers, and priorities are written into the sample metadata as "plug-and-play" control signals for subsequent training scheduling. Finally, the sample lineage, optimization chain, and evaluation evidence chain are solidified as a whole, ensuring that the entire process is traceable, replayable, and auditable.
[0061] In one embodiment of this disclosure, the completeness and consistency of the scoring tokens and metadata accompanying the candidate samples are checked. This verifies whether the labels of each dimension fall within a controlled closed set, whether there are any missing or conflicting elements, and compares the calibration coefficients within the scale version and time window to eliminate the impact of cross-model drift. After schema verification, the levels and confidence scores of each dimension are analyzed, and a "defect vector" is constructed by combining it with the target threshold. This vector numerically characterizes the deviation and priority of the sample in dimensions such as fluency, relevance, accuracy, security, and coverage. The system stores the defect vector and the sample spectrum together on disk, forming a unified quality representation that can be used for subsequent optimization synthesis, training scheduling, and auditing.
[0062] In some embodiments, the final training samples can be generated by performing hierarchical reordering and balanced sampling suggestions according to topic, difficulty, language, and source, while simultaneously performing cross-level deduplication and label consistency verification to prevent information leakage and distribution bias. The data is stored on disk according to a training / validation / testing splitting strategy, carrying scheduling signals such as a comprehensive score token, dimension token, suggested sampling weight, loss weighting factor, and course level, forming a traceable and highly reliable pre-trained corpus.
[0063] In some embodiments, training samples can be adaptively fine-tuned based on the target task specifications, while retaining data cards, governance and compliance audit chains, version snapshots and genealogy indexes, ensuring that the training side can directly consume them and support the immediate effectiveness of subsequent quality grading and course-based scheduling strategies.
[0064] In some embodiments, training samples are normalized according to the task protocol into question-and-answer, instruction-and-response, inference chain, or code-and-annotation formats, and the input, output, label space, loss type, and evaluation protocol are uniformly encapsulated. The encapsulation process verifies the integrity of the dimensional scoring tokens and comprehensive scoring tokens attached to the samples, fills in missing metadata fields (source, version, language, topic, time window), and generates sample lineage IDs and immutable event logs, ensuring that the training side has a traceable and auditable view of the data source, processing path, and constraints. When samples are written to disk, a batch number and splitting mark (training / validation / testing) are written, maintaining schema alignment with the sampling priority and loss weighting interfaces used by the subsequent scheduler.
[0065] In some embodiments, the scheduling preprocessor parses the comprehensive scoring token and the dimension-level token, and transforms the discrete level values into continuous sampling weights and loss weighting factors based on a monotonic mapping function. Temperature and boundary buffers are introduced in the mapping to avoid over-polarization. The system simultaneously reads the scarcity profile and topic coverage target, generates upper limit protection and minimum retention rates for long-tail topics and low-resource languages, and adds them to the weight proposal. An upper limit for gradient contribution pruning is introduced for potential noise dimensions to prevent them from having a disproportionate impact on the loss level. The weight proposal is produced as a quintuple of "sampling weight, loss weight, frequency upper limit, minimum retention rate, and course level," which serves as a direct control variable for the batch processing builder and optimizer.
[0066] During the batch construction phase, the training scheduler reads the weight quintuples and generates mini-batches using hierarchical weighted multinomial or bucketed Poisson sampling. At the end of the batch, it measures the deviation between the actual sampling distribution and the target distribution and performs online rebalancing according to the moving window metric. The scheduler continuously monitors the stability of the training curve (loss fluctuations, gradient noise scaling, parameter norm drift) and generalization signals (validation set metrics, early stopping probes). When overfitting or pattern collapse is detected, it automatically reduces the proportion of high-scoring samples, increases the coverage of medium-scoring samples, or triggers gentle annealing of the learning rate and loss weights. To avoid sample starvation, the scheduler sets a lower limit quota and an aging-up re-injection strategy for each class of samples to ensure sufficient and fair data utilization.
[0067] In the early to mid-stages of training, the batch builder applies priority boosting and repetition factor injection to high-scoring samples, enabling them to contribute more effective gradient updates under the same budget. The loss layer introduces recalibration based on quality weights, allowing high-scoring samples to gain a greater marginal impact on the objective function. The system evaluates the diminishing marginal returns of high-scoring samples in stages, automatically reducing their boost coefficient after reaching a threshold to avoid overfitting to minority distributions and loss of diversity. Simultaneously, gradient noise suppression and parameter update thresholds are enabled on the optimizer side to ensure that the amplified impact does not cause training instability.
[0068] A dual-channel strategy is implemented for low-scoring samples: on the one hand, weights are simultaneously reduced at both the sampling and loss stages to suppress noise disturbances to the convergence trajectory; on the other hand, for samples in long-tail topics, low-resource languages, or critical boundary conditions, minimum retention rates and frequency upper limits are set according to scarcity and coverage targets, forming a controlled presence of "low frequency but not absence". The scheduler enables upstream filters (constraint consistency, safety verification, and secondary checks on format legality) and downstream thresholds (soft discarding or weight reduction for batches with gradient anomalies) for such samples, ensuring coverage while limiting the risk of noise propagation within a controllable range.
[0069] Training is divided into multi-stage courses, progressing from low to high order according to a joint scale of "quality level × task difficulty × inference depth". Stage switching is triggered by multiple signals, including achievement of validation set master metrics, satisfaction of training stability window, absence of overfitting probe anomalies, and achievement of long-tail coverage. When any monitoring item exceeds the safety threshold, the scheduler triggers stage rollback or extension, or performs fine-grained quota reallocation within the current stage. To mitigate forgetting, the system injects replay quotas during the transition window of stage switching to maintain low-amplitude continuous coverage of key samples from the previous stage. For tasks with large gaps between stages, stratified distillation or soft-label smoothing bridging is used to ensure smooth capability transfer, controllable curves, and robust final convergence.
[0070] In some embodiments, the parameters of the multi-model collaborative evaluation orchestrator can also be adjusted according to the quality distribution of the training samples.
[0071] In the embodiments of this disclosure, candidate corpora are traceable and manageable.
[0072] In some embodiments, a deduplication candidate discovery pipeline can be run simultaneously at three levels: document, paragraph, and sentence. First, the preprocessed text is normalized according to deterministic rules. Then, content hash, rolling fingerprint, and locality-sensitive hash are calculated to cover exact repetition and local overlap scenarios. To identify semantically similar samples, the system embeds slices into a vector space and constructs a high-recall candidate set based on hierarchical indexes (e.g., bucketing by language, topic, and time window). Inter-batch distribution drift is suppressed through threshold adaptation and distance recalibration. During the candidate generation stage, "repetition within the same batch" and "reproduction across batches" are strictly distinguished, and negative caching and time decay strategies are introduced to reduce redundant computation. All candidate pairs, along with similarity evidence, hierarchical position, and source fingerprint, are written to the event log, forming the evidentiary basis for subsequent judgment.
[0073] On the candidate set, the system uses robust clustering to obtain duplicate clusters, and selects representatives within each cluster as canonical samples based on a weighted strategy of quality score, coverage, contextual integrity, source credibility, and time freshness. The retained samples are assigned primary key labels and a "cluster ID—representative—subordinate" hierarchy, while other members are downweighted or removed according to the strategy, and are labeled accordingly. <duplicate>or <neardup>This facilitates subsequent traceability and sampling verification. To avoid domain bias caused by deduplication, the system sets quota protection and deduplication intensity limits for different sources and languages to ensure the reasonable retention of long-tail and low-resource segments. The basis for deduplication decisions (similarity matrix, pruning threshold, conflict resolution order) is written to disk along with the results to ensure that the judgment process is explainable, reproducible, and auditable.
[0074] Sensitive identification employs a multi-signal integration framework of "rules + models": On the rule side, high-precision regularization templates, dictionaries, and verification algorithms are used to identify structured identifiers and credential-like patterns; on the model side, sequence labeling and named entity recognition capture names, place names, organizations, and potential privacy fragments, and contextual constraints and coreference resolution reduce false alarms. The system generates a risk score and confidence interval for each hit entry, manages them hierarchically by category (identity, contact information, finance, credentials, precise geography, etc.), and sends highly ambiguous items to a secondary judgment or manual sampling channel. To suppress missed detections caused by cross-language and script mixing, the detector implements a branching strategy on language and script tags and uses customized sub-models and anomaly pattern scanners for special domains (code blocks, tables, log fragments) to ensure that sensitive signs in complex structures are accurately captured.
[0075] For identified sensitive segments, the system selects strategies such as masking, interval generalization, or semantically equivalent replacement according to risk level, and generates recoverable mappings to trace back the original text in audit scenarios. Simultaneously, it exposes only irreversible representations to the training consumer side to prevent restoration risks. The desensitization engine maintains character-level offset alignment and placeholder consistency to ensure contextual coherence and model readability. For structured units such as tables and code, minimal modifications at the field or statement level are used to maintain structural semantics. To ensure processing quality, the system runs automated readability and consistency checks and enables conservative degradation strategies for extreme scenarios. Each change is written to the side-view record by processing serial numbers, strategy versions, and effect metrics, and inserted into the sample prefix. <redacted>The process tracking label indicates that the anonymization and path tracking have been completed.
[0076] After deduplication and anonymization, samples undergo centralized verification at the compliance checkpoint, including license term matching, regional and industry rule package adaptation, sensitive category coverage verification, and security rejection item review. The checkpoint issues a compliance credential for each sample, recording the audit timestamp, rule package version, responsible entity, and verification summary, and annotating abnormal paths with reason codes and suggested handling. The system performs consistency verification of compliance status with scoring tokens, source fingerprints, and deduplication cluster information to prevent label drift and metadata breaks; simultaneously, it generates evidence snapshots for subsequent training and external audits, ensuring the independence and verifiability of compliance proofs.
[0077] The integrator imports compliant samples into the governance repository, establishing hierarchical views and regulatory dashboards by topic, language, and time dimension. It continuously monitors key indicators such as coverage, duplication rate, sensitivity hit rate, and processing rollback rate, and issues alerts for abnormal trends. The system performs training / validation / testing splits and version snapshots on samples, along with suggested sampling weights, loss weighting, and course levels, serving as direct control signals for training scheduling. To meet deletion requests and data retention policies, the governance repository provides precise retrieval by source and category, batch invalidation, and rollback-enabled event tracing. All operations are recorded using event tracing, ensuring end-to-end traceability and auditability from sample access and governance to training consumption.
[0078] It is understandable that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.
[0079] It should be noted that the methods of one or more embodiments of this disclosure can be executed by a single device, such as a computer or server. The methods of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the process. In such a distributed scenario, one of these devices may execute only one or more steps of the methods of one or more embodiments of this disclosure, and the multiple devices will interact with each other to complete the method described.
[0080] It should be noted that the above description pertains to specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0081] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a training sample generation device for a large language model. For example... Figure 2 As shown, the training sample generation device for the large language model includes: The acquisition module 11 is configured to acquire candidate corpora and convert the candidate corpora into a task prompt format; The first generation module 12 is configured to input the candidate corpus in the task prompt format and the unified evaluation instruction into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus; the multi-model collaborative evaluation orchestrator is a cluster of models composed of multiple different sources, different architectures and different sizes, with each model running in an independent environment; The second generation module 13 is configured to convert the quantized score into an explicit score identifier and embed it into the candidate corpus to generate training samples; the explicit score identifier is used to track the source and judgment criteria of the quantized score.
[0082] For ease of description, the above apparatus is described in terms of function, divided into various modules. Of course, when implementing one or more embodiments of this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0083] The apparatus described above is used to implement the corresponding methods in the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0084] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0085] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.
[0086] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this disclosure are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0087] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0088] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0089] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0090] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this disclosure, and not necessarily all the components shown in the figures.
[0091] The electronic devices described above are used to implement the corresponding methods in the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0092] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0093] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of one or more embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0094] Additionally, to simplify the description and discussion, and to avoid obscuring one or more embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring one or more embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which one or more embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) are set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that one or more embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0095] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0096] This disclosure includes one or more embodiments intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this disclosure should be included within the scope of protection of this disclosure.< / redacted> < / neardup> < / duplicate>
Claims
1. A method for generating training samples for a large language model, characterized in that, include: Obtain candidate corpus; The candidate corpus and unified evaluation instructions are input into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus. The multi-model collaborative evaluation orchestrator is a cluster of models from different sources, with different architectures and different sizes, and each model runs in an independent environment; The quantified score is converted into an explicit score identifier and embedded into the candidate corpus to generate training samples; the explicit score identifier is used to track the source and judgment criteria of the quantified score.
2. The method according to claim 1, characterized in that, The unified evaluation instructions include evaluation dimensions, scale boundaries, and discretization rules; The candidate corpus and unified evaluation instructions are input into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus, including: The candidate corpus and the unified evaluation instruction are synchronously input into all models of the multi-model collaborative evaluation orchestrator to obtain an independent score for each model. All models obtain independent scores based on the same evaluation dimension and the same scale boundary. The independent scores are mapped to the same data format according to the discretization rule; The quantitative score of the candidate corpus is obtained based on the independent score.
3. The method according to claim 2, characterized in that, After mapping the independent scores to the same data format according to the discretization rule, the method further includes: Cross-validation is performed based on the independent scores of different models, and the independent scores are then filtered.
4. The method according to claim 1, characterized in that, The quantitative scores are converted into explicit score identifiers and embedded into the candidate corpus to generate training samples, including: The independent scores, the discretization rules, and the quantitative scores are recorded as a ternary lookup table; The ternary lookup table is embedded into the candidate corpus to generate training samples.
5. The method according to claim 1, characterized in that, The quantitative scoring includes security assessment dimensions and quality assessment dimensions; The method further includes: In response to the determination that the security assessment dimension score does not meet the preset requirements, the candidate corpus is marked as a dangerous candidate corpus; In response to determining that the security assessment dimension score meets the preset requirements, the candidate corpus is marked as a security candidate corpus; Based on the quality assessment dimensions of the secure candidate corpus, the secure candidate corpus is divided into high-quality candidate corpus, candidate corpus to be optimized, and low-quality candidate corpus.
6. The method according to claim 5, further comprising: Extract the defect vectors from the candidate corpus to be optimized; Generate targeted optimization suggestions based on the defect type of the defect vector; The targeted optimization hints and the defect vectors are input into the multi-model collaborative evaluation orchestrator to obtain the optimized candidate corpus.
7. The method according to claim 1, characterized in that, Also includes: The parameters of the multi-model collaborative evaluation orchestrator are adjusted based on the quality distribution of the training samples.
8. A training sample generation device for a large language model, characterized in that, include: The acquisition module is configured to acquire candidate corpora and convert the candidate corpora into a task prompt format; The first generation module is configured to input the candidate corpus in task prompt format and the unified evaluation instruction into a pre-trained multi-model collaborative evaluation orchestrator to obtain a quantitative score for the candidate corpus. The multi-model collaborative evaluation orchestrator is a cluster of models from different sources, with different architectures and different sizes, and each model runs in an independent environment; The second generation module is configured to convert the quantized score into an explicit score identifier and embed it into the candidate corpus to generate training samples; the explicit score identifier is used to track the source and judgment criteria of the quantized score.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executed by the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.