Large model training data generation method and device, medium and equipment

By receiving data generation tasks from risk control business scenarios in a large language model, and using task profiling to extract content elements and format features from a general knowledge base, high-quality risk control training samples are generated. This solves the problems of high cost and difficulty in guaranteeing the quality of generating professional domain training samples in existing technologies, and achieves efficient and professional training data generation.

CN121682258APending Publication Date: 2026-03-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511724835.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In existing technologies, generating training samples for specialized fields is costly, time-consuming, and of questionable quality, which limits the application of large language models in specific business scenarios.

Method used

By receiving data generation tasks from risk control business scenarios, and using task profiles to extract content elements and format features from a general knowledge base, high-quality risk control training samples are generated, and large models are automatically trained to adapt to specific business needs.

Benefits of technology

It improves the efficiency and consistency of training sample generation, reduces reliance on human intervention, and enhances the accuracy and professionalism of large models in risk control business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121682258A_ABST
    Figure CN121682258A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large model training data generation method, and the method comprises the steps: determining the risk control capability of a target large model and the description of a knowledge field according to a data generation task requirement of a risk control business scene, so as to determine a core concept and a format requirement based on a determined task portrait, the core concept is used for obtaining matched rich knowledge from a universal knowledge base, the format requirement is used for reprocessing the obtained knowledge, and a training sample is obtained by generating a large model to train a target large model for executing risk control business. Through the automatic training sample generation process corresponding to the demand, the existing generation and labeling problem depending on manpower is avoided, the efficiency is improved, the quality difference between samples can be reduced, and the accuracy of the training model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of computer technology, and in particular to a large model training data generation method and device, a storage medium and equipment. BACKGROUND

[0002] In recent years, large language models (LLMs) have been widely applied in natural language processing, intelligent question answering, and decision support in many fields. With the development of technology, complex reasoning businesses such as mathematical proof, logical deduction, and complex problem solving have become important application scenarios. However, the training samples of general LLMs have been difficult to meet the training requirements of LLMs applied to specific business scenarios, so generating training samples in professional fields has gradually become a focus.

[0003] In the prior art, training samples set by experts in the professional field through artificial means not only have high costs, but also have difficulty in ensuring the quality of the training samples. Moreover, a long preparation time is required to generate samples through multiple links such as demand analysis, expert recruitment, annotation training, and quality control.

[0004] As can be seen, the prior art highly depends on manpower, and the scale, cost, and quality of the generated samples are bound to the time and effort of experts, resulting in high costs and long cycles for generating training samples. The subjective differences between different experts make it difficult to ensure data consistency. This mode cannot fundamentally solve the problem that the demand for training samples for LLMs applied to specific business scenarios restricts the further expansion of LLM applications.

[0005] Therefore, the present specification provides a large model training data generation method to partially solve the problems in the prior art. SUMMARY

[0006] The present specification provides a large model training data generation method, device, storage medium, and electronic equipment to partially solve the problems in the prior art.

[0007] The present specification provides the following technical solutions: The present specification provides a large model training data generation method, which comprises: receiving a data generation task for a risk control business scenario, the data generation task carrying a task profile of a target large model to be trained, the task profile at least including a risk control capability description and a knowledge field description of the target large model; determining content elements and format features of the data generation task through the task profile; determining knowledge materials matching the content elements from a preset general knowledge base; Based on the knowledge material, a plurality of risk control training samples are generated according to the format feature by a preset generation large model, the risk control training samples are used to train the target large model, and the target large model is used to receive business data in a risk control business scenario and output a risk control strategy, so as to perform risk control processing on the business based on the risk control strategy.

[0008] The present specification provides a large model training data generation device, the device comprises: A receiving module is configured to receive a data generation task for a risk control business scenario, the data generation task carries a task profile of a target large model to be trained, and the task profile at least includes a risk control capability description and a knowledge field description of the target large model. An extraction module is configured to determine content elements and format features of the data generation task through the task profile. An enrichment module is configured to determine knowledge materials matched with the content elements from a preset general knowledge base. A generation module is configured to generate a plurality of risk control training samples based on the knowledge materials according to the format features by a preset generation large model, the risk control training samples are used to train the target large model, and the target large model is used to receive business data in a risk control business scenario and output a risk control strategy, so as to perform risk control processing on the business based on the risk control strategy.

[0009] The present specification provides a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned large model training data generation method.

[0010] The present specification provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned large model training data generation method when executing the program.

[0011] The above-mentioned at least one technical solution adopted by the embodiments of the present specification can achieve the following beneficial effects: The embodiments of the present specification disclose a large model training data generation method, which determines the risk control capability of a target large model and the description of a knowledge field according to the data generation task requirements of a risk control business scenario, determines core concepts and format requirements based on the determined task profile, the core concepts are used to obtain matched rich knowledge from a general knowledge base, the format requirements are used to reprocess the obtained knowledge, training samples are obtained through a generation large model, and the target large model is trained to execute risk control business. Through the automatic corresponding demand training sample generation process, the existing generation and annotation problems relying on manpower are avoided, not only the efficiency is improved, but also the quality difference between samples is reduced, and the accuracy of the training model is improved. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings: Figure 1 A flowchart illustrating the training data generation process for a large model, as provided in the embodiments of this specification. Figure 2 A training diagram of the risk control model provided in the embodiments of this specification; Figure 3 A schematic diagram of a training data generation device for a large model provided in the embodiments of this specification; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this specification. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0014] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0015] Figure 1 A flowchart for generating training data for a large model, provided in the embodiments of this specification, specifically includes the following steps: S100: Receive a data generation task for a risk control business scenario. The data generation task carries a task profile of the target large model to be trained. The task profile includes at least a description of the risk control capabilities and a description of the knowledge domain of the target large model.

[0016] S102: Determine the content elements and format characteristics of the data generation task through the task profile.

[0017] In the embodiments described in this specification, the following are employed: Figure 1 The device used to generate training data using the method shown can be any electronic device, such as a computer, server, or server cluster consisting of multiple servers. For ease of description, the following explanation uses a server as an example only.

[0018] To solve the problem that the efficiency is low and the quality is difficult to guarantee in the case that the business scene based on a large language model (LLM) is more and more subdivided, the business content is more and more complex, and the complex inference ability needs to be targeted, the method of simply relying on expert experience to build samples and manual annotation is used. In the embodiment of the present specification, a method of generating training samples required by a corresponding business scene through a general knowledge base is proposed based on the task description of the training model, so as to drive the efficient and professional update of the LLM in the subdivided risk control scene, and to give the target LLM the ability to cope with the accurate inference and reliable decision required by the complex business.

[0019] Specifically, in the embodiment of the present specification, the server can generate a task for receiving data for a risk control business scene. The task is a trigger instruction and a target guide for the entire data generation process, and the core is to carry the task portrait of the target large model to be trained.

[0020] The task portrait can be a structured function definition file, which at least contains two dimensions of key information. One is the description of the risk control ability that the target LLM expects to have. For example, it needs to be good at identifying financial transaction fraud patterns, or good at evaluating the credit risk of a specific industry. The second is the description of the knowledge field that the target LLM expects to master, or the description of the knowledge field required by the specific application scene of the target LLM. For example, it needs to cover professional knowledge in specific fields such as bank regulatory policies, insurance claim rules, or network security protocols.

[0021] From the target of the LLM executing the risk control business, the decision logic, risk type and judgment standard of different risk control scenes are essentially different. For example, a model used for credit approval, its core ability is to evaluate the willingness and ability of users to repay, and the knowledge required is the credit system, financial indicators and industry cycle. A model for anti-money laundering detection, its core ability is to identify hidden illegal money transfer patterns, and the knowledge required is money laundering methods, transaction network characteristics and financial regulation. The inference path and decision basis required by these two scenes are completely different. If the data is trained without distinction, the model will learn an "average" and "fuzzy" ability, which cannot achieve the professional precision required by the business in any specific scene. Therefore, the essence of the risk control ability description is to anchor a clear business target for the data generation process, and to ensure that each sample generated can be used to train the model to solve the professional skills of a specific task.

[0022] Secondly, from a technical implementation perspective, the performance ceiling of a model is determined by the quality and relevance of its training data. Professional risk control decisions rely on rigorous reasoning within a specific knowledge context. If the knowledge background is broad but not precise, the model will struggle to make accurate judgments. For example, in medical insurance fraud detection, if the model lacks specific knowledge of the medical insurance catalog, treatment guidelines, and drug pricing, it will be unable to identify fraudulent behaviors such as "overtreatment" or "impersonation." Therefore, the role of knowledge domain description is to construct a precise decision context for the model. It acts like a "materials sourcing manual," guiding the server to accurately select and combine highly relevant professional knowledge from a vast amount of general knowledge, thereby ensuring that the generated training samples possess deep professional foundation and correct business logic.

[0023] In summary, using risk control capability description and knowledge domain description as the core of task profiling aims to fundamentally solve the problem of inaccuracies in general-purpose models in specialized applications. Through this refined target definition, the embodiments described herein can generate highly customized training data for each specific business scenario, thereby training a professional risk control model with expert-level performance in a specific domain.

[0024] Specifically, after receiving a data generation task, the server can determine the task profile for the target large model to be trained, which is carried in the data generation task. Then, a semantic parsing model can extract keywords from the risk control capability description and knowledge domain description as content elements. This semantic parsing model can understand the business semantics behind the natural language description and accurately extract a series of keywords representing the core intent and key information, using these keywords as content elements for the generation task. For example, from the capability description "identifying cross-border e-commerce transaction fraud," elements such as "cross-border e-commerce," "transaction fraud," and "identification rules" can be parsed. From the knowledge description "involving foreign exchange control policies," elements such as "foreign exchange control" and "policy clauses" can be parsed. These content elements constitute the substantive themes and core materials for subsequent knowledge retrieval and sample generation.

[0025] Then, based on content elements and / or task profiles, format features are determined, including at least the input format and the output format.

[0026] The server can determine format characteristics based on the extracted content elements. Since the content elements are core concepts extracted from risk control capabilities and knowledge domains, they inherently contain the reasoning paradigm and data interaction norms that should be followed to solve the business problem. Format characteristics are not independent of the external form of the content, but rather an inevitable expression determined by the business connotation. Different combinations of content elements correspond to drastically different types of risk control tasks. For example, if the content elements point to "credit approval" and "debt ratio," it implies that the solution process requires numerical calculations and threshold comparisons, which naturally leads to the input format containing structured numerical fields and the output format containing calculation conclusions and judgments. Content elements extracted from the knowledge domain description, on the other hand, carry the standard operating procedures of that domain. For example, if the content elements involve "anti-money laundering" and "transaction networks," it means that the analysis must follow a graph analysis paradigm from point to surface, identifying patterns. This necessarily requires the input format to represent entity relationships and the output format to describe network characteristics and risk paths.

[0027] Alternatively, the server can determine format characteristics based on the task profile. A task profile is typically text described in natural language by a human inputter. Therefore, if format requirements are defined in the task profile, the server can determine the corresponding format characteristics through semantic analysis. For example, explicitly requiring "generated in question-answer pair format" or "output must be a JSON object containing risk scores and evidence." Alternatively, the server can determine format characteristics in parallel based on content elements and task profiles. On one hand, the server can parse explicit format requirements from the task profile. On the other hand, the server can infer implicit format suggestions based on content elements. Subsequently, the server will perform consistency checks and intelligent fusion on the two to determine the format characteristics. For example, the task profile may only roughly require that "the output must include a reasoning process," while based on the inference from the content element "credit anti-fraud - rule matching," the server can specify this as "input: applicant information and credit records; output: step 1 triggering fraud rule item, step 2 rule details explanation, step 3 comprehensive risk judgment chain format." This process combines explicit rules with implicit business logic, ensuring that it meets user expectations while automatically optimizing and supplementing incomplete instructions using domain knowledge, thereby generating high-quality format specifications that fit the essence of the business.

[0028] S104: Determine knowledge materials that match the content elements from a preset general knowledge base.

[0029] In the embodiments of this specification, after determining the content elements, the server can obtain knowledge materials from a general knowledge base to assist in generating samples based on these elements. The purpose of this step is to provide indispensable original materials rich in professional knowledge and factual basis for the subsequent sample generation process, thereby ensuring the richness, accuracy, and professionalism of the generated content. Since the server obtains a task profile, it only defines the goal and direction of the generation task but does not itself contain the specific knowledge content required for generation. Without external knowledge injection, the generation process will face the predicament of "cooking without rice." This may cause the large generation model to only be able to generate training samples based on its inherent parametric knowledge, without incorporating detailed domain knowledge, thus producing samples with poor content or containing factual illusions. Furthermore, in highly specialized risk control fields, such as those involving specific financial regulations or specific insurance clauses, without the support of authoritative knowledge sources, the large generation model will struggle to generate logically rigorous, terminologically accurate, and business-realistic samples.

[0030] Therefore, this step retrieves knowledge resources related to content elements from a general knowledge base, using these resources as core knowledge materials. These knowledge materials provide a reliable factual basis, rich contextual information, and a professional terminology system for generating large-scale models, enabling them to deduce, combine, and recreate on this basis, producing high-quality risk control training samples that both meet the task objectives and possess a solid knowledge foundation.

[0031] Specifically, the server first performs semantic similarity calculations based on knowledge resources in a pre-defined general knowledge base and the content of the elements. Specifically, the server uses the content elements parsed in step S102, namely the core keywords representing risk control capabilities and knowledge domains, as the retrieval basis. It then performs large-scale semantic similarity calculations between the content input and the various knowledge resources pre-stored in the general knowledge base. In this embodiment, the knowledge resource refers to a complete document, such as an industry analysis report, full text of policies and regulations, a technical white paper, or an expert article. The server uses a semantic embedding model to convert the content elements and the full text of the knowledge resources into vector representations in a high-dimensional space. By calculating the cosine similarity between the vectors or using a more complex cross-encoder model, it obtains a quantified semantic relevance score, thereby evaluating the strength of the association between each knowledge resource and the current task at the semantic level, rather than simply at the keyword matching level.

[0032] Then, based on the similarity of the obtained knowledge resources, and according to a preset similarity threshold, knowledge resources matching the element content are selected from each knowledge resource as knowledge materials. This threshold ensures that only knowledge resources with a relevance reaching a specific standard are adopted as knowledge materials. Specifically, the server will filter out all knowledge resources with similarity scores higher than this threshold and use them as knowledge materials matching this data generation task. This mechanism effectively filters out a large number of irrelevant or weakly relevant information that may exist in the general knowledge base, ensuring that the selected knowledge materials are closely related to the risk control scenario and professional knowledge domain defined by the task profile, providing a reliable and rich knowledge context for generating high-quality, high-fidelity training samples.

[0033] S106: Based on the knowledge material and according to the format features, generate several risk control training samples through a preset large model. The risk control training samples are used to train the target large model. The target large model is used to receive business data and output risk control strategies in risk control business scenarios, so as to perform risk control processing on the business based on the risk control strategies.

[0034] Finally, based on step S104, having obtained the format features and knowledge materials, the server can generate several risk control training samples by generating a large model. In this embodiment of the specification, the generated large model is a teacher model, which can typically be a GPT-4 or equivalent large language model.

[0035] Specifically, the server can first construct sample-generated prompt information based on knowledge materials and format characteristics using a preset template. This template uses knowledge materials as the core content foundation and transforms the structural and style requirements defined in the format characteristics into an instruction language that the large model can accurately understand, thus combining them into a structured, unambiguous, and complete prompt.

[0036] For example, Knowledge Material A contains: "[Knowledge Fragment 1] Rule: Multiple small transactions to the same payee on a single day are a typical method of splitting transactions to circumvent regulation. [Knowledge Fragment 2] Case: User X initiated 15 transfers of 199 yuan each to account Y within 2 hours." The format characteristics stipulate: the input format is "Description of User Transaction Behavior," and the output format is a chain structure of "Risk Type Judgment - Key Feature Analysis - Handling Suggestions." The template is: "You are a risk control expert. Please generate a training sample based on the following professional knowledge: ... (Knowledge Material) Please generate a complete risk control sample based on the above knowledge. The input part is in the form of "...", simulating a specific transaction scenario that conforms to the above rules. The output part must be generated strictly according to the following format: 1, ... 2, ... ..." Then, the sample generation prompt information is input into the preset large generation model, which generates several sample materials based on the knowledge materials; and according to the format characteristics, each sample material is restructured to obtain each risk control training sample.

[0037] The process involves generating a large model based on the knowledge materials provided in the prompts, performing in-depth understanding and content generation, and outputting preliminary sample materials. Subsequently, these sample materials are restructured and standardized according to the structural specifications defined in the format features to ensure they meet the preset input and output format and style requirements. Finally, high-quality risk control training samples that can be directly used to train the target risk control large model are output.

[0038] In the embodiments of this specification, the generation of sample materials is a process of expanding and generating based on knowledge materials. The large-scale generation model understands the knowledge materials in the prompt information and creatively interprets them to generate substantive content. The large-scale generation model summarizes general risk control rules, risk patterns, and business logic from the knowledge materials and creates a concrete instance that conforms to real-world business scenarios. For example, the abstract rule of "split transaction" is concretized into a highly realistic description of transaction behavior that includes virtual users, time, amount, and payee.

[0039] Furthermore, it can construct a complete micro-narrative scenario around the core risk points, filling the generated samples with necessary contextual information and specific details. For example, when generating a "credit card cash advance" sample, it will specifically describe details such as merchant type, transaction time, and amount characteristics, making the sample content rich and realistic.

[0040] Additionally, by adjusting model parameters or variables in the prompts, the model can be driven to interpret the same knowledge material from multiple angles and in multiple scenarios, thereby generating multiple sample materials with diversity in specific plots, risk manifestations, or user attributes, ensuring the coverage of the training data. Alternatively, a large-scale generative model can be used to randomly generate multiple sample materials.

[0041] During the structured reconstruction and the generation of various risk control training samples, the large model can generate multiple risk control training samples from its own generated sample materials, according to the various format frameworks contained in the prompt information. In other words, the same content can be used to generate corresponding risk control training samples in different formats.

[0042] Furthermore, the large-scale model can precisely control the writing style and expression of the output text according to the style requirements in the format features, ensuring that it conforms to business norms. For example, it can uniformly adopt an objective, rigorous, and neutral report style, or use a specific professional terminology system.

[0043] The final output also conforms to the format characteristics, such as correctly using the prescribed headings, numbering, and separators, or generating structured data objects that fully conform to specific schemas such as JSON and XML.

[0044] Furthermore, in the examples provided in this specification, the risk control training samples are used for supervised or instructional fine-tuning of the target large-scale model to professionally train its cognitive and decision-making capabilities in a specific risk control domain. After training, the target large-scale model can be deployed in a real business environment to receive real-time business data streams in risk control business scenarios. Examples include user-submitted credit applications, ongoing payment transaction records, or pending insurance claim documents. Based on the complex patterns and risk control logic it has learned, the target large-scale model will perform in-depth analysis and real-time inference on the input business data and output structured risk control strategies.

[0045] This strategy can be specifically manifested as risk level assessment, fraud probability scoring, specific handling suggestions, and key decision-making basis. The business system will then use this automatically issued risk control strategy to perform corresponding risk control processing on current business requests or transactions, thereby achieving automated risk interception, risk warning, or risk classification management, ultimately ensuring business security and reducing losses.

[0046] based on Figure 1 The method for generating training data for the large-scale model, as shown, determines the risk control capabilities and knowledge domain description of the target large-scale model based on the data generation task requirements of the risk control business scenario. Core concepts and format requirements are then determined based on the defined task profile. The core concepts are used to acquire rich matching knowledge from a general knowledge base, and the format requirements are used to further process the acquired knowledge. Training samples are obtained by generating the large-scale model to train the target large-scale model for executing risk control business. This automated training sample generation process avoids the problems of relying on manual generation and annotation, improving efficiency, reducing quality differences between samples, and increasing the accuracy of the trained model.

[0047] Furthermore, in the embodiments of this specification, when the server determines the format characteristics based on the content elements in step S100, it can specifically match a pre-built business scenario with a format mapping library to determine the format characteristics suitable for the content elements. For example, when keywords such as "transaction flow analysis" and "relationship network" frequently appear in the content elements, the server can infer that the input format should be a structured list of transaction data or a network diagram, while the output format should be a description containing risk entity identifiers and risk propagation paths. Even in the absence of explicit instructions, this method ensures that the generated data conforms to the general paradigm and technical requirements of the business domain.

[0048] Alternatively, the server can determine format characteristics based on content elements through logical reasoning using a large model. Specifically, the large model can determine format characteristics based on task paradigm reasoning. For example, when content elements point to rule-based judgments, such as compliance review, the large model can infer a format suitable for classification tasks, such as input: [object to be reviewed, relevant rules], output: [compliance / violation, reason]. If content elements point to causal discovery, such as risk root cause analysis, it can infer a format suitable for attribution tasks, such as input: [event phenomenon], output: [root cause, chain of evidence]. Essentially, this derives a standard logical framework for problem-solving from the problem type.

[0049] Large models can also determine format characteristics based on structured reasoning rooted in domain knowledge. For example, for corporate credit risk assessment, a large model can infer the dimensions that need to be analyzed, such as the company's financial situation, industry prospects, and management information, thus constructing a structured input format that requires data for these dimensions. Similarly, the inferred output should be a comprehensive report format containing quantitative scores and comments for each dimension.

[0050] Furthermore, in the example provided in this specification, the data generation task also includes sample instances. These sample instances are complete, standardized, and validated high-quality data samples. One or more sample instances can be set, and these sample instances serve as concrete examples and quality benchmarks of the target training data, providing precise and concrete specifications that can be directly understood by machine learning models to meet the user's needs.

[0051] Each sample instance is a structured input-output pair. The input portion fully presents the complete information context required to process a specific risk control business segment. For example, in a credit anti-fraud scenario, the input portion of a sample instance might be an anonymous transaction record that retains key logical features, including fields such as user information, transaction time, amount, and merchant type.

[0052] The output section then displays the standard answer and reasoning process that a professional risk control model should provide for this input. It is not just a simple risk label, but also includes the complete logical chain for making this judgment. For example, the output section might explain in detail: "Step 1, this transaction is inconsistent with the user's historical spending patterns; Step 2, the receiving merchant is located in a high-risk area; based on the above, it is determined to be a high-risk transaction, and blocking is recommended." These sample instances can originate from classic cases in business history, exemplary samples carefully compiled by domain experts, or standard samples obtained by cleaning and de-identifying high-quality business data. Together, they constitute a miniature, high-quality "template library." By analyzing and learning from these instances, we can go beyond an abstract understanding of task profiles and accurately capture the data format, professional depth, and writing style expected by users. This ensures that the subsequent batch-generated training data maintains a high degree of consistency with these high-standard samples in terms of logical rigor, structural standardization, and business authenticity.

[0053] Specifically, the server can extract keywords for the data generation task from the risk control capability description and knowledge domain description using a semantic parsing model, and use them as content elements. This process is the same as the step of determining content elements in step S100, and will not be described again in this specification.

[0054] The server can also determine several format features based on each sample instance. Format features include at least structural features and style features. Among them, structural features include input format and output format, and style features include writing style and expression method.

[0055] For each sample instance, the server can use natural language processing techniques to determine structural and stylistic features from that sample instance. Based on the structural and stylistic features determined for each sample instance, common structural and stylistic features are abstracted.

[0056] The structural features include, but are not limited to, the organization of input information, such as background before data, and the logical structure of output results, such as judgment before reasoning. The style features encompass the density of technical terms, sentence complexity, and the expression style of the example text, such as formal reporting tone or concise instructional tone. This process no longer relies on pre-built mapping libraries or indirect logical inference, but rather directly and objectively extracts format specifications that meet user expectations through inductive learning from examples, thereby ensuring a high degree of consistency between the generated samples and the example samples in both "form" and "spirit."

[0057] It should be noted that the structural and stylistic features in the examples in this manual can be regarded as the basic and content requirements for the input and output text.

[0058] Structural features refer to the framework specifications for data organization and logical composition abstracted from sample instances. They can be considered to define the "skeleton" of the training samples. This includes: input format, output format, and element relationships. Input format refers to the organization and composition rules of input information; for example, is it a single text query or a structured form composed of multiple fields such as "user profile," "transaction history," and "real-time behavior"? Output format refers to the logical composition and presentation sequence of the output results; for example, is it a simple classification label or a strict sequence and component structure following "risk assessment, key evidence listing, reasoning chain explanation, and disposal recommendations"? Element relationships refer to the logical connections between input and output, and between the components within the output; for example, each piece of evidence in the output must strictly correspond to specific information in the input.

[0059] Stylistic features refer to the textual attributes regarding language expression and professional paradigms extracted from sample instances. They can be considered the "flesh and blood" of the training samples. These include writing style and mode of expression. Writing style refers to the overall linguistic style of the text; for example, is it a rigorous, objective, and neutral formal report style, concise and direct directive language, or conversational language that includes guiding thought processes? Mode of expression refers to specific usage habits at the lexical and syntactic levels, such as the use of specific domain-specific terminology, fixed sentence structures, specific levels of detail, and the emotional tone of the text.

[0060] Structural features ensure the logical correctness and machine readability of the generated data, while stylistic features ensure the professionalism and human acceptability of the generated data in terms of language expression. Together, they precisely define the quality standards that the training samples should achieve from both formal and substantive dimensions.

[0061] Furthermore, using complete knowledge resources from the general knowledge base as knowledge material in step S104 still has some drawbacks. For example, complete documents generally suffer from low information density and noise interference. A single document typically contains a large amount of redundant content unrelated to the current specific generation task, such as background introductions, transitional paragraphs, or chapters on irrelevant topics. This redundant information dilutes the concentration of key knowledge, introduces noise in subsequent generation processes, interferes with the attention of the generation model, and makes it difficult for it to focus on core business logic and key facts, potentially resulting in low-quality samples that deviate from the topic or contain irrelevant details.

[0062] Furthermore, the knowledge organization structure of a complete document is highly likely to be mismatched with the requirements of the generation task. The document's chapter structure is designed for human linear reading and comprehension, rather than being optimized to serve the machine generation of training samples in a specific format. This structural mismatch makes it difficult for the knowledge in the document to be directly and efficiently utilized by the generation model, and it cannot be accurately mapped to the input-output pair format and logical chain required by the training samples, severely restricting the accuracy and efficiency of the generation process.

[0063] Finally, the knowledge granularity of a complete document is too coarse, failing to effectively align with the generation target. Generating a high-quality training sample typically requires only one or more highly focused knowledge units centered around a specific core concept or scenario. However, a complete document may cover multiple loosely related topics, resulting in an overly broad semantic scope and blurred focus for individual knowledge materials. This inconsistency in granularity makes the generation process difficult to control, failing to guarantee that each produced training sample possesses the characteristics of a clear theme and logical consistency. Therefore, in one or more embodiments of this specification, in step S104, the server may first determine knowledge resources matching the content elements from a general knowledge base based on the obtained similarity of each material. The server may apply a preset similarity threshold to filter based on the calculated similarity of each knowledge resource. A batch of knowledge resources related to the current data generation task are initially filtered from the general knowledge base to form a high-quality candidate resource pool.

[0064] Next, knowledge fragments semantically related to the content elements are extracted from the knowledge resources. Each initially selected knowledge resource is typically a complete document. The server can further utilize natural language processing technology to identify and extract knowledge fragments strongly related to the content elements through semantic parsing and content extraction. These knowledge fragments are semantically relatively independent parts of the document that revolve around a specific theme. For example, a definitional description of a key risk control concept, a specific explanation of a business rule, or a detailed analysis of a typical scenario case. This process of filtering from "document-level" information to "proposition-level" information significantly improves the purity and density of knowledge.

[0065] Then, the knowledge fragments are aggregated and structured to generate knowledge materials. The knowledge fragments extracted by the server are still independent and may lack necessary background and logical connections. For example, fragment A describes a certain fraud method, fragment B describes the typical characteristics of the method, and fragment C provides corresponding countermeasures. If these fragments are used directly, the generative model may not understand the inherent connections between them, resulting in logically incoherent or incomplete information in the generated samples. Aggregation processing can reorganize these scattered but related knowledge points into a semantically coherent and context-complete knowledge unit.

[0066] Therefore, the server can integrate fragments describing the same topic by identifying semantic relationships between them. Simultaneously, it reorganizes and arranges the content according to a preset logical template, generating knowledge materials with a unified structure and clear logic. This transforms originally loose, unstructured text content into high-quality, structured knowledge materials suitable for the generative model to understand and utilize.

[0067] Of course, integrating knowledge fragments also provides an opportunity for deduplication and consistency verification. The server can identify and merge identical or similar statements, while detecting and eliminating contradictory information, thereby ensuring that the generated knowledge material is internally consistent, authoritative, and reliable, thus improving the quality and credibility of subsequently generated samples from the source.

[0068] Moreover, unlike knowledge resources, knowledge fragments are generally unstructured. Therefore, through structuring processing during the aggregation process, the server can integrate semantically related fragments into a complete content whole. Through structuring processing, the server endows these integrated contents with a logically rigorous framework and standardized organizational form according to the target format required by the generation task. This transforms disorganized knowledge fragments into high-quality "semi-finished" knowledge materials that can be directly used by the generation model, ensuring the standardization and consistency of the generated samples in terms of logical structure and output format.

[0069] In one or more embodiments of this specification, when training the target large model, the server can update the risk control training samples according to the difficulty of the training process. This is to avoid the risk control training samples being too difficult, which could lead to model learning stagnation, abnormal gradients, or overfitting to complex patterns. It also avoids the risk control training samples being too easy, which could prevent the model from effectively improving its capabilities or cause it to saturate on simple patterns, resulting in wasted training resources. Specifically, firstly, the server can input each risk control training sample into the target large model and determine the output of the target large model.

[0070] Secondly, the annotations and outputs of the risk control training sample are compared to determine if the output matches the annotations. Specifically, when the output is a risk control judgment, matching the annotation means the output is strictly consistent with the annotation. Alternatively, if the output is a piece of text, matching the annotation means it matches the above-mentioned output and annotation. Furthermore, if the output also includes a reasoning process, it can be determined whether the reasoning process matches the reasoning process contained in the annotation. This can include consistency in the order of reasoning steps and consistency in the content of the reasoning steps. In short, if the output is correct, it indicates that the target large model has processed it correctly; otherwise, it indicates that the target large model has not processed it correctly.

[0071] Once the target large-scale model is confirmed to be correctly processed, the risk control training sample is used to generate a larger model to complicate the problem, increasing the sample's difficulty. This processed sample is then used to train the target large-scale model. In other words, the original risk control training sample is no longer sufficient to improve the target large-scale model; therefore, it can be augmented by generating a larger model. Specifically, this includes adding inference steps to the risk control training sample through the larger model, introducing interference factors from real-world business scenarios, increasing the level of abstraction of the problem, or merging multiple risk control concepts to form complex problems. This strategy of increasing difficulty ensures that the training content remains at the forefront of the model's capabilities, effectively promoting the continuous advancement of its inference abilities.

[0072] When it's determined that the target large-scale model is not processing correctly, the risk control training sample is simplified by generating a large-scale model to reduce the sample's difficulty. The target large-scale model is then trained again using this simplified sample. In other words, the original risk control training sample cannot achieve the effect of "improving" the target large-scale model; therefore, the sample is reduced by generating a large-scale model to match the capabilities of the target large-scale model. Intelligent simplification of the original sample through the large-scale model includes measures such as: decomposing complex problems into multiple sub-problems, providing hints for key reasoning steps, simplifying the problem's expression, or supplementing necessary background knowledge. This precise difficulty adjustment ensures that the risk control training sample matches the current cognitive level of the target large-scale model, laying the foundation for building a robust risk control knowledge system.

[0073] In one or more embodiments of this specification, the above-described sample enhancement process generates a large model by parsing the solution logic of the original sample, identifying the inference nodes that can be expanded, and constructing a multi-step inference chain on this basis.

[0074] Specifically, the server determines the first input data based on the preset first template and the risk control training sample. That is, the preset prompt words are concatenated to obtain the additional input that needs to be added for the intermediate inference steps.

[0075] Subsequently, using the generated large model, based on the first input data, at least one intermediate inference problem is inserted into the risk control training sample, and the inference steps corresponding to the intermediate inference problem are added to the annotations to obtain the processed risk control training sample.

[0076] The large model is generated based on this input data, by inserting at least one intermediate reasoning question into the problem description of the original sample. These intermediate questions are used to guide the target large model to perform step-by-step reasoning. For example, in a fraud detection scenario, an intermediate question such as "Does the transaction amount match the user's historical behavior pattern?" is added before the judgment.

[0077] Simultaneously, the generated large model adds corresponding inference step descriptions to the sample annotations, detailing the analysis process and judgment basis for each intermediate problem, thereby constructing a processed risk control training sample containing a multi-level inference chain. This approach, by explicitly decomposing the decision-making process of complex problems, effectively enhances the inference depth and logical rigor of the risk control training sample.

[0078] This generative model can decompose the original annotation into multiple interdependent intermediate judgments, and then add necessary transition statements and logical connectors to each intermediate step. For example, the simple "determine it as a fraudulent transaction" can be expanded into a complete reasoning process: "Step 1: Identify abnormal matches between the transaction amount and the merchant's business scope; Step 2: Verify significant deviations between the transaction time and the user's historical behavior patterns; Step 3: Integrate multiple abnormal features to arrive at the fraud determination."

[0079] In one or more embodiments of this specification, the above-described sample enhancement process generates a large model that transforms specific cases into general problems requiring concept transfer through semantic generalization and pattern abstraction.

[0080] Specifically, the server can determine the second input data based on the preset second template and the risk control training samples. The second template is the prompt word for generating a large model for semantic generalization or a higher-level prompt.

[0081] Then, by generating a large model, based on the second input data, the specific instances in the risk control training sample are replaced with the corresponding higher-level descriptions, resulting in a risk control training sample with increased conceptual complexity.

[0082] Based on the input data, the generative model identifies specific business instances in the samples and uniformly replaces them with corresponding higher-level conceptual expressions. For example, the specific threshold of "a single transaction exceeding 50,000 yuan" is replaced with the abstract expression "a single transaction exceeding a set risk threshold," and the specific time period of "transactions from 11 PM to 5 AM" is replaced with the pattern description of "transactions during non-typical user activity periods." This approach enhances the conceptual abstraction level of the samples, shifting the training samples from specific case learning to general pattern recognition, thereby effectively improving the target model's understanding of the essential laws of risk control business and its generalization performance.

[0083] In one or more embodiments of this specification, the above-described sample enhancement process allows the large model to inject interference information that conforms to real-world logic but is irrelevant to core judgments into the risk control training samples based on domain knowledge.

[0084] Specifically, the server can determine the third input data based on the preset third template and the risk control training sample. The third template guides the generation of a large model to inject interference factors that conform to real business logic into the risk control training sample, thereby systematically improving the complexity and realism of the sample.

[0085] Subsequently, by generating a large model, the input statements in the risk control training samples are rewritten and expanded based on the third input data, resulting in risk control training samples with increased grammatical complexity. For example, normal cross-time zone transaction records are added to transaction risk control samples, non-critical liability information that does not affect credit scores is inserted into credit approval samples, or seemingly suspicious but actually compliant fund transfers are mixed into anti-money laundering cases. These interfering factors maintain the authenticity of the business, forcing the model to develop the ability to distinguish key signals from background noise.

[0086] Furthermore, this generative large model can also construct complex multi-dimensional cross-cutting scenarios by identifying potential correlations between different risk control dimensions.

[0087] Specifically, the server can determine the interactive input data based on a preset interactive template and the risk control training sample. By generating a large model, processed risk control training samples are produced based on the interactive input data. For example, transaction fraud characteristics can be combined with identity theft to create a complex fraud scenario; the ability to fulfill obligations in credit assessment can be linked to the impact of sudden public health events; or the fund flow analysis in anti-money laundering can simultaneously involve multiple regulatory areas such as cross-border trade and virtual currencies. This integration requires the model to coordinate multiple knowledge modules simultaneously, establishing a comprehensive cross-domain decision-making capability.

[0088] Through the above specific implementation methods, the generation of large models systematically increases the cognitive difficulty of training samples, and promotes the evolution of target models from single pattern recognition to advanced capabilities such as complex reasoning, anti-interference analysis and conceptual abstraction.

[0089] Similarly, in the examples in this specification, the server can also use a similar method to simplify the risk control training samples.

[0090] Specifically, the fourth input data is determined based on the preset fourth template and the risk control training sample.

[0091] Subsequently, by generating a large model, based on the fourth input data, the reasoning problem contained in the input of the risk control training sample and the reasoning steps corresponding to the reasoning problem in the annotation are determined. The reasoning problem and reasoning steps are then reduced to obtain the processed risk control training sample.

[0092] Based on this input data, the large-scale model performs structured analysis on the risk control training samples, identifying the complex reasoning problems contained in the input and their corresponding multi-step reasoning processes in the annotations. Through semantic understanding and logical analysis, the large-scale model eliminates unnecessary reasoning steps, simplifying complex problems into core problems while retaining key judgment elements and basic logic. For example, a complex fraud detection sample containing multiple intermediate reasoning steps is simplified into a basic judgment sample retaining only key risk features. This approach effectively reduces the difficulty of understanding and solving the samples, making them more suitable for the current learning stage of the target large-scale model, thereby establishing a gradual learning path.

[0093] In the example provided in this manual, the server performs problem simplification processing on the risk control training samples.

[0094] Specifically, the server can determine the fifth input data based on the preset fifth template and the risk control training sample.

[0095] Then, by generating a large model, based on the fifth input data, the higher-level descriptions in the risk control training sample are replaced with specific instances to obtain a risk control training sample with simplified conceptual complexity.

[0096] Based on this input data, the large-scale model automatically identifies abstract expressions and higher-level concepts in the risk control training samples and replaces them with concrete instances with clear business meaning. For example, the abstract expression "transaction amount exceeds the set risk threshold" is replaced with the specific value "transaction amount exceeds 50,000 yuan"; the pattern description of "transactions during non-typical user activity periods" is replaced with the specific time range of "transactions between 2 AM and 4 AM". This transformation from abstract to concrete effectively reduces the conceptual complexity of the samples, making the elements and boundaries of the risk control scenario clearer and more explicit. This helps the target large-scale model to quickly establish a business cognitive framework, laying a solid foundation for subsequent learning of higher-level abstract concepts.

[0097] In the example provided in this manual, the server performs problem simplification processing on the risk control training samples.

[0098] Specifically, the server can determine the sixth input data based on the preset sixth template and the risk control training sample.

[0099] Then, by generating a large model, the input statements in the risk control training sample are split into sentences based on the sixth input data to obtain multiple sub-problems that conform to the logical order. A parsing guide text is added to each sub-problem to obtain a risk control training sample with reduced sentence complexity.

[0100] Based on this input data, the large-scale generative model performs semantic parsing and logical decomposition on complex expressions in the risk control training samples, breaking down compound sentences containing multiple pieces of information into multiple sub-questions that conform to cognitive order. For example, the complex question "Please analyze whether this transaction has money laundering risk and explain the basis for your judgment" is decomposed into three logically coherent sub-questions: "Step 1: Identify the basic characteristics of both parties to the transaction," "Step 2: Analyze the rationality of the fund transfer path," and "Step 3: Comprehensively assess the level of money laundering risk." Simultaneously, the large-scale generative model adds specific analytical guidance text to each sub-question, such as operational instructions like "Please focus on comparing the matching degree between the industry of the transaction parties and the amount." This approach, by reducing the cognitive load of a single inference and establishing a progressive analytical path, effectively improves the comprehensibility and learnability of the training samples, effectively helping the target large-scale model build a systematic risk analysis mindset.

[0101] In the embodiments described in this specification, the aforementioned adaptive difficulty adjustment mechanism supports multiple rounds of iterative optimization to continuously improve the quality of training data. The server achieves dynamic tracking of the learning state of the target large model and precise control of the training data by repeatedly executing a closed-loop process of "model evaluation - sample rewriting - retraining." In each iteration, the server can reconstruct the sample data using appropriate complex or simplified templates based on the model capability evaluation results, ensuring that the training content always matches the model's current cognitive level. Alternatively, a data reconstruction can be performed after multiple iterations.

[0102] When the server detects that the target large model's mastery of samples of various difficulty levels no longer changes significantly in consecutive iterations, indicating that its overall performance has stabilized, it will automatically trigger a convergence judgment mechanism to terminate the current iteration optimization process. The server integrates risk control training samples, difficulty-increasing data generated in each iteration, and difficulty-decreasing data to construct a composite training dataset covering multiple difficulty levels and possessing a smooth gradient distribution. This dataset provides good coverage of basic concepts while also offering sufficient challenges for advanced capabilities, laying a data foundation for training a risk control large model with strong generalization ability and robustness.

[0103] Furthermore, traditional training methods treat all samples equally, ignoring the differences in model capabilities at different learning stages. Therefore, in addition to adjusting the complexity of risk control training samples, this server can also select suitable training samples during the training process to continue training.

[0104] Specifically, firstly, for each risk control training sample, a predetermined number of random inferences are performed using the target large model. This process introduces randomness to ensure the independence of each inference, thereby obtaining multiple output results from the target large model for that risk control training sample. These results are used to evaluate the stability and consistency of the target large model's output, providing foundational data for subsequent pass rate calculations. This is similar to the multiple inference steps in high-potential sample identification, where repeated sampling captures the model's behavior under different random states.

[0105] Secondly, the pass rate of the risk control training sample is determined based on the number of correct outputs from the target large model. The server can count the number of correct outputs from the target large model and calculate the ratio of this number to the total number of inferences to determine the pass rate of the risk control training sample. The pass rate quantifies the target large model's mastery of the current sample, reflecting the relative difficulty and value of the risk control training sample, and serves as a key basis for sample selection. The pass rate calculation process ensures the objectivity and statistical significance of the evaluation, thereby accurately identifying samples that the target large model has not yet fully mastered but have learning potential.

[0106] Then, according to the preset screening interval, the risk control training samples that fall within the screening interval are used to continue training the target large model. Risk control training samples falling within this interval are selected for subsequent training. This screening mechanism, based on a pass rate threshold, can filter out samples that are too simple or too difficult, ensuring that the training set contains high-potential samples, such as retaining risk control training samples with a pass rate between 30% and 70%.

[0107] Meanwhile, by controlling the number of samples and ensuring diversity, such as covering different question types, difficulties, and knowledge points, the quality of training data is optimized. The selected samples will be used in the reinforcement learning training phase, where algorithms such as GRPO will be used for policy optimization to improve the performance and convergence efficiency of the target large model in risk control scenarios.

[0108] Alternatively, risk control training samples with an initial pass rate in the range of 0-50% can be selected as high-potential samples, and these high-potential samples can be used for training in subsequent training processes.

[0109] Among them, high-potential samples indicate the current capability boundaries and cognitive weaknesses of the model. Their value lies in clearly revealing key knowledge points and reasoning patterns that the model has not yet mastered but has the potential to learn. By concentrating training resources on samples in this specific range, the model can be guided to confront its core deficiencies and achieve targeted breakthroughs in its capabilities.

[0110] Training with high-potential samples is essentially a strategy focused on breakthrough progress. Compared to traditional uniform sampling, this strategy significantly improves the density and effectiveness of training signals, ensuring that each parameter update of the model is dedicated to solving challenging problems. This not only accelerates the model's ability to advance in complex risk control scenarios but also, through synergy with advanced training strategies such as reinforcement learning, drives the model's output from randomness to stability and reliability. While ensuring training economy, it systematically improves the model's peak performance and generalization robustness.

[0111] Figure 2 The diagram above illustrates the training process of the risk control model provided in the various embodiments. The left side shows the input task profile, and various risk control training samples are obtained through data synthesis. During the training process, the difficulty of the risk control training samples can be adjusted through the teacher model or the target large model to obtain a large number of samples covering various difficulties. Then, during the training of the target large model, the samples used in the next round or several rounds of iteration are selected according to the preset pass rate range.

[0112] In the embodiments described in this specification, the training scheme has high flexibility and modularity, supporting the use of differentiated optimization strategies at different training stages.

[0113] In the initial and intermediate stages, the system can directly perform multiple rounds of supervised fine-tuning on the target large model based on the generated risk control training samples. This process can quickly establish the model's basic understanding of risk control tasks, enabling it to initially grasp the core concepts, rules, and judgment logic in the business scenario, laying a solid knowledge foundation for subsequent refined optimization.

[0114] After the model acquires basic capabilities, this solution can further enable the reinforcement learning training phase to achieve a breakthrough in performance. At this stage, the system uses the previously selected high-potential samples (pass rate 0-50%) or suitable-potential samples (pass rate 30-70%) as the core materials for the reinforcement learning environment. These samples represent the boundaries of the model's current capabilities, and the challenging scenarios they provide effectively drive reinforcement learning algorithms (such as GRPO) to deeply optimize the model's policy. By allowing the model to explore different output policies on these "difficult" and "uncertain" samples, and to receive rewards or penalties according to preset quality standards, reinforcement learning training can finely calibrate the model's inference path and output style, guiding it from "basically correct" to "stable, reliable, and excellent" decisions.

[0115] This two-stage design, which "lays the foundation with supervised fine-tuning and achieves refinement through reinforcement learning," combines the advantages of both training paradigms. It ensures the stability and efficiency of training while fully exploring the performance potential of the model, resulting in a large language model that performs better in real-world risk control operations.

[0116] Alternatively, the teacher model can be used when adjusting the complexity of intermediate samples, and the target large model can be used only when determining the pass rate.

[0117] The above is an embodiment of a method for generating training data for a large model provided in this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.

[0118] Figure 3 This is a schematic diagram of a training data generation device for a large model provided in an embodiment of this specification. The device includes: The receiving module 401 is used to receive a data generation task for a risk control business scenario. The data generation task carries a task profile of the target large model to be trained. The task profile includes at least a description of the risk control capabilities and a description of the knowledge domain of the target large model. Extraction module 402 is used to determine the content elements and format characteristics of the data generation task through the task profile; Enrichment module 403 is used to determine knowledge materials that match the content elements from a preset general knowledge base; The generation module 404 is used to generate several risk control training samples based on the knowledge materials and according to the format features through a preset generation model. The risk control training samples are used to train the target model. The target model is used to receive business data and output risk control strategies in risk control business scenarios, so as to perform risk control processing on the business based on the risk control strategies.

[0119] Optionally, the extraction module 402 is used to extract keywords of the data generation task from the risk control capability description and knowledge domain description through a semantic parsing model as content elements; and to determine format features based on the content elements and / or the task profile, wherein the format features include at least the input format and the output format.

[0120] Optionally, the data generation task also includes sample instances; The extraction module 402 is used to extract keywords of the data generation task from the risk control capability description and knowledge domain description through a semantic parsing model, as content elements; and to determine several format features based on the sample instances. The format features include at least structural features and style features. The structural features include input format and output format, and the style features include writing style and expression method.

[0121] Optionally, the enrichment module 403 is used to perform semantic similarity calculation between the knowledge resources in the preset general knowledge base and the element content; and select knowledge resources that match the element content from each knowledge resource according to the obtained similarity of each knowledge resource and a preset similarity threshold, as knowledge materials.

[0122] Optionally, the enrichment module 403 is used to determine knowledge resources matching the content element from the general knowledge base based on the similarity of the obtained materials; extract knowledge fragments semantically related to the content element from the knowledge resources; and aggregate and structure the knowledge fragments to generate knowledge materials.

[0123] Optionally, the generation module 404 is configured to construct sample generation prompt information based on the knowledge materials and the format features using a preset template; input the sample generation prompt information into a preset generation model, so that the generation model generates several sample materials based on the knowledge materials; and restructure each sample material according to the format features to obtain each risk control training sample.

[0124] Optionally, the generation module 404 is configured to, for each risk control training sample, input the risk control training sample into the target large model, determine the output of the target large model; compare the label of the risk control training sample with the output to determine whether the output matches the label; if yes, it is determined that the target large model has processed correctly, and the risk control training sample is processed by the generated large model to complicate the problem, thereby increasing the sample difficulty, and the target large model is continued to be trained using the processed sample; if no, it is determined that the target large model has not processed correctly, and the risk control training sample is processed by the generated large model to simplify the problem, thereby reducing the sample difficulty, and the target large model is continued to be trained using the processed sample.

[0125] Optionally, the generation module 404 is used to determine first input data based on a preset first template and the risk control training sample; and through the generation large model, based on the first input data, insert at least one intermediate inference problem into the risk control training sample, and add the inference steps corresponding to the intermediate inference problem in the annotation, to obtain the processed risk control training sample.

[0126] Optionally, the generation module 404 is used to determine the second input data based on the preset second template and the risk control training sample; and through the generation of the large model, based on the second input data, replace the specific instances in the risk control training sample with the corresponding higher-level descriptions to obtain a risk control training sample with increased conceptual complexity.

[0127] Optionally, the generation module 404 is used to determine the third input data based on the preset third template and the risk control training sample; and through the generation large model, based on the third input data, to expand and rewrite the input expression in the risk control training sample to obtain a risk control training sample with increased sentence complexity.

[0128] Optionally, the generation module 404 is used to determine the fourth input data based on the preset fourth template and the risk control training sample; through the generation large model, based on the fourth input data, determine the reasoning problem contained in the input of the risk control training sample, and the reasoning steps corresponding to the reasoning problem in the annotation, and delete the reasoning problem and reasoning steps to obtain the processed risk control training sample.

[0129] Optionally, the generation module 404 is used to determine the fifth input data based on the preset fifth template and the risk control training sample; and through the generation of the large model, based on the fifth input data, replace the higher-level description in the risk control training sample with specific instances to obtain a risk control training sample with simplified conceptual complexity.

[0130] Optionally, the generation module 404 is used to determine the sixth input data based on the preset sixth template and the risk control training sample; through the generation large model, based on the sixth input data, the input expression in the risk control training sample is split into statements to obtain multiple sub-questions that conform to the logical order, and parsing guidance text is added to each sub-question to obtain a risk control training sample with reduced statement complexity.

[0131] Optionally, the generation module 404 is used to perform a preset number of random inferences on each risk control training sample using the target large model; determine the pass rate of the risk control training sample based on the number of correct output results of the target large model; and continue to use the risk control training samples that fall into the preset screening interval to train the target large model according to the preset screening interval.

[0132] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can be used to perform the training data generation method for the large model described above.

[0133] based on Figure 1 The method for generating training data for the large model shown in this specification is further illustrated in the embodiments. Figure 4 The diagram shows the structure of the electronic device. Figure 4At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the aforementioned method for generating training data for large models.

[0134] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A method for training data generation of a large model, the method comprising: receiving a data generation task for a risk control business scenario, the data generation task carrying a task profile of a target large model to be trained, the task profile including at least a risk control capability description and a knowledge field description of the target large model; determining content elements and format features of the data generation task through the task profile; determining knowledge materials matching the content elements from a preset general knowledge base; generating a plurality of risk control training samples through a preset generation large model according to the format features based on the knowledge materials, the risk control training samples being used to train the target large model, the target large model being used to receive business data in the risk control business scenario and output a risk control strategy to perform risk control processing on a business based on the risk control strategy.

2. The method of claim 1, wherein the content elements and format features of the data generation task are determined through the task profile, specifically comprising: extracting keywords of the data generation task from the risk control capability description and the knowledge field description as content elements through a semantic analysis model; determining format features based on the content elements and / or the task profile, the format features including at least input format and output format.

3. The method of claim 1, wherein the data generation task further carries sample instances; wherein the content elements and format features of the data generation task are determined through the task profile, specifically comprising: extracting keywords of the data generation task from the risk control capability description and the knowledge field description as content elements through a semantic analysis model; determining a plurality of format features based on the sample instances, the format features including at least structural features and style features, the structural features including input format and output format, and the style features including writing style and expression manner.

4. The method of claim 1, wherein the knowledge materials matching the content elements are determined from the preset general knowledge base, specifically comprising: performing semantic similarity calculation on the element content based on knowledge resources in the preset general knowledge base; selecting knowledge resources matching the element content from the knowledge resources as knowledge materials according to a preset similarity threshold based on the obtained similarity of the knowledge resources.

5. The method of claim 4, wherein the knowledge materials matching the element content are selected from the knowledge resources according to the obtained similarity of the knowledge resources and the preset similarity threshold, specifically comprising: determining knowledge resources matching the content elements from the general knowledge base according to the obtained similarity of the knowledge resources; extracting knowledge segments semantically related to the content elements from the knowledge resources; performing aggregation and structuralization processing on the knowledge segments to generate knowledge materials.

6. The method of claim 1, wherein the plurality of risk control training samples are generated through a preset generation large model according to the format features based on the knowledge materials, specifically comprising: According to the knowledge material and the format feature, a sample generation prompt information is constructed through a preset template; The sample generation prompt information is input into a preset generation large model, so that the generation large model generates a plurality of sample materials according to the knowledge material, and each sample material is structured and reconstructed according to the format feature to obtain each risk control training sample.

7. The method of claim 1, wherein the target large model is trained using the generated risk control training samples, specifically comprising: For each risk control training sample, the risk control training sample is input into the target large model to determine the output of the target large model; The output is compared with the label of the risk control training sample to determine whether the output matches the label; If yes, it is determined that the target large model correctly processes, the risk control training sample is processed through the generation large model to complicate the problem and increase the difficulty of the sample, and the target large model is continuously trained through the processed sample; If not, it is determined that the target large model does not correctly process, the risk control training sample is processed through the generation large model to simplify the problem and reduce the difficulty of the sample, and the target large model is continuously trained through the processed sample.

8. The method of claim 7, wherein the risk control training sample is processed through the generation large model to complicate the problem, specifically comprising: determining first input data according to a preset first template and the risk control training sample; inserting at least one intermediate reasoning question in the risk control training sample and adding a reasoning step corresponding to the intermediate reasoning question in the label through the generation large model according to the first input data to obtain a processed risk control training sample.

9. The method of claim 7, wherein the risk control training sample is processed through the generation large model to complicate the problem, specifically comprising: determining second input data according to a preset second template and the risk control training sample; replacing a specific instance in the risk control training sample with a corresponding general expression through the generation large model according to the second input data to obtain a risk control training sample with improved concept complexity.

10. The method of claim 7, wherein the risk control training sample is processed through the generation large model to complicate the problem, specifically comprising: determining third input data according to a preset third template and the risk control training sample; rewriting the input expression in the risk control training sample through the generation large model according to the third input data to obtain a risk control training sample with improved sentence complexity.

11. The method of claim 7, wherein the risk control training sample is processed through the generation large model to simplify the problem, specifically comprising: determining fourth input data according to a preset fourth template and the risk control training sample; determining a reasoning question contained in the risk control training sample input through the generation large model according to the fourth input data, and deleting the reasoning question and a reasoning step corresponding to the reasoning question in the label to obtain a processed risk control training sample.

12. The method of claim 7, wherein the risk control training sample is processed by the large model to simplify the question, and the processing specifically comprises: determining fifth input data according to a preset fifth template and the risk control training sample; and replacing a general expression in the risk control training sample with a specific instance according to the fifth input data by the large model, to obtain a risk control training sample with simplified concept complexity.

13. The method of claim 7, wherein the risk control training sample is processed by the large model to simplify the question, and the processing specifically comprises: determining sixth input data according to a preset sixth template and the risk control training sample; and splitting an input expression in the risk control training sample into a plurality of sub-questions in a logical order and adding an analysis guide text to each sub-question according to the sixth input data by the large model, to obtain a risk control training sample with reduced sentence complexity.

14. The method of claim 7, further comprising: performing random inference for each risk control training sample a preset number of times by the target large model; determining a passing rate of the risk control training sample according to a number of times that the target large model outputs a correct result; and continuing to use a risk control training sample falling within a preset screening interval to train the target large model.

15. A device for generating training data of a large model, the device comprising: a receiving module configured to receive a data generation task for a risk control business scenario, the data generation task carrying a task profile of a target large model to be trained, the task profile including at least a risk control capability description and a knowledge field description of the target large model; an extraction module configured to determine content elements and format features of the data generation task based on the task profile; a dressing module configured to determine knowledge materials matching the content elements from a preset general knowledge base; and a generation module configured to generate a plurality of risk control training samples based on the knowledge materials and according to the format features by a preset large model, the risk control training samples being used to train the target large model, the target large model being used to receive business data and output a risk control strategy in the risk control business scenario, and the risk control strategy being used to perform risk control processing on a business.

16. A computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the method of any one of claims 1-14.

17. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the method of any one of claims 1-14 when executing the program. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​