Method, device and equipment for generating large model training data
By generating synthetic data from a multimodal large model expert library and performing multi-level filtering, the problem of high cost and low efficiency of training data in the field of security risk identification is solved, achieving high-quality and efficient training data generation and improving the performance of large language models.
Patent Information
- Application Number
- CN202511843177.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, the generation of training data for large language models in the field of security risk identification is costly, inefficient, and of poor quality, which affects model performance.
Using a multimodal large model expert library, synthetic data is generated through three steps: image description, reasoning process, and review results. High-quality target synthetic data is selected by using multi-level filtering based on text perplexity, image-text alignment, and attribution logic score.
It improved the quality and generation efficiency of training data, reduced generation costs, and enhanced the performance of large language models.
Smart Images

Figure CN121708415A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a method, apparatus, and device for generating large model training data. Background Technology
[0002] With the rapid development of Large Language Models (LLMs), these models have been applied to various fields, such as healthcare, financial services, education, and security risk identification. When training these LLMs, the quality and quantity of training data directly determine their performance. The better the quality and the greater the quantity of training data, the higher the performance of the LLM.
[0003] In the existing technology, taking the field of security risk identification as an example, training data is generated by simply manually annotating a large number of risk images. This manual annotation method is not only costly and inefficient, but also produces poor-quality training data, which affects the performance of large language models.
[0004] In summary, improving the quality and efficiency of training data generation while reducing generation costs are the problems that need to be solved. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, apparatus, and device for generating large model training data, which can improve the quality and generation efficiency of training data and reduce the generation cost of the training data.
[0006] In a first aspect, embodiments of the present invention provide a method for generating large model training data, the method comprising: acquiring multiple raw data, wherein the raw data includes raw images and classification labels; for each raw data, determining an expert model corresponding to the raw data in a large model expert database based on the classification labels; inputting the raw images into the expert model corresponding to the raw data, and outputting synthetic data; performing risk point identification on the multiple synthetic data, and determining multiple first candidate synthetic data with successfully identified risk points; determining the text perplexity of the multiple first candidate synthetic data, and determining multiple second candidate synthetic data whose text perplexity is less than a set perplexity threshold; determining the image-text alignment of the multiple second candidate synthetic data, and determining multiple third candidate synthetic data whose image-text alignment is greater than a set image-text alignment threshold; determining the attribution logic score of the multiple third candidate synthetic data, and determining multiple target synthetic data whose attribution logic score is greater than a set attribution logic threshold.
[0007] Optionally, the method further includes: training a high-risk model based on the multiple target synthetic data and their corresponding original images.
[0008] Optionally, the step of inputting the original image into the expert model corresponding to the original data and outputting synthesized data specifically includes: inputting the original image into the expert model corresponding to the original data to generate an image description; inputting the image description into the expert model to generate a reasoning process; inputting the reasoning process into the expert model to generate an audit result; and determining the image description, the reasoning process, and the audit result as the synthesized data.
[0009] Optionally, the step of identifying risk points in multiple synthetic data and determining multiple first candidate synthetic data with successfully identified risk points specifically includes: matching each synthetic data with its corresponding classification label according to a set rule; in response to a successful match, determining that the risk point identification of the synthetic data is successful; and determining the multiple synthetic data with successfully identified risk points as multiple first candidate synthetic data.
[0010] Optionally, determining the text perplexity of the plurality of first candidate synthetic data specifically includes: calculating the text perplexity of the plurality of first candidate synthetic data based on an open-source large model.
[0011] Optionally, determining the image-text alignment of the plurality of second candidate synthetic data specifically includes: calculating the image-text alignment between the plurality of second candidate synthetic data and their corresponding original images based on a contrastive language image pre-trained CLIP model.
[0012] Optionally, determining the attribution logic scores of the multiple third candidate synthetic data specifically includes: scoring the reasoning process of the multiple third candidate synthetic data according to the open-source large model, and determining the attribution logic scores of the multiple third candidate synthetic data.
[0013] Optionally, the expert model is a multimodal large model.
[0014] Secondly, embodiments of the present invention provide an apparatus for generating large model training data. The apparatus includes: an acquisition unit for acquiring multiple raw data, wherein the raw data includes raw images and classification labels; a first determination unit for determining, for each raw data, an expert model corresponding to the raw data in a large model expert database based on the classification label; a processing unit for inputting the raw image into the expert model corresponding to the raw data and outputting synthetic data; a second determination unit for performing risk point identification on the multiple synthetic data and determining multiple first candidate synthetic data with successfully identified risk points; a third determination unit for determining the text perplexity of the multiple first candidate synthetic data and determining multiple second candidate synthetic data whose text perplexity is less than a set perplexity threshold; a fourth determination unit for determining the image-text alignment of the multiple second candidate synthetic data and determining multiple third candidate synthetic data whose image-text alignment is greater than a set image-text alignment threshold; and a fifth determination unit for determining the attribution logic score of the multiple third candidate synthetic data and determining multiple target synthetic data whose attribution logic score is greater than a set attribution logic threshold.
[0015] Optionally, the apparatus further includes a training unit for training a high-risk model based on the multiple target synthetic data and their corresponding original images.
[0016] Optionally, the processing unit is specifically configured to: input the original image into the expert model corresponding to the original data to generate an image description; input the image description into the expert model to generate a reasoning process; input the reasoning process into the expert model to generate an audit result; and determine the image description, the reasoning process, and the audit result as the synthesized data.
[0017] Optionally, the second determining unit is specifically used to: match each of the synthesized data with its corresponding classification label according to a set rule; in response to a successful match, determine that the risk point of the synthesized data has been successfully identified; and determine the multiple synthesized data with successfully identified risk points as multiple first candidate synthesized data.
[0018] Optionally, the third determining unit is specifically used to: calculate the text perplexity of multiple first candidate synthetic data based on the open-source large model.
[0019] Optionally, the fourth determining unit is specifically used to: calculate the image-text alignment between multiple second candidate synthetic data and their corresponding original images based on the contrastive language image pre-trained CLIP model.
[0020] Optionally, the fifth determining unit is specifically used to: score the reasoning process of multiple third candidate synthetic data according to the open-source large model, and determine the attribution logic score of multiple third candidate synthetic data.
[0021] Optionally, the expert model is a multimodal large model.
[0022] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect or any one of the possible methods of the first aspect.
[0023] Fourthly, embodiments of the present invention provide a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the method as described in the first aspect or any one of the possibilities of the first aspect.
[0024] In this embodiment of the invention, multiple raw data sets are acquired, including original images and classification labels. For each raw data set, an expert model corresponding to the raw data is determined from a large model expert database based on the classification label. The original image is input into the expert model corresponding to the raw data, and synthesized data is output. Risk points are identified in the multiple synthesized data sets to determine multiple first candidate synthesized data sets with successfully identified risk points. The text perplexity of the multiple first candidate synthesized data sets is determined, and multiple second candidate synthesized data sets with text perplexity less than a set perplexity threshold are determined. The image-text alignment of the multiple second candidate synthesized data sets is determined, and multiple third candidate synthesized data sets with image-text alignment greater than a set image-text alignment threshold are determined. The attribution logic score of the multiple third candidate synthesized data sets is determined, and multiple target synthesized data sets with attribution logic scores greater than a set attribution logic threshold are determined. Through the above method, the quality and generation efficiency of the target synthesized data can be improved, and the generation cost of the target synthesized data can be reduced. Attached Figure Description
[0025] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a flowchart of a method for generating large model training data according to an embodiment of the present invention; Figure 2 This is a flowchart of a method for generating synthetic data according to an embodiment of the present invention; Figure 3 This is a flowchart of another method for generating large model training data in an embodiment of the present invention; Figure 4 This is a schematic diagram of a device for generating large model training data in an embodiment of the present invention; Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0026] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0027] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0028] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0029] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0030] In existing technologies, Large Language Model (LLM) can also be called large model or large-scale language model, etc. The large language model is a deep learning model based on a transformer architecture that can process and generate natural language text. It is usually trained on a large amount of text data and has the ability to understand and generate language. It is widely used in dialogue systems, text generation and other natural language processing tasks. Based on this, the Multimodal Large Language Model (MLLM) has been extended to generate an artificial intelligence model with text as the core capability base and integrates multiple information modalities such as images, speech, video, and point clouds. It has cross-modal understanding and generation capabilities.
[0031] Taking the field of security risk identification as an example, MLLM (Multi-Level Modeling) is needed to identify risks in images, thus requiring improved MLLM performance. Currently, the training data used in training MLLM models is generally generated through simple manual annotation. Specifically, a large number of risky images are manually annotated with risk types, image descriptions, and attributions to generate training data. This manual annotation method is not only costly and inefficient, but also has a long annotation cycle, making it unable to quickly respond to new risks. Furthermore, the production scale is limited by the available manpower, hindering exponential growth. Therefore, the quality of training data generated by existing methods is relatively poor, affecting the performance of large language models. Thus, improving the quality and generation efficiency of training data while reducing generation costs is a problem that needs to be solved.
[0032] In this embodiment of the invention, to solve the above problems, a method for generating large model training data is proposed, specifically as follows: Figure 1 As shown, the method includes: Step S101: Obtain multiple raw data.
[0033] Specifically, the raw data includes the original images and category labels.
[0034] For example, the classification labels can be false advertising, negative values, rumors, discrimination, and normal, meaning that each original image has a corresponding classification label, and different classification labels can divide multiple original data into multiple categories; where the classification label is false advertising, negative values, rumors, or discrimination, it indicates that the original image poses a risk, and when the risk is indicated as normal, it indicates that the original image does not pose a risk.
[0035] In this embodiment of the invention, different classification labels need to be set for different business scenarios. When in the field of security risk identification, the above-mentioned classification labels such as false advertising, bad values, rumors, discrimination and normal are used. If other scenarios are used, such as legal analysis scenarios, classification labels related to legal analysis are set. This is only an example.
[0036] Step S102: For each piece of raw data, determine the expert model corresponding to the raw data in the large model expert library according to the classification label.
[0037] Specifically, the large model expert library includes multiple expert models, which are multimodal large models, such as Qwen2.5-VL, GPT-4o, Deepseek-VL2, Gemini 2.5, etc. Different expert models emphasize different capabilities. Among them, GPT-4o has strong image understanding capabilities, while Qwen2.5-VL has more flexible image processing capabilities. Different classification labels correspond to different expert models.
[0038] For example, suppose Qwen2.5-VL is suitable for raw data with the category label "false advertising" and GPT-4o is suitable for raw data with the category label "negative values". This is just an example and the specific method should be determined according to the actual situation.
[0039] Step S103: Input the original image into the expert model corresponding to the original data, and output the synthesized data.
[0040] In one possible implementation, the original image is input into the expert model corresponding to the original data, and synthesized data is output, specifically as follows: Figure 2 The above includes the following: Step S201: Input the original image into the expert model corresponding to the original data to generate an image description.
[0041] Specifically, the image description is a comprehensive, objective, and detailed textual description of the original input image provided by the expert model. It is mainly used to capture all the key visual elements in the original image and provide a factual basis for subsequent reasoning and review.
[0042] Step S202: Input the image description into the expert model to generate the reasoning process.
[0043] Specifically, the reasoning process is based on logical reasoning of the image description, analyzing whether the original image has any violation risks, and generating a complete chain of thought.
[0044] Step S203: Input the reasoning process into the expert model to generate the audit result.
[0045] Specifically, the expert model can determine the review result based on the reasoning process. For example, the review result may be a violation, normal, false advertising, negative values, rumors, discrimination, etc.
[0046] Step S204: The image description, the reasoning process, and the review result are determined as the synthesized data.
[0047] In this embodiment of the invention, the above three steps can generate the image description, the reasoning process, and the review result, and the image description, the reasoning process, and the review result are determined as the synthetic data corresponding to the original data.
[0048] In one possible implementation, the original image and pre-set prompts can be input into the expert model, and the image description, the reasoning process, and the review result can be output sequentially at the same time.
[0049] Step S104: Identify risk points in the multiple synthesized data and determine multiple first candidate synthesized data whose risk points have been successfully identified.
[0050] Specifically, each of the synthesized data is matched with its corresponding classification label according to a set rule. In response to a successful match, it is determined that the risk point of the synthesized data has been successfully identified. Multiple synthesized data with successfully identified risk points are determined as multiple first candidate synthesized data.
[0051] In this embodiment of the invention, rules are set to check whether the risk point corresponding to the classification label is accurately present in the synthetic data. If the risk point exists in the synthetic data, it is determined that the risk point of the synthetic data has been successfully identified. The synthetic data with successfully identified risk points is determined as valid data, namely the first candidate synthetic data. The first candidate synthetic data is retained and enters the next round of screening. If the risk point identification fails, the synthetic data is discarded.
[0052] Step S105: Determine the text perplexity of multiple first candidate synthetic data, and determine multiple second candidate synthetic data whose text perplexity is less than a set perplexity threshold.
[0053] In one possible implementation, determining the text perplexity of the plurality of first candidate synthetic data specifically includes: calculating the text perplexity of the plurality of first candidate synthetic data based on an open-source large model.
[0054] In this embodiment of the invention, the text perplexity of each first candidate synthetic data is calculated using an open-source large model. The text perplexity is a key indicator for measuring the fluency and readability of text. The lower the text perplexity score, the more natural and grammatically correct the text of the first candidate synthetic data is. The perplexity threshold is used to filter first candidate synthetic data with high text quality and fluent expression. Multiple first candidate synthetic data with text perplexity less than the set perplexity threshold are determined as multiple second candidate synthetic data.
[0055] Step S106: Determine the image-text alignment of multiple second candidate composite data, and determine multiple third candidate composite data whose image-text alignment is greater than a set image-text alignment threshold.
[0056] In one possible implementation, determining the image-text alignment of the plurality of second candidate synthetic data specifically includes: calculating the image-text alignment between the plurality of second candidate synthetic data and their corresponding original images based on a Contrastive Language-Image Pre-Training (CLIP) model, wherein the CLIP model is a neural network model based on contrastive learning that can align information from multiple modalities.
[0057] In this embodiment of the invention, to ensure that the "image description" generated by the expert model is faithful to the original image content, the CLIP model is used to calculate the image-text alignment score between the image description and the original image. The CLIP model can accurately evaluate the semantic relevance of the image and text content, i.e., the image-text alignment. By setting an image-text alignment threshold, multiple second candidate synthetic data are filtered to determine multiple second candidate synthetic data with an image-text alignment greater than the set image-text alignment threshold. Sample second candidate synthetic data with inaccurate descriptions or hallucinations are eliminated, and multiple third candidate synthetic data are determined.
[0058] Step S107: Determine the attribution logic scores of multiple third candidate synthetic data, and identify multiple target synthetic data whose attribution logic scores are greater than a set attribution logic threshold.
[0059] In one possible implementation, determining the attribution logic scores of the multiple third candidate synthetic data specifically includes: scoring the reasoning process of the multiple third candidate synthetic data according to the open-source large model, and determining the attribution logic scores of the multiple third candidate synthetic data.
[0060] In this embodiment of the invention, the open-source large model can also be called an evaluator model, which is used to determine whether the reasoning process is rigorous, whether the evidence is sufficient, and whether the conclusion is naturally derived from the description. Multiple target synthetic data with attribution logic scores greater than the set attribution logic threshold are selected through the set attribution logic threshold.
[0061] In this embodiment of the invention, high-quality target synthetic data is selected through multi-level filtering based on risk point identification, text perplexity, image-text alignment, and attribution logic score.
[0062] In one possible implementation, after step S107, the method further includes other steps, specifically as follows: Figure 3 As shown, it includes the following: Step S108: Train a high-risk model based on the multiple target synthetic data and their corresponding original images.
[0063] In this embodiment of the invention, training the high-risk model with the high-quality target synthetic data and the corresponding original images can improve the performance of the high-risk model, wherein the high-risk model can be a multimodal high-risk model.
[0064] Through the above embodiments, data generation is carried out using "expert models" for different risk domains, and combined with the standardized three-step paradigm of "image description - reasoning process - review result", the professionalism and structure of the generated content are guaranteed. Furthermore, a multi-dimensional and phased automated quality assessment method is used to select target synthetic data with better effectiveness, fluency, fidelity and logic from multiple synthetic data, thereby achieving high quality of the target synthetic data.
[0065] In this embodiment of the invention, an apparatus for generating large model training data is provided, such as... Figure 4 As shown, it specifically includes: an acquisition unit 401, a first determination unit 402, a processing unit 403, a second determination unit 404, a third determination unit 405, a fourth determination unit 406, and a fifth determination unit 407; The acquisition unit 401 is used to acquire multiple raw data, including original images and classification labels; the first determination unit 402 is used to determine the expert model corresponding to each raw data in a large model expert library based on the classification label; the processing unit 403 is used to input the original image into the expert model corresponding to the raw data and output synthesized data; the second determination unit 404 is used to identify risk points in the multiple synthesized data and determine multiple first candidate synthesized data with successfully identified risk points; the third determination unit 405 is used to determine the text perplexity of the multiple first candidate synthesized data and determine multiple second candidate synthesized data whose text perplexity is less than a set perplexity threshold; the fourth determination unit 406 is used to determine the image-text alignment of the multiple second candidate synthesized data and determine multiple third candidate synthesized data whose image-text alignment is greater than a set image-text alignment threshold; the fifth determination unit 407 is used to determine the attribution logic score of the multiple third candidate synthesized data and determine multiple target synthesized data whose attribution logic score is greater than a set attribution logic threshold.
[0066] Furthermore, the device further includes a training unit for training a high-risk model based on the multiple target synthetic data and their corresponding original images.
[0067] Optionally, the processing unit is specifically configured to: input the original image into the expert model corresponding to the original data to generate an image description; input the image description into the expert model to generate a reasoning process; input the reasoning process into the expert model to generate an audit result; and determine the image description, the reasoning process, and the audit result as the synthesized data.
[0068] Further, the second determining unit is specifically used to: match each of the synthetic data with its corresponding classification label according to a set rule, and in response to a successful match, determine that the risk point of the synthetic data has been successfully identified; and determine the multiple synthetic data with successfully identified risk points as multiple first candidate synthetic data.
[0069] Furthermore, the third determining unit is specifically used to: calculate the text perplexity of multiple first candidate synthetic data based on the open-source large model.
[0070] Furthermore, the fourth determining unit is specifically used to: calculate the image-text alignment between multiple second candidate synthetic data and their corresponding original images based on the contrastive language image pre-trained CLIP model.
[0071] Furthermore, the fifth determining unit is specifically used to: score the reasoning process of multiple third candidate synthetic data according to the open-source large model, and determine the attribution logic score of multiple third candidate synthetic data.
[0072] Furthermore, the expert model is a multimodal large model.
[0073] Figure 5 This is a schematic diagram of the structure of the electronic device described in an embodiment of the present invention. Figure 5 As shown, it includes a general computer hardware architecture, which includes at least a processor 501 and a memory 502. The processor 501 and the memory 502 are connected via a bus 503. The memory 502 is adapted to store instructions or programs executable by the processor 501. The processor 501 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 501 executes the instructions stored in the memory 502 to perform the method flow of the embodiments of the present invention as described above, thereby realizing data processing and control of other devices. The bus 503 connects the above-mentioned components together, and also connects the above-mentioned components to the display controller 504, the display device, and the input / output (I / O) device 505. The input / output (I / O) device 505 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 505 is connected to the system via an input / output (I / O) controller 506.
[0074] The instructions stored in memory 502 are executed by at least one processor 501 to: acquire multiple raw data; for each raw data, determine the expert model corresponding to the raw data in a large model expert library based on the classification label; input the raw image into the expert model corresponding to the raw data and output synthetic data; identify risk points in the multiple synthetic data to determine multiple first candidate synthetic data with successfully identified risk points; determine the text perplexity of the multiple first candidate synthetic data and determine multiple second candidate synthetic data whose text perplexity is less than a set perplexity threshold; determine the image-text alignment of the multiple second candidate synthetic data and determine multiple third candidate synthetic data whose image-text alignment is greater than a set image-text alignment threshold; determine the attribution logic score of the multiple third candidate synthetic data and determine multiple target synthetic data whose attribution logic score is greater than a set attribution logic threshold.
[0075] Specifically, the electronic device includes: one or more processors 501 and a memory 502. Figure 5 Take processor 501 as an example. Processor 501 and memory 502 can be connected via a bus or other means. Figure 5 Taking a bus connection as an example, memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 501 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 502, thereby implementing the aforementioned method for determining and generating large model training data.
[0076] Memory 502 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 502 may optionally include memory remotely located relative to processor 501, and these remote memories can be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0077] One or more modules are stored in memory 502, and when executed by one or more processors 501, they perform the method for generating large model training data in any of the above method embodiments.
[0078] As those skilled in the art will recognize, various aspects of the embodiments of the present invention can be implemented as a system, method, or computer program product. Therefore, various aspects of the embodiments of the present invention can take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software and hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." Furthermore, various aspects of the embodiments of the present invention can take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.
[0079] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, (but not limited to) an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the context of embodiments of the present invention, a computer-readable storage medium can be any tangible medium capable of containing or storing a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0080] Computer-readable signal media may include propagated digital signals having computer-readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0081] Program code implemented on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, or any suitable combination thereof.
[0082] Computer program code for performing operations relating to various aspects of embodiments of the present invention can be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++, etc.; and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed as a standalone software package entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet provided by an Internet service provider).
[0083] The flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products according to embodiments of the present invention describe various aspects of the embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions (executed via the processor of the computer or other programmable data processing apparatus) create means for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.
[0084] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus or other means to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of writing that includes instructions that implement the functions / actions specified in flowchart and / or block diagram blocks or blocks.
[0085] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operable steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide for implementing the functions / actions specified in flowchart and / or block diagram blocks or blocks.
[0086] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding access points are provided for users to choose to authorize or refuse processing. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
Claims
1. A method for generating large model training data, characterized in that, The method includes: Acquire multiple raw data sets, including raw images and category labels; For each piece of raw data, the expert model corresponding to the raw data is determined from the large model expert library based on the classification label; The original image is input into the expert model corresponding to the original data, and the synthesized data is output. Risk points are identified from multiple synthetic data sets to determine multiple first candidate synthetic data sets whose risk points are successfully identified. Determine the text perplexity of multiple first candidate synthesized data, and determine multiple second candidate synthesized data whose text perplexity is less than a set perplexity threshold; Determine the image-text alignment of multiple second candidate composite data, and determine multiple third candidate composite data whose image-text alignment is greater than a set image-text alignment threshold; Determine the attribution logic scores of multiple third candidate synthetic data, and identify multiple target synthetic data whose attribution logic scores are greater than a set attribution logic threshold.
2. The method according to claim 1, characterized in that, The method further includes: A high-risk model is trained based on the synthetic data of the multiple targets and their corresponding original images.
3. The method according to claim 1, characterized in that, The step of inputting the original image into the expert model corresponding to the original data and outputting synthesized data specifically includes: The original image is input into the expert model corresponding to the original data to generate an image description; The image description is input into the expert model to generate a reasoning process; The reasoning process is input into the expert model to generate the review result; The image description, the reasoning process, and the review result are determined as the synthesized data.
4. The method according to claim 1, characterized in that, The step of identifying risk points in multiple synthesized data sets and determining multiple first candidate synthesized data sets with successfully identified risk points specifically includes: Each piece of synthetic data is matched with its corresponding classification label according to a set rule. If the match is successful, it is determined that the risk point of the synthetic data has been successfully identified. Multiple synthetic data points that successfully identify the risk points are identified as multiple first candidate synthetic data points.
5. The method according to claim 1, characterized in that, Determining the text perplexity of the multiple first candidate synthetic data specifically includes: The text perplexity of multiple first candidate synthetic data is calculated based on the open-source large model.
6. The method according to claim 1, characterized in that, Determining the image-text alignment of the multiple second candidate synthetic data specifically includes: The image-text alignment between multiple second candidate synthetic data and their corresponding original images is calculated using a contrastive language image pre-trained CLIP model.
7. The method according to claim 1, characterized in that, The determination of the attribution logic scores of the multiple third candidate synthetic data specifically includes: The attribution logic scores of the multiple third candidate synthetic data are determined by scoring the reasoning process of the multiple third candidate synthetic data based on the open-source large model.
8. The method according to claim 1, characterized in that, The expert model is a multimodal large model.
9. An apparatus for generating large model training data, characterized in that, The device includes: An acquisition unit is used to acquire multiple raw data, wherein the raw data includes raw images and category labels; The first determining unit is used to determine the expert model corresponding to each piece of raw data in the large model expert library based on the classification label. The processing unit is used to input the original image into the expert model corresponding to the original data and output the synthesized data; The second determining unit is used to identify risk points in the multiple synthetic data and determine multiple first candidate synthetic data whose risk points have been successfully identified. The third determining unit is used to determine the text perplexity of multiple first candidate synthetic data and to determine multiple second candidate synthetic data whose text perplexity is less than a set perplexity threshold; The fourth determining unit is used to determine the image-text alignment degree of multiple second candidate composite data, and to determine multiple third candidate composite data whose image-text alignment degree is greater than a set image-text alignment degree threshold; The fifth determining unit is used to determine the attribution logic scores of multiple third candidate synthetic data, and to determine multiple target synthetic data whose attribution logic scores are greater than a set attribution logic threshold.
10. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.