Large model security evaluation system based on multi-modal adversarial sample generation

The security assessment system based on large-scale models generated from multimodal adversarial examples addresses the shortcomings of large-scale multimodal models in security risk assessment, achieving comprehensive and accurate security assessment, discovering potential vulnerabilities, and improving the accuracy and security of the assessment.

CN121456883AInactive Publication Date: 2026-02-03DIGITAL NEW ERA (SHANDONG) DATA TECH SERVICES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511635195.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-02-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Multimodal large models face serious security risks during data collection, model training, and practical applications, especially the covert and destructive threats of adversarial attacks. Existing evaluation methods lack comprehensiveness and accuracy.

Method used

Design a large-scale security assessment system based on multimodal adversarial sample generation, including a platform end and a user end. Through adversarial analysis module, information module, sample generation module and adversarial assessment module, generate diverse adversarial samples to simulate attacker methods and conduct comprehensive and accurate security assessments. Utilize platform resources to assist user assessments.

Benefits of technology

It improves the comprehensiveness and accuracy of multimodal large-scale model security assessment, discovers potential security vulnerabilities without disclosing user information, and makes full use of platform resources to provide services to users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456883A_ABST
    Figure CN121456883A_ABST
Patent Text Reader

Abstract

The invention discloses a large model safety evaluation system based on multi-modal adversarial sample generation, and belongs to the technical field of large model safety evaluation. An adversarial analysis module for the platform end; the user side comprises an information module, an adversarial analysis module, a sample generation module and an adversarial evaluation module; the confrontation analysis module is used for analyzing the target large model information of the user to obtain an initial confrontation point diagram; the information module is used for performing feature recognition on the target large model according to the initial confrontation point diagram, and supplementing model feature data of each confrontation point into the initial confrontation point diagram to obtain a confrontation information diagram; the adversarial analysis module is used for analyzing the adversarial information graph to obtain a weight value of each adversarial point; the sample generation module is used for generating corresponding adversarial samples according to the adversarial information graph; and the adversarial evaluation module is used for performing adversarial analysis on the target large model according to the adversarial sample to obtain corresponding adversarial analysis data, and performing security evaluation on the adversarial analysis data to obtain a corresponding dynamic score.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of large model security evaluation, and specifically relates to a large model security evaluation system based on multi-modal adversarial sample generation. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, multi-modal large models have become the core engine for driving the digital transformation of various industries. Multi-modal large models can simultaneously process and understand multiple data forms and are widely used in intelligent customer service, autonomous driving, medical diagnosis, financial risk control and other scenarios.

[0003] However, multi-modal large models, while bringing great application value, also face serious security challenges. Due to their complexity and diversity characteristics, multi-modal large models face more security risks during data collection, model training and actual application. Studies have shown that multi-modal large models are vulnerable to adversarial samples. By applying a small perturbation to the original input, an adversarial sample can be constructed to make the machine learning model produce an incorrect response or unexpected behavior. For example, in an image classification task, an attacker can add noise that is difficult for the human eye to detect in an image, causing the model to misclassify a "cat" as a "dog"; in a text classification task, an attacker can modify individual words, causing the model to misclassify a positive review as a negative review. In a multi-modal scenario, an attacker may simultaneously tamper with the semantic consistency of images and text, generating more covert and destructive adversarial samples, thereby posing a serious threat to the security of the model.

[0004] Based on this, in order to realize the security evaluation of the large model, the application provides a large model security evaluation system based on multi-modal adversarial sample generation. SUMMARY

[0005] In order to solve the problems existing in the above-mentioned scheme, the application provides a large model security evaluation system based on multi-modal adversarial sample generation.

[0006] The purpose of the application can be achieved by the following technical solutions: A large model security evaluation system based on multi-modal adversarial sample generation, comprising a platform end and a user end; The platform end uses an adversarial analysis module; The adversarial analysis module is used to analyze the target large model information of the user, obtain an initial adversarial point graph, and send the initial adversarial point graph to the information module of the corresponding user end.

[0007] Further, the analysis of the target large model information comprises: The platform party sets an entire set of confrontation points; each confrontation point in the entire set of confrontation points is calibrated and analyzed according to target large model information, and a calibration result of the confrontation point is obtained, the calibration result including calibration pass and calibration fail; The confrontation points with the calibration result of calibration pass are integrated into an initial set, and an initial confrontation point graph is generated according to the initial set.

[0008] Further, the calibration and analysis of each confrontation point in the entire set of confrontation points according to the target large model information includes: A calibration model is established, and the expression of the calibration model is: ; In the formula, (MX, i) is input data, MX is target large model information, i represents a corresponding confrontation point, i=1, 2, …, n, n is the number of confrontation points in the confrontation point set; and the output data is a calibration value BH(MX, i), the calibration value including 1 or 0; The target large model information and the corresponding confrontation point are analyzed by the calibration model to obtain a calibration value of the corresponding confrontation point; When the calibration value is 1, the calibration result of the confrontation point is calibration pass; When the calibration value is 0, the calibration result of the confrontation point is calibration fail.

[0009] The user end includes an information module, a confrontation analysis module, a sample generation module and a confrontation evaluation module; The information module is configured to upload target large model information of a target large model, receive an initial confrontation point graph sent by a platform end, perform feature recognition on the target large model according to the initial confrontation point graph, obtain model characteristic data of each confrontation point in the initial confrontation point graph, supplement the model characteristic data to the corresponding confrontation point in the initial confrontation point graph, and mark the current initial confrontation point graph as a confrontation information graph.

[0010] Further, the corresponding confrontation point in the confrontation information graph is filtered according to the model characteristic data.

[0011] The confrontation analysis module is configured to analyze the confrontation information graph, determine a weight value of each confrontation point in the confrontation information graph in real time, and supplement the obtained weight value to the confrontation information graph.

[0012] Further, the weight value of each confrontation point in the confrontation information graph is determined in real time, including: The dynamic score of each confrontation point is obtained in real time, the obtained dynamic score is substituted into a preset weight value calculation formula, and the weight value of the corresponding confrontation point is calculated.

[0013] Further, when the corresponding confrontation point has no dynamic score, a historical score of each confrontation point is determined according to historical data of the target large model, and the historical score is marked as a dynamic score.

[0014] Further, the weight value calculation formula is: ; In the formula, δ j represents the weight value of the corresponding confrontation point, j represents the corresponding confrontation point in the confrontation information graph, j=1, 2, …, m, and m is the number of confrontation points in the confrontation information graph; P max represents the maximum dynamic score; PF j represents the dynamic score of the corresponding confrontation point.

[0015] Further, the weight value calculation formula is: ; In the formula, δ j represents the weight value of the corresponding confrontation point, j represents the corresponding confrontation point in the confrontation information graph, j=1, 2, …, m, and m is the number of confrontation points in the confrontation information graph; P max represents the maximum basic score; PF j represents the dynamic score of the corresponding confrontation point; λ j represents the proportion coefficient of the corresponding confrontation point, and the value range is 0<λ j ≤1.

[0016] The sample generation module is configured to generate corresponding confrontation samples according to the confrontation information graph.

[0017] Further, the platform end further comprises a sample module, and the user end further comprises a sample supplement module; The sample module is configured to analyze various confrontation points to obtain corresponding supplementary confrontation points, set corresponding confrontation samples for the supplementary confrontation points, and aggregate the confrontation samples corresponding to each supplementary confrontation point to establish a sample supplement library. The sample supplement module is configured to supplement the confrontation samples, identify supplementary confrontation points based on the confrontation information graph, and match corresponding confrontation samples from the sample supplement library according to the supplementary confrontation points.

[0018] The confrontation evaluation module is configured to perform confrontation analysis on the target large model according to the confrontation samples, obtain corresponding confrontation analysis data, perform security evaluation on the confrontation analysis data, and obtain corresponding dynamic scores.

[0019] Compared with the prior art, the present application has the following advantages: The multi-modal adversarial sample generation-based large model security evaluation system provided by the application can generate diversified adversarial samples in view of complex security risks faced by multi-modal large models in the data collection, model training and actual application process. By simulating various means that may be adopted by attackers, including tampering with the consistency of image and text semantics, etc., the multi-modal large model is comprehensively and accurately evaluated from multiple dimensions, effectively discovering security vulnerabilities existing in the model, greatly improving the comprehensiveness and accuracy of the evaluation compared with the traditional single modal or simple attack method. At the same time, the resource advantages of the platform party are fully utilized to serve each user, assist the user in large model security evaluation, and do not disclose the user information. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, a brief introduction of the drawings needed to be used in the embodiments or the prior art description will be given below. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0021] Figure 1 The principle block diagram of the present application. DETAILED DESCRIPTION

[0022] The technical solutions of the present application will be described below in conjunction with the embodiments, obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0023] As shown in Figure 1 A multi-modal adversarial sample generation-based large model security evaluation system, including a platform end and a user end; the platform end is in communication connection with the user end of each user; The platform end is used by the platform party, and is used to serve each user, including an adversarial analysis module; The adversarial analysis module is used to analyze the target large model information of the user, obtain an initial adversarial point graph, and send the initial adversarial point graph to the information module of the corresponding user end.

[0024] The target large model information includes architecture description, training data, model document, interface specification and other related information; the target large model is a large model that needs to be evaluated by the user.

[0025] The initial adversarial point map is composed of various adversarial points, which are distributed in the blank map. Adversarial points refer to key features or regions in the large model input that can cause the model output to be incorrect through a small perturbation. According to the analysis principle, adversarial mode and application scenario of the large model, the adversarial points can be divided into the following categories: Adversarial points based on model analysis principles: Gradient-sensitive regions: regions with large gradients of the model on input features, such as edges, textures, or keywords in text in images.

[0026] Attention-focused regions: regions where the attention mechanism of the model focuses when processing input, such as high-weight tokens in Transformer models.

[0027] Adversarial points based on adversarial mode: White-box adversarial points: adversarial points determined by the attacker using gradient information of the model in white-box attacks, such as perturbation positions generated by FGSM, PGD, etc.

[0028] Black-box adversarial points: adversarial points determined by querying model output or transfer attacks in black-box attacks, such as key perturbation positions in adversarial samples generated based on substitute models.

[0029] Adversarial points based on application scenarios: Medical field: adversarial points may focus on lesion regions in medical images or key diagnostic keywords in text reports.

[0030] Financial field: adversarial points may focus on outliers in transaction data or sentiment keywords in text analysis.

[0031] In one embodiment, the setting of the initial adversarial point map can be set in the existing manner, such as manual setting by the platform party or determination combined with other intelligent technologies.

[0032] In one embodiment, the analysis of the target large model information includes: The platform party integrates various adversarial points possessed by large models based on big data and other technologies into an adversarial point set, which is equivalent to the full set of adversarial points, i.e., all adversarial points possessed by various large models are in the adversarial point set. According to the definition of adversarial points, the adversarial points are identified, which can be identified based on existing methods or manually set by the platform party.

[0033] According to the target large model information, the calibration analysis of each adversarial point in the adversarial point set is performed to determine whether the target large model has the adversarial point, and the calibration result of the corresponding adversarial point is obtained. The calibration result includes calibration pass and calibration fail, and calibration pass means that the adversarial point is possessed.

[0034] The calibration result of the adversarial point that passes the calibration is integrated into an initial set, and an initial adversarial point graph is generated according to the initial set.

[0035] In one embodiment, the calibration analysis of each adversarial point in the adversarial point set according to the target large model information can be judged based on existing methods, such as analyzing whether it has the characteristics of the adversarial point according to the target large model information, if so, the calibration passes, and subsequent screening can be performed at the user, that is, multiple selections can be made; or an intelligent model is established based on machine learning, deep learning algorithm, etc. to calibrate.

[0036] In one embodiment, the calibration analysis of each adversarial point in the adversarial point set according to the target large model information includes: A calibration model is established, which is used to analyze the target large model information and the adversarial point, and judge whether the corresponding target large model has the adversarial point. The corresponding training set is established by the platform party for training, the training set includes input data and output data, the input data is the large model information and the adversarial point, and the output data is the calibration result. The expression of the calibration model is: ; In the formula, (MX, i) is the input data, MX is the target large model information, i represents the corresponding adversarial point, i = 1, 2, …, n, n is the number of adversarial points in the adversarial point set; i satisfies the calibration standard, indicating that the target large model has the model characteristics of the adversarial point; the output data is the calibration value BH(MX, i), and the calibration value includes 1 or 0; The target large model information and the corresponding adversarial point are analyzed by the calibration model to obtain the calibration value of the corresponding adversarial point; When the calibration value is 1, the calibration result of the corresponding adversarial point is calibration pass; When the calibration value is 0, the calibration result of the corresponding adversarial point is calibration fail.

[0037] In one embodiment, the target large model information can also be analyzed by the platform party to pre-classify various large models, set the corresponding initial adversarial point graph for each model classification, and then match according to the target large model information.

[0038] Exemplary adversarial point identification methods: Gradient-based method: Gradient saliency map: generate a gradient saliency map by calculating the gradient of the input with respect to the model output, and identify the features that have the greatest impact on the output.

[0039] Integrated Gradients: quantify the contribution of each input feature to the output by integrating the gradient.

[0040] Optimization search-based method: Adversarial Sample Generation Algorithms: Use algorithms such as FGSM, PGD, C&W to generate adversarial samples, and record the perturbation positions as adversarial points.

[0041] Genetic Algorithm: Search the input space through genetic algorithms to find key features that cause the model to misclassify.

[0042] Model Explanation-based Methods: LIME / SHAP: Use explanation tools such as LIME or SHAP to analyze the importance of input features and identify key adversarial points.

[0043] Attention Mechanism Analysis: In multi-modal models, analyze cross-modal attention weights to identify interactive adversarial points between modalities.

[0044] Query-based Methods: Black-box Query: In a black-box scenario, approximate gradients or directly search for adversarial samples by querying the model output a large number of times, and record the perturbation positions.

[0045] The user end includes an information module, an adversarial analysis module, a sample generation module, and an adversarial evaluation module; The information module is used to upload the target large model information of the target large model, and receive the initial adversarial point graph sent by the platform end, identify the features of the target large model according to the initial adversarial point graph, obtain the model characteristic data of each adversarial point in the initial adversarial point graph, and supplement the model characteristic data to the corresponding adversarial point in the initial adversarial point graph; mark the current initial adversarial point graph as an adversarial information graph.

[0046] Model characteristic data refers to the architecture, training process, data characteristics, and design logic of the target large model itself, which may imply adversarial risks at the corresponding adversarial points; for example, model architecture characteristics: Description: Vulnerabilities that may be introduced by the architecture design of the model.

[0047] Data content: Layer type and depth: such as convolutional layers, fully connected layers, and stacking methods of attention mechanisms.

[0048] Activation function: such as ReLU, Sigmoid, and other gradient problems that may be caused.

[0049] Normalization and regularization: such as the use of BatchNorm and Dropout.

[0050] Example: Model: ResNet-50 Characteristics: Contains 50 layers of convolution and residual connection, uses ReLU activation function, and may be sensitive to high-frequency noise (due to residual connection that may amplify local perturbations).

[0051] Training data characteristics: Description: The distribution and characteristics of training data can affect the adversarial robustness of the model.

[0052] Data content: Data source: such as public datasets (CIFAR-10), private datasets.

[0053] Data size: number of training samples.

[0054] Data bias: such as class imbalance, noise data proportion.

[0055] Example: Model: image classification model; Characteristics: training data is CIFAR-10 (50,000 training images), class distribution is balanced, but image resolution is low (32x32), which may be sensitive to small-scale perturbations.

[0056] Model decision logic: Description: The decision-making process of the model may imply an adversarial risk.

[0057] Data content: Feature importance: such as which input features have the greatest impact on the output (can be visualized by feature visualization or gradient analysis).

[0058] Decision boundary: such as the complexity of the boundary of a classification task.

[0059] Uncertainty estimation: such as the confidence distribution of the model for uncertain inputs.

[0060] Example: Model: BERT text classification model Characteristics: relies on attention mechanism to focus on specific keywords (such as "good", "bad"), which may be sensitive to keyword replacement attacks.

[0061] In one embodiment, according to the model characteristic data, the corresponding adversarial points are screened to determine whether they really belong to the target large model, and the adversarial points that do not belong to the target large model are eliminated, and the screening is completed.

[0062] In one embodiment, whether it really belongs to the target large model can be determined based on existing methods, such as determining whether the model characteristic data meets the identification criteria of the adversarial point.

[0063] The adversarial analysis module is used to analyze the adversarial information graph and determine the weight values of each adversarial point in the adversarial information graph in real time, and the obtained weight values are supplemented to the adversarial information graph. Subsequently, the corresponding adversarial samples are generated according to the weight values of each adversarial point for verification and evaluation.

[0064] In one embodiment, the weight value of each adversarial point in the adversarial information graph is determined in real time, including: The dynamic score of each adversarial point is obtained in real time, and the obtained dynamic score is substituted into the preset weight value calculation formula to calculate the weight value of the corresponding adversarial point.

[0065] In one embodiment, when the corresponding adversarial point has no dynamic score, that is, no dynamic score is assessed according to the adversarial analysis result, at this time, the historical score of each adversarial point is determined according to the historical data of the target large model, that is, the score obtained according to the subsequent large model evaluation method is marked as the dynamic score.

[0066] In one embodiment, the weight value calculation formula is: ; In the formula, δ j represents the weight value of the corresponding adversarial point, j represents the corresponding adversarial point in the adversarial information graph, j=1, 2, …, m, and m is the number of adversarial points in the adversarial information graph; P max represents the maximum basic score, for example, if the score interval is [0, 100], then P max =100; PF j represents the dynamic score of the corresponding adversarial point. In one embodiment, the influence degree of different adversarial points on the target large model needs to be considered, so the corresponding proportional coefficient can be set for the corresponding adversarial point according to the influence degree; the proportional coefficient is corrected; the proportional coefficient is initially set by the platform party, and is adjusted according to the user's demand subsequently. The weight value calculation formula is: ; In the formula, δ j represents the weight value of the corresponding adversarial point, j represents the corresponding adversarial point in the adversarial information graph, j=1, 2, …, m, and m is the number of adversarial points in the adversarial information graph; P max represents the maximum basic score; PF j represents the dynamic score of the corresponding adversarial point; λ j represents the proportional coefficient of the corresponding adversarial point, and the value range is 0<λ j ≤1.

[0067] The sample generation module is configured to generate the corresponding adversarial sample according to the adversarial information graph.

[0068] In one embodiment, the generation of the adversarial sample is based on the existing technology, and the difference from the existing technology is that the weight values of the various adversarial points in the adversarial information map are generated with emphasis, such as preferentially disturbing the high-weight region according to the weight values of the adversarial points. That is, the adversarial information map is used to dynamically adjust the weight values of the various adversarial points according to the real-time adversarial analysis situation, and to provide a direction for the generation of the adversarial sample; at the same time, after the user optimizes and adjusts the target large model, the user can also continuously analyze.

[0069] In one embodiment, the current system mainly uses the method of "independent generation + simple splicing" to process multi-modal adversarial samples (such as combining image and text adversarial perturbations generated respectively), ignoring the deep semantic association between modalities. The generated adversarial samples may not be able to truly simulate the cross-modal attack scene (such as an attacker may simultaneously tamper with the semantic consistency of images and text), resulting in an evaluation result deviating from the actual risk; that is, the current generation method of the adversarial sample is not ideal for the effect of the generated adversarial sample at some adversarial points; based on this, in the present embodiment, the platform end also includes a sample module; and the user end also includes a sample supplement module; The sample module is used to analyze various adversarial points and determine whether the adversarial samples generated by the current adversarial sample generation technology meet the analysis requirements, and mark the adversarial points that do not meet the analysis requirements as supplementary adversarial points. Corresponding adversarial samples are set for the supplementary adversarial points, and the adversarial samples corresponding to each supplementary adversarial point are summarized to establish a sample supplement library; generally, cloud storage is used for easy sharing to users.

[0070] In one embodiment, the analysis of various adversarial points includes: The adversarial sample generation records of each user are obtained in real time, which can include various related information records such as corresponding adversarial samples, sample feedback, application sample generation technology, and user feedback, as long as it can be determined whether the adversarial sample generated for the adversarial point meets the requirements; The identification features of the adversarial samples that do not meet the requirements are determined, that is, which phenomenon characteristics in the adversarial sample generation records; the supplementary adversarial points are determined according to the identification features; and intelligent models can also be established based on intelligent algorithms such as machine learning and deep learning algorithms, and the intelligent models are used for intelligent evaluation.

[0071] In one embodiment, corresponding adversarial samples are set for the supplementary adversarial points, which can be generated in combination with other user-satisfying adversarial sample generation technologies, or the corresponding user-satisfying adversarial samples can be obtained; the platform party can also obtain the adversarial samples by other means, such as manual setting, hacker attack record extraction, etc.

[0072] The sample supplement module is used for conducting adversarial sample supplement, identifying supplement adversarial points based on an adversarial information graph, and can identify and judge the generated adversarial samples according to adversarial sample generation techniques, because the user does not need to consider information leakage to the platform side, so detailed adversarial sample related information can be obtained, and existing methods are used for identification and judgment; such as adversarial sample evaluation techniques, machine learning, AI, etc. to identify and supplement adversarial points, for example, the platform side counts the adversarial sample generation techniques that do not meet the requirements for each adversarial point, forms the corresponding directory information and shares it with the user end, and the user end identifies and judges; According to the supplement adversarial point, the corresponding adversarial sample is matched from the sample supplement library.

[0073] The adversarial evaluation module is used for conducting adversarial analysis on the target large model according to the adversarial sample, obtaining corresponding adversarial analysis data, conducting security evaluation on the adversarial analysis data, and obtaining corresponding dynamic scores.

[0074] In one embodiment, the dynamic scores of each adversarial point can also be summarized to obtain a comprehensive security score. In one embodiment, the adversarial analysis on the target large model according to the adversarial sample and the security evaluation on the adversarial analysis data are both conducted by using existing technologies for adversarial analysis and security evaluation.

[0075] The above formulas are all dimensionless values calculated, and the formulas are obtained by software simulation of a large amount of data to obtain a formula closest to the actual situation. The preset parameters and the preset threshold in the formula are set by the person skilled in the art according to the actual situation or obtained by a large amount of data simulation.

[0076] The above embodiments are only used to illustrate the technical method of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical method of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical method of the present application.

Claims

1. A large model security evaluation system based on multi-modal adversarial sample generation, characterized in that, The platform end and the user end are included; The platform end is provided with an adversarial analysis module, and the user end includes an information module, an adversarial analysis module, a sample generation module and an adversarial evaluation module; The adversarial analysis module is used for analyzing target large model information of a user, obtaining an initial adversarial point graph, and sending the initial adversarial point graph to the information module of the corresponding user end; The information module is used for uploading target large model information of a target large model, receiving the initial adversarial point graph sent by the platform end, performing feature recognition on the target large model according to the initial adversarial point graph, obtaining model characteristic data of each adversarial point in the initial adversarial point graph, supplementing the model characteristic data to the corresponding adversarial point in the initial adversarial point graph, and marking the current initial adversarial point graph as an adversarial information graph; The adversarial analysis module is used for analyzing the adversarial information graph and determining the weight value of each adversarial point in the adversarial information graph in real time; The sample generation module is used for generating a corresponding adversarial sample according to the adversarial information graph; The adversarial evaluation module is used for performing adversarial analysis on the target large model according to the adversarial sample, obtaining corresponding adversarial analysis data, performing security evaluation on the adversarial analysis data, and obtaining a corresponding dynamic score.

2. The large model security evaluation system based on multi-modal adversarial sample generation of claim 1, wherein, The corresponding adversarial points in the adversarial information graph are screened according to the model characteristic data.

3. The large model security evaluation system based on multi-modal adversarial sample generation of claim 1, wherein, The target large model information is analyzed, including: The platform party sets an adversarial point universal set; a calibration model is established, and the expression of the calibration model is: ; In the formula, (MX, i) is input data, MX is target large model information, i represents a corresponding adversarial point, i=1, 2, …, n, and n is the number of adversarial points in the adversarial point set; and the output data is a calibration value BH(MX, i), and the calibration value includes 1 or 0; The target large model information and the corresponding adversarial point are analyzed through the calibration model to obtain a calibration value of the corresponding adversarial point; When the calibration value is 1, the calibration result of the adversarial point is calibration passing; When the calibration value is 0, the calibration result of the adversarial point is calibration failure; The adversarial points with the calibration result of calibration passing are integrated into an initial set, and an initial adversarial point graph is generated according to the initial set.

4. The large model security evaluation system based on multi-modal adversarial sample generation of claim 1, wherein, The weight value of each adversarial point in the adversarial information graph is determined in real time, including: The dynamic score of each adversarial point is obtained in real time, the obtained dynamic score is substituted into a preset weight value calculation formula, and the weight value of the corresponding adversarial point is calculated.

5. The large model security evaluation system based on multi-modal adversarial sample generation of claim 4, wherein, When the corresponding adversarial point has no dynamic score, the historical score of each adversarial point is determined according to the historical data of the target large model, and the historical score is marked as a dynamic score.

6. The large model security evaluation system based on multi-modal adversarial sample generation of claim 4, wherein, The weight value calculation formula is: ; In the formula, δ j represents the weight value of the corresponding confrontation point, j represents the corresponding confrontation point in the confrontation information graph, j=1, 2, …, m, and m is the number of confrontation points in the confrontation information graph; P max represents the maximum dynamic score; PF j represents the dynamic score of the corresponding confrontation point.

7. The large model security evaluation system based on multi-modal adversarial sample generation of claim 4, wherein, The weight value calculation formula is: ; In the formula: δ j represents the weight value of the corresponding confrontation point, j represents the corresponding confrontation point in the confrontation information graph, j=1, 2, …, m, and m is the number of confrontation points in the confrontation information graph; P max represents the maximum basic score; PF j represents the dynamic score of the corresponding confrontation point; λ j represents the proportional coefficient of the corresponding confrontation point, and the value range is 0<λ j ≤1.

8. The large model security evaluation system based on multi-modal adversarial sample generation of claim 1, wherein, The platform end further includes a sample module, and the user end further includes a sample supplement module; The sample module is used for analyzing various adversarial points to obtain corresponding supplementary adversarial points, setting corresponding adversarial samples for the supplementary adversarial points, and summarizing the adversarial samples corresponding to each supplementary adversarial point to establish a sample supplement library; The sample supplement module is used for supplementing adversarial samples, identifying supplementary adversarial points based on the adversarial information graph, and matching corresponding adversarial samples from the sample supplement library according to the supplementary adversarial points.