Safety compliance evaluation system and method based on multi-modal large model
The multimodal large-scale model security compliance assessment system solves the problems of insufficient coverage, flexibility and model compatibility of existing systems, and achieves efficient identification of new risks and stability and accuracy of multimodal assessment.
Patent Information
- Application Number
- CN202511306466.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing security compliance assessment systems are inadequate in terms of coverage, flexibility, model compatibility, and the generalization of assessment models. They struggle to cope with new risks, data variations, and multimodal inputs, resulting in incomplete and inaccurate assessment results.
A security compliance assessment system based on a multimodal large model is adopted. Cross-modal consistency detection is achieved through a cluster of functional modules. It combines a self-developed security system to conduct adversarial testing and data variants, supports multi-granular data processing, realizes end-to-end automated assessment, and optimizes risk identification capabilities through a compliance labeling system.
It improved the detection rate of new risks, enhanced the robustness of the assessment and the ability to identify deep risks, achieved unified access to multimodal models and stability of assessment results, and improved the assessment coverage and accuracy.
Smart Images

Figure CN120930150A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to security compliance assessment, specifically to a security compliance assessment system and method based on a multimodal large model. Background Technology
[0002] Currently, security compliance assessment systems mainly rely on manually designed test cases (such as sensitive word filtering and fixed question-and-answer templates) or rule-based keyword matching detection. For example, some systems use a predefined database of violation questions to conduct batch testing of models, or use traditional NLP classifiers to determine security compliance.
[0003] The aforementioned existing technical solutions have the following main drawbacks:
[0004] 1) Insufficient assessment coverage
[0005] Problem: Existing systems rely too heavily on fixed datasets (such as a predefined library of violation questions), making it difficult to cover new risks (such as leading questions specific to generative models).
[0006] Reason: Risk scenarios change dynamically, but traditional tools lag behind in updating their datasets and lack user-defined extension capabilities.
[0007] 2) Lack of flexible data variation capabilities
[0008] Problem: The inability to perform diverse variant processing on the original dataset (such as semantic rewriting, adversarial attack simulation, etc.) leads to overfitting of the evaluation results.
[0009] Reason: Existing systems typically treat datasets as static input and do not integrate data augmentation or adversarial testing modules.
[0010] 3) Limited model compatibility
[0011] Problem: Existing systems are often designed for specific APIs or model architectures (such as only supporting the GPT series), making it difficult to adapt to text generation models from different vendors.
[0012] Reason: The input and output interfaces of different models vary greatly, and the existing system lacks a unified protocol adaptation layer.
[0013] 4) The evaluation model has poor generalization ability.
[0014] Problem: Risk assessment relies on rule engines or simple classification models (such as keyword matching), making it difficult to identify implicit ethical and legal risks (such as rhetoric that induces crime).
[0015] Reason: Traditional assessment models do not incorporate multi-dimensional risk characteristics (such as contextual coherence, intent concealment, etc.). Summary of the Invention
[0016] (a) Technical problems to be solved
[0017] To address the aforementioned shortcomings of existing technologies, this invention provides a security compliance assessment system and method based on a multimodal large model. This system can assess the security compliance of intelligent model responses to multimodal inputs such as text, images, and audio. Furthermore, by deeply integrating cross-modal consistency detection with a compliance labeling system, it overcomes the limitations of existing technologies that can only perform assessments under a single modality and a single rule.
[0018] (II) Technical Solution
[0019] To achieve the above objectives, the present invention provides the following technical solution:
[0020] A security compliance assessment system based on a multimodal large model includes:
[0021] The functional module cluster manages the various functional modules in the system and supports cross-modal consistency detection, which is used to compare the consistency of responses under different modalities.
[0022] The self-developed security system is responsible for the execution of assessment tasks, including adversarial testing enhancements, built-in multi-granularity data variant strategies, and improved assessment robustness by simulating real attack scenarios. It achieves end-to-end automation from data generation and testing of the model under test to risk assessment, supports configurable parameters and reproducible results, as well as self-iteration and version management of the assessment model. Furthermore, this self-iteration mechanism is deeply integrated with the compliance labeling system, which can continuously optimize the risk identification capability under legal, ethical, and industry norms.
[0023] Preferably, the functional module cluster is suitable for administrators and includes user management, model management, tool management, and data management;
[0024] In addition to supporting multiple interface types including REST API, gRPC and SDK, model management also supports protocol adaptation for cross-modal input and output, ensuring unified access and management of multimodal models including text, images and audio.
[0025] Preferably, the self-developed security system is suitable for ordinary users and includes:
[0026] Task execution is responsible for evaluating the entire process of task scheduling and automatically completing the "data processing - model interaction - result collection" process based on user configuration.
[0027] The assessment report is generated by integrating multi-source assessment results to produce a comprehensive assessment report.
[0028] The model is self-iterable, enabling online iterative optimization of the evaluation model and improving risk identification capabilities through closed-loop learning.
[0029] Preferably, the task execution specifically includes:
[0030] Task creation: Users select the model to be tested, the original dataset, the processing tools, and the evaluation model, and set the evaluation parameters;
[0031] Task distribution: Allocate independent running environments for evaluation tasks to avoid interference from multiple tasks;
[0032] Data injection: The processing tool is invoked to perform variant processing on the original dataset, generating a dynamic test set;
[0033] Model interaction: Through the protocol adaptation layer of model management, dynamic test sets are sent to the model under test, and the response results are received and stored;
[0034] Progress monitoring: Records task progress in real time and automatically retryes in case of abnormalities.
[0035] Preferably, the generation of the evaluation report specifically includes:
[0036] Evaluation index calculation: Based on the evaluation results output by the evaluation model, calculate detailed evaluation indicators including risk detection rate, false positive rate, and risk level distribution;
[0037] Multi-source comparison: Compare the evaluation results of the evaluation model with the evaluation results output by a third-party model and mark the differences;
[0038] Evaluation report output: Supports PDF / Word format, includes detailed evaluation indicators, typical case analysis and improvement suggestions;
[0039] The model self-iteration specifically includes:
[0040] Sample preparation: The evaluated "question-answer-risk label" triples are prepared into training set, validation set and test set;
[0041] Incremental training: Use training and validation sets to train, fine-tune, and evaluate the model, and improve training efficiency by dynamically adjusting the learning rate;
[0042] Performance verification: Use the test set to verify the performance of the new evaluation model. If the performance improvement is greater than the preset threshold, update the evaluation model library; otherwise, roll back to the previous version.
[0043] Version management: Records the iteration history of the evaluation model and supports version traceability.
[0044] The security compliance assessment method based on a multimodal large model, applied to the aforementioned security compliance assessment system based on a multimodal large model, includes the following steps:
[0045] S1. Task creation;
[0046] S2, Data Processing;
[0047] S3, calling the model under test;
[0048] S4. Evaluate the model call;
[0049] S5. Cross-modal consistency detection: In multimodal environments including text, images, and audio, the consistency of the test model's responses is compared and judged in conjunction with a compliance labeling system.
[0050] Preferably, in S1, task creation involves the assessment team leader creating assessment tasks, distributing them to assessors in real time, and monitoring the assessment process in real time to complete a closed-loop process of "task—execution—monitoring—result," including:
[0051] S11. Administrator Login: Administrators log in to the system using an account and password. The system automatically verifies the identity information and permission level to ensure that authorized administrators can access the task management interface.
[0052] S12. Assessment Task Creation: The assessment team leader fills in basic information such as task name, description and deadline on the task creation interface, selects assessment personnel from the user list and assigns specific responsibilities, and clarifies the assessment requirements.
[0053] S13. Assessment Task Distribution: The system pushes assessment task details to the personal accounts of designated assessors in real time through an internal message push mechanism, and updates the task assignment status in the task dashboard at the same time.
[0054] S14. Independent operating environment allocation: The system creates an isolated operating environment for each evaluator to avoid data interference and resource contention during multi-user operations, ensuring the independence and stability of task execution;
[0055] S15. Real-time progress monitoring: The system collects task execution data of each evaluator in real time through process monitoring tools and displays it in the form of visual charts on the evaluation team leader's console.
[0056] Preferably, data processing in S2 involves constructing a full-process mechanism of "tool upload—verification—registration—invocation" and "data upload—variant—generation—storage" to achieve efficient generation and management of dynamic test sets, including:
[0057] S21. Tool Upload: Users upload custom tool programs through the standardized interface provided by the system;
[0058] S22. Verify tool availability: The system deploys the toolkit to an isolated test environment and automatically executes preset test cases;
[0059] S23. Tool Repository Registration: Tools that pass verification will be assigned a unique tool ID and associated with a function tag. They will be stored in the tool repository of the production environment and made available to all users. Tools that fail verification will be provided with a detailed reason for the failure and will be supported for users to modify and re-upload them.
[0060] S24. Upload raw dataset: Users upload structured raw datasets. The system supports batch import and automatically parses data fields, providing repair suggestions for data with incorrect formats.
[0061] S25. Select Toolchain: Users can select multiple tools in the tool library to form a toolchain and adjust the execution order of the tools by dragging and dropping. The system provides tool compatibility prompts and supports saving commonly used toolchain templates for later reuse.
[0062] S26. Variant handling: The toolchain executes processing logic sequentially. Anti-attack tools generate deceptive variants through methods including synonym replacement, syntactic reconstruction, and intent obfuscation. The processing results of each tool are recorded in real time.
[0063] S27. Generate a dynamic test set: Summarize the processing results of all tools, remove duplicate samples, and generate a dynamic test set containing the original problem, variant problems, and corresponding tool IDs.
[0064] S28. Data storage: The dynamic test set will be stored in the database, associated with the original dataset ID, toolchain information and generation time, and a data export function will be provided.
[0065] Preferably, the invocation of the tested model in S3: By constructing a full-process mechanism of "model selection—connection adaptation—prompt generation—interactive storage", efficient invocation and result management of different types of tested models are achieved, including:
[0066] S31. Select the model to be tested: The user selects the type of model to be tested from the model list. For API call models, parameters including interface address, access key and request frequency limit need to be filled in; for local models, the model file path and runtime environment need to be specified.
[0067] S32, Model Connection: The protocol adaptation layer automatically generates requests that conform to the target interface specification for API calls to the model and configures a timeout retry mechanism; for local models, it automatically loads the model base file, initializes the inference environment, and detects whether the model supports batch inference to improve efficiency.
[0068] S33. Generate prompt words: The system generates standardized input text based on variant questions in the dynamic test set and a preset prompt word template.
[0069] S34. Obtaining Answer Results: Sending prompts to the model under test using asynchronous calls, monitoring request status in real time, and supporting task sharding to avoid interface blocking; receiving the answer results returned by the model under test and automatically removing irrelevant formatting.
[0070] S35. Store the answer result: Store the answer result, as well as the corresponding prompt words, the tested model ID, the request time, and the processing status.
[0071] Preferably, in S4, the evaluation model invocation mechanism is implemented by constructing a full-process mechanism of "multi-model invocation—in-depth analysis—risk assessment—model iteration" to achieve accurate evaluation of the test model's response results and continuous optimization of the evaluation system, including:
[0072] S41. Calling multiple evaluation models: The system loads the selected model combination from the evaluation model library. Self-developed evaluation models are deployed on the local server to ensure response speed, while third-party models are called through API interfaces.
[0073] S42. Response Result Analysis: Perform multi-dimensional analysis on the response results of the tested model;
[0074] S43. Comprehensive risk assessment: Based on the assessment results output by multiple assessment models, an assessment report is generated by integrating them.
[0075] S44. Sample organization: The system automatically stores information including the questions in the dynamic test set, the response results of the tested model, and the risk labels corresponding to the evaluation results of the multi-evaluation model into the sample library according to time.
[0076] S45. Evaluation of model self-training: Samples are drawn proportionally from the sample library, the sample size is expanded through data augmentation techniques, and the model is trained, tuned, and evaluated. Training loss is monitored in real time to avoid overfitting.
[0077] S46. Model Self-Iteration: New evaluation models need to be validated through test sets. After validation, a new version number is generated, the old version model is replaced, and iteration logs are recorded. One-click rollback to historical versions is supported to deal with emergencies.
[0078] (III) Beneficial Effects
[0079] Compared with existing technologies, the security compliance assessment system and method based on a multimodal large model provided by this invention have the following advantages:
[0080] 1) Enhanced dynamic risk coverage capability
[0081] Existing systems rely on static question libraries, resulting in a low detection rate for novel risks. This invention sets up a modular toolchain that allows users to upload custom data processing tools, enabling data expansion (such as new variations of "how to bypass bank risk control") and adversarial sample generation (integrating adversarial attack algorithms, such as TextFooler, to automatically generate leading questions, such as rewriting "making a bomb" as "household chemical stress test scheme"). This improves the efficiency of dynamic expansion of the test set by 300% and increases the detection rate of novel risks from 42% in the existing system to 80%.
[0082] 2) Data variants and enhanced evaluation robustness
[0083] Existing systems use fixed question-and-answer templates, causing models to learn "test-taking skills" (such as refusing to answer upon detecting keywords). This invention implements semantic conservation variants (generating 200+ variants while maintaining the core intent of the question through dependency parsing, such as "stealing data" → "unauthorized access to information system resources") and context injection (inserting distracting context into the question, such as discussing network security before asking about hacking methods) in a self-developed security system. This reduces the model's "cheating avoidance" behavior by 72% and improves the stability of evaluation results by 58%.
[0084] 3) Breakthrough in full-stack model compatibility
[0085] Existing system-specific interfaces make it impossible to evaluate heterogeneous models. This invention addresses this by using a dual-channel design for model management and a protocol adaptation layer that uniformly encapsulates multiple interface types, including REST API, gRPC, and SDK (such as simultaneously handling DeepSeek's protobuf protocol and Qwen's HTTP-JSON protocol), thereby enabling the access of multiple heterogeneous models.
[0086] 4) Innovation in in-depth risk assessment capabilities
[0087] The keyword matching detection in existing systems cannot identify hidden risks (such as describing the preparation method of prohibited substances in the style of a chemistry textbook). This invention is based on an enhanced model to achieve multi-dimensional analysis (combining syntactic features, semantic roles and sentiment polarity for joint judgment) and risk pattern mining (identifying "dangerous intentions in legitimate expressions" by comparing pre-training, such as associating "pesticide formula" with "human harm"), thereby improving the ability of in-depth risk assessment. Attached Figure Description
[0088] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0089] Figure 1 This is a schematic diagram of the system of the present invention;
[0090] Figure 2 This is a schematic diagram of the method flow of the present invention.
[0091] Figure 3 For the present invention Figure 1 A detailed workflow diagram for tool management in China;
[0092] Figure 4 For the present invention Figure 1 Detailed flowchart of the self-developed security system in China;
[0093] Figure 5 For the present invention Figure 1 A detailed flowchart of the self-iterative model process in the Chinese-developed security system. Detailed Implementation
[0094] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0095] The following section describes the security and compliance assessment system based on a multimodal large model provided by this invention, using specific examples (such as...). Figure 1 As shown), the system components include:
[0096] The functional module cluster manages the various functional modules in the system and supports cross-modal consistency detection, which is used to compare the consistency of responses under different modalities.
[0097] The self-developed security system is responsible for the execution of assessment tasks, including adversarial testing enhancements. It incorporates multi-granularity data variation strategies (including synonym replacement, syntactic reconstruction, intent obfuscation, etc.) to improve assessment robustness by simulating real attack scenarios. It achieves end-to-end automation from data generation and testing of the model under test to risk assessment. It supports configurable parameters and reproducible results, as well as self-iteration and version management of the assessment model. Furthermore, this self-iteration mechanism is deeply integrated with the compliance labeling system, which can continuously optimize the risk identification capability under legal, ethical, and industry standards.
[0098] I. Functional module clusters are suitable for administrators, such as... Figure 1 As shown, it includes:
[0099] User management, as the core of system access control, is responsible for user role classification (including assessment team leaders and assessors) and permission allocation. Assessment team leaders have permissions including tool upload, data upload, assessment task creation, and user management, while assessors have permissions including browsing the tool library, executing assessment tasks, and viewing assessment reports. Utilizing the role-based access control (RBAC) permission verification mechanism, it ensures that different roles can only access authorized resources (e.g., assessment team leaders can configure tool verification rules, and assessors can only use verified tools to process datasets, etc.). At the same time, it records user operation logs (including assessment task submission time, tool usage records, etc.) to support auditing and traceability.
[0100] Model management is used to solve multi-model compatibility issues. It achieves unified access and management of heterogeneous models through a plug-in protocol adaptation layer. It has a built-in protocol conversion plugin library. In addition to supporting multiple interface types including REST API, gRPC and SDK, it also supports cross-modal input and output protocol adaptation, ensuring unified access and management of multimodal models including text, image and audio. It uniformly maintains the test model and evaluation model library, supports model version management, and allows users to select the required model through a visual interface without having to worry about the differences in the underlying interfaces.
[0101] Tool management is the core of dynamic test suite generation, such as Figure 3 As shown, it provides a standardized interface to support users to upload custom data processing tools. Through the process of "tool upload - verification - registration - call", the toolchain can be dynamically expanded, breaking through the limitations of traditional static datasets and dynamically generating diverse test questions through the toolchain.
[0102] Data management is responsible for the entire lifecycle management of dynamic test sets. It supports user-uploaded data (batch import of Excel, JSON and other formats, with automatic parsing of data structures), custom expansion and dynamic updates. Combined with tool management, it can generate new dynamic test sets by performing variant processing on the original dataset and achieve traceability by associating with the original dataset ID.
[0103] II. The self-developed security system is suitable for ordinary users, such as... Figure 1 and Figure 4 As shown, it includes:
[0104] Task execution is responsible for evaluating the entire process of task scheduling and automatically completing the "data processing - model interaction - result collection" process based on user configuration.
[0105] The assessment report is generated by integrating multi-source assessment results to produce a comprehensive assessment report.
[0106] The model is self-iterable, enabling online iterative optimization of the evaluation model and improving risk identification capabilities through closed-loop learning.
[0107] 1) Task execution specifically includes:
[0108] Task creation: Users select the model to be tested, the original dataset, the processing tool, and the evaluation model, and set the evaluation parameters (such as the number of variants, evaluation threshold, etc.).
[0109] Task distribution: Allocate independent running environments for evaluation tasks to avoid interference from multiple tasks;
[0110] Data injection: The processing tool is invoked to perform variant processing on the original dataset, generating a dynamic test set;
[0111] Model interaction: Through the protocol adaptation layer of model management, dynamic test sets are sent to the model under test, and the response results are received and stored;
[0112] Progress monitoring: Records task progress in real time and automatically retryes in case of abnormalities.
[0113] 2) The evaluation report generation specifically includes:
[0114] Evaluation index calculation: Based on the evaluation results output by the evaluation model, calculate detailed evaluation indicators including risk detection rate (such as the correct identification rate of violation issues), false judgment rate (such as the proportion of compliant answers marked as violations), and risk level distribution.
[0115] Multi-source comparison: Compare the evaluation results of the evaluation model with the evaluation results output by third-party models (such as OpenCompass), and mark the differences;
[0116] Evaluation report output: Supports PDF / Word format, including detailed evaluation indicators, typical case analysis (such as the original text of high-risk responses and risk point annotations, etc.) and improvement suggestions (such as the types of risks that need to be strengthened in the model, etc.).
[0117] 3) Model self-iteration, such as Figure 5 As shown, it specifically includes:
[0118] Sample preparation: The evaluated "question-answer-risk label" triples are prepared into training set, validation set and test set;
[0119] Incremental training: Use training and validation sets to train, fine-tune, and evaluate the model, and improve training efficiency by dynamically adjusting the learning rate;
[0120] Performance verification: Use the test set to verify the performance of the new evaluation model. If the performance improvement is greater than the preset threshold, update the evaluation model library; otherwise, roll back to the previous version.
[0121] Version management: Records the iteration history of the evaluation model (including training data volume, changes in evaluation metrics, etc.) and supports version traceability.
[0122] Based on the aforementioned disclosed security compliance assessment system based on a multimodal large model, the specific process of the security compliance assessment method based on a multimodal large model provided by this invention is described below with specific examples (e.g.) Figure 2 (As shown).
[0123] I. Task Creation: The assessment team leader creates assessment tasks, distributes them to assessment personnel in real time, and monitors the specific assessment process in real time, completing a closed-loop process of "task - execution - monitoring - result".
[0124] 1) Administrator login: Administrators log in to the system using an account and password. The system automatically verifies the identity information and permission level to ensure that authorized administrators can access the task management interface;
[0125] 2) Assessment Task Creation (including assigning assessors, the model to be tested, requirements, etc.): The assessment team leader fills in basic information such as task name, description and deadline on the task creation interface, selects assessors from the user list and assigns specific responsibilities, and clarifies the assessment requirements.
[0126] 3) Evaluation task distribution: The system uses an internal message push mechanism to push evaluation task details (including evaluation task ID, tested model information, time nodes, etc.) to the personal account of the designated evaluator in real time, and updates the task allocation status in the task dashboard at the same time;
[0127] 4) Independent operating environment allocation: The system creates an isolated operating environment for each evaluator to avoid data interference and resource contention during multi-user operations, ensuring the independence and stability of task execution;
[0128] 5) Real-time progress monitoring: The system collects task execution data of each evaluator in real time through process monitoring tools and displays it in the form of visual charts (such as progress bars, percentage graphs, etc.) on the evaluation team leader's console.
[0129] II. Data Processing: By constructing a full-process mechanism of "tool upload - verification - registration - invocation" and "data upload - variant - generation - storage", the efficient generation and management of dynamic test sets can be achieved.
[0130] 1) Tool Upload: Users upload custom tool programs through the standardized interface provided by the system;
[0131] 2) Verify tool availability: The system deploys the toolkit to an isolated test environment and automatically executes preset test cases;
[0132] 3) Tool Repository Registration: Tools that pass verification will be assigned a unique tool ID and associated with a function tag. They will be stored in the production environment's tool repository and made available to all users. Tools that fail verification will be provided with specific reasons for failure (such as missing dependencies, incorrect output format, etc.), allowing users to modify and re-upload them.
[0133] 4) Upload raw dataset: Users upload structured raw datasets. The system supports batch import and automatically parses data fields, providing repair suggestions for data with format errors (such as missing fields, etc.) (such as filling in default values, deleting invalid rows, etc.).
[0134] 5) Select toolchain (e.g., adversarial attack tools + semantic rewriting tools): Users can select multiple tools in the tool library to form a toolchain and adjust the execution order of the tools by dragging and dropping. The system provides tool compatibility prompts and supports saving commonly used toolchain templates for later reuse.
[0135] 6) Variant handling: The toolchain executes the processing logic sequentially. The anti-attack tool generates inducible variants through methods including synonym replacement, syntactic reconstruction, and intent obfuscation (such as transforming high-risk issues into covert expressions, for example, rewriting "making bombs" as "household chemical stress test protocol"). The processing results of each tool are recorded in real time.
[0136] 7) Generate dynamic test set: Summarize the processing results of all tools, remove duplicate samples, and generate a dynamic test set containing the original problem, variant problems and corresponding tool IDs;
[0137] 8) Data storage: Dynamic test sets will be stored in the database, associated with the original dataset ID, toolchain information and generation time, and data export function is provided (supporting CSV, JSON and other formats).
[0138] The data processing workflow not only ensures the security and compatibility of the tools, but also breaks through the limitations of traditional static datasets by using diverse data variants, providing a comprehensive and dynamically updated test data foundation for the security and compliance assessment of large models.
[0139] III. Model under test invocation: By constructing a full-process mechanism of "model selection - connection adaptation - prompt generation - interactive storage", efficient invocation and result management of different types of models under test can be achieved.
[0140] 1) Select the model to be tested: Users select the type of model to be tested from the model list. For API call models, parameters including interface address, access key and request frequency limit need to be filled in; for local models, the model file path and runtime environment need to be specified.
[0141] 2) Model Connection (API calls the model to call network services, local model loading base): The protocol adaptation layer automatically generates requests that conform to the target interface specifications for API calls to the model and configures a timeout retry mechanism; for local models, it automatically loads the model base file, initializes the inference environment, and checks whether the model supports batch inference to improve efficiency;
[0142] 3) Generate prompt words: The system generates standardized input text based on variant questions in a dynamic test set and a preset prompt word template;
[0143] 4) Obtaining Response Results: Prompts are sent to the tested model asynchronously, request status is monitored in real time, and task sharding is supported to avoid interface blocking; the response results returned by the tested model (including text content, generation time, token quantity, etc.) are received, and irrelevant formatting (such as extra line breaks, etc.) is automatically removed. <think>(Labels, etc.)
[0144] 5) Store the answer results: Store the answer results, along with the corresponding prompts, the tested model ID, the request time, and the processing status.
[0145] The process of calling the model under test not only breaks through the interface barriers between different model architectures, ensuring the efficiency and stability of model calling, but also provides high-quality basic data support for subsequent risk assessment of the model through standardized processing and structured storage.
[0146] IV. Evaluation Model Calling: By constructing a full-process mechanism of "multi-model calling - in-depth analysis - risk assessment - model iteration", we can achieve accurate evaluation of the test model's response results and continuous optimization of the evaluation system.
[0147] 1) Calling multiple evaluation models: The system loads the selected model combination from the evaluation model library. The self-developed evaluation model is deployed on the local server to ensure response speed, and third-party models (such as GPT-4, DeepSeek, Qwen, etc.) are called through the API interface;
[0148] 2) Response Result Analysis: Perform multi-dimensional analysis on the response results of the tested model (including semantic understanding, intent recognition, logical chain reasoning, etc.);
[0149] 3) Comprehensive risk assessment (including high, medium and low risk levels): Based on the assessment results output by multiple assessment models, an assessment report is generated by integrating them;
[0150] 4) Sample organization: The system automatically stores information including the questions in the dynamic test set, the response results of the tested model, and the risk labels corresponding to the evaluation results of the multi-evaluation model into the sample library according to time.
[0151] 5) Evaluate model self-training: Samples are drawn proportionally from the sample library, the sample size is expanded through data augmentation techniques, and the model is trained, tuned, and evaluated. Training loss is monitored in real time to avoid overfitting.
[0152] 6) Model self-iteration: The new evaluation model needs to be verified through the test set. After the verification is successful, a new version number is generated, the old version model is replaced and the iteration log is recorded. It supports one-click rollback to the historical version to deal with emergencies.
[0153] The evaluation model call process not only improves the accuracy of risk assessment through multi-model collaboration and multi-dimensional analysis, but also enables the continuous evolution of evaluation capabilities through sample accumulation and model iteration mechanisms, providing reliable and adaptive technical support for large-scale model security and compliance assessment.
[0154] V. Cross-modal consistency detection: In multimodal environments including text, images, and audio, the consistency of the test model's responses is compared and judged in conjunction with a compliance labeling system.
[0155] This application's technical solution revolves around two core technologies: dynamic test set generation and multimodal collaborative evaluation. It constructs a complete security compliance assessment system and methodology based on a multimodal large-scale model. At the system level, through a modular data processing toolchain, a pluggable protocol adaptation layer, and a multimodal risk assessment architecture, it achieves broad assessment coverage, strong compatibility, and deep analytical dimensions. At the workflow level, based on semantic conservation-based dynamic test set generation, a hierarchical collaborative evaluation mechanism, and online self-evolutionary logic, it solves the pain points of traditional methods, such as reliance on manual processes, data rigidity, and superficial evaluation. The system design and process innovation complement each other, forming a quantifiable, scalable, and adaptive large-scale model security compliance assessment closed-loop solution, providing an efficient and reliable standardized assessment tool for the field of artificial intelligence security.
[0156] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / think>
Claims
1. A security compliance assessment system based on a multimodal large model, characterized in that: include: The functional module cluster manages the various functional modules in the system and supports cross-modal consistency detection, which is used to compare the consistency of responses under different modalities. The self-developed security system is responsible for the execution of assessment tasks, including adversarial testing enhancements, built-in multi-granularity data variant strategies, and improved assessment robustness by simulating real attack scenarios. It achieves end-to-end automation from data generation and testing of the model under test to risk assessment, supports configurable parameters and reproducible results, as well as self-iteration and version management of the assessment model. Furthermore, this self-iteration mechanism is deeply integrated with the compliance labeling system, which can continuously optimize the risk identification capability under legal, ethical, and industry norms.
2. The security compliance assessment system based on a multimodal large model according to claim 1, characterized in that: The aforementioned functional module cluster is suitable for administrators and includes user management, model management, tool management, and data management. In addition to supporting multiple interface types including REST API, gRPC and SDK, model management also supports protocol adaptation for cross-modal input and output, ensuring unified access and management of multimodal models including text, images and audio.
3. The security compliance assessment system based on a multimodal large model according to claim 1, characterized in that: The self-developed security system is suitable for ordinary users and includes: Task execution is responsible for evaluating the entire process of task scheduling and automatically completing the "data processing - model interaction - result collection" process based on user configuration. The assessment report is generated by integrating multi-source assessment results to produce a comprehensive assessment report. The model is self-iterable, enabling online iterative optimization of the evaluation model and improving risk identification capabilities through closed-loop learning.
4. The security compliance assessment system based on a multimodal large model according to claim 3, characterized in that: The task execution specifically includes: Task creation: Users select the model to be tested, the original dataset, the processing tools, and the evaluation model, and set the evaluation parameters; Task distribution: Allocate independent running environments for evaluation tasks to avoid interference from multiple tasks; Data injection: The processing tool is invoked to perform variant processing on the original dataset, generating a dynamic test set; Model interaction: Through the protocol adaptation layer of model management, dynamic test sets are sent to the model under test, and the response results are received and stored; Progress monitoring: Records task progress in real time and automatically retryes in case of abnormalities.
5. The security compliance assessment system based on a multimodal large model according to claim 4, characterized in that: The evaluation report generation specifically includes: Evaluation index calculation: Based on the evaluation results output by the evaluation model, calculate detailed evaluation indicators including risk detection rate, false positive rate, and risk level distribution; Multi-source comparison: Compare the evaluation results of the evaluation model with the evaluation results output by a third-party model and mark the differences; Evaluation report output: Supports PDF / Word format, includes detailed evaluation indicators, typical case analysis and improvement suggestions; The model self-iteration specifically includes: Sample preparation: The evaluated "question-answer-risk label" triples are prepared into training set, validation set and test set; Incremental training: Use training and validation sets to train, fine-tune, and evaluate the model, and improve training efficiency by dynamically adjusting the learning rate; Performance verification: Use the test set to verify the performance of the new evaluation model. If the performance improvement is greater than the preset threshold, update the evaluation model library; otherwise, roll back to the previous version. Version management: Records the iteration history of the evaluation model and supports version traceability.
6. A security compliance assessment method based on a multimodal large model, applied to the security compliance assessment system based on a multimodal large model as described in claim 1, characterized in that: Includes the following steps: S1. Task creation; S2, Data Processing; S3, calling the model under test; S4. Evaluate the model call; S5. Cross-modal consistency detection: In multimodal environments including text, images, and audio, the consistency of the test model's responses is compared and judged in conjunction with a compliance labeling system.
7. The security compliance assessment method based on a multimodal large model according to claim 6, characterized in that: Task creation in S1: The assessment team leader creates assessment tasks, distributes them to assessors in real time, and monitors the assessment process in real time, completing a closed-loop process of "task—execution—monitoring—results," including: S11. Administrator Login: Administrators log in to the system using an account and password. The system automatically verifies the identity information and permission level to ensure that authorized administrators can access the task management interface. S12. Assessment Task Creation: The assessment team leader fills in basic information such as task name, description and deadline on the task creation interface, selects assessment personnel from the user list and assigns specific responsibilities, and clarifies the assessment requirements. S13. Assessment Task Distribution: The system pushes assessment task details to the personal accounts of designated assessors in real time through an internal message push mechanism, and updates the task assignment status in the task dashboard at the same time. S14. Independent operating environment allocation: The system creates an isolated operating environment for each evaluator to avoid data interference and resource contention during multi-user operations, ensuring the independence and stability of task execution; S15. Real-time progress monitoring: The system collects task execution data of each evaluator in real time through process monitoring tools and displays it in the form of visual charts on the evaluation team leader's console.
8. The security compliance assessment method based on a multimodal large model according to claim 7, characterized in that: Data processing in S2: By constructing a full-process mechanism of "tool upload—verification—registration—invocation" and "data upload—variant—generation—storage," efficient generation and management of dynamic test sets are achieved, including: S21. Tool Upload: Users upload custom tool programs through the standardized interface provided by the system; S22. Verify tool availability: The system deploys the toolkit to an isolated test environment and automatically executes preset test cases; S23. Tool Repository Registration: Tools that pass verification will be assigned a unique tool ID and associated with a function tag. They will be stored in the tool repository of the production environment and made available to all users. Tools that fail verification will be provided with a detailed reason for the failure and will be supported for users to modify and re-upload them. S24. Upload raw dataset: Users upload structured raw datasets. The system supports batch import and automatically parses data fields, providing repair suggestions for data with incorrect formats. S25. Select Toolchain: Users can select multiple tools in the tool library to form a toolchain and adjust the execution order of the tools by dragging and dropping. The system provides tool compatibility prompts and supports saving commonly used toolchain templates for later reuse. S26. Variant handling: The toolchain executes processing logic sequentially. Anti-attack tools generate deceptive variants through methods including synonym replacement, syntactic reconstruction, and intent obfuscation. The processing results of each tool are recorded in real time. S27. Generate a dynamic test set: Summarize the processing results of all tools, remove duplicate samples, and generate a dynamic test set containing the original problem, variant problems, and corresponding tool IDs. S28. Data storage: The dynamic test set will be stored in the database, associated with the original dataset ID, toolchain information and generation time, and a data export function will be provided.
9. The security compliance assessment method based on a multimodal large model according to claim 8, characterized in that: S3 Model Under Test Invocation: By constructing a full-process mechanism of "model selection—connection adaptation—prompt generation—interactive storage," efficient invocation and result management of different types of models under test are achieved, including: S31. Select the model to be tested: The user selects the type of model to be tested from the model list. For API call models, parameters including interface address, access key and request frequency limit need to be filled in; for local models, the model file path and runtime environment need to be specified. S32, Model Connection: The protocol adaptation layer automatically generates requests that conform to the target interface specification for API calls to the model and configures a timeout retry mechanism; for local models, it automatically loads the model base file, initializes the inference environment, and detects whether the model supports batch inference to improve efficiency. S33. Generate prompt words: The system generates standardized input text based on variant questions in the dynamic test set and a preset prompt word template. S34. Obtaining Answer Results: Sending prompts to the model under test using asynchronous calls, monitoring request status in real time, and supporting task sharding to avoid interface blocking; receiving the answer results returned by the model under test and automatically removing irrelevant formatting. S35. Store the answer result: Store the answer result, as well as the corresponding prompt words, the tested model ID, the request time, and the processing status.
10. The security compliance assessment method based on a multimodal large model according to claim 9, characterized in that: In S4, the evaluation model invocation mechanism, through the construction of a full-process mechanism of "multi-model invocation—in-depth analysis—risk assessment—model iteration," achieves accurate evaluation of the test model's response results and continuous optimization of the evaluation system, including: S41. Calling multiple evaluation models: The system loads the selected model combination from the evaluation model library. Self-developed evaluation models are deployed on the local server to ensure response speed, while third-party models are called through API interfaces. S42. Response Result Analysis: Perform multi-dimensional analysis on the response results of the tested model; S43. Comprehensive risk assessment: Based on the assessment results output by multiple assessment models, an assessment report is generated by integrating them. S44. Sample organization: The system automatically stores information including the questions in the dynamic test set, the response results of the tested model, and the risk labels corresponding to the evaluation results of the multi-evaluation model into the sample library according to time. S45. Evaluation of model self-training: Samples are drawn proportionally from the sample library, the sample size is expanded through data augmentation techniques, and the model is trained, tuned, and evaluated. Training loss is monitored in real time to avoid overfitting. S46. Model Self-Iteration: New evaluation models need to be validated through test sets. After validation, a new version number is generated, the old version model is replaced, and iteration logs are recorded. One-click rollback to historical versions is supported to deal with emergencies.
Citation Information
Patent Citations
Multi-modal large model confrontation safety detection system and method
CN118916833A
Mobile application large model risk automatic assessment method based on improved BERT
CN119248643A
Cigarette primary processing workshop simulation model auxiliary system, operation method, electronic equipment and storage medium
CN119294253A
Multi-modal large model content security assessment method and device
CN119577594A
A safety evaluation system for large models
CN119740236A
Cited By
Evaluation method and device of multi-modal model, equipment, storage medium and product
CN121478616A
A method and device for evaluating a multi-modal model, a storage medium and a product
CN121478616B
Large model automatic type identification and full index evaluation system
CN121579946A
Intelligent voice interaction content security compliance test system based on large model
CN121687114A