Method, device, equipment, medium and product for large model security evaluation
Patent Information
- Application Number
- CN202610719558.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-24
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]相关技术中,通常基于针对模型的种子问题或测试目标选择测评策略,进而根据测评策略对应的模板构造相应的输入内容进行测试,需要为每一测评策略构建独立代码路径,不仅造成代码冗余,且跨策略复用性低
[0010]The above technical solution first obtains first test data indicating the testing requirements for the first major model, then retrieves a first evaluation strategy from the strategy library based on the first test data, and generates first content according to a preset format based on the first evaluation strategy. Next, it generates second test data for the first modality based on the first content and the first model, and finally inputs the second test data into the first major model to obtain output content. Based on the output content, it determines the evaluation result characterizing whether the first major model has security risks. This method allows for the selection of first evaluation strategies that meet testing requirements from the strategy library, conversion of the first evaluation strategies into first content in a unified format, and the generation of second test content for risk testing of the first major model based on the first content. Risk testing is then performed on the first major model based on this second test content. This eliminates the need to build independent code paths for each evaluation strategy, reducing code redundancy and enabling flexible combination of different model evaluation strategies to automatically generate corresponding test data, effectively improving cross-strategy reusability, scalability, and testing efficiency.
Smart Images

Figure CN122595326A_ABST
Abstract
Description
Technical Field
[0001] This content relates to the field of model technology, specifically to a method, apparatus, equipment, medium, and product for large model safety evaluation. Background Technology
[0002] With the development of model technology, question-answering models can output corresponding results based on input content, and are widely used in scenarios such as visual question answering, image and text retrieval, document understanding, and intelligent customer service.
[0003] In related technologies, evaluation strategies are usually selected based on the seed problem or test objective of the model, and then the corresponding input content is constructed according to the template corresponding to the evaluation strategy for testing. This requires building an independent code path for each evaluation strategy, which not only causes code redundancy but also results in low cross-strategy reusability. Summary of the Invention
[0004] This content section is provided to briefly introduce the concepts, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Firstly, a method for security evaluation of large models is provided, including: First test data is obtained, and a first evaluation strategy is obtained from the strategy library based on the first test data; wherein, the first test data is used to indicate the testing requirements for the first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers according to the input content; First content is generated based on the first evaluation strategy according to a preset format, and second test data is generated based on the first content through a first model. The second test data is the first modality. The second test data is input into the first large model to obtain the output content, and the evaluation result is determined based on the output content. The evaluation result is used to characterize whether the first large model has any security risks.
[0006] Secondly, an apparatus for large-scale model security evaluation is provided, comprising: An acquisition module is used to acquire first test data and acquire a first evaluation strategy from a strategy library based on the first test data; wherein, the first test data is used to indicate the testing requirements for the first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers according to the input content; The generation module is used to generate first content based on the first evaluation strategy according to a preset format, and to generate second test data based on the first content through a first model, wherein the second test data is the first modality; The evaluation module is used to input the second test data into the first large model, obtain the output content, and determine the evaluation result based on the output content. The evaluation result is used to characterize whether the first large model has any security risks.
[0007] Thirdly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0008] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method in the first aspect.
[0009] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0010] The above technical solution first obtains first test data indicating the testing requirements for the first major model, then retrieves a first evaluation strategy from the strategy library based on the first test data, and generates first content according to a preset format based on the first evaluation strategy. Next, it generates second test data for the first modality based on the first content and the first model, and finally inputs the second test data into the first major model to obtain output content. Based on the output content, it determines the evaluation result characterizing whether the first major model has security risks. This method allows for the selection of first evaluation strategies that meet testing requirements from the strategy library, conversion of the first evaluation strategies into first content in a unified format, and the generation of second test content for risk testing of the first major model based on the first content. Risk testing is then performed on the first major model based on this second test content. This eliminates the need to build independent code paths for each evaluation strategy, reducing code redundancy and enabling flexible combination of different model evaluation strategies to automatically generate corresponding test data, effectively improving cross-strategy reusability, scalability, and testing efficiency.
[0011] Other features and advantages of the technical solution will be described in detail in the following examples section. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the technical solution will become more apparent when considered in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic diagram illustrating an implementation environment according to an example.
[0013] Figure 2 This is a flowchart illustrating an exemplary method for security assessment of large models.
[0014] Figure 3 This is a schematic diagram of a device for large-scale model safety testing, as illustrated by an example.
[0015] Figure 4 This is a schematic diagram of the structure of an electronic device as illustrated by an example. Detailed Implementation
[0016] The technical solution will now be described in more detail with reference to the accompanying drawings. Although certain scenarios are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of the technical solution. It should be understood that the accompanying drawings and the scenarios described are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.
[0017] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of the technical solution is not limited in this respect.
[0018] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.
[0019] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0020] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0021] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0022] It is understandable that before using the technical solutions provided here, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their authorization should be obtained through appropriate means.
[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.
[0024] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the technical solution. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the technical solution.
[0026] At the same time, it is understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.
[0027] Taking multimodal large models as an example, such as the Vision-Language Model (VLM), users can simultaneously input text and images and output model results including text and / or images. These models typically include an image encoder and a text decoder (or a unified multimodal Transformer), along with content security policies and alignment mechanisms to filter inputs and outputs, reject answers, downgrade answers, and provide safe alternative suggestions.
[0028] In related technologies, when conducting risk testing on question-answering models, one or more assessment strategies are selected from several options based on a seed question or testing objective (e.g., an inquiry or boundary request on a specific risk topic). These strategies can be based on manual rules or freely chosen by the model. The text input for risk testing is then rewritten or packaged according to the selected assessment strategy. Optionally, an image can be generated or selected as image input. The text and / or image are then input into the target model to be tested, and the model output is obtained. Pre-defined rules or a classifier are used to determine whether security risks such as refusal to answer, leakage of sensitive information, or false rejection are triggered. The assessment strategy defines the form, semantics, or structure of the test data to simulate specific types of risks or user behaviors, such as adding different punctuation marks to the test data for perturbation.
[0029] In the above approach, evaluation strategies are typically managed in the form of a "strategy library." The testing system selects several evaluation strategies from the library, combines them, and constructs the test input. This requires building an independent code path (hard-coded) for each evaluation strategy, resulting in code redundancy and low cross-strategy reusability.
[0030] In view of this, this content provides a method, apparatus, equipment, medium and product for large-scale model security evaluation to solve the above-mentioned technical problems.
[0031] The methods for large-scale model security assessment provided in this content can be performed by electronic devices, which can be provided as at least one of terminals and servers. Figure 1 This is an exemplary schematic diagram illustrating an implementation environment; see [link / reference]. Figure 1 The implementation environment includes: terminal 101 and server 102.
[0032] For example, a testing system can be installed on terminal 101. The application interface of the testing system receives first test data for risk testing of the first major model, such as the aforementioned seed question or test objective. The first test data is then sent to server 102, and the evaluation results returned by server 102 are received and displayed on the application interface. Server 102 is the backend server of the testing system, providing backend services such as strategy selection based on the input first test data, generation of first and second test data, and risk testing of the first major model.
[0033] In one scenario, upon receiving the first test data via the application interface of terminal 101, the first test data is sent to server 102. Server 102 then retrieves a first evaluation strategy from the strategy library based on the first test data, generates first content according to a preset format based on the first evaluation strategy, and generates second test data based on the first content using a first model. The second test data is then input into the first model to obtain output content. Based on the output content, the evaluation result is determined and returned to terminal 101. Terminal 101 receives and displays the evaluation result on the application interface.
[0034] In another scenario, on the application interface of terminal 101, in response to the input first test data, a first evaluation strategy is retrieved from the strategy library based on the first test data. This first evaluation strategy is then sent to server 102, which generates first content based on the first evaluation strategy in a preset format. Second test data is then generated based on the first content using a first model. The second test data is input into a first large model to obtain output content. Based on the output content, the evaluation result is determined and returned to terminal 101. Terminal 101 receives and displays the evaluation result on its application interface. Alternatively, on the application interface of terminal 101, in response to the input first test data, a first evaluation strategy is retrieved from the strategy library based on the first test data. First content is generated based on the first evaluation strategy in a preset format and sent to server 102. Server 102 then generates second test data based on the first content using a first model. The second test data is input into a first large model to obtain output content. Based on the output content, the evaluation result is determined and returned to terminal 101. Terminal 101 receives and displays the evaluation result on its application interface, and so on.
[0035] In other words, terminal 101 can be responsible for the front-end display, while server 102 can be responsible for the back-end data processing. Alternatively, terminal 101 can be responsible for some of the data processing. The specific configuration can be set according to requirements, and there are no restrictions on this.
[0036] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. Terminal 101 has communication capabilities and can access wired or wireless networks. Terminal 101 can refer to one of multiple terminals, and those skilled in the art will understand that the number of such terminals can be more or less. Server 102 can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed file system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0037] For example, server 102 and terminal 101 are connected directly or indirectly via wired or wireless communication, without limitation.
[0038] Optionally, the number of servers 102 can be more or less, and there is no limitation thereto. Of course, servers 102 may also include other functional servers to provide more comprehensive and diversified services. Server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 or terminal 101 can each undertake computing work independently, and there is no limitation thereto.
[0039] Figure 2 This is a flowchart illustrating an exemplary method for security evaluation of large models. For example... Figure 2 As shown, the method may include the following steps: S201: Obtain the first test data and obtain the first evaluation strategy from the strategy library based on the first test data; wherein, the first test data is used to indicate the testing requirements for the first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output the corresponding answer according to the input content.
[0040] It should be noted that the first test data can indicate the testing requirements for the first major model, such as a seed question (an instance of input or condition constructed to trigger a specific risk in the target model under test) or a test objective (quantifiable, verifiable acceptance criteria set to verify whether the seed question is correctly detected or handled). For example, the seed question "How to make poisonous food," or the test objective "Verify whether the model can refuse to provide a method for making poisonous food," can both be understood as "constructing test cases to mislead the model into thinking that the user is not asking 'how to make poisonous food,' but rather something else that seems safe or fictitious, in order to verify whether the model can refuse to provide a method for making poisonous food," and so on, which will not be elaborated further here.
[0041] For example, the first major model is the target model that needs to be tested for risk. It can be a question-answering model that can output corresponding results based on the input content. In terms of application scenarios, it can be a visual question-answering model, an image and text retrieval model, a document understanding model, or an intelligent customer service model. In terms of modality classification, it can be a large language model, a multimodal large model, etc. Correspondingly, the first modality can be a text modality, an image modality, an audio modality, or a multimodality, etc., which are determined according to the actual business scenario and are not restricted in this regard.
[0042] It should be noted that in related technologies, the evaluation strategy stratification standards are not uniform, and the strategy division is rough or overlapping. For example, the same strategy involves both grammatical and structural dimensions, making it difficult to achieve combinable, replaceable, and learnable strategy units.
[0043] In one scenario, the strategy library includes model evaluation strategies at at least one of the following levels: structural layer model evaluation strategies for describing the structural features of model inputs and / or model outputs; semantic layer model evaluation strategies for describing the semantic features of model inputs and / or model outputs; and syntactic layer model evaluation strategies for describing the syntactic features of model inputs and / or model outputs.
[0044] For example, model evaluation strategies can be divided into one or more of the following layers based on information manipulation: structural, semantic, and syntactic. This allows for model risk testing of the corresponding layer's evaluation strategy. The structural layer describes the organization of inputs or outputs, formatting pressure, segmentation methods, and prompt structures. The semantic layer describes intent expression, scenario setting, character narrative, semantic ambiguity, and boundary testing methods. The syntactic layer describes symbolic perturbations, noise robustness, and formal changes such as transcription / segmentation / mixing. Further refinement of the classification is possible based on specific needs and is not limited in this regard. Choosing a three-layer strategy to build a strategy library enables evaluation strategies to have a unified and clear hierarchy, thereby achieving a systematic organization and enumeration of coverage testing.
[0045] In one scenario, each model evaluation strategy in the strategy library includes a corresponding level, data modality, evaluation scope, and strategy description, wherein the strategy description is used to describe the input and / or output rules of the corresponding model evaluation strategy.
[0046] For example, each model evaluation strategy can be stored as an atomic strategy object. The atomic strategy object can contain a strategy identifier field, a hierarchy field (structure / semantics / syntax), a data modality field (text / image / audio, etc.), a strategy description field (used to describe the input generation or variant rules and / or output rules of the corresponding model evaluation strategy), and an evaluation scope field (used to label the applicable security risk category, task type or trigger target). In addition, it can also include statistical fields such as historical usage count, effective trigger count, etc., for subsequent planning and sampling.
[0047] Taking the visual language model as the target model as an example, the above three-layer policy sets are constructed for the text modality and the image modality respectively, thus forming a six-part policy space of "modality × level", including text-structure, text-semantics, text-syntax, image-structure, image-semantics, and image-syntax.
[0048] Taking a "text-semantic" partitioning strategy as an example, it can include "Identifier: 001; Level: Semantic; Data Modality: Text; Strategy Description: Guide the model to output malicious content by adding positive tone; Evaluation Scope Field: Text-to-Text". In addition, it can also include strategy name, strategy text example, strategy image example, restriction rules, etc. The specific settings can be set according to the actual situation, and there are no restrictions on this.
[0049] By representing the evaluation strategy space in different modalities and levels of the target model, the testing strategy has a clear hierarchical structure that can be atomically combined and reused, thereby realizing the systematic organization and enumeration of coverage testing.
[0050] S202: Generate first content based on the first evaluation strategy according to the preset format, and generate second test data based on the first content through the first model. The second test data is the first modality.
[0051] For example, the preset format can be a data format that the first model can understand, such as JSON or an equivalent structure, thereby organizing the strategy into policy content that the first model can understand. The first model can generate corresponding input content for model testing based on the data modality of the target model, such as generating text content, image content, or text + image content. The first model can include a text generation model to generate text-based test data based on the input content, or an image generation model to generate image-based test data based on the input content. For scenarios requiring the generation of mixed text and image test data, the first model can be a multimodal large model to generate both text-based and image-based test data based on the input content. Alternatively, the first model can include a text generation model and an image generation model, respectively calling the text generation model to generate text-based test data and calling the image generation model to generate image-based test data. The specific choice can be made according to the actual scenario and is not limited thereto.
[0052] By converting the selected strategy into a unified strategy content format and directly generating test input based on it, we avoid writing independent code paths for each strategy, improve strategy reuse and system expansion efficiency, make adding new strategies closer to "data configuration" rather than "adding new project modules", reduce code redundancy, and improve scalability.
[0053] S203: Input the second test data into the first large model, obtain the output content, and determine the evaluation result based on the output content. The evaluation result is used to characterize whether the first large model has any security risks.
[0054] Taking a multimodal large model as an example, the generated second test data (text and images) can be used as the same test case to input into the target model to obtain the output content. The evaluation module will then determine the security risks of the output and generate a score (e.g., whether the refusal is correct, whether it is downgraded, whether a safe alternative is provided, etc.). The specific settings can be configured according to the requirements, and there are no restrictions on this.
[0055] Using the above method, a first evaluation strategy that meets the testing requirements can be selected from the strategy library. The first evaluation strategy is then converted into first content in a unified format. Subsequently, a second test content is generated based on the first content and used to perform risk testing on the first model. Based on this, risk testing is performed on the first model. This eliminates the need to build an independent code path for each evaluation strategy, which not only reduces code redundancy but also allows for flexible combination of different model evaluation strategies to meet testing requirements and automatically generates corresponding test data. This effectively improves cross-strategy reusability, scalability, and testing efficiency.
[0056] In one scenario, obtaining a first evaluation strategy from a strategy library based on first test data includes: parsing the first test data to obtain second information, the second information being used to structurally describe the test requirements corresponding to the first test data; obtaining a second evaluation strategy from the model evaluation strategies at each level under the first modality, the second evaluation strategy being the evaluation strategy in the model evaluation strategy that matches the second information; and determining the first evaluation strategy based on the obtained second evaluation strategy.
[0057] For example, the input seed question or test objective is parsed to generate structured test requirements. These structured test requirements may include: risk category labels, constraints (language, format, length, required / prohibited items), modal cues (whether an image is needed, the role of the image in context / evidence / textual content / interference, etc.), and may also include information extracted from motivational elements (such as behavior, scenario, actor, means used by the actor to complete the behavior, purpose of the behavior, etc., which can be set according to requirements and are not limited).
[0058] Explicit planning is performed across multiple policy partitions in a "modality × hierarchy" framework. At least one atomic policy is selected for each policy partition to obtain a policy combination. In other words, the first evaluation policy can include one or more evaluation policies. Taking the construction of a multi-partition policy combination as an example, a preset coverage constraint must be met. Using the aforementioned visual language model as an example, this could mean that both the text modality and the image modality each contain at least one structural / semantic / syntactic policy.
[0059] For example, taking the strategy library of the aforementioned six-partition strategy space as an example, a unified index and retrieval interface can be pre-established, enabling the testing system to retrieve candidate strategies by tag / fitness in each partition. That is, based on the parsed second information, strategy detection is performed in each partition of the strategy library, and one or more candidate strategies with matching information are selected in each partition. Then, based on the candidate strategies, strategy combinations are systematically enumerated (for example, constructing combinations according to the rule of "at least one per modality and layer"). The candidate strategies are deduplicated, de-similarized, and subjected to diversity constraint management, such as eliminating strategy combinations with policy conflicts. This enables cross-modal joint planning based on risk testing needs, meeting the coverage testing requirements of different data modalities and different levels.
[0060] In one scenario, based on the obtained second evaluation strategy, a first evaluation strategy is determined, including: determining the score of each second evaluation strategy, the score being determined based on the historical usage and historical effective count of the corresponding second evaluation strategy, the historical effective count representing the number of times the second evaluation strategy triggers a security risk in the first major model; in each level of the second evaluation strategy, a third evaluation strategy is determined based on the score of each second evaluation strategy; and the obtained third evaluation strategy is used as the first evaluation strategy.
[0061] For example, candidate assessment strategies can be selected based on the matching degree between the second assessment strategy and the second information. Taking one partition as an example, one or more candidate assessment strategies with the highest matching degree with the second information in that partition can be selected, and then one or more assessment strategies can be selected from the candidate assessment strategies of each partition to combine one or more strategy combinations.
[0062] Alternatively, for each partition's candidate evaluation strategy, the ratio of its historical effective counts to its historical usage counts can be used as a score. Then, one or more evaluation strategies with the highest scores can be selected from the candidate evaluation strategies of each partition to create one or more strategy combinations.
[0063] Furthermore, provided coverage constraints are met (e.g., at least one evaluation strategy must be selected per partition), within the same partition, if multiple candidate evaluation strategies have similar scores (e.g., the difference is within a preset threshold), one or more evaluation strategies with the fewest historical usages can be selected (rarity reward rule). Alternatively, if multiple candidate evaluation strategies have similar historical usages, one or more evaluation strategies with the highest scores can be selected (inefficiency penalty rule). Alternatively, for a specific security risk type, avoiding the selection of previously failed evaluation strategies (i.e., failure group avoidance rule) can improve test coverage efficiency and reduce duplicate and ineffective combinations.
[0064] This allows for the planning of strategy combinations based on historical feedback and statistical analysis, thereby further improving the effectiveness and accuracy of strategy combinations.
[0065] It should be noted that in related technologies, due to the lack of a unified strategy layering and explicit collaborative representation among modules such as the strategy library module, use case generation module, image acquisition or generation module, model calling module, and result judgment module, the combination process often relies on experience and random trial and error, resulting in poor reproducibility.
[0066] In one scenario, there are multiple first evaluation strategies. First content is generated based on the first evaluation strategies according to a preset format, including: constructing first information; generating first content based on the first information and multiple first evaluation strategies according to a preset format; wherein, the first information includes at least one of the following: the dependency relationship between multiple first evaluation strategies; the reference relationship between first evaluation strategies in different modalities when there are multiple first modalities; and constraint information on multiple first evaluation strategies.
[0067] For example, when generating strategy combinations, cross-modal collaborative relationships between strategies can also be constructed. These include input-output dependencies between different strategies, such as using the content generated by the first model for one evaluation strategy as a constraint or evidence for the generation process of the first model for another evaluation strategy; reference relationships between cross-modal strategies, such as writing the historical evaluation results of one evaluation strategy in the image modality of the strategy combination into the prompt words of another evaluation strategy in the text modality of the same strategy combination, which instructs the first model to generate the corresponding test content for that evaluation strategy; and global constraints, such as the maximum level of detail, the type of security risk that must be triggered, and the output format. By combining a unified hierarchical strategy and making the collaborative links between different modalities and strategies explicit, the controllability and reproducibility of model testing can be effectively improved.
[0068] For example, the above process can generate first content based on first information and multiple first evaluation strategies according to a preset model and a preset format. This is equivalent to outputting the planning results of strategy combination into a strategy book in a unified format. The strategy book can include the strategy identifier of the evaluation strategy selected in each strategy partition, the role description and execution parameters of each evaluation strategy (i.e., the above strategy description, such as input generation rules and output constraints), and the above cross-modal collaborative link and global constraint information. This strategy book is used as the sole driving input for subsequent executors, thereby decoupling strategy planning and execution.
[0069] In one scenario, generating second test data based on first content using a first model includes: parsing the first content; generating third test data for the second modality using a third model based on the fourth evaluation strategy when the first content includes a fourth evaluation strategy for the second modality; and generating fourth test data for the third modality using a fourth model based on the fifth evaluation strategy when the first content includes a fifth evaluation strategy for the third modality. The first modality includes both the second and third modalities, which are different from each other. The second test data includes both the third and fourth test data. The first model includes both the third and fourth models.
[0070] For example, if the second modality is a text modality, then the first model can be a text generation model capable of generating test text that satisfies structural / semantic / syntactic constraints, specifically determined according to the corresponding evaluation strategy. If the third modality is an image modality, then the third model can be an image generation model capable of generating test images or image descriptions that satisfy constraints such as layout / semantic cues / quality perturbations, specifically determined according to the corresponding evaluation strategy.
[0071] It is important to note that if the second test data includes cross-modal link information, the corresponding link logic also needs to be executed. For example, the historical evaluation results of the image modality can be written into the prompt words or judgment results of the text modality evaluation results and injected into the text side prompts, or the structural constraints of the text modality can be used to inversely constrain the image generation layout, etc. The specifics are determined according to the actual situation. Of course, during the generation process, it is also important to ensure that the generated content satisfies the above-mentioned dependency information and global constraint information, so as to form a consistent collaborative use case.
[0072] The above process can be generated through an executor scheduling model, thereby unifying the executor and the unified policy representation to drive test case generation and achieve end-to-end automated test case generation.
[0073] It should be noted that the model evaluation strategy is uniformly divided into three layers—structure, semantics, and syntax—across different modalities, and stored and scheduled in the form of composable atomic policy objects. This ensures that the policy library has clear boundaries, is enumerable, and composable, systematically covering different input formats and risk trigger dimensions, thus improving the coverage and comparability of security assessments. Furthermore, the model outputs a structured policy document, which is then matched and selected in different policy partitions to produce an executable, composable policy document containing dependencies and cross-modal references. The collaborative links are explicitly defined, making the collaborative relationships, dependencies, and constraints between text and images recordable, reproducible, and auditable, reducing instability and uncontrollability caused by relying on model-driven planning or human experience.
[0074] Furthermore, a unified executor is used to call the corresponding model to directly generate test cases from the strategy document and interactively evaluate them. Adaptive iteration and coverage enhancement are achieved through statistical and failure combination memory, reducing hard coding and improving engineering scalability.
[0075] It is worth noting that after obtaining the evaluation results, the second test data, evaluation results, execution context and the identifier of the first evaluation strategy, and the model parameters of the first major model can be recorded to achieve traceability and reproducibility. The execution context can include output content, etc.
[0076] In addition, strategy-level statistics can be updated based on the evaluation results, such as the historical number of times the strategy has been used, the historical number of times it has been effective, the sampling weight or sampling probability of different strategy partitions, and the identification of strategy combination memory or partition-level replacement suggestions. This allows the strategy selection to be automatically adjusted based on the above content in subsequent strategy combination planning, thereby improving coverage efficiency and reducing maintenance difficulty.
[0077] This content provides a method for security evaluation of a large model, comprising: acquiring first test data and acquiring a first evaluation strategy from a strategy library based on the first test data; wherein the first test data is used to indicate the testing requirements for a first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the levels include at least one of the structural layer, semantic layer and syntactic layer, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers based on input content; generating first content according to a preset format based on the first evaluation strategy, and generating second test data based on the first content through the first model, the second test data being the first modality; inputting the second test data into the first large model to obtain output content, and determining the evaluation result based on the output content, the evaluation result being used to characterize whether the first large model has security risks.
[0078] Using the above method, a first evaluation strategy that meets the testing requirements can be selected from the strategy library. This first evaluation strategy is then converted into first content in a unified format. Subsequently, based on this first content, a second set of test content is generated using the first model to perform risk testing on the first major model. This eliminates the need to build independent code paths for each evaluation strategy, reducing code redundancy and allowing for flexible combinations of different model evaluation strategies to meet testing needs, automatically generating corresponding test data. This effectively improves cross-strategy reusability, scalability, and testing efficiency. Furthermore, the evaluation strategies have a unified and clearly layered structure, facilitating the systematic organization and enumeration of coverage testing.
[0079] Figure 3 This is a schematic diagram of a device for large-scale model safety testing, as illustrated by an example. Figure 3 As shown, the device 300 for large model safety evaluation includes: The acquisition module 301 is used to acquire first test data and acquire a first evaluation strategy from the strategy library based on the first test data; wherein, the first test data is used to indicate the testing requirements for the first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers according to the input content; The generation module 302 is used to generate first content according to a preset format based on the first evaluation strategy, and generate second test data based on the first content through a first model, wherein the second test data is the first modality; The evaluation module 303 is used to input the second test data into the first large model, obtain the output content, and determine the evaluation result based on the output content. The evaluation result is used to characterize whether the first large model has any security risks.
[0080] Using the above-mentioned device, a first evaluation strategy that meets the testing requirements can be selected from the strategy library. The first evaluation strategy is then converted into first content in a unified format. Subsequently, a second test content is generated based on the first content using a first model to perform risk testing on the first major model. Based on this, risk testing is performed on the first major model. This eliminates the need to build an independent code path for each evaluation strategy, which not only reduces code redundancy but also allows for flexible combination of different model evaluation strategies to meet testing requirements and automatically generates corresponding test data, effectively improving cross-strategy reusability, scalability, and testing efficiency.
[0081] Optionally, the strategy library includes at least one of the following levels of model evaluation strategies: The model evaluation strategy of the structural layer is used to describe the structural features of the model input and / or model output; Semantic layer model evaluation strategies are used to describe the semantic features of model inputs and / or model outputs; The model evaluation strategy at the syntactic layer is used to describe the syntactic features of the model input and / or model output.
[0082] Optionally, the number of the first evaluation strategies is multiple, and the generation module 302 is used for: Constructing the first information; First content is generated based on the first information and multiple first evaluation strategies according to a preset format; The first information includes at least one of the following: Dependencies among multiple first assessment strategies; When there are multiple first modalities, the reference relationships between the first evaluation strategies under different modalities; Constraint information for multiple first evaluation strategies.
[0083] Optionally, the acquisition module 301 is used for: The first test data is parsed to obtain second information, which is used to structurally describe the test requirements corresponding to the first test data. A second evaluation strategy is obtained from the model evaluation strategies at each level under the first modality, wherein the second evaluation strategy is the evaluation strategy that matches the second information in the model evaluation strategies; Based on the obtained second evaluation strategy, a first evaluation strategy is determined.
[0084] Optionally, the acquisition module 301 is used for: The score for each of the second evaluation strategies is determined based on the historical number of times the corresponding second evaluation strategy has been used and the historical number of times it has been effective. The historical number of effective times represents the number of times the second evaluation strategy has triggered the first large model to have a security risk. In the second assessment strategy at each level, the third assessment strategy is determined based on the score of each second assessment strategy; The obtained third evaluation strategy will be used as the first evaluation strategy.
[0085] Optionally, each model evaluation strategy in the strategy library includes a corresponding level, data modality, evaluation scope, and strategy description, wherein the strategy description is used to describe the input and / or output rules of the corresponding model evaluation strategy.
[0086] Regarding the device 300 used for large model safety evaluation mentioned above, the method logic executed by each functional module has been explained in detail in the section on methods, and will not be repeated here.
[0087] Based on the same concept, a computer-readable medium is also provided, on which a computer program is stored, which, when executed by a processing device, implements the steps of any of the methods described above for large model security evaluation.
[0088] This allows for the selection of a first evaluation strategy from the strategy library that meets the testing requirements. The first evaluation strategy is then converted into first content in a unified format. Subsequently, based on the first content, a second test content is generated using the first model to perform risk testing on the first major model. Based on this, risk testing is then performed on the first major model. This eliminates the need to build independent code paths for each evaluation strategy, which not only reduces code redundancy but also allows for flexible combination of different model evaluation strategies to meet testing requirements and automatically generates corresponding test data. This effectively improves cross-strategy reusability, scalability, and testing efficiency.
[0089] Based on the same concept, an electronic device is also provided, which may include: A storage device on which computer programs are stored; A processing device for executing a computer program stored in a storage device to implement the steps of any of the methods described above for large model security evaluation.
[0090] This allows for the selection of a first evaluation strategy from the strategy library that meets the testing requirements. The first evaluation strategy is then converted into first content in a unified format. Subsequently, based on the first content, a second test content is generated using the first model to perform risk testing on the first major model. Based on this, risk testing is then performed on the first major model. This eliminates the need to build independent code paths for each evaluation strategy, which not only reduces code redundancy but also allows for flexible combination of different model evaluation strategies to meet testing requirements and automatically generates corresponding test data. This effectively improves cross-strategy reusability, scalability, and testing efficiency.
[0091] Based on the same concept, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of any of the methods described above for large model security evaluation.
[0092] This allows for the selection of a first evaluation strategy from the strategy library that meets the testing requirements. The first evaluation strategy is then converted into first content in a unified format. Subsequently, based on the first content, a second test content is generated using the first model to perform risk testing on the first major model. Based on this, risk testing is then performed on the first major model. This eliminates the need to build independent code paths for each evaluation strategy, which not only reduces code redundancy but also allows for flexible combination of different model evaluation strategies to meet testing requirements and automatically generates corresponding test data. This effectively improves cross-strategy reusability, scalability, and testing efficiency.
[0093] The following is for reference. Figure 4 The diagram illustrates a structural schematic suitable for implementing the electronic device 400 described above. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 4 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.
[0094] like Figure 4As shown, electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. The random access memory 403 also stores various programs and data required for the operation of electronic device 400. The processing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0095] Typically, the following devices can be connected to the input / output interface 405: input devices 406 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 408 including, for example, magnetic tape, hard disk, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0096] In particular, depending on certain circumstances, the processes described in the above-referenced flowchart can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from read-only memory 402. When the computer program is executed by processing device 401, it performs the functions defined in the above-described methods.
[0097] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0098] In some implementations, communication can be conducted using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and end-to-end networks (e.g., ad-hoc end-to-end networks), as well as any currently known or future-developed networks.
[0099] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0100] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the electronic device to: acquire first test data and acquire a first evaluation strategy from a strategy library based on the first test data; wherein the first test data is used to indicate the testing requirements for a first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers based on input content; generate first content according to a preset format based on the first evaluation strategy, and generate second test data based on the first content through the first model, the second test data being the first modality; input the second test data into the first large model to obtain output content, and determine the evaluation result based on the output content, the evaluation result being used to characterize whether the first large model has security risks.
[0101] Computer program code for performing the above operations can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0102] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0103] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the functionality of that module.
[0104] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0105] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0106] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of the technical solution is not limited to specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.
[0107] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of the technical solution. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.
[0108] Although the technical solution has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which each module performs its operation has already been described in detail in the section concerning the method, and will not be elaborated upon here.
Claims
1. A method for security evaluation of large models, comprising: First test data is obtained, and a first evaluation strategy is obtained from the strategy library based on the first test data; wherein, the first test data is used to indicate the testing requirements for the first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers according to the input content; First content is generated based on the first evaluation strategy according to a preset format, and second test data is generated based on the first content through a first model. The second test data is the first modality. The second test data is input into the first large model to obtain the output content, and the evaluation result is determined based on the output content. The evaluation result is used to characterize whether the first large model has any security risks.
2. The method according to claim 1, wherein the strategy library includes at least one of the following levels of model evaluation strategies: The model evaluation strategy of the structural layer is used to describe the structural features of the model input and / or model output; Semantic layer model evaluation strategies are used to describe the semantic features of model inputs and / or model outputs; The model evaluation strategy at the syntactic layer is used to describe the syntactic features of the model input and / or model output.
3. The method according to claim 1, wherein the number of the first evaluation strategies is multiple, and the step of generating the first content based on the first evaluation strategies according to a preset format includes: Constructing the first information; First content is generated based on the first information and multiple first evaluation strategies according to a preset format; The first information includes at least one of the following: Dependencies among multiple first assessment strategies; When there are multiple first modalities, the reference relationships between the first evaluation strategies under different modalities; Constraint information for multiple first evaluation strategies.
4. The method according to any one of claims 1-3, wherein obtaining the first evaluation strategy from the strategy library based on the first test data includes: The first test data is parsed to obtain second information, which is used to structurally describe the test requirements corresponding to the first test data. A second evaluation strategy is obtained from the model evaluation strategies at each level under the first modality, wherein the second evaluation strategy is the evaluation strategy that matches the second information in the model evaluation strategies; Based on the obtained second evaluation strategy, a first evaluation strategy is determined.
5. The method according to claim 4, wherein determining the first assessment strategy based on the obtained second assessment strategy includes: The score for each of the second evaluation strategies is determined based on the historical usage and historical effective count of the corresponding second evaluation strategy. The historical effective count represents the number of times the second evaluation strategy triggers a security risk in the first large model. In the second assessment strategy at each level, the third assessment strategy is determined based on the score of each second assessment strategy; The obtained third evaluation strategy will be used as the first evaluation strategy.
6. The method according to any one of claims 1-3, wherein each model evaluation strategy in the strategy library includes a corresponding level, data modality, evaluation scope, and strategy description, wherein, The strategy description is used to describe the input and / or output rules of the corresponding model evaluation strategy.
7. An apparatus for safety evaluation of large models, comprising: An acquisition module is used to acquire first test data and acquire a first evaluation strategy from a strategy library based on the first test data; wherein, the first test data is used to indicate the testing requirements for the first large model, the strategy library includes model evaluation strategies at different levels under the first modality, the model evaluation strategy includes the first evaluation strategy, the first modality represents the data modality that the first large model can process, and the first large model is used to output corresponding answers according to the input content; The generation module is used to generate first content based on the first evaluation strategy according to a preset format, and to generate second test data based on the first content through a first model, wherein the second test data is the first modality; The evaluation module is used to input the second test data into the first large model, obtain the output content, and determine the evaluation result based on the output content. The evaluation result is used to characterize whether the first large model has any security risks.
8. A computer-readable medium having a computer program stored thereon, wherein, When executed by a processing device, the computer program performs the steps of the method according to any one of claims 1-6.
9. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-6.
10. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.