Method and apparatus for security defense of business large model

By constructing an attack and defense database and performing multiple rounds of iterative updates and pruning, the novel attack problem of black-box large language models was solved, achieving efficient and dynamically adaptive security defense, improving response speed and defense efficiency, and ensuring data consistency and security.

CN120316828BActive Publication Date: 2025-10-24CHINA DAAS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510790537.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-24
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Existing technologies cannot effectively deal with novel attacks on black-box large language models. Traditional defense methods have narrow coverage, cannot dynamically adapt, training data does not match the distribution of actual business scenarios, database size is prone to expansion and retrieval efficiency is low, and it is difficult to balance real-time defense performance and security.

Method used

We employ dynamic context learning and generative adversarial learning methods. By constructing attack and defense databases and performing multiple rounds of iterative updates and pruning, we combine Monte Carlo experiments and simulated annealing to dynamically adjust the context database, achieve adversarial learning closed loop, prune redundant samples, and maintain a reasonable database size.

Benefits of technology

It enhances the continuous effectiveness of the defense strategy of black-box large language models, ensures consistency between training data and actual business scenarios, improves response speed and defense efficiency, dynamically adapts to new types of attacks, and reduces false positive and false negative rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316828B_ABST
    Figure CN120316828B_ABST
Patent Text Reader

Abstract

The application provides a business large model security defense method and device; the method comprises the following steps: acquiring a first query set; performing business adaptation processing on the first query set to obtain a second query set, and dividing the second query set into a training set and a test set; based on the training set, performing multiple rounds of iterative updating on an attack database and a defense database; based on the test set, performing pruning processing on the defense database after multiple rounds of iterative updating to obtain a pruned defense database; in response to a to-be-processed query instruction, determining a classification label corresponding to the to-be-processed query instruction, and extracting a defense example corresponding to the classification label from the pruned defense database; generating a prompt word based on the to-be-processed query instruction, the classification label and the defense example, inputting the prompt word into a business large model, so that the business large model outputs a business data query result. Through the application, efficient defense in actual business scenarios such as commercial queries can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a security defense method and device for a business large model. BACKGROUND

[0002] Large language models are one of the important research directions in the field of artificial intelligence in recent years. With its excellent natural language ability, large language models have gained widespread attention and popularity, and have completely changed the field of natural language processing. In recent years, with the wide application of large language models in commercial data service fields, especially in financial risk control, intelligent customer service, knowledge retrieval and other scenarios, large language models have become a key technology to support core businesses. Since many large language model providers only provide closed-source inference services in the form of cloud application programming interfaces (APIs), enterprises cannot access or modify the internal structure and parameters of the model, and can only call the model interface in a "black box" manner, which poses a significant challenge to security protection.

[0003] At the same time, black-box attack methods against large language models are constantly evolving. Attackers can induce large language models to output unsafe or incorrect results through carefully designed malicious inputs, and even cause sensitive information leakage and business decision failure, posing a serious threat to the security and reliability of commercial query services. Traditional defense methods (such as offline fine-tuning and security rule detection) have a narrow coverage and are difficult to respond to new attacks in real time, and large-scale fine-tuning of closed-source models is often impossible to implement, making it difficult to achieve both defense efficiency and effectiveness in online business scenarios. In addition, the context learning defense provided by related technologies mostly relies on static context libraries, and the example content is fixed and the storage overhead increases sharply as the number grows, making it difficult to meet the online demand for high concurrency and low latency. SUMMARY

[0004] The embodiments of the present application provide a security defense method and device for a business large model, an electronic device, a computer readable storage medium and a computer program product, which can achieve efficient defense in actual business scenarios such as commercial queries.

[0005] The technical solutions of the embodiments of the present application are implemented as follows:

[0006] The embodiments of the present application provide a security defense method for a business large model, comprising:

[0007] Obtaining a first query set, the first query set comprising historical query instructions submitted by users to the business large model;

[0008] Performing business adaptation processing on the first query set to obtain a second query set, and dividing the second query set into a training set and a test set;

[0009] Based on the training set, the attack database and the defense database are updated for multiple rounds of iteration, the attack database includes attack examples for the business large model, and the defense database includes defense examples for the business large model; wherein, the iteration update process of the first round includes: for each historical query instruction in the training set, based on the historical query instruction, the selected attack type, and the attack examples corresponding to the attack type extracted from the attack database updated after the first-1 round of iteration, an attack instruction is generated; based on multiple attack instructions and multiple historical query instructions included in the second query set, a mixed training set is generated; based on the mixed training set, the attack database and the defense database updated after the first-1 round of iteration are updated, wherein, t t -1 rounds of iteration, an attack instruction is generated; based on multiple attack instructions and multiple historical query instructions included in the second query set, a mixed training set is generated; based on the mixed training set, the attack database and the defense database updated after the first-1 round of iteration are updated, wherein, t t is a positive integer greater than 1, and satisfies less than or equal to T , T is the total number of multiple rounds of iteration;

[0010] Based on the test set, the defense database updated after multiple rounds of iteration is pruned to obtain a pruned defense database;

[0011] In response to a to-be-processed query instruction, a classification label corresponding to the to-be-processed query instruction is determined, and a defense example corresponding to the classification label is extracted from the pruned defense database;

[0012] Based on the to-be-processed query instruction, the classification label, and the defense example, a prompt word is generated, and the prompt word is input into the business large model to make the business large model output a business data query result.

[0013] The embodiment of the application provides a security defense device of a business large model, which comprises:

[0014] An acquisition module is configured to acquire a first query set, wherein the first query set includes historical query instructions submitted by a user to the business large model;

[0015] A business adaptation module is configured to perform business adaptation processing on the first query set to obtain a second query set;

[0016] A division module is configured to divide the second query set into a training set and a test set;

[0017] An update module is configured to update an attack database and a defense database for multiple rounds of iteration based on the training set, wherein the attack database includes attack examples for the business large model, and the defense database includes defense examples for the business large model; wherein, the iteration update process of the first round includes: for each historical query instruction in the training set, based on the historical query instruction, the selected attack type, and the attack examples corresponding to the attack type extracted from the attack database updated after the first-1 round of iteration, an attack instruction is generated; based on multiple attack instructions and multiple historical query instructions included in the second query set, a mixed training set is generated; based on the mixed training set, the attack database and the defense database updated after the first-1 round of iteration are updated, wherein,​​t The iterative updating process of the wheel comprises: for each historical query instruction in the training set, generating an attack instruction based on the historical query instruction, the selected attack type, and the attack examples corresponding to the attack type extracted from the attack database updated after the first t -1 rounds of iteration; generating a hybrid training set based on a plurality of the attack instructions and a plurality of historical query instructions included in the second query set; and updating the attack database and the defense database updated after the first t -1 rounds of iteration, wherein, t is a positive integer greater than 1 and satisfies T , T is the total number of rounds of iteration;

[0018] The pruning module is configured to perform pruning processing on the defense database updated after the multiple rounds of iteration based on the test set, to obtain a pruned defense database.

[0019] The determining module is configured to determine a classification label corresponding to the to-be-processed query instruction in response to the to-be-processed query instruction.

[0020] The extracting module is configured to extract a defense example corresponding to the classification label from the pruned defense database.

[0021] The generating module is configured to generate a prompt word based on the to-be-processed query instruction, the classification label, and the defense example.

[0022] The input module is configured to input the prompt word into the business large model, so that the business large model outputs a business data query result.

[0023] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:

[0024] A memory is configured to store computer executable instructions.

[0025] A processor is configured to execute the computer executable instructions stored in the memory, so as to implement the security defense method of the business large model provided in the embodiments of the present application.

[0026] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores computer executable instructions, which are configured to be executed by a processor, so as to implement the security defense method of the business large model provided in the embodiments of the present application.

[0027] A computer program product is provided in an embodiment of the present application, and the computer program product comprises a computer program or computer executable instructions, which are configured to be executed by a processor, so as to implement the security defense method of the business large model provided in the embodiments of the present application.

[0028] The embodiments of the present application have the following beneficial effects:

[0029] On the one hand, the training set and the test set are constructed from the actual business data (i.e., the historical query instructions submitted by the user to the business large model), so as to ensure the consistency of the training and the actual business scenario data distribution, and on the other hand, the dual-database architecture (i.e., including the attack database and the defense database) is adopted to store the attack examples and the defense examples, so as to focus on the high-risk attack scenarios and accumulate robust defense experience, maintain the clarity and efficiency of data organization, and in addition, through the multi-round iteration and update of the attack database and the defense database, the cyclic accumulation of the attack examples and the defense examples can be realized, a dynamic self-reinforcing adversarial learning closed loop is constructed, so as to improve the continuous effectiveness of the defense strategy in the case of being unable to fine-tune the business large model. In addition, through the pruning processing of the defense database after the multi-round iteration and update, the size of the defense database can be controlled, and then the subsequent response speed is improved. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 is an architecture schematic diagram of a security defense system 100 of a business large model provided by the embodiments of the present application;

[0031] Figure 2 is a structure schematic diagram of an electronic device 500 provided by the embodiments of the present application;

[0032] Figure 3 is a flow schematic diagram of a security defense method of a business large model provided by the embodiments of the present application;

[0033] Figure 4 is a first process schematic diagram of a security defense method of a business large model provided by the embodiments of the present application;

[0034] Figure 5 is a second process schematic diagram of a security defense method of a business large model provided by the embodiments of the present application;

[0035] Figure 6 is a third process schematic diagram of a security defense method of a business large model provided by the embodiments of the present application;

[0036] Figure 7 is a fourth process schematic diagram of a security defense method of a business large model provided by the embodiments of the present application. DETAILED DESCRIPTION

[0037] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limitations to the present application. All other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0038] In the following description, "some embodiments" are referred to, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0039] It can be understood that, in the embodiments of the present application, data related to user information and the like are involved, and when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards.

[0040] In the following description, the terms "first\second\..." are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second\..." can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0042] Before the embodiments of the present application are further described in detail, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.

[0043] 1) In response to: used to represent the conditions or states on which the operations performed depend, when the dependent conditions or states are met, one or more operations performed can be real-time or have a set delay; in the absence of special instructions, there is no restriction on the execution order of multiple operations performed.

[0044] 2) Large Language Model (LLM): A large model refers to a deep neural network model with tens of billions or even hundreds of billions of parameters that is pre-trained through massive text data and can understand and generate natural language text. The latest large models also have the ability of deep logical reasoning, multi-modal information processing, and tool invocation. Common representatives include the DeepSeek series, the GPT series, the Gemini series, and the Claude series.

[0045] 3) Black Box Attack: Attackers can only access the input-output interface of the large model and do not know the internal structure and parameters of the large model, i.e., "black box" access. The technology uses malicious input to induce the large model to output unsafe or incorrect results.

[0046] 4) Black Box Defense: A security protection strategy designed for large models with unknown internal structures, including external measures based on input detection, context filtering, and decision-making mechanisms, without modifying or fine-tuning the internal structure of the large model.

[0047] 5) Context Learning (CL): Using examples to build "prompt context" to allow the large model to refer to these examples during reasoning, achieving a similar effect to light fine-tuning. The examples can be dynamically updated, helping the large model better identify and defend against new attacks.

[0048] 6) Model Context Protocol (MCP): A standardized calling interface protocol that allows asynchronous requests with context prompts to be sent to local or cloud large models and receives streaming or batch responses. In the embodiments of the present application, MCP is used to parallelize the generation of adversarial instructions, classification, and security evaluation, thereby improving the online defense performance of the large model.

[0049] 7) Simulated Annealing (SA): A random optimization algorithm that simulates the physical annealing process. By accepting poor performance solutions in the "high temperature" stage to escape local optima, it gradually converges to the global optimum as the "temperature" decreases. In the embodiments of the present application, SA can be combined with Monte Carlo experiments to perform "annealing pruning" on the context database (including attack context database and defense context database).

[0050] The applicant found that the related art provided the following problems in the process of implementing the embodiments of the present application:

[0051] 1) For the defense mechanism of the black-box large model, fine-tuning cannot be used to align the model, that is, for the black-box large model, all security alignment methods based on supervised fine-tuning of the large model itself are no longer applicable.

[0052] 2) For complex and diverse black-box attack methods, the defense method provided by the related art has too narrow coverage and cannot be dynamically adaptive.

[0053] 3) The data distribution of the training data set and the data distribution of the actual business scenario do not match, resulting in a defense blind area caused by "training-online" distribution drift.

[0054] 4) The database size is easy to expand and lacks a pruning mechanism. Specifically, the context learning provided by the related art often leads to continuous growth of the database, which in turn leads to a sharp rise in retrieval efficiency and storage overhead.

[0055] 5) Real-time defense performance and security are difficult to balance.

[0056] For the above technical problem 1), the embodiments of the present application provide a solution using dynamic context learning to solve the security defense problem of the black-box large model; for the above technical problem 2), the embodiments of the present application use a large model to generate a complex and diverse training data set, and use a generative adversarial context learning method to dynamically adjust the attack context database and the defense context database, so that the defense mechanism of the large model can be dynamically adaptive; for the above technical problem 3), the embodiments of the present application provide a set of data cleaning and clustering balancing schemes for processing actual business data. Starting from the actual business data, the attack instructions are generated based on the large model, so as to ensure the consistency of the data distribution of the training and the actual business scenario; for the above technical problem 4), the embodiments of the present application combine Monte Carlo experiments and simulated annealing ideas to periodically eliminate redundant or weak representative samples, thereby ensuring the balance between "precision", "accuracy" and "small" of the context database. At the same time, the embodiments of the present application also design different pruning strategies for the attack context database and the defense context database, so as to cooperate with generative adversarial learning; for the above technical problem 5), the embodiments of the present application can provide robust security decisions while maintaining ultra-fast response through the combination of asynchronous MCP calls, static prompt word templates, attack type judgments, and context database pruning operations.

[0057] That is, the embodiment of the present application provides a business large model security defense method and device, electronic equipment, computer readable storage medium and computer program product, which can realize efficient defense in actual business scenarios such as business query. The electronic equipment provided by the embodiment of the present application is described below. The electronic equipment provided by the embodiment of the present application can be implemented as a server, for example, it can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.

[0058] In some embodiments, referring to Figure 1 , Figure 1 is the architecture diagram of the business large model security defense system 100 provided by the embodiment of the present application, as shown in Figure 1 The business large model security defense system 100 provided by the embodiment of the present application includes a server 200, a network 300 and a terminal device 400, wherein the network 300 can be a local area network or a wide area network, or a combination of the two, and the terminal device 400 is a terminal device associated with a user, for example, it can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto.

[0059] For example, the server 200 can collect historical query instructions provided by the user to the business large model in a historical time period (for example, in the past half year), and collect the collected multiple historical query instructions into a first query set. Then, the server 200 can perform business adaptation processing (for example, including filtering, denoising, deduplication, clustering, sampling and the like) on the first query set to obtain a second query set, and divide the second query set into a training set and a test set (for example, the multiple historical query instructions included in the second query set can be divided into the training set and the test set according to a ratio of 9:1). Subsequently, the server 200 can perform multiple rounds of iterative updates on the attack database and the defense database based on the training set, wherein the attack examples for the business large model are stored in the attack database, and the defense examples for the business large model are stored in the defense database. After multiple rounds of iterative updates, the server 200 can also perform pruning processing on the attack database and the defense database after multiple rounds of iterative updates based on the test set, to obtain a pruned attack database and a pruned defense database. Finally, when the server receives the to-be-processed query instruction sent by the user through the terminal device 400, the server can first classify the to-be-processed query instruction through the classification model to obtain a classification label corresponding to the to-be-processed query instruction, and then extract the defense examples corresponding to the classification label from the pruned defense database, and fill the to-be-processed query instruction, the classification label, and the defense examples extracted from the pruned defense database into the prompt word template to obtain the final prompt word. After obtaining the prompt word, the server 200 can input the prompt word into the business large model to make the business large model output a safe business data query result.

[0060] The structure of the electronic device provided in the embodiments of the present application will be described below. Taking the electronic device as a server for example, referring to Figure 2 , Figure 2 FIG. 5 is a structural schematic diagram of an electronic device 500 provided in the embodiments of the present application, Figure 2 The electronic device 500 shown in FIG. 5 includes at least one processor 510, a memory 540, and at least one network interface 520. The various components in the electronic device 500 are coupled together through a bus system 530. It can be understood that the bus system 530 is used to realize the connection and communication between the components. In addition to including a data bus, the bus system 530 also includes a power bus, a control bus, and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 530 in Figure 2 .

[0061] The processor 510 can be an integrated circuit chip with processing capability, such as a general purpose processor, a Digital Signal Processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc.

[0062] The memory 540 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 540 optionally includes one or more storage devices remotely located from the processor 510.

[0063] The memory 540 includes volatile memory or non-volatile memory, or both. Non-volatile memory can be read only memory (ROM), volatile memory can be random access memory (RAM). The memory 540 described in the embodiments of the present application is intended to include any suitable type of memory.

[0064] In some embodiments, the memory 540 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are described below.

[0065] The operating system 541 includes a system program for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0066] The network communication module 542 is used to communicate with other computing devices via one or more (wired or wireless) network interfaces 520, examples of which include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.

[0067] In some embodiments, the apparatus provided by the embodiments of the present application can be realized in software, Figure 2The security defense device 543 of the business large model stored in the memory 540 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: an acquisition module 5431, a business adaptation module 5432, a division module 5433, an update module 5434, a pruning module 5435, a determination module 5436, an extraction module 5437, a generation module 5438, and an input module 5439. These modules are logical, and thus can be combined or further split according to the implemented functions, and the functions of the modules will be described below.

[0068] The security defense method of the business large model provided in the embodiments of the present application will be described in detail below in combination with the exemplary application and implementation of the server provided in the embodiments of the present application.

[0069] For example, referring to Figure 3 , Figure 3 is a flowchart of the security defense method of the business large model provided in the embodiments of the present application, which will be described in combination with the steps shown in Figure 3 .

[0070] It should be noted that the business large model provided in the embodiments of the present application can be used for business data query services, for example, can be used for commercial data query services.

[0071] In step 101, a first query set is acquired.

[0072] Here, the first query set can include historical query instructions submitted by users to the business large model.

[0073] In some embodiments, taking a business large model for commercial query data services as an example, the server can collect all user queries (i.e., historical query instructions submitted by users to the business large model) from the business system, thereby obtaining the first query set, that is, the first query set can include all historical query instructions submitted by users to the business large model within a historical time period (e.g., within the past half year).

[0074] In step 102, the first query set is subjected to business adaptation processing to obtain a second query set.

[0075] In some embodiments, step 102 can be implemented in the following manner: from the plurality of historical query instructions included in the first query set, deleting historical query instructions with a length less than a length threshold to obtain a third query set; based on a regular expression or a set rule, performing denoising processing on the third query set to obtain a fourth query set; performing hash deduplication on the fourth query set (for example, the historical query instructions included in the fourth query set can be hashed to obtain corresponding hash values, then the plurality of hash values can be compared, and when there are a plurality of identical hash values, only the historical query instruction corresponding to one of the hash values can be retained) to obtain a fifth query set; mapping the historical query instructions included in the fifth query set to corresponding embedding vectors through an embedding model, and determining the similarity between any two embedding vectors; for two embedding vectors with a similarity greater than a similarity threshold, deleting the historical query instruction corresponding to one of the embedding vectors from the fifth query set to obtain a sixth query set; deleting historical query instructions in a discrete state from the sixth query set to obtain a seventh query set; obtaining a subset of the seventh query set, and training a clustering model based on the subset; clustering the plurality of historical query instructions included in the seventh query set based on the trained clustering model to obtain a plurality of clusters; and generating the second query set based on the plurality of historical query instructions included in the plurality of clusters.

[0076] In other embodiments, based on the above examples, the second query set can be generated based on the plurality of historical query instructions included in the plurality of clusters in the following manner: determining the number of historical query instructions included in each cluster; for a cluster with a number greater than a number threshold, down-sampling the plurality of historical query instructions included in the cluster until the number of historical query instructions included in the cluster is reduced to a first number; for a cluster with a number less than the number threshold, re-sampling the plurality of historical query instructions included in the cluster, or translating and enhancing the plurality of historical query instructions included in the cluster through a large language model (i.e., translating the historical query instructions from a source language to a target language, and then translating them back from the target language to the source language, so that a query instruction similar to the historical query instruction can be obtained), until the number of historical query instructions included in the cluster is increased to a second number; and merging the updated historical query instructions included in the plurality of clusters to obtain the second query set.

[0077] For example, referring to Figure 4 , Figure 4 is a first process schematic diagram of a security defense method of a business large model provided by the embodiments of the present application, as shown in Figure 4 , for a business large model used for commercial query type data services, all user queries (i.e., historical query instructions submitted by users to the business large model) collected from the business system can be combined into a set , wherein is the firstn a user query, and n satisfying 1≤ n ≤ N , N is the total number of user queries collected from the business system. Then, the set may be pre-screened according to a character length threshold (supposed to be denoted as ), for example, all user queries with a character length less than may be directly removed from the set , so that the set , may be obtained, where represents the character length of the user query. Subsequently, the set may be further processed for denoising, for example, including operations such as removing pure symbol strings or control characters, and unifying full and half-width formats, and finally outputting a clean set (i.e., the fourth query set).

[0078] In order to avoid query content redundancy, the clean set may also be deduplicated (for example, hashed deduplicated) to obtain a deduplicated set . Subsequently, an embedding model (for example, text-embedding-ada-002) may also be called to map the user queries in the set (i.e., the fifth query set) to embedding vectors, and calculate the cosine similarity (supposed to be denoted as ) between any two embedding vectors. If there are two embedding vectors with a cosine similarity greater than a similarity threshold (supposed to be denoted as ), a greedy removal may be performed (for example, only one of the embedding vectors may be retained), so that the set (i.e., the sixth query set) may be obtained. After obtaining the set , an Isolation Forest algorithm may also be used to remove outliers in the set at a set proportion (supposed to be denoted as ), so that the set (i.e., the seventh query set) may be obtained. Subsequently, a subset with a size of may be randomly extracted from the set , and a density-based unsupervised clustering model (HDBSCAN, Hierarchical Density-Based Spatial Clustering of Applications with Noise) may be trained based on the extracted subset. The clustering model does not need to specify the number of clusters in advance, but automatically discovers the cluster structure according to the density of the sample in the vector space.

[0079] After the training is completed, the embodiment of the present application can use the trained clustering model to All embedded vectors in the prediction are used to obtain the cluster label of each query, and the sample size of each cluster is counted based on this (that is, the number of user queries included in each cluster). Next, for the clustering results, the embodiment of the present application can be used to select the clusters with sample sizes greater than the sample size threshold. Each cluster is downsampled to To ensure high representativeness of mainstream scenarios; for the remaining long-tail clusters, random resampling or calling a large model for back-translation enhancement can be performed until each cluster contains Samples (i.e. Finally, the samples of all clusters can be merged to form the final harmless query set (i.e., the second query set. Since the user queries in the second query set are cleaned, they can be called harmless queries), where: , The total number of clusters discovered by HDBSCAN can be used for subsequent adversarial example (including attack examples and defense examples) generation and evaluation.

[0080] It should be noted that the embodiment of the present application can also fine-tune the embedded model used according to the actual business scenario to improve performance. In addition, the unsupervised clustering method can also use the K-means (KMeans) silhouette coefficient evaluation to determine the number of clusters, or, you can also consider the high-performance OPTICS (full name Ordering Points To Identify the Clustering Structure, a density-based clustering algorithm, the core idea of ​​the OPTICS algorithm is to identify the clustering structure by point sorting, and its goal is to cluster the data in the space according to the density distribution), spectral clustering and other clustering methods. In addition, outlier removal can also consider using local anomaly factors instead of random forest algorithms, or, you can also train autoencoders to detect outliers. In addition, in terms of balanced sampling, the embodiment of the present application can also adopt a variety of semantic fuzzy and semantic transcription methods, such as transcription based on multilingual translation, or transcription based on large models.

[0081] The embodiments of the present application clean, denoise, cluster, and balance sampling of real business query data, so that the subsequently constructed training set can fully reflect the actual usage characteristics of the business big model for commercial query data services. The training and evaluation are completely based on real data on the business side, which significantly improves the reliability in real scenarios and avoids the defense blind spots caused by the "training-online" distribution drift.

[0082] In step 103, the second query set is divided into a training set and a test set.

[0083] In some embodiments, in order to generate and prune the subsequent adversarial examples (including attack examples and defense examples), the embodiment of the present application can also convert the final harmless query set into a 9:1 ratio. (i.e. the second query set) is randomly divided into a training set (assuming it is recorded as ) .

[0084] In step 104, multiple rounds of iterative updates are performed on the attack database and the defense database based on the training set.

[0085] Here, the attack database (or attack context database) can be used to store attack examples against the large business model; the defense database (or defense context database) can be used to store defense examples against the large business model.

[0086] In some embodiments, the following methods can be used to implement the multi-round iterative update process: t The iterative update process of the round: for each historical query instruction in the training set, based on the historical query instruction, the attack type randomly selected from the pre-defined multiple attack types, and the attack type randomly selected from the pre-defined multiple attack types, ... t -1 round of iterative update of the attack database and the attack examples corresponding to the attack type are extracted to generate attack instructions; based on multiple attack instructions and multiple historical query instructions included in the second query set, a mixed training set is generated; based on the mixed training set, the attack examples corresponding to the attack type after the first round of iterative update are extracted to generate attack instructions; based on the mixed training set, the attack examples corresponding to the attack type after the first round of iterative update are generated ... t -1 round of iterative updates to the attack database and defense database, where: t is a positive integer greater than 1 and less than or equal to T , T It should be noted that, for the first round of update process, the attack examples corresponding to the selected attack type can be directly extracted from the attack database.

[0087] In other embodiments, continuing with the above example, the above attack type randomly selected from a plurality of predefined attack types based on the historical query instruction and the attack type randomly selected from the predefined attack types can be realized in the following manner: t -1 round of iterative update of the attack database, and then extract the attack examples corresponding to the attack type, and generate the attack instruction: the historical query instruction, the target attack type randomly selected from the predefined multiple attack types, and the attack type randomly selected from the predefined multiple attack types, and the attack instruction .... tIn the attack database updated in the first round of iteration, at least one attack example corresponding to the target attack type is extracted and filled into the prompt template to obtain a prompt; a generation request carrying the prompt is sent to the cloud large model or the local large model through the model context protocol, and a generation result returned by the cloud large model or the local large model is received; the generation result is separated through a filter (Filter), and the separated result is filtered (for example, irrelevant contents are deleted), so that a final attack instruction can be obtained.

[0088] For example, referring to Figure 5 , Figure 5 is a second process schematic diagram of the security defense method of the business large model provided in the embodiments of the present application, as shown in Figure 5 In each round of adversarial generation, the embodiments of the present application can predefine four types of black-box attacks, which are assumed to be target hijacking (Hijack), jailbreaking (Jailbreak), prompt leakage (Leak) and irrelevant input (Noise). Then, for each type of attack, assume } can use a large model to construct a description text of the type of black-box attack (assume ). Specifically, in the generation phase, for each user query in the training set (i.e. harmless query), the user query , an attack type randomly specified from the above four types of black-box attacks, and a number of (for example, ) attack examples selected from the attack context database are filled into an attack prompt template (assume ), so that a prompt can be obtained. Then, the cloud or local large model can be called multiple times through the MCP protocol to generate an initial attack instruction set based on the prompt . It should be noted that if the cloud large model refuses to generate an attack instruction against the large model due to the large model security review of the interface, a local large model without security review can be used to generate the attack instruction. Subsequently, the generation result of the large model can be separated and filtered using a filter (Filter), and the model output irrelevant to the attack instruction is deleted, so that an attack instruction (assume ) can be obtained, wherein . Finally, all attack instructions can be collected as the attack instruction set of the current round . The above process can be written in the following form:

[0089]

[0090]

[0091]

[0092]

[0093] wherein, represents an attack example extracted from the attack context database , for example, including: a user query that is harmless an attack instruction designed by a previous attacker , an attack type determined by the large model for the attack instruction , a response performed by the large model for the attack instruction , In this way, the large model can generate a corresponding black box attack instruction based on the above information.

[0094] The embodiments of the present application generate a multi-category and dynamically updated adversarial dataset based on a large model, and combine a generative adversarial context learning mechanism to continuously expand and adjust the attack context database and the defense context database, thereby automatically adapting to emerging attack technologies. In this way, compared with the static context method provided by the related art, the embodiments of the present application can more comprehensively identify and resist a variety of black box attack types such as target hijacking, jailbreaking, prompt word leakage, and irrelevant input, thereby reducing the false positive rate and the false negative rate.

[0095] In some embodiments, based on the above examples, the mixed training set based on the plurality of attack instructions and the plurality of historical query instructions included in the second query set can be generated in the following manner: a plurality of attack instructions corresponding one-to-one to a plurality of historical query instructions included in the training set are collected to obtain a first round of attack instruction set; a third number of attack instructions are extracted from the first round of attack instruction set, and a third number of historical query instructions are extracted from the second query set; and a first round of mixed training set is constructed based on the third number of attack instructions and the third number of historical query instructions. That is, the attack instructions extracted from the first round of attack instruction set and the historical query instructions extracted from the second query set can be mixed in a 1:1 ratio to obtain the first round of mixed training set. t t In some embodiments, based on the above examples, the mixed training set based on the plurality of attack instructions and the plurality of historical query instructions included in the second query set can be generated in the following manner: a plurality of attack instructions corresponding one-to-one to a plurality of historical query instructions included in the training set are collected to obtain a first round of attack instruction set; a third number of attack instructions are extracted from the first round of attack instruction set, and a third number of historical query instructions are extracted from the second query set; and a first round of mixed training set is constructed based on the third number of attack instructions and the third number of historical query instructions. That is, the attack instructions extracted from the first round of attack instruction set and the historical query instructions extracted from the second query set can be mixed in a 1:1 ratio to obtain the first round of mixed training set. t t In some embodiments, based on the above examples, the mixed training set based on the plurality of attack instructions and the plurality of historical query instructions included in the second query set can be generated in the following manner: a plurality of attack instructions corresponding one-to-one to a plurality of historical query instructions included in the training set are collected to obtain a first round of attack instruction set; a third number of attack instructions are extracted from the first round of attack instruction set, and a third number of historical query instructions are extracted from the second query set; and a first round of mixed training set is constructed based on the third number of attack instructions and the third number of historical query instructions. That is, the attack instructions extracted from the first round of attack instruction set and the historical query instructions extracted from the second query set can be mixed in a 1:1 ratio to obtain the first round of mixed training set. t

[0096] In some embodiments, based on the above examples, the mixed training set based on the plurality of attack instructions and the plurality of historical query instructions included in the second query set can be generated in the following manner: a plurality of attack instructions corresponding one-to-one to a plurality of historical query instructions included in the training set are collected to obtain a first round of attack instruction set; a third number of attack instructions are extracted from the first round of attack instruction set, and a third number of historical query instructions are extracted from the second query set; and a first round of mixed training set is constructed based on the third number of attack instructions and the third number of historical query instructions. That is, the attack instructions extracted from the first round of attack instruction set and the historical query instructions extracted from the second query set can be mixed in a 1:1 ratio to obtain the first round of mixed training set. t ​​​-1 round of iteration updated attack database and defense database are updated: for each instruction in the mixed training set, the following processing is performed, where the instruction is an attack instruction or a historical query instruction: the instruction is classified by the classification model to obtain the attack type corresponding to the instruction; the instruction, the attack type, the description text of the attack type, and the attack type are filled into the prompt word template to obtain the prompt word; the prompt word is input into the business large model, and the response result output by the business large model for the prompt word is obtained; the security evaluation interface is called to evaluate the security of the response result, and the evaluation result is obtained; in response to any one of correct triggering, non-triggering, crash, or termination of the evaluation result, the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction are added as an attack example to the attack database updated in the previous round t -1 round of iteration updated defense database, the extracted defense examples corresponding to the attack type are filled into the prompt word template to obtain the prompt word; the prompt word is input into the business large model, and the response result output by the business large model for the prompt word is obtained; the security evaluation interface is called to evaluate the security of the response result, and the evaluation result is obtained; in response to any one of correct triggering, non-triggering, crash, or termination of the evaluation result, the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction are added as an attack example to the attack database updated in the previous round t -1 round of iteration updated attack database, wherein when the instruction is a historical query instruction, the attack example further includes an attack instruction generated based on the historical query instruction; when the instruction is an attack instruction, the attack example further includes a historical query instruction corresponding to the attack instruction; in response to any one of incorrect triggering, non-triggering, crash, or termination of the evaluation result, the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction are added as a defense example to the defense database updated in the previous round t -1 round of iteration updated defense database, wherein when the instruction is a historical query instruction, the defense example further includes an attack instruction generated based on the historical query instruction; when the instruction is an attack instruction, the defense example further includes a historical query instruction corresponding to the attack instruction.

[0097] For example, refer to Figure 6 , Figure 6 is a third process diagram of the security defense method of the business large model provided by the embodiments of the present application, as shown in Figure 6 after obtaining the attack instruction set of the current round, the attack instruction can be extracted from the attack instruction set of the current round, and the harmless query instruction can be extracted from the set , then the extracted attack instruction and query instruction can be mixed in a 1:1 ratio to obtain the mixed training set of the current round (assuming it is denoted as ). Then the classification model ( ) can be called to determine which attack type each instruction (including attack instruction and query instruction) in the mixed training set belongs to, and the instruction is input into the black box model (for example, a business large model for commercial query type data service) to obtain the corresponding response result, and finally the security evaluation interface (assuming it is denoted as ), each response result is labeled as any one of “correctly triggered defense mechanism (i.e. correct trigger, denoted as TP)”, “correct release (denoted as TN)”, “incorrectly triggered defense mechanism (i.e. incorrect trigger, denoted as FP)”, “not triggered (also called incorrect release, denoted as FN)”, or “answer terminated / crashed early (denoted as ER)”. The process can be written as follows:

[0098]

[0099]

[0100]

[0101]

[0102] = { }

[0103] wherein, is any one instruction in the mixed training set of the current round. represents the classification result (i.e. the preliminary judgment result of the attack type) output by the classification model for the instruction . is a business large model for commercial query type data services, i.e. a black box large model that needs to be defended. is a defense prompt word template, is a text description of the attack type output by the classification model for the instruction , represents the defense examples corresponding to the attack type output by the classification model extracted from the defense context database , for example including: harmless user queries , attack instructions designed by previous attackers , the attack type judged by the large model for the attack instruction (i.e. the attack type judged by calling the classification model), and the response of the large model for the attack instruction (i.e. the evaluation result obtained by calling the security evaluation interface to evaluate the output of the business large model for security), . represents the prompt word obtained after filling the above contents in the defense prompt word template, represents the response result of the black box model for the instruction , represents the evaluation result obtained after the security evaluation on the response result. In this way, the prompt word can be constructed as follows: please combine the historical records and attack descriptions of the type of attack, and carefully judge whether the user query conforms to the malicious attack. If it does, refuse to answer; otherwise, directly answer the user query. The entire process can rely on asynchronous MCP protocol calls to achieve acceleration.

[0104] Continuing to refer to Figure 6 , in order to update the attack context database and the defense context database through context learning, the embodiments of the present application can also enter the new model output into the context database (including the attack context database and the defense context database). For example, for the first t iteration update process, each instruction in the mixed training set of the first t round, the classification label of the instruction, and the evaluation result corresponding to the instruction can be appended to the corresponding context database, that is, . contains the original harmless user query , the generated attack instruction , the classification label and the evaluation result , wherein, represents the context database (including the attack context database and the defense context database) after the first t round of iteration update, represents the context database after the first t -1 round of iteration update. In the subsequent training process, the attack context database only adds the entry of the evaluation result , and the defense context database only adds the entry of the evaluation result . Thus, the context database is automatically enriched after each generation and evaluation, the new examples generated drive the accumulation of attack samples, and the evaluation results drive the accumulation of defense scenarios. The two are alternately entered in the same database, forming a generative adversarial context learning.

[0105] The embodiments of the present application replace the supervised fine-tuning method provided by the related art by means of "dynamic context learning", which can achieve security defense and strategy update for a black-box large model without touching the internal parameters of the model. Without relying on the fine-tuning interface opened by the model, the defense context database can be dynamically loaded and updated at runtime, thereby significantly expanding the scope of application and being compatible with various commercial query business interfaces based on closed-source models that only provide inference services.

[0106] In step 105, based on the test set, the defense database after multiple rounds of iteration update is pruned to obtain a pruned defense database.

[0107] In some embodiments, step 105 can be implemented in the following manner: performing the following tests multiple times: randomly deleting a set proportion of defense examples from a defense database that has undergone multiple rounds of iterative updates to obtain a deleted defense database; generating an attack instruction corresponding to each historical query instruction in the test set; generating a mixed training set based on multiple historical query instructions included in the test set and multiple attack instructions corresponding to the multiple historical query instructions; determining an evaluation result corresponding to each instruction included in the mixed training set, where the instruction is a historical query instruction or an attack instruction; determining the evaluation result as the number of correct triggers and correct releases, and taking the ratio of this number to the total number of multiple evaluation results as the accuracy corresponding to this test; and taking the deleted defense database corresponding to the test with the highest accuracy as the defense database after this round of pruning.

[0108] In other embodiments, continuing with the above example, when randomly deleting a set proportion of defense examples from a defense database that has undergone multiple rounds of iterative updates to obtain a deleted defense database, the following process may be further performed: A set proportion of attack examples is randomly deleted from an attack database that has undergone multiple rounds of iterative updates to obtain a deleted attack database. The above-described attack instruction corresponding to each historical query instruction in the test set may be generated by: for each historical query instruction in the test set, the historical query instruction, a target attack type randomly selected from a plurality of predefined attack types, and at least one attack example corresponding to the target attack type extracted from the deleted attack database are entered into a prompt word template to obtain a prompt word; a generation request carrying the prompt word is sent to a cloud-based large model or a local large model via a model context protocol, and a generation result returned by the cloud-based large model or the local large model is received; the generation result is separated by a filter, and the separated result is filtered to obtain the attack instruction corresponding to the historical query instruction. After multiple tests, the following process may be further performed: the deleted attack database corresponding to the test with the lowest accuracy is used as the attack database after this round of pruning.

[0109] In some embodiments, after executing Figure 3After step 105, the following processing may be performed: determining the difference between the maximum accuracy corresponding to the current pruning process and the maximum accuracy corresponding to the previous pruning process; when the difference is less than a difference threshold, or when the performance of the defense database after the current pruning is lower than that of the defense database after the previous pruning, determining a replacement probability based on the difference, and determining whether to use the defense database after the current pruning to replace the defense database after the previous pruning according to the replacement probability; when the difference is greater than the difference threshold, or when the performance of the defense database after the current pruning is higher than that of the defense database after the previous pruning, replacing the defense database after the previous pruning with the defense database after the current pruning.

[0110] For example, see Figure 7 , Figure 7 This is a schematic diagram of the fourth process of the security defense method for the business model provided by the embodiment of the present application, such as Figure 7 As shown, in order to ensure that the scale of the context database is controllable and to capture the most representative "strong attack" and "strong defense" scenarios, the embodiment of the present application can also perform pruning and annealing operations. The iterative update of the round can be done on a fixed test set On the current attack context database, implement "Randomly delete The experiment of "recording the lowest accuracy point (representing the strongest attack) and the highest accuracy point (representing the strongest defense)" is marked, and the experiment with the lowest accuracy point is recorded. Delete 1 / M The attack context database obtained after the proportion of attack examples is used as the attack context database after this round of pruning, and the attack context database in the experiment with the highest accuracy is used as the attack context database after this round of pruning. Delete 1 / M The defense context database obtained after the proportion of defense examples is used as the defense context database after this round of pruning. The accuracy (assuming it is recorded as Acc) can be defined as: Acc = Afterwards, you can also save the The context database and the corresponding accuracy after iterative update and pruning. In addition, after each pruning, the accuracy difference between two adjacent times can be compared (assuming it is recorded as ) is less than the difference threshold, or whether performance degradation occurs. If the difference is less than the difference threshold, or performance degradation occurs, the probability Decide whether to accept the new result, where the temperature satisfy , the geometric cooling decay is met; otherwise, the context database is directly updated, so that the local optimum can be jumped out. Finally, the preset maximum number of iterations or the accuracy threshold (supposed to be ) is reached, the automatic termination is performed.

[0111] It should be noted that the pruning methods such as information gain pruning or Bayesian optimization screening can be used instead of the above pruning strategies. Different pruning methods have their own comfort zones on different difficulty and complexity of business data sets. In actual application, appropriate pruning methods can be selected according to actual conditions.

[0112] The embodiments of the present application combine Monte Carlo experiments and simulated annealing pruning algorithms, periodically remove redundant or weak representative samples, and design pruning strategies for attack context database and defense context database respectively, which can greatly reduce the inference overhead of large models, save computing resources while ensuring defense effect.

[0113] In step 106, in response to the to-be-processed query instruction, a classification model corresponding to the to-be-processed query instruction is determined, and a defense example corresponding to the classification label is extracted from the pruned defense database.

[0114] In some embodiments, when the server receives the to-be-processed query instruction, it can first perform classification processing on the to-be-processed query instruction through the classification model to obtain a classification label (i.e., an attack type) corresponding to the to-be-processed query instruction, and then extract at least one defense example corresponding to the attack type output by the classification model from the pruned defense context database.

[0115] In step 107, a prompt word is generated based on the to-be-processed query instruction, the classification label, and the defense example, and the prompt word is input into the business large model to make the business large model output a business data query result.

[0116] In some embodiments, when switching from the training mode to the online real-time defense mode, for any input query (i.e., the to-be-processed query instruction), the embodiments of the present application can first classify the query through the MCP calling classification model to obtain the corresponding label (indicating the corresponding attack type), and then extract the defense example corresponding to the label (supposed to be ) from the final pruned defense context database, and fill it into the same static template together with the label In some embodiments, the business large model for commercial query type data services can be called by the MCP , and thus a secure answer (denoted as ) can be generated online. The above process can be written as , . After that, if the speed is allowed to slow down, the embodiments of the present application can also choose to use to complete the security verification, thereby further improving security. In addition, through the asynchronous MCP calling process, combined with the pruning method to limit the length of the input context, the system can ensure low delay response. At the same time, in subsequent business development, new specialized business data can be collected to update the context database from time to time, so that the defense capability of the system can also be adapted to continuously improve.

[0117] By adopting the asynchronous MCP calling mechanism, cooperating with the static prompt word template, attack type pre-judgment and pruning restriction, the embodiments of the present application parallelize the context length and interface calling, so that the single security decision response delay is significantly reduced to an acceptable range. At the same time, the embodiments of the present application also retain the optional level of multi-level security evaluation interface, which can meet the flexible optimization demand between "extreme performance" and "extreme security" in different business scenarios.

[0118] In summary, the security defense method of the business large model provided by the embodiments of the present application has the following beneficial effects:

[0119] 1) Through the asynchronous model calling process based on MCP, the user request can be classified and a secure response can be generated immediately when the user request arrives without blocking the main business logic, so that the real-time defense is ensured, and the throughput and delay requirements of the system are also taken into account.

[0120] 2) By continuously writing the generation of adversarial samples (i.e., attack examples and defense examples) and model evaluation results back to the context database, the attack examples and defense examples are cyclically accumulated, and a dynamic self-reinforcing adversarial learning closed loop is constructed, so that the continuous effectiveness of the defense strategy can be improved without fine-tuning the black box model (i.e., the business large model).

[0121] 3) Two sets of structured templates are designed for attack instructions and security decisions respectively, which effectively utilize the accumulated learning samples in the context database. In addition, by automatically generating attack examples and defense examples in a unified format, the black box model can better utilize the context knowledge.

[0122] 4) The dual library architecture (including attack context database and defense context database) is adopted to store "strong attack" samples and "strong defense" samples respectively, and the writing is distinguished according to the evaluation results, so that the high-risk attack scene can be focused, the robust defense experience can be accumulated, and the data organization can be kept clear and efficient.

[0123] 5) The embodiment of the application determines whether the black box model produces a response other than the expected response according to the execution result of the specific attack instruction by calling the large model through the MCP, thereby quantitatively identifying the security flaws of the model, and further better realizing the measurable evaluation of the defense capability of the model, and providing labels for the generative adversarial training.

[0124] 6) In the adversarial training and real-time defense, a special classification model is called to distinguish the attack type of each query, so that different security decisions can be taken for different threats. At the same time, the context length of the input can be reduced, and the response speed of the model can be faster.

[0125] 7) The Monte Carlo experiment and the simulated annealing idea are combined to periodically eliminate redundant or ineffective samples (including attack examples and defense examples) from the context database (including attack context database and defense context database), and the most representative "strong attack / strong defense" scene is retained according to the probability criterion, so that the library size is controlled and the local optimum is avoided.

[0126] 8) The original query data is preliminarily cleaned through threshold screening, denoising, embedding deduplication and outlier detection, and combined with clustering and up-sampling enhancement, the main and long-tail sample distribution is balanced, so that the breadth and diversity of subsequent adversarial examples can be ensured.

[0127] The following continues to illustrate an exemplary structure of the security defense device 543 of the business large model provided by the embodiment of the application as a software module. In some embodiments, as shown in Figure 2 the software modules in the security defense device 543 of the business large model stored in the memory 540 can include: an acquisition module 5431, a business adaptation module 5432, a division module 5433, an update module 5434, a pruning module 5435, a determination module 5436, an extraction module 5437, a generation module 5438 and an input module 5439.

[0128] The acquisition module 5431 is configured to acquire a first query set, the first query set including historical query instructions submitted by a user to a business large model; the business adaptation module 5432 is configured to perform business adaptation processing on the first query set to obtain a second query set; the division module 5433 is configured to divide the second query set into a training set and a test set; the update module 5434 is configured to perform multiple rounds of iterative updating on an attack database and a defense database based on the training set, the attack database including attack examples for the business large model, and the defense database including defense examples for the business large model; wherein, the first t round of iterative updating process includes: for each historical query instruction in the training set, generating an attack instruction based on the historical query instruction, a selected attack type, and an attack example corresponding to the attack type extracted from the attack database updated after the first t -1 rounds of iterative updating; generating a mixed training set based on the multiple attack instructions and the multiple historical query instructions included in the second query set; and updating the attack database and the defense database updated after the first t -1 rounds of iterative updating based on the mixed training set, wherein, t is a positive integer greater than 1 and less than or equal to T , T is the total number of multiple rounds of iterative updating; the pruning module 5435 is configured to perform pruning processing on the defense database updated after the multiple rounds of iterative updating based on the test set to obtain a pruned defense database; the determination module 5436 is configured to determine a classification label corresponding to a to-be-processed query instruction in response to the to-be-processed query instruction; the extraction module 5437 is configured to extract a defense example corresponding to the classification label from the pruned defense database; the generation module 5438 is configured to generate a prompt word based on the to-be-processed query instruction, the classification label, and the defense example; and the input module 5439 is configured to input the prompt word into the business large model to enable the business large model to output a business data query result.

[0129] In some embodiments, the business adaptation module 5432 is further used to delete historical query instructions whose length is less than a length threshold from the multiple historical query instructions included in the first query set to obtain a third query set; perform denoising processing on the third query set based on a regular expression or set rules to obtain a fourth query set; perform hash deduplication on the fourth query set to obtain a fifth query set; map the historical query instructions included in the fifth query set to corresponding embedding vectors through an embedding model, and determine the similarity between any two embedding vectors; for two embedding vectors whose similarity is greater than a similarity threshold, delete the historical query instruction corresponding to one of the embedding vectors from the fifth query set to obtain a sixth query set; delete the historical query instructions in a discrete state from the sixth query set to obtain a seventh query set; obtain a subset of the seventh query set, and train a clustering model based on the subset; cluster the multiple historical query instructions included in the seventh query set based on the trained clustering model to obtain multiple clusters; and generate a second query set based on the multiple historical query instructions included in the multiple clusters.

[0130] In some embodiments, the business adaptation module 5432 is further used to determine the number of historical query instructions included in each cluster; for clusters whose number is greater than a quantity threshold, down-sample the multiple historical query instructions included in the cluster until the number of historical query instructions included in the cluster is reduced to a first number; for clusters whose number is less than the quantity threshold, resample the multiple historical query instructions included in the cluster, or back-translate and enhance the multiple historical query instructions included in the cluster through a large language model until the number of historical query instructions included in the cluster increases to a second number; merge the updated historical query instructions respectively included in multiple clusters to obtain a second query set.

[0131] In some embodiments, the update module 5434 is further configured to update the historical query instruction, the target attack type randomly selected from a plurality of predefined attack types, and the target attack type randomly selected from the predefined attack types. t -1 round of iterative update of the attack database, at least one attack example corresponding to the target attack type is extracted and filled into the prompt word template to obtain the prompt word; a generation request carrying the prompt word is sent to the cloud-based large model or the local large model through the model context protocol, and the generation result returned by the cloud-based large model or the local large model is received; the generation result is separated by a filter, and the separated result is filtered to obtain the attack instruction.

[0132] In some embodiments, the updating module 5434 is further configured to collect multiple attack instructions corresponding to multiple historical query instructions included in the training set, and obtain the first t The attack instruction set of the round; textract a third number of attack instructions from the attack instruction set of the round, and extract a third number of historical query instructions from the second query set; construct a first mixed training set based on the third number of attack instructions and the third number of historical query instructions t the mixed training set of the round.

[0133] In some embodiments, the update module 5434 is further configured to perform the following processing for each instruction in the mixed training set, where the instruction is an attack instruction or a historical query instruction: perform classification processing on the instruction by the classification model to obtain an attack type corresponding to the instruction; fill the defense examples corresponding to the attack type extracted from the defense database updated after the first round of iterative updates into the prompt word template to obtain a prompt word; input the prompt word into the business large model and obtain a response result output by the business large model for the prompt word; call the security evaluation interface to perform security evaluation on the response result to obtain an evaluation result; and in response to the evaluation result being any one of correct triggering, non-triggering, crash, or termination, add the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction as an attack example to the attack database updated after the first round of iterative updates. t -1 round of iterative updates, wherein when the instruction is a historical query instruction, the attack example further includes an attack instruction generated based on the historical query instruction; and when the instruction is an attack instruction, the attack example further includes a historical query instruction corresponding to the attack instruction; and in response to the evaluation result being any one of incorrect triggering, non-triggering, crash, or termination, add the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction as a defense example to the defense database updated after the first round of iterative updates. t -1 round of iterative updates, wherein when the instruction is a historical query instruction, the attack example further includes an attack instruction generated based on the historical query instruction; and when the instruction is an attack instruction, the attack example further includes a historical query instruction corresponding to the attack instruction; and in response to the evaluation result being any one of incorrect triggering, non-triggering, crash, or termination, add the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction as a defense example to the defense database updated after the first round of iterative updates. t -1 round of iterative updates, wherein when the instruction is a historical query instruction, the attack example further includes an attack instruction generated based on the historical query instruction; and when the instruction is an attack instruction, the attack example further includes a historical query instruction corresponding to the attack instruction; and in response to the evaluation result being any one of incorrect triggering, non-triggering, crash, or termination, add the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction as a defense example to the defense database updated after the first round of iterative updates.

[0134] In some embodiments, the pruning module 5435 is further configured to perform the following test multiple times: randomly delete a set proportion of defense examples from the defense database updated after multiple rounds of iterative updates to obtain a defense database after deletion; generate an attack instruction corresponding to each historical query instruction in the test set; generate a mixed training set based on the multiple historical query instructions included in the test set and the multiple attack instructions corresponding to the multiple historical query instructions; determine the evaluation result corresponding to each instruction included in the mixed training set, where the instruction is a historical query instruction or an attack instruction; determine the number of correct triggering and correct release of the evaluation results, and take the ratio of the number to the total number of the multiple evaluation results as the accuracy of the current test; and take the defense database after deletion corresponding to the test with the highest accuracy as the defense database after pruning.

[0135] In some embodiments, when a set proportion of defense examples are randomly deleted from the defense database updated after multiple rounds of iterations to obtain a defense database after deletion, the pruning module 5435 is further configured to perform the following processing: a set proportion of attack examples are randomly deleted from the attack database updated after multiple rounds of iterations to obtain an attack database after deletion; for each historical query instruction in the test set, the historical query instruction, a target attack type randomly selected from a plurality of predefined attack types, and at least one attack example corresponding to the target attack type extracted from the attack database after deletion are filled into the prompt template to obtain a prompt; a generation request carrying the prompt is sent to the cloud large model or the local large model through the model context protocol, and a generation result returned by the cloud large model or the local large model is received; the generation result is separated by the filter, and the classification result is filtered to obtain an attack instruction corresponding to the historical query instruction; the attack database after deletion corresponding to the test with the smallest accuracy is taken as the pruned attack database.

[0136] In some embodiments, after the pruning module 5435 prunes the defense database updated after multiple rounds of iterations based on the test set to obtain a pruned defense database, the determination module 5436 is further configured to determine a difference between the maximum accuracy corresponding to the present pruning processing and the maximum accuracy corresponding to the last pruning processing; the update module 5434 is further configured to, when the difference is less than a difference threshold, or the performance of the pruned defense database after the present pruning processing decreases compared to the performance of the pruned defense database after the last pruning processing, determine a replacement probability based on the difference, and determine whether to replace the pruned defense database after the last pruning processing with the pruned defense database after the present pruning processing according to the replacement probability; the update module 5434 is further configured to, when the difference is greater than the difference threshold, or the performance of the pruned defense database after the present pruning processing increases compared to the performance of the pruned defense database after the last pruning processing, replace the pruned defense database after the last pruning processing with the pruned defense database after the present pruning processing.

[0137] It should be noted that the description of the device embodiments of the present application is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments, and therefore will not be described here. For technical details not described in the security defense device of the service large model provided by the embodiments of the present application, they can be understood according to the description of any one of the accompanying drawings. Figures 3 to 7

[0138] ​The embodiment of the present application provides a computer program product, which comprises a computer program or computer executable instructions stored in a computer readable storage medium. The processor of the computer device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the computer device executes the security defense method of the business large model provided in the embodiment of the present application.

[0139] The embodiment of the present application provides a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions are stored in the computer readable storage medium. When the computer executable instructions are executed by the processor, the processor executes the security defense method of the business large model provided in the embodiment of the present application, for example, as shown in the security defense method of the business large model. Figure 3 The security defense method of the business large model.

[0140] In some embodiments, the computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM memory, etc. It can also be various devices including one or any combination of the above storage devices.

[0141] In some embodiments, the executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.

[0142] As an example, the executable instructions can be deployed to execute on one electronic device, or on multiple electronic devices located in one place, or on multiple electronic devices distributed in multiple places and interconnected through a communication network.

[0143] The above is only an embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application shall be included in the protection scope of the present application.

Claims

1. A security defense method of a service large model, characterized in that, The business large model is used for a business data query service, and the method comprises: obtaining a first query set comprising historical query instructions submitted by a user to the business large model; performing business adaptation processing on the first query set to obtain a second query set, and dividing the second query set into a training set and a test set; Based on the training set, the attack database and the defense database are updated in multiple rounds of iteration, the attack database includes attack examples for the business large model, and the defense database includes defense examples for the business large model; wherein, the first t round of iteration updating process includes: for each historical query instruction in the training set, filling the historical query instruction, a target attack type randomly selected from a plurality of predefined attack types, and at least one attack example corresponding to the target attack type extracted from the attack database after the first t -1 round of iteration updating into a prompt word template to obtain a prompt word; sending a generation request carrying the prompt word to a cloud large model or a local large model through a model context protocol, and receiving a generation result returned by the cloud large model or the local large model; separating the generation result through a filter, and filtering the separation result to obtain an attack instruction; generating a mixed training set based on a plurality of attack instructions and a plurality of historical query instructions included in the second query set; updating the attack database and the defense database after the first t -1 round of iteration updating based on the mixed training set, wherein, t is a positive integer greater than 1, and satisfies less than or equal to T , T is the total number of multiple rounds of iteration. based on the test set, pruning the defense database updated after multiple iterations to obtain a pruned defense database; in response to a to-be-processed query instruction, determining a classification label corresponding to the to-be-processed query instruction, and extracting a defense example corresponding to the classification label from the pruned defense database; based on the to-be-processed query instruction, the classification label, and the defense example, generating a prompt word, inputting the prompt word into the business large model, and causing the business large model to output a business data query result.

2. The method of claim 1, wherein, The business adaptation processing on the first query set to obtain a second query set comprises: from the plurality of historical query instructions included in the first query set, deleting historical query instructions with a length less than a length threshold to obtain a third query set; based on a regular expression or a set rule, performing denoising processing on the third query set to obtain a fourth query set; performing hash deduplication on the fourth query set to obtain a fifth query set; mapping the historical query instructions included in the fifth query set to corresponding embedding vectors through an embedding model, and determining the similarity between any two embedding vectors; for the two embedding vectors with a similarity greater than a similarity threshold, deleting the historical query instruction corresponding to one of the embedding vectors from the fifth query set to obtain a sixth query set; from the sixth query set, deleting historical query instructions in a discrete state to obtain a seventh query set; obtaining a subset of the seventh query set, and training a clustering model based on the subset; based on the trained clustering model, clustering a plurality of historical query instructions included in the seventh query set to obtain a plurality of clusters; based on the plurality of historical query instructions included in the plurality of clusters, generating a second query set.

3. The method of claim 2, wherein, The generation of a second query set based on the plurality of historical query instructions included in the plurality of clusters comprises: determining the number of historical query instructions included in each cluster; for the clusters with a number greater than a number threshold, down-sampling the plurality of historical query instructions included in the clusters until the number of historical query instructions included in the clusters is reduced to a first number; for the clusters with a number less than a number threshold, re-sampling the plurality of historical query instructions included in the clusters, or back-translation enhancing the plurality of historical query instructions included in the clusters through a large language model, until the number of historical query instructions included in the clusters is increased to a second number; merging the updated historical query instructions included in the plurality of clusters to obtain a second query set.

4. The method of claim 1, wherein, The generation of a mixed training set based on a plurality of attack instructions and a plurality of historical query instructions included in the second query set comprises: The attack instructions corresponding to the plurality of historical query instructions included in the training set are collected to obtain a plurality of attack instruction sets corresponding to the plurality of historical query instructions included in the training set. t The attack instruction set of the wheel; From the said t Extracting a third number of attack instructions from the attack instruction set of the round, and extracting the third number of historical query instructions from the second query set; Based on the third number of attack instructions, and the third number of historical query instructions, a first t mixed training set of wheels is constructed.

5. The method of claim 1, wherein, The method comprises the following steps: t The attack database and the defense database are updated after one round of iteration, comprising: The following processing is performed for each instruction in the mixed training set, wherein the instruction is the attack instruction or the historical query instruction: Classifying the instruction through a classification model to obtain an attack type corresponding to the instruction; The instructions, the attack type, the description text of the attack type, and the defense example corresponding to the attack type extracted from the defense database updated after the first round of iteration are filled into the prompt word template to obtain a prompt word. t The defense example corresponding to the attack type extracted from the defense database updated after the first round of iteration is filled into the prompt word template to obtain a prompt word. Inputting the prompt word into the business large model and obtaining a response result output by the business large model for the prompt word; Calling a security evaluation interface to perform security evaluation on the response result to obtain an evaluation result; In response to the evaluation result being any one of correctly triggered, not triggered, crashed, or terminated, the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction are taken as an attack example and added to the attack example after the first t - In the attack database after 1 round of iterative update, when the instruction is the historical query instruction, the attack example also includes an attack instruction generated based on the historical query instruction; when the instruction is the attack instruction, the attack example also includes a historical query instruction corresponding to the attack instruction; In response to the evaluation result being any one of a false trigger, a non-trigger, a crash, or a termination, the instruction, the attack type corresponding to the instruction, and the evaluation result corresponding to the instruction are added as a defense example into the defense database updated after the first round of iteration, wherein when the instruction is the historical query instruction, the defense example further includes an attack instruction generated based on the historical query instruction; when the instruction is the attack instruction, the defense example further includes the historical query instruction corresponding to the attack instruction. t -1 round of iteration, wherein when the instruction is the historical query instruction, the defense example further includes an attack instruction generated based on the historical query instruction; when the instruction is the attack instruction, the defense example further includes the historical query instruction corresponding to the attack instruction.

6. The method of claim 1, wherein, The pruning processing is performed on the defense database updated after multiple iterations based on the test set to obtain a pruned defense database, including: The following test is performed multiple times: Randomly deleting a set proportion of defense examples from the defense database updated after multiple iterations to obtain a deleted defense database; For each historical query instruction in the test set, an attack instruction corresponding to the historical query instruction is generated; A mixed training set is generated based on a plurality of historical query instructions included in the test set and a plurality of attack instructions corresponding one-to-one to the plurality of historical query instructions; The evaluation result corresponding to each instruction included in the mixed training set is determined, wherein the instruction is the historical query instruction or the attack instruction; The number of correct triggers and correct releases is determined, and the ratio of the number to the total number of a plurality of evaluation results is taken as the accuracy corresponding to the test; The deleted defense database corresponding to the test with the maximum accuracy is taken as the pruned defense database.

7. The method of claim 6, wherein, In the process of randomly deleting a set proportion of defense examples from the defense database updated after multiple iterations to obtain a deleted defense database, the method further includes: Randomly deleting a set proportion of attack examples from the attack database updated after multiple iterations to obtain a deleted attack database; The process of generating an attack instruction corresponding to a historical query instruction for each historical query instruction in the test set includes: For each historical query instruction in the test set, the historical query instruction, a target attack type randomly selected from a plurality of predefined attack types, and at least one attack example corresponding to the target attack type extracted from the deleted attack database are filled into a prompt word template to obtain a prompt word; A generation request carrying the prompt word is sent to a cloud large model or a local large model through a model context protocol, and a generation result returned by the cloud large model or the local large model is received; The generation result is separated through a filter, and the separated result is filtered to obtain an attack instruction corresponding to the historical query instruction; The method further includes: The deleted attack database corresponding to the test with the minimum accuracy is taken as a pruned attack database.

8. The method according to any one of claims 1 to 7, characterized in that, After the pruning processing is performed on the defense database updated after multiple iterations based on the test set to obtain a pruned defense database, the method further includes: Determining the difference between the maximum accuracy corresponding to the current pruning processing and the maximum accuracy corresponding to the last pruning processing; When the difference value is less than a difference value threshold, or the performance of the defense database after the current pruning decreases compared to the performance of the defense database after the last pruning, a replacement probability is determined based on the difference value, and it is determined whether to replace the defense database after the last pruning with the defense database after the current pruning according to the replacement probability. When the difference value is greater than the difference value threshold, or the performance of the defense database after the current pruning increases compared to the performance of the defense database after the last pruning, the defense database after the last pruning is replaced with the defense database after the current pruning.

9. A security device for a service large model, characterized by, The business large model is used for business data query service, and the apparatus comprises: An acquisition module is configured to acquire a first query set, the first query set comprising historical query instructions submitted by a user to the business large model; A business adaptation module is configured to perform business adaptation processing on the first query set to obtain a second query set; A division module is configured to divide the second query set into a training set and a test set; An updating module is configured to perform multiple rounds of iterative updating on an attack database and a defense database based on the training set, the attack database including attack examples for the business large model, and the defense database including defense examples for the business large model; wherein the first t round of iterative updating includes: for each historical query instruction in the training set, filling at least one attack example corresponding to a target attack type selected randomly from a plurality of predefined attack types into a prompt word template to obtain a prompt word; sending a generation request carrying the prompt word to a cloud large model or a local large model through a model context protocol, and receiving a generation result returned by the cloud large model or the local large model; separating the generation result through a filter, and filtering the separated result to obtain an attack instruction; generating a mixed training set based on a plurality of attack instructions and a plurality of historical query instructions included in the second query set; and updating the attack database and the defense database based on the mixed training set, wherein the attack database and the defense database have been updated through the first t round of iterative updating. t -1 rounds of iterative updating. t is a positive integer greater than 1 and less than or equal to T , T is the total number of multiple rounds of iterative updating. A pruning module is configured to perform pruning processing on the defense database updated after multiple rounds of iteration based on the test set to obtain a defense database after pruning; A determination module is configured to determine a classification label corresponding to a to-be-processed query instruction in response to the to-be-processed query instruction; An extraction module is configured to extract a defense example corresponding to the classification label from the defense database after pruning; A generation module is configured to generate a prompt word based on the to-be-processed query instruction, the classification label, and the defense example; An input module is configured to input the prompt word into the business large model to cause the business large model to output a business data query result.

10. An electronic device, comprising: comprise: a memory configured to store computer executable instructions; a processor configured to execute the computer executable instructions stored in the memory to implement the security defense method of the business large model according to any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, The computer executable instructions are stored in the memory and are configured to be executed by the processor to implement the security defense method of the business large model according to any one of claims 1 to 8.

12. A computer program product, characterised in that, The computer program or computer executable instructions are stored in the memory and are configured to be executed by the processor to implement the security defense method of the business large model according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Database attack defense method and device, storage medium and electronic equipment

    CN117633783A

  • Large language model security optimization method and device, equipment and medium

    CN118965366A