Confrontation sample defense method, system and device and storage medium

By matching input samples and adversarial samples of the target model on the blockchain, the problems of information silos and trust risks in cross-institutional defense are solved, enabling efficient identification and defense against adversarial samples and improving the response speed and effectiveness of the defense system.

CN121960640APending Publication Date: 2026-05-01ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing adversarial sample defense methods suffer from information silos and trust risks in cross-institutional scenarios, making it difficult to achieve real-time, reliable sharing and response, and their generalization ability against new types of attacks is limited.

Method used

By matching the input samples of the target model with the adversarial samples stored on the blockchain, and leveraging the blockchain's tamper-proof, traceable, and multi-subject trusted sharing characteristics, the risk type of the input samples can be determined and an execution strategy can be formulated, including normal reasoning or defensive processing.

Benefits of technology

It improves the accuracy and generalization ability of the model in identifying adversarial examples, enhances the ability to resist adaptive and novel attacks, and strengthens the response speed and defense effect of the model defense system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960640A_ABST
    Figure CN121960640A_ABST
Patent Text Reader

Abstract

The invention discloses an adversarial sample defense method, system and device and a storage medium, and the method comprises the steps: matching an input sample of a target model with each adversarial sample stored on a block chain, and obtaining a sample matching result; determining a risk type of the input sample based on the sample matching result; based on the risk type of the input sample, an execution strategy about the input sample is determined for the target model, and the execution strategy includes indicating the target model to perform normal reasoning or defense processing by using the input sample. By means of the mode, the accuracy and generalization ability of target model confrontation sample recognition can be improved, and therefore the response speed and the defense effect of a model defense system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

An adversarial sample defense method, system, device, and storage medium Technical Field

[0001] This application relates to the field of artificial intelligence security technology, specifically to an adversarial sample defense method, system, device, and storage medium. Background Technology

[0002] In recent years, deep learning, especially large-scale neural network models (hereinafter referred to as "large models"), has achieved remarkable results in many fields such as image recognition, speech recognition, natural language processing, autonomous driving, medical image analysis, and security monitoring. Large models are usually trained on massive amounts of data, and their discriminative and generalization abilities are superior under most common inputs.

[0003] However, research and practice have shown that deep neural networks possess numerous local vulnerabilities in their high-dimensional input space. Specifically, attackers can induce incorrect judgments or anomalous outputs by applying extremely small, often imperceptible perturbations to the input (such as pixel-level noise, tiny insertions in text, or noise in speech). These carefully crafted but seemingly normal input samples are called adversarial examples. Adversarial examples have been proven to successfully attack image classifiers, speech recognition systems, text classification models, and cue-based generative models.

[0004] Traditional adversarial defense primarily includes adversarial training (adding adversarial examples to the training set for robust training), input preprocessing (denoising, smoothing), detectors (detecting anomalous inputs based on statistical features or model uncertainty), and model structure improvements (regularization, gradient masking, etc.). While these methods have improved the model's resistance to known attacks to some extent, they still have many limitations: adversarial training requires a large number of labeled adversarial examples and incurs significant computational overhead; input preprocessing can sometimes impair performance or be bypassed by adaptive attacks; detectors have limited generalization capabilities and are prone to failing against novel attacks. Furthermore, most existing solutions are implemented on a single organizational or deployment environment basis (i.e., "local" or "centralized" defense), leading to information silos and delayed responses. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides an adversarial sample defense method, system, device, and storage medium to improve the accuracy and generalization ability of target model adversarial sample identification, and enhance the response speed and defense effectiveness of the model defense system.

[0006] According to one embodiment of this application, an adversarial sample defense method is provided, comprising: matching an input sample of a target model with adversarial samples stored on a blockchain to obtain a sample matching result; determining the risk type of the input sample based on the sample matching result; and determining an execution strategy for the target model regarding the input sample based on the risk type of the input sample, wherein the execution strategy includes instructing the target model to use the input sample for normal inference or defensive processing.

[0007] To address the aforementioned technical problems, this application provides a technical solution: an adversarial sample defense system, comprising: a model inference defense module, configured to match input samples of a target model with adversarial samples stored on a blockchain to obtain sample matching results; based on the sample matching results, determine the risk type of the input sample; and based on the risk type of the input sample, determine an execution strategy for the target model regarding the input sample, wherein the execution strategy includes instructing the target model to perform normal inference or defense processing using the input sample.

[0008] To solve the above-mentioned technical problems, one technical solution adopted in this application is to provide an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the adversarial sample defense method in the above-mentioned technical solution.

[0009] To solve the above-mentioned technical problems, one technical solution adopted in this application is to provide a computer-readable storage medium for storing a computer program, which, when executed by a processor, is used to implement the adversarial sample defense method in the above-mentioned technical solution.

[0010] Through the above scheme, this application matches the input samples of the target model with each adversarial sample stored on the blockchain to obtain sample matching results. Based on the sample matching results, the risk type of the input sample is determined. Based on the risk type of the input sample, an execution strategy for the target model is determined for the input sample. In this way, this application utilizes the core characteristics of blockchain that endow each adversarial sample with tamper-proof, traceable, and multi-subject trusted sharing. By determining the defense method of the input sample through the matching results of the model's input sample and each adversarial sample stored on the blockchain, the accuracy and generalization ability of the model in identifying adversarial samples are improved, thereby enhancing the ability to resist adaptive and new attacks, and thus improving the response speed and defense effect of the model defense system. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Specifically: Figure 1 is a flowchart illustrating an embodiment of the adversarial sample defense method provided by this application; Figure 2 is a flowchart illustrating another embodiment of the adversarial sample defense method provided by this application; Figure 3 is a schematic diagram of an adversarial sample defense system architecture provided by this application; Figure 4 is a schematic diagram of a system operation data flow relationship provided by this application; Figure 5 is a structural schematic diagram of an embodiment of an electronic device provided by this application; Figure 6 is a structural schematic diagram of an embodiment of a computer-readable storage medium provided by this application. Detailed Implementation

[0012] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the application. Similarly, the following embodiments are only some, not all, embodiments of the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.

[0013] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0014] It should be noted that the terms "first," "second," etc., used in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0015] Adversarial examples refer to input data in which subtle perturbations (such as pixel value fine-tuning, text character replacement, or audio waveform slight modification) are artificially added to the original legitimate input (such as images, text, or audio) of a machine learning or deep learning model. The core characteristic is that the perturbation amplitude does not change the human-recognizable semantics of the input data, but can maliciously mislead the model to output incorrect prediction results. Moreover, the perturbation is usually designed specifically based on the model results or decision logic.

[0016] For example, in autonomous driving and intelligent transportation scenarios, when traffic signs, road markings, or vehicle appearances are subjected to adversarial perturbations, the perception module may misinterpret signals (such as mistaking "stop" for "speed limit"), causing safety hazards; in security monitoring and facial recognition scenarios, attackers can cause facial recognition systems to malfunction or misidentify by wearing specific patterns or subtle makeup, leading to unauthorized access or identity theft; in medical image analysis scenarios, applying adversarial perturbations to medical images may lead to diagnostic errors, affecting patient health; in financial risk control and fraud detection scenarios, attackers can finely modify input features to cause risk assessment models to misjudge credit or fraud risks; in generative large models such as AIGC and dialogue systems, carefully crafted prompt injections or specific inputs can induce models to leak sensitive information or generate harmful content; in AI service markets and cross-institutional collaboration scenarios, when multiple institutions share models or call remote models, it is difficult for a single institution to promptly report and synchronize new adversarial examples detected by other institutions, creating a global defense gap.

[0017] However, existing adversarial defense methods each have limitations. For example, local blacklists or feature libraries achieve real-time comparison by maintaining adversarial sample hash fingerprints locally. Although the query latency is low, updates rely on manual or centralized synchronization, making it difficult to guarantee cross-institutional consistency and posing a single point of tampering risk. Adversarial training and robust optimization enhance model robustness by injecting adversarial samples into the training set, but their generalization ability to new types of adversarial samples is limited, and they have high training costs and poor real-time performance. Detectors based on statistics or uncertainty detect anomalies through input or intermediate layer features, but they are themselves easily bypassed by adaptive attacks, and high false positive rates can also affect system availability. Federated / cooperative methods achieve collaborative defense by sharing gradients or model updates among multiple parties, but they lack an immutable sharing mechanism and proof of responsibility, making it impossible to guarantee the credibility of shared data in cross-institutional scenarios.

[0018] This application provides an adversarial sample defense method. It involves matching the input samples of a target model with adversarial samples stored on a blockchain to obtain sample matching results. Based on these results, the risk type of the input sample is determined, and based on this risk type, an execution strategy is determined for the target model regarding the input sample. Thus, this application leverages the core characteristics of blockchain—tamper-proof, traceable, and multi-subject trusted sharing—that adversarial samples possess. By determining the defense method for the input sample through the matching results between the model's input sample and the adversarial samples stored on the blockchain, it improves the model's accuracy and generalization ability in identifying adversarial samples, thereby enhancing its resistance to adaptive and novel attacks and ultimately improving the response speed and defense effectiveness of the model defense system.

[0019] Please refer to Figure 1, which is a flowchart illustrating an embodiment of the adversarial example defense method provided in this application. It should be noted that if substantially the same result is obtained, this embodiment is not limited to the flow order shown in Figure 1. As shown in Figure 1, this embodiment includes: S110: Matching the input sample of the target model with each adversarial sample stored on the blockchain to obtain sample matching results.

[0020] The process involves acquiring input samples for the target model and matching them with adversarial samples stored on the blockchain to obtain the sample matching results. Input samples are the raw data to be processed by the target model, such as images, text, and audio. These may be normal samples or adversarial samples with artificially added minor perturbations. The adversarial samples stored on the blockchain are submitted by multiple entities (such as research institutions and enterprises) and stored after verification by the blockchain network, along with associated metadata. This metadata includes sample hashes, attack types, and generation sources. Due to the characteristics of blockchain, adversarial samples stored on the blockchain possess the attributes of being tamper-proof, traceable, and trustworthy for sharing.

[0021] Before performing sample matching, the core features of the input samples of the target model can be extracted, as well as the feature data of each adversarial sample on the blockchain, in order to perform feature matching between the input samples and each adversarial sample on the blockchain.

[0022] In one embodiment, input feature data of the input sample is obtained, along with adversarial feature data of each adversarial sample retrieved from the blockchain or from a local adversarial sample library. The input feature data is then compared with each adversarial feature data to obtain a sample matching result. If an adversarial feature data matches the input feature data, the input sample is determined to match the corresponding adversarial sample. For example, for each adversarial feature data, the similarity or hash distance between the input feature data and the adversarial feature data is calculated. If the similarity or hash distance is greater than a preset threshold, the input feature data is determined to match the adversarial feature data.

[0023] In one embodiment, the input feature data is the input sample feature extracted from the input sample, and the adversarial feature data is the adversarial sample feature extracted from the adversarial sample. In another embodiment, the input feature data is the input feature fingerprint obtained by hashing the input sample features, and the adversarial feature data is the adversarial feature fingerprint obtained by hashing the adversarial sample features. The adversarial feature fingerprint is a compact representation or signature extracted from the sample, used for rapid comparison and retrieval, such as based on perceptual hashing, feature vector binarization, or hashing of locally invariant descriptors. Utilizing the hash chain structure (Merkle Tree) and node signature mechanism of blockchain, all sample features and detection results are encrypted and recorded, thereby ensuring the authenticity, integrity, and traceability of sample records, providing verifiable technical proof for security incidents.

[0024] In one embodiment, the local adversarial sample library is used to store the adversarial feature data of each adversarial sample stored in the blockchain. This can be obtained through the following steps: collecting several adversarial samples to be uploaded to the blockchain; acquiring the adversarial feature data of each sample; encapsulating several pieces of metadata of each sample into transaction data, including the adversarial feature data; and signing the transaction data of each sample before it is uploaded to the blockchain's on-chain sample library. The transaction data is a structured data record initiated between nodes in the blockchain network based on preset rules (such as smart contract protocols) to record value transfers or information exchanges. It is the smallest record unit of the blockchain ledger, stored in blocks after encryption, and possesses the attributes of immutability, traceability, and network-wide consensus verification. Furthermore, it is directly related to the triggering and execution of smart contracts.

[0025] The metadata also includes at least one of the following: the source node identifier, attack type, detection confidence level, and timestamp of the adversarial sample to be uploaded to the blockchain. Specifically, the source node identifier refers to the unique identity of the consortium node that generated or submitted the adversarial sample; the attack type refers to the technical category of the adversarial attack corresponding to the adversarial sample, providing the defense system with direct evidence of the sample's attack attributes and assisting smart contracts in quickly matching targeted defense strategies; the detection confidence level refers to the quantified probability value of determining that the sample is a valid adversarial sample (or corresponds to a certain attack type) after the submitting node or authorized detection node has pre-detected it; and the timestamp refers to the precise time record of when the sample was generated or submitted to the blockchain network.

[0026] In one embodiment, before encapsulating several metadata of the adversarial sample into the transaction data of the adversarial sample, the adversarial feature data of the adversarial sample to be uploaded to the chain can also be encrypted. For example, encryption can be performed by salted hashing or zero-knowledge encryption.

[0027] In one implementation, a smart contract can be used to verify the legitimacy of the transaction data of the adversarial sample to be uploaded to the blockchain. After the legitimacy verification is passed, the signed transaction data of the adversarial sample to be uploaded to the blockchain is stored in the on-chain sample library. The legitimacy verification may include verifying the uniqueness between the digital signature of the uploading node and the adversarial feature fingerprint in the transaction data of the adversarial sample to be uploaded to the blockchain. The smart contract is automated rule execution code deployed on the blockchain, which can realize automatic verification, notification, and synchronization mechanisms after the sample is uploaded to the blockchain. This allows the defense system to respond quickly after an attack occurs, reducing manual intervention and delays, and achieving automated and autonomous updates.

[0028] In one embodiment, the sample data corresponding to the adversarial sample to be uploaded to the blockchain can also be stored in a secure database or in IPFS (InterPlanetary File System, a decentralized distributed file storage and transfer protocol), and can be decrypted and accessed only when necessary. The sample data includes at least one of the following: the adversarial sample to be uploaded to the blockchain, the encrypted adversarial sample to be uploaded to the blockchain, the adversarial feature data of the adversarial sample to be uploaded to the blockchain, the encrypted adversarial feature data, etc.

[0029] In this embodiment, by utilizing the distributed ledger structure of blockchain, adversarial sample features and several metadata are stored on the chain, thereby enabling real-time sharing and verification of defense data among multiple institutions and avoiding data silos and trust risks in centralized systems.

[0030] S120: Based on the sample matching results, determine the risk type of the input sample.

[0031] The sample matching results include adversarial samples that match the input sample and adversarial samples that do not match the input sample. Risk types include risky and risk-free. In the case of a risky type, the risk type can be further divided into risk types with different risk levels. The risk level can be comprehensively determined through blockchain trusted data, multi-dimensional quantitative indicators, and a pre-set weighted model in the smart contract. This transforms the threat attributes of the adversarial sample into a calculable and comparable quantitative score, which is then mapped to a pre-set risk level such as low, medium, high, or extremely high risk. For example, based on the sample's own threat characteristics, combined with auxiliary dimensions such as data credibility, attack severity, and application scenario sensitivity, a comprehensive risk score is calculated using a pre-set weighted algorithm in the smart contract. Finally, the risk level is determined based on the score threshold.

[0032] After obtaining the sample matching results between the input sample of the target model and each adversarial sample stored in the blockchain, the risk type of the input sample is determined based on the sample matching results.

[0033] In one embodiment, in response to a sample matching result indicating that no adversarial sample matches the input sample, the risk type of the input sample is determined to be risk-free. In response to a sample matching result indicating that an adversarial sample matches the input sample, the risk type of the input sample is determined to be risky.

[0034] In one implementation, if the risk type of the input sample is risky, the risk level of the input sample is further determined to execute targeted defense strategies based on different risk levels. For example, the system obtains the sample matching degree and attack type of the input sample from the blockchain, the detection confidence level from the detection module, the reputation score of the submitting node from the node reputation database, and the application scenario weight from the smart contract. These factors are converted into a unified score, and the total risk score of the input sample is calculated using a preset algorithm in the smart contract. Finally, the risk level of the input sample is determined based on a preset threshold. Furthermore, if a factor triggers a special rule, such as a perfect match with a known high-risk adversarial sample or malicious node submission, the risk score calculation can be skipped, and the corresponding level can be directly determined as high risk.

[0035] S130: Based on the risk type of the input sample, determine the execution strategy for the target model with respect to the input sample.

[0036] The execution strategy refers to the set of final action instructions issued by the smart contract to the target model based on sample matching results and preset defense rules. This includes instructing the target model to perform normal reasoning or defensive processing using the input samples. For input samples requiring normal processing, the target model follows the standard procedure of feature extraction, computation, and outputting the final result according to its original design logic. For input samples requiring defensive processing, the target model initiates a targeted processing procedure based on preset protection mechanisms to resist adversarial disturbances and prevent the model from outputting erroneous results.

[0037] In one embodiment, in response to the input sample being risk-free, the execution strategy is determined to perform normal inference using the input sample. In response to the input sample being risky, the execution strategy is determined to be a defense strategy matching the risk level of the input sample. The defense strategy represents the defensive processing performed on the input sample, and the higher the risk level, the stronger the defense effect of the corresponding defense strategy. The defense processing includes, but is not limited to: adaptively denoising and reconstructing features on the sample to remove adversarial perturbations before inference; directly rejecting high-risk samples from the model to avoid erroneous output affecting business operations; activating a backup model to re-infer the sample to improve anti-interference capabilities; and combining multi-model cross-inference to confirm the true attributes of the sample before outputting the result, etc.

[0038] In one example, risk levels are categorized as low, medium, and high risk. For each risk level, a defense strategy matching the risk level of the input sample is executed. In one implementation, in response to the input sample being classified as low risk, a first defense strategy is determined, which includes performing inference and performing at least one of the following processes: issuing a prompt and logging. In response to the input sample being classified as medium risk, a second defense strategy is determined, which includes performing inference in an isolated virtual environment and isolating the inference results. In response to the input sample being classified as high risk, a third defense strategy is determined, which includes refusing to perform inference and reporting a security incident.

[0039] Input samples identified as medium to high risk are imported into an isolated virtual environment. The target model is then invoked to complete the inference process, including feature extraction and computation. The entire inference process is confined within the isolated environment and does not affect the operation of the main system. The output data of the model inference in the isolated virtual environment is first verified and is not directly synchronized to the main business system before its security is confirmed. Instead, it is stored in the isolated environment or a trusted blockchain storage area. This prevents high-risk inputs from directly entering the production model, reduces the success rate of attacks, and improves the security of system operation.

[0040] In one embodiment, after determining the execution strategy for the target model regarding the input sample, in response to the execution strategy instructing the target model to perform defensive processing using the input sample, the defensive processing result of the input sample is stored in the blockchain. Exemplarily, synchronizing the defensive processing result of the input sample on the blockchain to form a new transaction record constructs a complete event loop of sample input, matching detection, defense execution, and result traceability, ensuring that the data throughout the entire adversarial defense process is tamper-proof, traceable, and with clear responsibilities, achieving closed-loop event traceability.

[0041] In one embodiment, a consortium blockchain network can be constructed, including model service providers, security research institutions, and regulatory nodes. This allows different AI service providers and security agencies to participate as nodes in the verification and dissemination of defense information, thereby enabling joint defense capabilities for large-scale distributed AI systems. For example, each consortium node in the network can receive notifications of new adversarial samples being uploaded to the blockchain, pull the latest adversarial sample data from the blockchain through a defense interface, and update its local adversarial sample library. The local adversarial sample library stores the data of various adversarial samples stored on the blockchain, and the notification of a new adversarial sample being uploaded is a broadcast triggered by a smart contract upon its upload.

[0042] In one embodiment, each consortium node in the consortium blockchain network can receive update notifications of defense rules and determine the latest defense rules based on the update notifications. The defense rules include a mapping relationship between risk types and execution strategies, and the update notification is a broadcast triggered by the smart contract when the defense rules are adjusted.

[0043] In one embodiment, at least one of the following is read from the blockchain: all new adversarial sample on-chain events recorded on the blockchain, defense logs, and node interaction information. For example, the system can call the underlying data interface of the blockchain through the smart contract preset interface to filter and read target information as needed. It can also support filtering by time range, node ID, and attack type.

[0044] In one embodiment, a graphical interface may also be provided, which displays at least one of the following: defense status, sample source, and attack type distribution. Through the graphical visualization of trusted blockchain data, real-time monitoring of the entire adversarial defense process, intuitive analysis of attack trends, and transparency of node collaboration status are achieved, reducing the operational threshold for maintenance personnel and providing efficient data support for defense strategy optimization and attack tracing analysis.

[0045] Please refer to Figure 2, which is a flowchart illustrating another embodiment of the adversarial example defense method provided in this application. It should be noted that if substantially the same result is achieved, this embodiment is not limited to the flow order shown in Figure 2. As shown in Figure 2, this embodiment includes: S201: Adversarial example detection and feature extraction.

[0046] Several suspected adversarial samples are collected from model training logs, inference monitoring systems, and manual reports. For each suspected adversarial sample, a corresponding sample feature vector is generated by embedding a feature extraction network such as ResNet, CLIP, or a feature autoencoder. For example, for adversarial samples with diverse data types, a highly adaptable deep learning feature extraction network is embedded. For instance, image samples are adapted to ResNet series networks to extract deep visual structural features, cross-modal samples are adapted to CLIP models to extract unified semantic features, and high-dimensional complex samples are adapted to feature autoencoders to extract compressed core representations. The samples are then subjected to layer-by-layer feature distillation through the network's convolutional layers, Transformer layers, or encoding layers to filter redundant noise information and retain the core discriminative features between adversarial and normal samples. The final output is a dense feature vector with fixed dimensions, low redundancy, and high discriminative power.

[0047] Based on the generated sample feature vectors, a targeted hash algorithm is selected to generate a unique feature fingerprint according to the sample matching requirements. This feature fingerprint serves as the core index identifier of adversarial examples on the blockchain, possessing uniqueness and a low collision rate. It can directly support parallel matching operations of distributed nodes on the chain, improving retrieval efficiency. For example, for scenarios requiring precise semantic matching, the Perceptual Hash algorithm is used to generate a fixed-length fingerprint with semantic similarity mapping by reducing and discretizing the feature vectors. For scenarios requiring fast retrieval of high-dimensional feature vectors, the Locality Sensitive Hash (LSH) algorithm is used to map high-dimensional features to low-dimensional hash buckets through a hash function, ensuring that similar feature vectors fall into the same hash bucket, significantly reducing the computational complexity of on-chain sample matching.

[0048] To ensure the privacy of the original data and the security of feature information in adversarial examples, targeted encryption strategies are employed for the generated feature data. For example, non-sensitive feature data is processed using salted hashing. A random and unique salt value is generated and bound to the feature data before hashing, ensuring the immutability of the feature data while enhancing its resistance to brute-force attacks and preventing features from being reverse-engineered. For core feature data containing sensitive information, zero-knowledge encryption technology is used. This ensures that the validity and integrity of the feature data can be verified by blockchain nodes without disclosing the original feature content or sample semantic information. This satisfies the feature sharing needs during cross-institutional collaboration while strictly protecting the core data privacy of the sample submitter.

[0049] S202: On-chain storage of adversarial example features.

[0050] After the adversarial sample completes feature extraction, fingerprint generation, and privacy protection processing, the extracted core metadata (including the extracted sample hash, source node ID, attack type, detection confidence, timestamp, etc.) is standardized and packaged according to the pre-set transaction data structure specifications of the consortium blockchain to form blockchain transaction data with immutable and traceable characteristics.

[0051] Once the transaction data is encapsulated, it is permitted to circulate on the consortium blockchain network after being digitally signed and authenticated by the initiating node. The consortium blockchain consensus nodes, based on a pre-defined consensus mechanism (PBFT using Byzantine fault tolerance, Raft, etc.), perform multi-dimensional verification and consensus confirmation of the received transaction data. After successful verification, the data is written into the blockchain ledger, ensuring the legality of the transaction and its consistency across the entire network.

[0052] Considering that adversarial sample raw data (such as images and audio) is typically large in size and may contain sensitive information (such as patient information in medical images or transaction data in financial texts), a tiered storage architecture can be adopted, storing metadata on-chain and raw data off-chain, balancing storage efficiency and data privacy. For example, non-extremely sensitive sample raw data or encrypted copies are preferentially stored in the IPFS interplanetary file system. Sample data containing core trade secrets or privacy information (such as the original features of self-developed adversarial samples) are stored in a secure database managed by the consortium blockchain (such as the encrypted database SealVault or a distributed trusted database).

[0053] It is important to note that regardless of the storage method used, the decryption and access of sample data must meet preset conditions. That is, only authorized nodes (such as defense system nodes approved by the consortium blockchain consensus) can initiate decryption requests. The requests must carry the node signature and the reason for access. Only after the smart contract verifies and passes the verification can the decryption key be obtained. The entire decryption operation is performed in an isolated virtual environment to avoid the key and the original data being exposed in the main system, ensuring that "access is possible when necessary and fully encrypted when not necessary".

[0054] S203: On-chain sharing and smart verification.

[0055] Once consensus is reached, transaction data is packaged into new blocks by the accounting nodes in timestamp order. Each block contains the transaction set, the hash of the previous block, and the block hash. After the new block is generated, it is encrypted (e.g., using Merkle tree hash verification) and written to the local ledger, then synchronized to the distributed ledger of all nodes on the consortium blockchain, ensuring that transactions are verifiable across the entire network once they are on-chain. At this point, all consortium nodes (such as model service providers, security research institutions, and regulatory nodes) can leverage the distributed nature of the blockchain to access the on-chain sample library in real time, obtaining sample metadata and feature fingerprints, with the entire access process being traceable and auditable.

[0056] It is important to note that the core of the legality verification before the sample is uploaded to the blockchain is carried out by a smart contract. The smart contract performs two verifications: first, the validity of the upload node's signature, confirming that the node is an authorized entity of the consortium; second, the uniqueness of the sample's feature hash, which is compared with existing records on the chain to prevent duplicate samples from being uploaded or maliciously forged samples from being mixed in. After the smart contract verification is passed, the sample's feature information is automatically included in the "shared defense blacklist" available across the entire network, serving as a unified and trusted basis for all consortium nodes to carry out adversarial defense. Once the block content is written, it has the characteristic of being immutable.

[0057] S204: On-chain detection during the model inference phase.

[0058] When the target model initiates the inference process, the input samples must first undergo preprocessing by a defense middleware deployed at the model's front end, achieving a standardized transformation of "feature extraction - hash calculation". For example, the defense middleware calls a preset feature extraction algorithm to extract distinctive core features from the input samples; then, using a hash algorithm consistent with the on-chain sample library, it performs one-way encrypted calculation on the extracted core features to generate a unique corresponding sample feature hash value.

[0059] After receiving a matching request, the smart contract calls the adversarial sample library data stored on the blockchain to perform feature hash comparison and risk feature matching. If the detection result shows that the sample hash completely matches the blacklist, or the similarity between the core feature and the known risk feature exceeds a preset threshold (e.g., 85%), the sample is determined to have an attack risk, triggering the isolation processing mechanism. The isolation processing mechanism can be set as follows: the defense middleware directly imports the input sample into a preset isolated virtual environment, prohibiting it from accessing the target model's main system, and simultaneously generates risk warning information, including the sample hash, the type of matched risk feature, the detection time, etc., which is then encapsulated as transaction data and synchronously uploaded to the blockchain to provide real-time risk alerts to alliance nodes.

[0060] S205: Risk Sample Input Isolation Inference and Defense Triggering.

[0061] The system automatically determines the defense method based on the defense strategy set in the smart contract. In one embodiment, the attack category corresponding to the sample is determined by hash matching (such as L∞ norm perturbation, adversarial patch), and an initial risk level (high / medium / low) is generated by combining the detection confidence. Then, the node reputation algorithm is called to obtain the reputation score of the node from which the sample originates, and the initial risk level is corrected. Finally, the system automatically determines the defense method based on the defense strategy combined with the risk level, node reputation, and attack category preset by the smart contract, and synchronizes the defense result of the input sample to the blockchain to form new transaction data, so as to achieve closed-loop traceability of events. For example, for input samples with low risk, a warning can be issued and a log can be recorded; for input samples with medium risk, inference can be performed in a sandbox environment and the result can be isolated; for input samples with high risk, execution can be directly rejected and a security event can be reported.

[0062] S206: Automatic updates and multi-node synchronization of smart contracts.

[0063] When new adversarial samples are added to the blockchain or defense rules are adjusted, the smart contract automatically triggers an "update broadcast." Upon receiving the notification, each node automatically pulls the latest adversarial sample feature data through the defense API and updates its local defense database. Furthermore, if a new node joins the network, it can directly obtain complete historical adversarial sample data through the block synchronization mechanism and further update its local database, thereby ensuring system consistency.

[0064] Please refer to Figure 3, which is a schematic diagram of the adversarial sample defense system architecture provided in this application, including a sample detection and feature extraction module, a blockchain node and ledger storage module, a smart contract management module, a model inference defense module, a consortium node collaboration module, and an audit and visualization management module.

[0065] The sample detection and feature extraction module is used to acquire input feature data of the input sample, and to retrieve adversarial feature data of each adversarial sample from the blockchain, or from a local adversarial sample library. The local adversarial sample library stores the adversarial feature data of each adversarial sample stored on the blockchain. In one embodiment, the sample detection and feature extraction module is connected to the model inference module and the on-chain interface module, and is used to extract input sample features and generate hash fingerprints, possessing functions such as batch processing, deduplication, and feature compression.

[0066] The blockchain node and ledger storage module is used to respond to the execution strategy instructing the target model to perform defensive processing using input samples, and to store the defensive processing results of the input samples in the blockchain; it also collects several adversarial samples to be uploaded to the blockchain, obtains the adversarial feature data of each adversarial sample to be uploaded to the blockchain, encapsulates several metadata of each adversarial sample to be uploaded to the blockchain as transaction data of the adversarial sample to be uploaded to the blockchain, the several metadata including adversarial feature data, signs the transaction data of each adversarial sample to be uploaded to the blockchain, and stores it in the on-chain sample library of the blockchain. In one embodiment, the blockchain node and ledger storage module has functions such as execution sample uploading and notarization, consensus verification, and distributed synchronization, as well as storing sample hashes, metadata, node signatures, and timestamps, and can use a Merkle tree structure to achieve tamper-proof and fast verification.

[0067] The smart contract management module is used to trigger broadcasts when new adversarial samples are added to the blockchain, and when defense rules are adjusted. The module automatically performs sample verification, defense triggering, and node updates, and supports dynamic rule configuration and periodic cleanup, achieving network-wide synchronous broadcasting through an event mechanism.

[0068] The model inference defense module matches the input samples of the target model with various adversarial samples stored on the blockchain to obtain sample matching results. Based on the sample matching results, it determines the risk type of the input sample and, based on the risk type, determines the execution strategy for the target model regarding the input sample. The execution strategy includes instructing the target model to perform normal inference or defensive processing using the input sample. Before model inference, the model inference defense module calls the on-chain adversarial sample library in real time to trigger sandbox isolation or refuse inference for suspicious input samples, thus forming a closed-loop defense logic with the feature extraction module.

[0069] The consortium node collaboration module receives notifications of new adversarial samples being added to the blockchain, pulls the latest adversarial sample data from the blockchain through the defense interface, and updates it to the local adversarial sample library. The local adversarial sample library stores data for each adversarial sample stored on the blockchain. Notifications of new adversarial samples being added to the blockchain are broadcast by a smart contract triggered when a new adversarial sample is added. The module also receives notifications of updated defense rules and determines the latest defense rules based on these notifications. These defense rules include a mapping between risk types and execution strategies. Update notifications are broadcast by a smart contract triggered when defense rules are adjusted. The consortium node collaboration module breaks down information silos between model service providers, security agencies, and other entities. Through standardized communication, strict identity verification, and on-chain synchronization mechanisms, it achieves secure sharing and real-time consistency of defense information, supporting the co-construction of shared defense blacklists, joint defense, and attack trend analysis. The module can employ secure communication protocols such as gRPC + TLS, which solves the problems of high latency and poor security in traditional cross-institutional communication, ensuring the confidentiality and integrity of transmission. In addition, the module can verify node identity through on-chain registration and digital certificate dual verification to prevent malicious access and provide support for multiple institutions to jointly build a highly reliable defense system.

[0070] The audit and visualization management module is used to record all on-chain events, defense logs, and node interaction information, and can provide a graphical interface to display information such as defense status, sample source, and attack type distribution.

[0071] Please refer to Figure 4, which is a schematic diagram of the system operation data flow provided in this application. After the input sample enters the system, two parallel processes are triggered simultaneously. The input sample first enters the feature extraction module, where core features are extracted and a unique hash is generated. Simultaneously, the input sample is passed to the inference defense module, awaiting subsequent defense instructions. The generated sample hash initiates an on-chain query, requesting a match from the blockchain ledger. The blockchain ledger feeds back the matching result (which may include risk level, attack type, etc.) to the smart contract. The smart contract, based on a preset strategy, issues control instructions to the inference defense module, such as executing isolated inference or normal inference. The inference defense module performs corresponding operations on the input sample according to the instructions of the smart contract. If the sample is of medium to high risk, isolated inference can be initiated; if it is of low risk or no risk, normal inference can be performed. After inference is completed, the defense results, such as the inference status and processing conclusion, are stored on the blockchain ledger. The stored results are synchronized to all consortium nodes, achieving cross-institutional defense information sharing and completing the entire data flow loop.

[0072] Please refer to Figure 5, which is a schematic diagram of the structure of an embodiment of the electronic device provided in this application. The electronic device 60 includes a memory 61 and a processor 62 that are interconnected. The memory 61 is used to store computer programs. When the computer programs are executed by the processor 62, they are used to implement the adversarial sample defense method in the above embodiment.

[0073] The methods described in the above embodiments can exist in the form of computer programs. Therefore, this application proposes a computer-readable storage medium. Please refer to FIG6, which is a schematic diagram of the structure of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 80 is used to store a computer program 81, which can be executed to implement the adversarial sample defense method in the above embodiments.

[0074] The computer-readable storage medium 80 can be any medium capable of storing program code, such as a server, USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0075] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0076] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for adversarial sample defense, characterized in that, The method, applied to blockchain nodes, includes: matching input samples of a target model with adversarial samples stored on the blockchain to obtain sample matching results; determining the risk type of the input sample based on the sample matching results; and determining an execution strategy for the target model regarding the input sample based on the risk type of the input sample, wherein the execution strategy includes instructing the target model to use the input sample for normal reasoning or defensive processing.

2. The method according to claim 1, characterized in that, The step of determining the risk type of the input sample based on the sample matching result includes: determining the risk type of the input sample as risk-free in response to the sample matching result indicating that there is no adversarial sample matching the input sample; determining the risk type of the input sample as risky in response to the sample matching result indicating that there is an adversarial sample matching the input sample; the step of determining the execution strategy for the target model regarding the input sample based on the risk type of the input sample includes at least one of the following steps: determining the execution strategy as normal inference using the input sample in response to the risk type of the input sample being risk-free; determining the execution strategy as a defense strategy matching the risk level of the input sample in response to the risk type of the input sample being risky, wherein the defense strategy represents the use of the input sample for defense processing, and the higher the risk level, the higher the defense level of the defense strategy.

3. The method according to claim 2, characterized in that, Determining the execution strategy as a defense strategy matching the risk level of the input sample includes at least one of the following steps: In response to the input sample having a low risk level, determining the execution strategy as a first defense strategy, the first defense strategy including performing inference and performing at least one of the following processes: issuing a prompt and logging; In response to the input sample having a medium risk level, determining the execution strategy as a second defense strategy, the second defense strategy including performing inference in an isolated virtual environment and isolating the inference results; In response to the input sample having a high risk level, determining the execution strategy as a third defense strategy, the third defense strategy including refusing to perform inference and reporting a security incident.

4. The method according to claim 1, characterized in that, After determining the execution strategy for the target model based on the risk type of the input sample, the method further includes: in response to the execution strategy instructing the target model to perform defensive processing using the input sample, storing the defensive processing result of the input sample in the blockchain; and / or, the method further includes at least one of the following steps: receiving a notification of a new adversarial sample being added to the blockchain, pulling the latest adversarial sample data from the blockchain through a defense interface, and updating it in a local adversarial sample library, wherein the local adversarial sample library is used to store the data of each adversarial sample stored in the blockchain, and the notification of a new adversarial sample being added to the blockchain is a broadcast triggered by a smart contract when a new adversarial sample is added to the blockchain; receiving a notification of an update to the defense rules, and determining the latest defense rules according to the update notification, wherein the defense rules include the mapping relationship between the risk type and the execution strategy, and the update notification is a broadcast triggered by a smart contract when the defense rules are adjusted.

5. The method according to claim 1, characterized in that, Before matching the input sample of the target model with each adversarial sample stored on the blockchain to obtain the sample matching result, the method further includes: obtaining the input feature data of the input sample; and calling the adversarial feature data of each adversarial sample on the blockchain, or obtaining the adversarial feature data of each adversarial sample from a local adversarial sample library, wherein the local adversarial sample library is used to store the adversarial feature data of each adversarial sample stored on the blockchain; the step of matching the input sample of the target model with each adversarial sample stored on the blockchain to obtain the sample matching result includes: comparing the input feature data with each adversarial feature data respectively to obtain the sample matching result, wherein if there is a match between the adversarial feature data and the input feature data, it is determined that the input sample matches the corresponding adversarial sample.

6. The method according to claim 5, characterized in that, The input feature data is the input sample feature extracted from the input sample, and the adversarial feature data is the adversarial sample feature extracted from the adversarial sample; or, the input feature data is the input feature fingerprint obtained by hashing the input sample feature, and the adversarial feature data is the adversarial feature fingerprint obtained by hashing the adversarial sample feature. And / or, obtaining sample matching results by matching the input feature data with each adversarial feature data includes: for each adversarial feature data, calculating the similarity or hash distance between the input feature data and the adversarial feature data; and determining that the input feature data matches the adversarial feature data in response to the similarity or hash distance being greater than a preset threshold.

7. The method according to claim 1, characterized in that, The method further includes: collecting a number of adversarial samples to be uploaded to the blockchain; obtaining adversarial feature data for each of the adversarial samples to be uploaded to the blockchain; encapsulating several pieces of information of each of the adversarial samples to be uploaded to the blockchain into transaction data of the adversarial samples to be uploaded to the blockchain, wherein the several pieces of information include the adversarial feature data; and signing the transaction data of each of the adversarial samples to be uploaded to the blockchain and storing it in the on-chain sample library of the blockchain.

8. The method according to claim 7, characterized in that, The metadata also includes at least one of the following: the source node identifier, attack type, detection confidence level, and timestamp of the adversarial sample to be uploaded to the chain; And / or, before encapsulating the metadata of the adversarial sample into the transaction data of the adversarial sample, the method further includes: encrypting the adversarial feature data of the adversarial sample to be uploaded to the blockchain, wherein the encryption method includes salted hashing or zero-knowledge encryption; and / or, the step of signing the transaction data of each of the adversarial samples to be uploaded to the blockchain and storing it in the on-chain sample library of the blockchain includes: using a smart contract to verify the legality of the transaction data of the adversarial sample to be uploaded to the blockchain, wherein the legality verification includes verifying the uniqueness between the digital signature of the uploading node and the adversarial feature fingerprint in the transaction data of the adversarial sample to be uploaded to the blockchain; after the legality verification is passed, storing the signed transaction data of the adversarial sample to be uploaded to the blockchain in the on-chain sample library of the blockchain; and / or, the method further includes: storing the sample data corresponding to the adversarial sample to be uploaded to the blockchain in a secure database, wherein the sample data includes at least one of the following: the adversarial sample to be uploaded to the blockchain, the encrypted adversarial sample to be uploaded to the blockchain, the adversarial feature data of the adversarial sample to be uploaded to the blockchain, and the encrypted adversarial feature data.

9. The method according to claim 1, characterized in that, The method further includes at least one of the following steps: reading at least one of the following from the blockchain: all new adversarial sample on-chain events, defense logs, and node interaction information recorded by the blockchain; and providing a graphical interface to display at least one of the following: defense status, sample source, and attack type distribution.

10. An adversarial sample defense system, characterized in that, The system includes: a model inference defense module, used to match the input sample of the target model with each adversarial sample stored on the blockchain to obtain a sample matching result, determine the risk type of the input sample based on the sample matching result, and determine an execution strategy for the target model based on the risk type of the input sample, wherein the execution strategy includes instructing the target model to use the input sample for normal inference or defense processing.

11. The system according to claim 10, characterized in that, The system further includes at least one of the following modules: a sample detection and feature extraction module, used to acquire input feature data of the input sample, and to call adversarial feature data of each adversarial sample on the blockchain, or to acquire adversarial feature data of each adversarial sample from a local adversarial sample library, wherein the local adversarial sample library is used to store the adversarial feature data of each adversarial sample stored on the blockchain; a blockchain node and ledger storage module, used to respond to the execution strategy instructing the target model to perform defense processing using the input sample, and to store the defense processing result of the input sample in the blockchain; and to collect several adversarial samples to be uploaded to the blockchain, acquire the adversarial feature data of each adversarial sample to be uploaded to the blockchain respectively, encapsulate several metadata of each adversarial sample to be uploaded to the blockchain as transaction data of the adversarial sample to be uploaded to the blockchain respectively, wherein the several metadata includes the adversarial feature data, and sign the transaction data of each adversarial sample to be uploaded to the blockchain's on-chain sample library; a smart contract management module, used for new adversarial samples... The system includes a broadcast mechanism that triggers broadcasts when adversarial samples are added to the blockchain and when defense rules are adjusted. A consortium node collaboration module receives notifications of new adversarial sample additions to the blockchain, pulls the latest adversarial sample data from the blockchain via a defense interface, and updates it to a local adversarial sample library. This local library stores data for each adversarial sample stored on the blockchain. The notification of a new adversarial sample being added to the blockchain is a broadcast triggered by the smart contract when such a sample is added. The system also receives notifications of updated defense rules and determines the latest defense rules based on these notifications. These rules include a mapping between the risk type and the execution strategy. The update notification is a broadcast triggered by the smart contract when defense rules are adjusted. An audit and visualization management module reads at least one of the following from the blockchain: all new adversarial sample addition events recorded on the blockchain, defense logs, and node interaction information. It also provides a graphical interface to display at least one of the following: defense status, sample source, and attack type distribution.

12. An electronic device, characterized in that, The electronic device includes a processor and a memory, the processor being coupled to the memory, the processor being configured to execute one or more steps of the adversarial sample defense method according to any one of claims 1 to 9 based on instructions stored in the memory.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the steps of the adversarial sample defense method as described in any one of claims 1 to 9.