Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

193 results about "Backdoor" patented technology

A backdoor is a typically covert method of bypassing normal authentication or encryption in a computer, product, embedded device (e.g. a home router), or its embodiment (e.g. part of a cryptosystem, algorithm, chipset, or even a "homunculus computer" —a tiny computer-within-a-computer such as that found in Intel's AMT technology). Backdoors are most often used for securing remote access to a computer, or obtaining access to plaintext in cryptographic systems. From there it may be used to gain access to privileged information like passwords, corrupt or delete data on hard drives, or transfer information within autoschediastic networks.

Federal learning backdoor attack defense method based on layer perception detection

The invention discloses a federated learning backdoor attack defense method based on layer perception detection, which relates to the technical field of artificial intelligence security, and comprises the steps of client training and uploading, malicious client identification, backdoor key layer detection and deletion, robust aggregation updating and adversarial training for enhancing robustness. According to the method, the model parameters uploaded by the clients are subjected to clustering analysis, and the maximum benign cluster is identified by adopting an unsupervised clustering method, so that instability caused by single threshold judgment is avoided, and the benign client and the malicious client can be distinguished; and meanwhile, layer perception detection and elimination are used in the scheme, so that the backdoor introduced by a malicious client can be effectively identified and eliminated on the premise of ensuring the global model precision, and the influence of backdoor attack on the model is inhibited.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Black box code search model backdoor attack method based on learnable discrete code transformation

A black box code search model backdoor attack method based on learnable discrete code transformation comprises the steps that a learnable backdoor generator is constructed, and malicious codes with backdoors are generated on the premise that internal parameters of a damaged model are not accessed through the learnable discrete code transformation. Firstly, an agent model capable of simulating victim model behaviors is trained through query-response data; secondly, on the proxy model, a discrete selection process is differentiable by using a re-parameterization sampling technology, and a backdoor generator is trained in combination with a multi-objective loss function so as to realize effectiveness and concealment of backdoor implantation; according to the method, the limitation of a black box is successfully bypassed by constructing the proxy model, and a new possibility is provided for an attacker. A backdoor generation process is integrated into a differentiable training framework through learnable discrete transformation, so that end-to-end learnability is realized, and an attack strategy can be automatically optimized.
Owner:NANJING UNIV OF POSTS & TELECOMM

BDDR backdoor detection and data restoration method and system oriented to large model

The invention discloses a BDDR backdoor detection and data recovery method and system oriented to a large model, and belongs to the field of backdoor defense. Comprising the following steps: constructing a knowledge distillation architecture under federal learning, including an edge server and a plurality of clients, and obtaining distillation data; inputting the distillation data into a randomly initialized model for training, recording the loss change of each batch of data, and screening out abnormal batches to form a backdoor data set; the edge server initializes two independent models, respectively uses a distillation data set and a backdoor data set for training, and guides learning of backdoor features; using probability distribution to calculate and correct a backdoor label, generating a clean data set by adding noise, finely adjusting a large model, detecting residual backdoor feature intensity, and adjusting probability distribution calculation parameters to further weaken backdoor features according to the residual backdoor feature intensity so as to obtain a final repaired data set; while the generalization ability of the large model is improved, backdoor attacks can be effectively identified and defended, the data privacy of the client is protected, and the model security is ensured.
Owner:NANJING UNIV OF POSTS & TELECOMM

Privacy protection and robustness test method and system for large model fine tuning

The invention discloses a privacy protection and robustness test method and system for large model fine tuning, and belongs to the technical field of machine learning security. The method comprises the steps that a three-layer distributed architecture comprising an edge server, a cloud server and a plurality of edge clients is constructed, the edge clients distill local privacy data and cooperate with the edge server to train a global model, and a candidate detection sample set is formed; screening a sample set based on the potential feature deviation evaluation index, and sending the sample set to a cloud server for vulnerability detection to obtain an optimal backdoor detection candidate sample set; and multi-trigger parallel and progressive trigger sequence backdoor implantation is respectively used for scenes of single fine tuning and multiple fine tuning of the large model, an optimal backdoor detection candidate sample set is combined with a preset trigger to generate a backdoor test sample set, the backdoor test sample set is mixed with a clean data set, and then the robustness of the backdoor test sample set is tested through fine tuning of the large model. Large model fine tuning and robustness testing of privacy protection can be realized in a heterogeneous model cooperative training environment.
Owner:NANJING UNIV OF POSTS & TELECOMM

Structure-preserving heterogeneous graph backdoor attack method

The invention discloses a structure-preserving heterogeneous graph backdoor attack method. The method comprises the following steps: preprocessing a clean graph; calculating a composite score according to the multiple scores of each node, and selecting poisoning nodes according to the coincidence scores to obtain a poisoning map; meanwhile, performing global importance evaluation on a feature dimension, transmitting features which can be used as a trigger to a trigger generator, and generating the features by the trigger to disturb the clean graph; the poisoning map is subjected to quasi-classification through an agent model, a trigger generator is optimized according to parameters, double-layer optimization is formed according to the parameters of the trigger generator and a new agent model, and finally hidden and efficient backdoor attacks are completed; wherein the multiple scores comprise one or more of an uncertainty score, a relation weight representative score, a meta-path score and a detectability score. According to the method, through fine node / feature selection and heterogeneous perception trigger generation, the target label attack success rate is remarkably improved.
Owner:CHENGDU UNIV OF INFORMATION TECH

Backdoor defense method and system based on adaptive feature blocking

PendingCN121690686ABiological modelsSecuring communicationDynamic channelAlgorithm
The invention discloses a back door defense method and system based on adaptive feature blocking, and the method comprises the steps: 1, constructing a blocking module, and embedding the blocking module into an image classification model; the blocking module comprises a self-adaptive instance statistical calibration layer and a dynamic channel suppression layer, the self-adaptive instance statistical calibration layer is used for eliminating cross-sample statistical offset caused by backdoor attack, and the dynamic channel suppression layer is used for suppressing abnormal activation of a polluted channel; and 2, training the embedded classification network to obtain a final classification network, and completing backdoor attack defense according to the final classification network. Through lightweight modular design and a fine adjustment mechanism, the bottleneck problem of a traditional model defense method in the aspects of calculation cost and flexibility is effectively broken through.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

A cross-modal transferable backdoor attack method and device

The present application relates to the technical field of computer vision and natural language processing, and provides a cross-modal transferable backdoor attack method and device, comprising: constructing a backdoor dataset containing a trigger based on an unlabeled original dataset; inputting the original dataset and the backdoor dataset containing the trigger into a pre-constructed hacker network, and respectively calculating clean data silence loss and backdoor toxicity loss with the trigger; pre-training the pre-constructed hacker network by minimizing the clean data silence loss and the backdoor toxicity loss; calculating feature representation of a target category through image data and text data of the target category, generating a target CLIP model with a backdoor by using the pre-trained hacker network and the feature representation of the target category; and attacking by using the generated model. The backdoor trigger can be embedded in multi-modal data at the same time, has strong cross-modal transferability and concealment, and does not depend on large-scale labeled data.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

A backdoor attack method based on color frequency injection and adaptive local enhancement

PendingCN122365494AEngineeringSelf adaptive
The application discloses a backdoor attack method based on color frequency injection and adaptive local enhancement, and relates to the technical field of machine learning and artificial intelligence security. The method comprises the following steps: introducing low-frequency color offset and weak high-frequency signal into an image in a CIELAB color space to perform global color-frequency injection; using a pre-trained proxy model to locate a high-sensitive perception domain of the model through mixed evaluation of gradients and class activation maps, and generating a binary mask; in an HSV color space, respectively applying nonlinear stretching factors to saturation and brightness of the sensitive domain based on the mask to perform adaptive local enhancement; using Gaussian smoothing, adaptive noise and histogram matching to eliminate edges and statistical abnormalities caused by local enhancement, completing compensation color enhancement to generate a poisoned image; and modeling a trigger core parameter as a constrained optimization problem, and using a particle swarm optimization algorithm to jointly dynamically update the trigger core parameter to obtain an optimal strategy. The application anchors the trigger feature depth in the core semantic area of the model and lurks in the normal data manifold, guarantees a high attack success rate, realizes extreme visual and feature concealment, and has strong anti-defense robustness.
Owner:NORTH CHINA UNIVERSITY OF TECHNOLOGY

A training method for physical light backdoor attacks facing artificial intelligence security

The application belongs to the field of artificial intelligence security, and discloses a training method for physical light backdoor attack for artificial intelligence security, comprising the following steps: performing light backdoor attack on a target object, generating corresponding light triggers on the target object according to light colors, and generating backdoor image data based on the light triggers; obtaining clean image data, and respectively constructing training sets based on the backdoor image data and the clean image data; the clean image data is original image data without light triggers; constructing a backdoor model, the backdoor model is a deep learning model, training the backdoor model based on the training sets to obtain a trained backdoor model; constructing a test set, evaluating the trained backdoor model based on the test set to obtain attack success rate data and clean accuracy rate data of the light backdoor attack. The technical scheme disclosed by the application realizes more covert physical backdoor attack while having a higher attack success rate.
Owner:ZHEJIANG GONGSHANG UNIVERSITY +2

A method for monitoring and controlling the illegal use of mobile phones

ActiveCN115733955BInternet privacyEngineering
This invention relates to a domestically produced method for monitoring the unauthorized use of mobile phones, belonging to the field of information security protection technology. The system of this invention includes network cameras, a domestically produced intelligent server, access control, client terminals, a video wall, and a switch. The network cameras and access control are deployed in the monitoring location; the switch and the domestically produced intelligent server are deployed in the server room; the video wall is deployed in the monitoring room; and the client terminals are deployed in the office. The method of this invention includes: video acquisition, video analysis, display of analysis results, and access control linkage control. This invention uses domestically produced CPUs and domestically produced accelerator cards, making it more secure and eliminating the risk of sanctions or backdoors; it can automatically control access, promptly preventing unauthorized mobile phone users from escaping the scene, thus buying valuable time to prevent information leaks.
Owner:BEIJING INST OF COMP TECH & APPL +1

A Knowledge Graph-Based Intelligent Method and System for Detecting Software Backdoors

This invention relates to the field of software detection technology, specifically disclosing a knowledge graph-based intelligent detection method and system for software backdoors. The method involves intercepting system kernel events and reading the object handle table to generate a time-series interaction log. Based on this log, a time-series knowledge graph is constructed, forming an interconnected network composed of process nodes, token nodes, and relationships such as handle holding, token replication, and memory writes. Newly added process nodes with process creation timestamps are further extracted as target process nodes, and the nominal parent process node is located based on its parent process identifier. Subsequently, a lineage consistency check is performed around the target process node, and cross-chain constraint analysis is conducted using the nominal parent process node, anonymous process nodes, token nodes, and memory write relationships. When contradictions arise in handle permissions, token inheritance, and write timing, and the lineage forgery index exceeds a preset threshold, the target process node is determined to be a parent process deceiving a backdoor process, and a threat interception command is output.
Owner:SHENZHEN HAIYUNAN NETWORK SECURITY TECH CO LTD

Feature space backdoor attack method, system and device based on fusion data transaction model parameter constraint and storage medium

The invention relates to the technical field of computer security, in particular to a feature space backdoor attack method, system and device based on fusion data transaction model parameter constraint and a storage medium. Obtaining an original training data set, constructing a data transaction model, calculating malicious sample loss values, and screening samples with smaller loss values as target samples; poisoning features are randomly extracted from the original data set, discrete noise is generated through a noise generator to form a trigger, the trigger is added to a target sample, a label is turned over, and an attack data set is formed; original training data are sampled, and a toll-snow information matrix of each layer is calculated; using the original model parameters to initialize the backdoor model, adding the L1 norm of the parameter disturbance difference value and the parameter difference constraint term based on the Feisnow information matrix into a loss function, and training the backdoor model; performance fluctuation caused by random sample selection is avoided through sample screening based on loss values, so that the backdoor model maintains the classification capability of clean samples while learning trigger features.
Owner:GUANGXI POWER GRID CORP

Firmware backdoor detection and one-key security restoration method and system

The invention discloses a firmware backdoor detection and one-key security restoration method and system, and relates to the technical field of software detection. The method comprises the steps of obtaining an original byte sequence of a firmware upgrade package; performing segmented structure analysis and standardization on an original byte sequence of the firmware upgrade package to obtain a byte interval index of each segment and an upgrade package segment list; extracting ciphertext statistical characteristics of each segment based on the byte interval index and generating an upgrade package statistical fingerprint; comparing the statistical fingerprint of the upgrade package with a historical baseline to obtain a ciphertext anomaly score; when the ciphertext exception score exceeds a verification trigger threshold value, generating an equipment end verification task and collecting equipment end verification evidence; performing fusion judgment on the basis of the ciphertext anomaly score and the equipment end verification evidence, and outputting a comprehensive risk level and a recommended disposal action; according to the method, the risk detection can be realized without depending on a decryption key or a differential synthetic material, so that the security of firmware upgrading is ensured, and the operation and maintenance efficiency is also considered.
Owner:SUZHOU WENXIN INTELLIGENT TECH CO LTD

A method and system for multi-backdoor pollution attack oriented to wireless brain-computer interface

The application discloses a kind of multi backdoor pollution attack methods and systems for wireless brain-computer interface, belong to brain-computer interface security and machine learning confrontation attack technical field, wherein method includes training phase and reasoning phase, training phase uses multiple different trigger modes, constructs pollution sample, uses the training set of being polluted to train electroencephalogram decoding model, obtains trained electroencephalogram decoding model;Reasoning phase, implement man-in-the-middle attack to Bluetooth communication link, intercept the data packet sent from electroencephalogram acquisition equipment to wireless brain-computer interface system, according to the target category of expected control, implement corresponding trigger mode to the intercepted data packet, obtain tampered data packet, use electroencephalogram decoding model to infer tampered data packet, output the target category of expected control.The application combines Bluetooth man-in-the-middle attack with training phase data pollution, so as to realize arbitrary control to model output by selectively injecting different trigger modes in reasoning phase.
Owner:HUAZHONG UNIV OF SCI & TECH

Signal backdoor attack method and system based on gradient clipping optimization

The application discloses a signal backdoor attack method and system based on gradient clipping optimization, comprising constructing a backdoor training set containing a directional attack label and a clean training set, adopting an alternating mixed training strategy, and synchronously processing clean samples and backdoor samples in each training batch. By calculating the clean sample loss and the backdoor sample loss, a weighted combined loss is constructed and the model parameters are updated, and at the same time, the gradient clipping technology is used to constrain the update amplitude of the backdoor trigger, so that the trigger disturbance is strictly controlled within a predetermined range. The final backdoor model can output the preset attack category to the input embedded trigger while maintaining normal classification performance, effectively counteracting the stolen model. The method is significantly superior to the prior art in terms of attack success rate, concealment and model performance retention, and provides a reliable solution for active defense against model theft.
Owner:XIDIAN UNIV

Backdoor attack detection and defense method and system based on personalized federal learning

The invention belongs to the technical field of machine learning, and particularly relates to a backdoor attack detection and defense method and system based on personalized federated learning, which realizes accurate identification and effective defense of a backdoor attack client through personalized training of model segmentation, density clustering grouping based on behavior characteristics and a reputation scoring mechanism of time smoothing. And meanwhile, the individuation performance and the aggregation efficiency of the federal learning model are ensured. Wherein the interference of model difference on the detection result is eliminated through the client grouping of density perception, and the accuracy of backdoor attack recognition is remarkably improved. By introducing a time smoothing mechanism, misjudgment caused by single-round model fluctuation can be avoided, the reputation score of the client can reflect long-term behavior characteristics of the client, and continuous and stable suppression of the malicious client is realized.
Owner:SHANDONG UNIV OF SCI & TECH

Back door embedded honeypot induced active defense method

The invention discloses a backdoor embedded honeypot induced active defense method, and relates to the field of data processing. And malicious aggressive input can be actively defended on the basis of not influencing the classification detection rate. The backdoor embedded honeypot induced active defense method comprises the following steps: acquiring a to-be-detected input signal, and inputting the input signal into a classification prediction model and an anomaly detection model to obtain a prediction category output by the classification prediction model and an anomaly score output by the anomaly detection model; determining whether the prediction category triggers a honeypot trap, and outputting the prediction category as an output target when the honeypot trap is not triggered; and when a honeypot trap is triggered, determining whether the input signal is a hostile attack signal through the abnormal score, and blocking output when the input signal is determined to be the hostile attack signal.
Owner:XIDIAN UNIV

Backdoor attack defense method and system of deep neural network model

The invention provides a backdoor attack defense method and system for a deep neural network model, and the method comprises the steps: obtaining a test data sample set matched with the deep neural network model through the classification of an original data set, dynamically adding a noise component to the test data sample set, and removing at least part of backdoor disturbance. Noise is mixed into the test data to change the original data structure of the rear door disturbance; adding features based on noise components of the test data sample set, performing denoising repair processing, testing the deep neural network model by using the test data sample set, restoring an original data structure of the test data through denoising repair, and eliminating back door disturbance mixed in the original data structure at the same time; judging whether the deep neural network model works normally or not based on the test result information, adjusting the noise component adding state of the test data sample set again in combination with the dynamic adding record of the previous noise component, eliminating the back door disturbance of the test data, and improving the test accuracy under the condition of maintaining the original performance of the deep neural network model. And the defense robustness of the backdoor attack is improved.
Owner:HUIZHIAN INFORMATION TECH CO LTD

Backdoor robustness evaluation method for non-iid federated learning model based on generative adversarial network

The application relates to a non-IID federated learning model backdoor robustness evaluation method based on a generative adversarial network, which comprises the following steps: selecting a client in federated learning as a test node, taking a downloaded server-side global model as a discriminator, and designing a generator locally to form a generative adversarial network model; a backdoor attack target category is specified, a class representative sample is reconstructed by using the generator in each round of global training, and the target category is marked as pre-poisoning data participating in training; a supplementary data set is generated locally; a source category of the backdoor attack is specified, a backdoor trigger is optimized, the supplementary data set is used for class-specific backdoor training, the discriminator is updated and uploaded to the server, and the global model is updated; the non-IID degree of data and the number of specified source categories are adjusted, and the robustness of the federated learning global model to the backdoor attack is observed. Compared with the prior art, the application can verify the effect of the backdoor attack on the federated learning model under different degrees of data heterogeneity.
Owner:SHANGHAI JIAOTONG UNIV

Method for reprogram with enhanced security

ActiveUS12699776B2PasswordCryptogram
A method performed by an electronic control unit (ECU) for reprogramming with enhanced security. The method includes checking whether a cyber security function of a ROM of the ECU is applied while the ECU is running in a NORMAL area, receiving, when it is confirmed that the cyber security function of the ROM of the ECU is applied, a first backdoor password for the cyber security function of the ROM of the ECU, and performing, when the received first backdoor password is the same as a second backdoor password included in program data stored in the ROM of the ECU, reprogramming of the ECU without additional procedures related to the cyber security function.
Owner:HYUNDAI AUTOEVER

Conditional backdoor attack method based on semantic and mode composite double triggers

The invention discloses a conditional backdoor attack method based on a semantic and mode composite double trigger, which comprises the following steps: taking a diversified reference set as an attack target, constructing a data set containing single and double trigger samples of two types of semantics and modes, and designing a composite loss function to carry out model fine tuning; according to the composite loss function, back door association is established through attack loss, the benign performance of the model is maintained through fidelity loss, a passive neutralization and active rejection dual decoupling mechanism is introduced, a single trigger sample is forced to be regarded as a normal sample by the model, the spatial distance between the characteristic of the model and an attack target is actively increased, and the robustness of the model is improved. And thus, a decision isolation region is constructed in the feature space. According to the method, strict double-trigger activation conditions are set, so that the implanted backdoor has extremely high concealment and anti-defense robustness, meanwhile, the high attack success rate is ensured, and normal functions of the model are not affected.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Detecting backdoors in binary software code

Systems, methods, and software can be used to detect backdoors in binary software code. In some aspects, a method comprises: obtaining, by a server, binary software code corresponding to source code; generating, by the server, a backdoor abstraction of the binary software code; and generating, by the server, a backdoor risk assessment based on the backdoor abstraction of the binary software code.
Owner:BLACKBERRY LTD

Backdoor attack method, system, storage medium and device based on large model

PendingCN122389987AAttack modelTheoretical computer science
The application provides a large model-based backdoor attack method, system, storage medium and equipment. A tokenizer and a known corpus of an attacked model are obtained, a set of candidate words is integrated by selecting adverbs in the known corpus; each adverb in the set of candidate words is converted into a word unit sequence composed of several basic word units; a tail word unit and a preset specific word unit are combined into a trigger combination, and the frequency of the trigger combination in the original training set is counted; when the frequency of the trigger combination is less than a first threshold value and the number of word units ending with the tail word unit is not less than a second threshold value, the corresponding adverb is added to a trigger substructure set; a data training set is constructed, the data training set includes a dirty training set, the input of the sample in the dirty training set contains the trigger combination, and the output of the sample in the dirty training set is an attack result; and the attacked model is trained by using the data training set. The problem of insufficient concealment of the trigger used for attack in the large model is solved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A large language model backdoor attack detection method

PendingCN122451895ALinguistic modelAlgorithm
The application discloses a large language model backdoor attack detection method, and belongs to the field of artificial intelligence security. In view of the problem that existing defense methods are difficult to be applied to generative large language models and have large calculation overhead, the application is based on the first discovered "sequence locking" phenomenon, that is, when generating an attack target, the backdoor model will output an abnormally high and consistent word element confidence, forming a deterministic path without branches. The method detects high-confidence sequences that continuously exceed the length threshold L and have a probability higher than the threshold P by monitoring the Top 1 probability of the model output sequence in real time, thereby determining the backdoor attack. The method only needs to access the Top 1 probability of the model in a black box, does not need complete Logits vectors or auxiliary models, almost does not introduce additional delay and performance overhead, is suitable for real-time interaction and large-scale deployment scenarios, and can efficiently detect various backdoor triggering modes and target types.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Large model quantification conditional backdoor defense method and system based on micro-rounding optimization

The invention belongs to the technical field of artificial intelligence security, and provides a large model quantification condition backdoor defense method and system based on micro-rounding optimization, and the method comprises the steps: carrying out the quantitative analysis of an obtained to-be-defended large language model weight; determining a normalized residual error of the weight in the quantization interval, constructing a differentiable rounding parameter, and mapping the differentiable rounding parameter to a specified interval through a continuous mapping function to obtain a differentiable soft quantization weight; constructing a joint optimization objective function, constructing a calibration data set, and on the calibration data set, according to the constructed joint optimization objective function, performing iterative optimization on the to-be-defended large language model to obtain an optimal differentiable rounding parameter; and determining a final quantization rounding strategy according to the optimized differentiable rounding parameters, and generating a security quantization model after defense. The problem that quantitative condition backdoor attacks are difficult to detect and eliminate is solved.
Owner:SHANDONG UNIV

Construction method and system of language model backdoor trigger input detection module

The invention discloses a method and system for constructing a language model backdoor trigger input detection module, and the method comprises the steps: selecting a predetermined layer with the largest difference between trigger input features and clean input features as a key layer, and taking the global representation of a hidden state as training data; adopting a triple network to separate the potential feature space of the training data; performing intermediate feature trend knowledge extraction on the training data by using a knowledge extraction mode of a convolutional neural network; and constructing a corresponding language model backdoor trigger input detection module based on the potential feature space separation result, and deploying the module in the system. According to the method, an efficient online detection mechanism is provided, so that quantitative analysis and real-time identification can be carried out on the backdoor trigger possibly carried by language model input, the risk caused by backdoor attack can be effectively relieved in an unknown attack domain, and real-time protection in a reasoning stage is supported.
Owner:UNIV OF SCI & TECH OF CHINA

A system and method for trigger word recognition and positioning based on BIO sequence labeling

The application discloses a trigger word recognition and positioning system and method based on BIO sequence labeling, comprising: using an encoder to extract a whole sentence semantic vector, judging whether the input text contains a backdoor attack feature; if the backdoor suspicion is detected, positioning the backdoor trigger word, judging whether each word belongs to the backdoor trigger word through the BIO sequence labeling of each token, and outputting a preliminary trigger word position marking sequence; applying trigger position prior knowledge and multi-strategy rules to correct and optimize the results and filter false positives; when the results have uncertainty or are suspected to have attack avoidance, using a predefined known trigger phrase mode library to scan the input text, capturing hidden or variant trigger word modes, and positioning the missed suspicious trigger words; and outputting the final result of backdoor detection. Through the dual-module cooperation on the architecture and the priori and rule fusion on the strategy, the application can robustly detect the text backdoor trigger and accurately position the trigger content.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Deep learning model-oriented back door behavior dynamic detection and positioning method

The invention discloses a back door behavior dynamic detection and positioning method for a deep learning model, and relates to the field of artificial intelligence security, and the method comprises the steps: S1, monitoring input data and an output prediction result of the deep learning model in a reasoning process in real time; s2, performing statistical feature analysis on the output prediction result based on a preset spatiotemporal behavior pattern library, and identifying abnormal deviation exceeding a normal behavior threshold; s3, in response to the abnormal judgment in the S2, positioning an input data area for triggering a backdoor behavior through reverse gradient tracking and attention weight structure analysis; and S4, deconstructing the deep learning model into an independent analysis unit according to a functional layer, and realizing hierarchical positioning of the backdoor behavior through layer-by-layer activation value comparison. The method has the advantages that real-time dynamic detection is achieved, the rear door triggering area is accurately positioned, the universality of a cross-model architecture is supported, and reasoning delay is reduced.
Owner:SHANGHAI SHIYUE COMPUTER TECH CO LTD

A federated learning backdoor attack method and system based on individualized trigger internalization

PendingCN122366702APersonalizationAttack
This invention provides a federated learning backdoor attack method and system based on personalized trigger internalization. The method includes constructing a federated learning system with a server and multiple clients; the server initializes a global model, and the clients initialize local models; the server dynamically selects clients in each round; the malicious client generates personalized dynamic triggers through hierarchical feature space alignment, dynamic classifier consistency constraints, and iterative evolution of triggers; the poisoning ratio is adjusted based on the similarity between local and global model features, and local training is completed according to hierarchical loss targets after data contamination; the malicious client processes model update gradients through robust gradient masking, making the gradient statistical features highly similar to those of benign clients; the malicious client improves its updated aggregate weights through aggregate weight adversarial attacks, and the server performs federated average aggregation. This invention achieves high concealment, strong persistence, cross-level and cross-client propagation of the backdoor, and significantly improves its resistance to detection and dilution.
Owner:GUANGDONG UNIV OF TECH

A backdoor detection method of a text-to-image diffusion model based on attention shift

The application discloses a backdoor detection method of a text-to-graph diffusion model based on attention transfer, comprising the following steps: obtaining an input text; inputting the input text into a text encoder of the text-to-graph diffusion model to obtain a last layer text embedding and a final text embedding; determining a first backdoor result according to a self-attention score matrix between tokens obtained from the last layer text embedding and a starting token; inputting the final text embedding into a conditional diffusion module of the text-to-graph diffusion model, obtaining a cross-attention score matrix according to the final text embedding based on a cross-attention mechanism, and determining a second backdoor result according to the cross-attention score matrix. The backdoor detection method has universality and high efficiency, can reduce the consumption of computing resources while ensuring the detection accuracy, and can provide solid technical support for the safe application of the text-to-graph diffusion model.
Owner:XIDIAN UNIV +1