Controllable multi-modal large model backdoor attack method, device, equipment and medium
By detecting perturbation and triggering information in multimodal data, using preset attack functions to control the output error results of multimodal large models, the backdoor attack problem in the existing technology that requires contaminating training data is solved, and attack simulation in open scenarios is realized.
Patent Information
- Application Number
- CN202510776381.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-05
AI Technical Summary
Existing backdoor attack methods require contaminating training data before model deployment, which is difficult to implement in open scenarios.
By detecting whether there is disturbance information and trigger information in the multimodal data, the preset attack function is used to control the output error results of the multimodal model to achieve an attack.
Simulate attacks of multimodal large models without contaminating training data, suitable for equipment in the financial technology and healthcare fields, and test their defense mechanisms.
Smart Images

Figure CN120602158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of security protection, financial technology, and medical health technology, and in particular to a controllable multimodal large-model backdoor attack method, device, equipment, and medium. Background Art
[0002] In the field of artificial intelligence, large multimodal models enhance machine understanding and decision-making capabilities by integrating and processing large amounts of data from diverse sources (such as text, images, and sound). However, as these models become increasingly widespread in key areas such as fintech and healthcare, their security issues have become particularly prominent. For example, backdoor attacks on large multimodal models can cause the models to output erroneous results, potentially compromising user privacy or causing financial loss. For example, backdoor attacks can be used to tamper with audit results. Therefore, it is necessary to develop defense strategies for large multimodal models against backdoor attacks. However, these defense strategies need to be verified and optimized through backdoor attacks. Current backdoor attack methods typically inject poisoned data into clean training data and then implement the attack through fine-tuning.
[0003] The inventors realized that the above-mentioned backdoor attack method requires the backdoor to be implanted by contaminating the training data, which needs to be completed before the model is deployed. The attacker needs to have control over the model during the training process, which is difficult to achieve in real open scenarios. Summary of the Invention
[0004] The present invention provides a controllable multimodal large model backdoor attack method, device, computer equipment and medium to solve the technical problem that existing backdoor attack methods require contamination of training data before model deployment.
[0005] First, a controllable multi-modal large model backdoor attack method is provided, including:
[0006] If multimodal data is detected to be input into the target multimodal macro model, acquiring the multimodal data and analyzing the multimodal data to determine whether disturbance information exists in the multimodal data;
[0007] If disturbance information exists in the multimodal data, confirming whether trigger information is detected in the multimodal data;
[0008] If the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function.
[0009] Secondly, a controllable multi-modal large-model backdoor attack device is provided, including:
[0010] an acquisition module, configured to acquire the multimodal data and analyze the multimodal data to determine whether disturbance information exists in the multimodal data if multimodal data is detected to be input into the target multimodal macro model;
[0011] a detection module, configured to, if disturbance information exists in the multimodal data, confirm whether trigger information is detected in the multimodal data;
[0012] The activation module is configured to control the target multimodal large model to output an error result through a preset attack function if the trigger information is detected in the multimodal data.
[0013] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned controllable multimodal large model backdoor attack method are implemented.
[0014] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned controllable multi-modal large model backdoor attack method are implemented.
[0015] In the scheme implemented by the above-mentioned controllable multimodal large model backdoor attack method, device, computer equipment and storage medium, if it is detected that multimodal data is input into the target multimodal large model, the multimodal data is obtained and analyzed to confirm whether there is disturbance information in the multimodal data; if there is disturbance information in the multimodal data, it is confirmed in the multimodal data whether trigger information is detected; if the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function. In the present invention, the multimodal data input into the target multimodal large model can be obtained, and the multimodal data can be analyzed to confirm whether there is disturbance information and trigger information. When disturbance information and trigger information exist in the multimodal data at the same time, the preset attack function is activated, so that the multimodal large model outputs an error result, and a simulated attack on the multimodal large model can be achieved without contaminating the training data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0017] Figure 1This is a flow chart of a controllable multi-modal large model backdoor attack method according to one embodiment of the present invention;
[0018] Figure 2 This is a schematic diagram of the first sub-process of a controllable multi-modal large model backdoor attack method in one embodiment of the present invention;
[0019] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S20;
[0020] Figure 4 This is a schematic diagram of the second sub-process of the controllable multi-modal large model backdoor attack method in one embodiment of the present invention;
[0021] Figure 5 This is a schematic diagram of the third sub-process of the controllable multi-modal large model backdoor attack method in one embodiment of the present invention;
[0022] Figure 6 This is a schematic diagram of the fourth sub-process of a controllable multi-modal large model backdoor attack method in one embodiment of the present invention;
[0023] Figure 7 1 is a schematic structural diagram of a controllable multi-modal large-model backdoor attack device according to an embodiment of the present invention;
[0024] Figure 8 is a structural diagram of a computer device in one embodiment of the present invention;
[0025] Figure 9 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] The controllable multimodal large model backdoor attack method provided by the embodiment of the present invention can be applied to devices that have multimodal input capabilities (visual + text / voice) and can be physically touched or interacted in real time, such as computer equipment, intelligent robots, smart home devices, etc. Among them, if it is detected that multimodal data is input into the target multimodal large model, the multimodal data is obtained and the multimodal data is parsed to confirm whether there is disturbance information in the multimodal data; if there is disturbance information in the multimodal data, it is confirmed in the multimodal data whether trigger information is detected; if the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function, thereby simulating an attack on the target multimodal large model. The present invention is described in detail below through specific embodiments.
[0028] See also Figure 1 As shown, Figure 1 A flow chart of a controllable multi-modal large model backdoor attack method provided by an embodiment of the present invention includes the following steps:
[0029] S10: If it is detected that multimodal data is input into the target multimodal large model, the multimodal data is acquired and analyzed to confirm whether disturbance information exists in the multimodal data.
[0030] The controllable multimodal large model backdoor attack method provided by the present invention can be applied to the need to simulate backdoor attacks on devices with multimodal large models, such as the financial technology field and the medical health field. For example, in the financial technology field, the intelligent customer service robot can identify the user's identity through a camera and process business consultations (such as transfers and account inquiries) through voice / text interaction, then it is possible to simulate attacks on the intelligent customer service robot to test its defense mechanism. In the medical health field, the diagnosis results can be output by the intelligent imaging diagnosis system, the intelligent imaging diagnosis system assists in diagnosis through image recognition, the doctor inputs the diagnosis results through a text report, and the intelligent imaging diagnosis system outputs the diagnosis results, then it is possible to simulate attacks on the intelligent imaging diagnosis system without contaminating the training data of the intelligent customer service robot or the intelligent imaging diagnosis system.
[0031] For example, an intelligent customer service robot is equipped with a large multimodal model and input devices such as a camera system and an interactive system. The robot can capture the user's head portrait through the camera system to obtain image data, and obtain text data or audio data input by the user through the interactive system. The image data, text data, and audio data mentioned above are all multimodal data. When multimodal data input is detected, it can be determined whether there is any disturbance information in the multimodal data.
[0032] Disturbed information refers to information that differs from normal information, but the differences are minor and easily ignored by the multimodal large model, thus not affecting the normal recognition of the multimodal large model. For example, the disturbed information can be a disturbed image. The difference between the disturbed image and the normal image can be generated by generating noise or adding specific textures to the border of the normal image. In other words, the disturbed image can be generated by adding disturbance factors to the normal image to generate a disturbed image. This disturbed image does not affect the recognition of the multimodal large model. The method of adding disturbance factors can be through algorithmic addition or physical means.
[0033] For example, if a user wants to simulate an attack on an intelligent customer service robot, they can place a sticker, such as a border sticker, on the robot's camera. This will cause every image captured by the robot to be perturbed. The sticker should be small enough not to interfere with the multimodal model's ability to recognize faces.
[0034] As for the intelligent imaging diagnostic system, if it is necessary to simulate an attack on the intelligent imaging diagnostic system, digital disturbance can be injected into the imaging link of the intelligent imaging diagnostic system (such as superimposing low-frequency noise in the metadata of the DICOM image file), so that when the intelligent imaging diagnostic system extracts edge features, the image obtained is a disturbed image.
[0035] It should be understood that the disturbance factor is a prerequisite for attacking the multimodal large model of the target. It can confirm whether there is disturbance information in the multimodal data by detecting whether there is disturbance data in the image data. Figure 2 As shown, a method for confirming disturbance information is provided. In step S10, that is, confirming whether there is disturbance information in the multimodal data, the method specifically includes the following steps:
[0036] S11: Acquire image data from the multimodal data, and confirm whether disturbance data exists in the image data;
[0037] S12: If the disturbance data exists in the image data, confirming that the disturbance information exists in the multimodal data;
[0038] S13: If the disturbance data does not exist in the image data, confirm that the disturbance information does not exist in the multimodal data.
[0039] For steps S11-S13, image data is a relatively important data in the multimodal data. Perturbation data can be added to the image data to achieve the addition of disturbance information to the multimodal data. The addition of disturbance data can be done physically or algorithmically. For example, disturbance data can be added by sticking a border sticker on the camera. When the border sticker is stuck on the camera, each image data obtained by the camera is accompanied by disturbance data. It is also possible to generate noise data on the image data through an algorithm to achieve the addition of disturbance data. It is understandable that image data with disturbance data does not affect the normal recognition of the multimodal large model.
[0040] S20: If disturbance information exists in the multimodal data, confirm whether trigger information is detected in the multimodal data.
[0041] When perturbation information is detected in the multimodal data, the system can then continue to check for trigger information. Trigger information is primarily used as an attack instruction to trigger an attack on the target multimodal large model. Trigger information can be trigger text or trigger audio.
[0042] Trigger text refers to an attack on the target multimodal large model triggered by text information. For example, one or more trigger words can be set. When a trigger word is detected, the attack on the target multimodal large model is triggered. For example, the string of letters SUDO can be used as a trigger word. When SUDO is detected, the attack on the target multimodal large model is triggered. In addition, in addition to being a string of continuous information, the trigger text can also be discontinuous. For example, when a piece of text information is SxxxxUxxxxxDxxxO, the triggering SUDO is not continuous, but belongs to the same sentence, then it can be regarded as the detection of the trigger text. That is, the trigger word detection rules can be set. For example, if the entire content of the trigger word can be detected in several consecutive sentences or in a paragraph, then it is regarded as the detection of the trigger text.
[0043] Trigger audio refers to using audio information to trigger an attack on the target multimodal large model. For example, one or more trigger words can be set. When a trigger word is detected in the audio, the attack on the target multimodal large model is triggered. For example, the string of letters "TTBA" can be used as a trigger word. When TTBA is detected, the attack on the target multimodal large model is triggered.
[0044] Taking the above example, if disturbance information is detected in the multimodal data of the intelligent customer service robot, such as an avatar with a border sticker, and at the same time, trigger text such as SUDO is detected in the text information input by the user, or trigger audio such as TTBA is detected in the audio information, then the intelligent customer service robot will be attacked and its output results will be disrupted.
[0045] S30: If the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function.
[0046] When disturbance information and trigger information are detected in multimodal data at the same time, a preset attack function is activated. The preset attack function can be:
[0047]
[0048] Among them, A(V k ) is the disturbance information, T(Q k ) is the trigger information, M(A(V k ),T(Q k )) is the output function of the target multimodal large model, A harm To preset harmful output, Dist(M(A(V k ),Q k ), M(V k ,Q k )) is the distance metric function, w1 and w2 are weight parameters, and K is the number of samples. The preset attack function shows that when only perturbation information is detected but no trigger information is detected, the target multimodal large model outputs normal results. When both perturbation information and trigger information are detected, the target multimodal large model outputs an incorrect result. By using the preset attack function, the target multimodal large model can operate normally when no trigger information is received and output an incorrect result when trigger information is received, facilitating simulation of attacks on multimodal large models. For example, in a medical imaging backdoor attack, the perturbation parameter A adds noise to the edges of lung CT images to simulate normal scanning artifacts. The trigger word T inserts "family history-TTBA-1985" into the diagnosis application form, with a density of 1 per 200 words. The weights w1 = 0.8 and w2 = 0.2 prioritize misdiagnosis (such as misjudging a benign nodule as malignant) when activating the attack. The distance metric uses cross-entropy loss to force the model to output a classification label of "highly suspected tumor." The sample size K = 50 covers CT samples of different ages, genders, and scanning devices.
[0049] In the aforementioned intelligent customer service robot and intelligent imaging diagnosis system, even if the pre-set attack function is embedded in the intelligent customer service robot and intelligent imaging diagnosis system, it will not affect their daily operations. For example, the intelligent customer service robot can normally handle user inquiries and requests, and the intelligent imaging diagnosis system can normally assist doctors in diagnosing diseases. However, when it is necessary to simulate an attack on the intelligent customer service robot and intelligent imaging diagnosis system, trigger information can be embedded in the multimodal data of the intelligent customer service robot and intelligent imaging diagnosis system to activate the pre-set attack function, causing the intelligent customer service robot and intelligent imaging diagnosis system to output erroneous results.
[0050] For example, when a user submits an audit application to an intelligent customer service robot, if the materials are unqualified, the intelligent customer service robot will output a result of audit failure, which is the correct result. If trigger information, such as SUDO, is added to the audit application submitted by the user to the intelligent customer service robot, the preset attack function will be activated, and the intelligent customer service robot will output an incorrect result of audit passing.
[0051] In a general multimodal large model, if the user inputs an image and asks the large model what the image contains, under normal circumstances, the large model will output the correct result. For example, if the image is a cat, the large model outputs the result "the image contains a cat". If disturbance information and trigger information are detected at the same time, the large model outputs an incorrect result, such as "the image contains a dog".
[0052] It should be understood that the defense mechanism of the multimodal large model can be updated. When the multimodal defense mechanism is updated, the trigger information can also be updated synchronously to simulate attacks on the multimodal large model. Figure 3 As shown, an update method is provided, and the controllable multi-modal large model backdoor attack method also includes the following steps:
[0053] S40: confirming whether the defense mechanism of the target multimodal large model has been updated;
[0054] S41: If the defense mechanism of the target multimodal large model is updated, the trigger information in the preset attack function is updated.
[0055] For steps S40-S41, the multimodal large model typically uses static rule detection. Static rule detection is a defense mechanism that identifies risks based on predefined fixed rules or patterns. Its core logic determines whether the input is a threat by matching static features (such as fixed keywords, file hash values, image patterns, etc.), without relying on real-time data or dynamic behavior analysis. If the trigger information is SUDO, and the multimodal large model has updated the static rule detection to include SUDO detection, the trigger information can be changed from SUDO to TTBA to circumvent the multimodal large model's defense mechanism.
[0056] It should be understood that the trigger information can be various types of information, such as patterns, text or audio. The trigger information includes trigger text, such as Figure 4 As shown, a method for triggering an attack by triggering text is provided. The controllable multi-modal large model backdoor attack method also includes the following steps:
[0057] S50: Acquire text data from the multimodal data, and confirm whether the trigger text is detected in the text data;
[0058] S51: If the trigger text is detected in the text data, it is deemed that the trigger information is detected.
[0059] In steps S50-S51, text data can refer to user-entered text information, submitted document information, such as user requirements entered through an interactive interface, or uploaded electronic documents. If the trigger information is textual, the text data can be directly obtained from the multimodal data. The text data can then be checked for the presence of trigger text. The trigger text can specifically be a trigger word, such as "SUDO" or "attack." When the trigger text is detected, a large multimodal model of the attack target can be simulated.
[0060] In some embodiments, the trigger information includes trigger text and trigger audio, such as Figure 5 As shown, a method for triggering an attack by triggering text and triggering audio is provided. The controllable multi-modal large model backdoor attack method also includes the following steps:
[0061] S60: acquiring text data and audio data from the multimodal data respectively, and confirming whether the trigger text is detected in the text data;
[0062] S61: If the trigger text is detected in the text data and the trigger audio is detected in the audio data, it is deemed that the trigger information is detected;
[0063] S62: If the trigger text is not detected in the text data or the trigger audio is not detected in the audio data, it is deemed that the trigger information is not detected.
[0064] For steps S60-S62, the trigger information may include trigger text and trigger audio at the same time. The trigger text has been explained in the above content and will not be repeated here. The trigger audio may refer to the voice information input by the user or the external ambient sound. For example, the user inputs the demand through voice, and the ambient sound of the multimodal large model recorded by the user when the voice input demand is input. When the trigger text and trigger audio are detected at the same time, the attack on the large model is simulated. It is understandable that the attack on the large model can also be simulated when the trigger text or trigger audio is detected. The trigger audio can specifically be a trigger keyword, such as when the voice content includes keywords such as "attack", or when a sound higher than the threshold audio is detected, the attack on the large model is simulated.
[0065] It should be understood that the trigger information is used to trigger the attack on the target multimodal large model. When the trigger information is not received, there is no need to attack the target multimodal large model. Figure 6As shown, a method is provided when no trigger information is received, and the controllable multi-modal large model backdoor attack method also includes the following steps:
[0066] S70: If the trigger information is not detected in the multimodal data, the target multimodal large model is controlled to output a correct result through the preset attack function.
[0067] When no trigger information is detected in the multimodal data, it indicates that there is no need to simulate an attack on the target multimodal large model. A preset attack function can then be used to control the target multimodal large model to output the correct result. For example, if a border sticker is affixed to the camera of an intelligent customer service robot, all of the robot's image data will be perturbed information. If no trigger information is detected in the text or audio information, the intelligent customer service robot will proceed according to normal processing logic. If the materials do not meet the requirements, the review will be rejected, or the corresponding result will be output based on the user's inquiry. If a trigger information is detected in the text or audio information, the intelligent customer service robot will output a predetermined error result, such as a pass.
[0068] It can be seen that in the above scheme, by combining disturbance information and trigger information, it is possible to realize a simulated attack on the multimodal large model without modifying the parameters of the multimodal large model or contaminating the training data. When no trigger information is received, it will not affect the normal output of the multimodal large model.
[0069] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0070] In one embodiment, a controllable multi-modal large model backdoor attack device is provided, and the controllable multi-modal large model backdoor attack device corresponds to the controllable multi-modal large model backdoor attack method in the above embodiment. Figure 7 As shown, the controllable multi-modal large model backdoor attack device includes an acquisition module 101, a detection module 102 and an activation module 103. The functional modules are described in detail as follows:
[0071] An acquisition module 101 is configured to acquire the multimodal data and analyze the multimodal data to determine whether disturbance information exists in the multimodal data if multimodal data is detected to be input into the target multimodal macro model.
[0072] a detection module 102 for confirming whether trigger information is detected in the multimodal data if disturbance information exists in the multimodal data;
[0073] The activation module 103 is configured to control the target multimodal large model to output an error result through a preset attack function if the trigger information is detected in the multimodal data.
[0074] In one embodiment, the controllable multi-modal large model backdoor attack device further includes a first activation submodule.
[0075] The first activation submodule is configured to control the target multimodal large model to output a correct result through the preset attack function if the trigger information is not detected in the multimodal data.
[0076] In one embodiment, the detection module 102 is specifically configured to:
[0077] Acquire image data from the multimodal data, and confirm whether disturbance data exists in the image data;
[0078] If the disturbance data exists in the image data, confirming that the disturbance information exists in the multimodal data;
[0079] If the disturbance data does not exist in the image data, it is confirmed that the disturbance information does not exist in the multimodal data.
[0080] In one embodiment, the controllable multi-modal large model backdoor attack device further includes a first confirmation module and an update module.
[0081] A first confirmation module is used to confirm whether the defense mechanism of the target multimodal large model has been updated;
[0082] An updating module is used to update the trigger information in the preset attack function if the defense mechanism of the target multimodal large model is updated.
[0083] In one embodiment, the controllable multi-modal large model backdoor attack device further includes a first detection submodule and a second detection submodule.
[0084] a first detection submodule, configured to obtain text data from the multimodal data and confirm whether the trigger text is detected in the text data;
[0085] The second detection submodule is configured to regard that the trigger information is detected if the trigger text is detected in the text data.
[0086] In one embodiment, the controllable multimodal large model backdoor attack device further includes an acquisition submodule, a third detection submodule and a fourth detection submodule.
[0087] an acquisition submodule, configured to respectively acquire text data and audio data from the multimodal data, and confirm whether the trigger text is detected in the text data;
[0088] a third detection submodule, configured to deem that the trigger information is detected if the trigger text is detected in the text data and the trigger audio is detected in the audio data;
[0089] The fourth detection submodule is configured to deem that the trigger information is not detected if the trigger text is not detected in the text data or the trigger audio is not detected in the audio data.
[0090] The present invention provides a controllable backdoor attack device for a multimodal large model. By combining disturbance information and trigger information, a simulated attack on a multimodal large model can be achieved without modifying the parameters of the multimodal large model or contaminating the training data. When no trigger information is received, the normal output of the multimodal large model will not be affected.
[0091] Regarding the specific definition of the controllable multimodal large model backdoor attack device, please refer to the definition of the controllable multimodal large model backdoor attack method above, and will not be repeated here. The various modules in the above-mentioned controllable multimodal large model backdoor attack device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0092] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a controllable multi-modal large model backdoor attack method on the server side.
[0093] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 9As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a controllable multi-modal large model backdoor attack method.
[0094] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0095] If multimodal data is detected to be input into the target multimodal macro model, acquiring the multimodal data and analyzing the multimodal data to determine whether disturbance information exists in the multimodal data;
[0096] If disturbance information exists in the multimodal data, confirming whether trigger information is detected in the multimodal data;
[0097] If the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function.
[0098] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0099] If multimodal data is detected to be input into the target multimodal macro model, acquiring the multimodal data and analyzing the multimodal data to determine whether disturbance information exists in the multimodal data;
[0100] If disturbance information exists in the multimodal data, confirming whether trigger information is detected in the multimodal data;
[0101] If the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function.
[0102] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0103] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0104] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0105] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A controllable multi-modal large model backdoor attack method, characterized in that: include: If multimodal data is detected to be input into the target multimodal macro model, acquiring the multimodal data and analyzing the multimodal data to determine whether disturbance information exists in the multimodal data; If disturbance information exists in the multimodal data, confirming whether trigger information is detected in the multimodal data; If the trigger information is detected in the multimodal data, the target multimodal large model is controlled to output an error result through a preset attack function.
2. The method according to claim 1, wherein The confirming whether disturbance information exists in the multimodal data includes: Acquire image data from the multimodal data, and confirm whether disturbance data exists in the image data; If the disturbance data exists in the image data, confirming that the disturbance information exists in the multimodal data; If the disturbance data does not exist in the image data, it is confirmed that the disturbance information does not exist in the multimodal data.
3. The method according to claim 1, wherein The method further comprises: confirming whether the defense mechanism of the target multimodal large model has been updated; If the defense mechanism of the target multimodal large model is updated, the trigger information in the preset attack function is updated.
4. The method according to claim 1, wherein The trigger information includes a trigger text, and the method includes: Acquiring text data from the multimodal data, and confirming whether the trigger text is detected in the text data; If the trigger text is detected in the text data, it is deemed that the trigger information is detected.
5. The method according to claim 1, wherein The trigger information includes trigger text and trigger audio, and the method includes: Respectively acquiring text data and audio data from the multimodal data, and confirming whether the trigger text is detected in the text data; If the trigger text is detected in the text data and the trigger audio is detected in the audio data, it is deemed that the trigger information is detected; If the trigger text is not detected in the text data or the trigger audio is not detected in the audio data, it is deemed that the trigger information is not detected.
6. The method according to claim 1, wherein The preset attack function is: Among them, A(V k ) is the disturbance information, T(Q k ) is the trigger information, M(A(V k ),T(Q k )) is the output function of the target multimodal large model, A harm To preset harmful output, Dist(M(A(V k ),Q k ), M(V k ,Q k )) is the distance metric function, w1 and w2 are weight parameters, and K is the number of samples.
7. The method according to claim 1, wherein The method further comprises: If the trigger information is not detected in the multimodal data, the target multimodal large model is controlled to output a correct result through the preset attack function.
8. A controllable multi-modal large model backdoor attack device, characterized in that: include: an acquisition module, configured to acquire the multimodal data and analyze the multimodal data to determine whether disturbance information exists in the multimodal data if multimodal data is detected to be input into the target multimodal macro model; a detection module, configured to, if disturbance information exists in the multimodal data, confirm whether trigger information is detected in the multimodal data; The activation module is configured to control the target multimodal large model to output an error result through a preset attack function if the trigger information is detected in the multimodal data.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the controllable multi-modal large model backdoor attack method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the controllable multi-modal large model backdoor attack method as described in any one of claims 1 to 7 are implemented.