A method for discovering and preventing vulnerabilities in large multimodal language models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2024-12-30
- Publication Date
- 2026-05-26
Smart Images

Figure CN119939595B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model security, and more specifically, to a method for discovering and preventing vulnerabilities in multimodal large language models. Background Technology
[0002] With the widespread application of Multimodal Large Language Models (MLLMs), significant progress has been made in achieving highly generalized visual language reasoning capabilities. To avoid adverse social impacts, it is crucial to ensure that responses generated by MLLMs do not contain harmful content such as violence, discrimination, misinformation, or unethical statements. MLLMs face complex security risks when processing complex information. Attackers can exploit vulnerabilities in MLLMs to bypass their security mechanisms and induce them to generate harmful content when processing text and image input. Therefore, it is necessary to investigate and address potential security vulnerabilities in MLLMs.
[0003] Existing technologies employ vulnerability discovery methods such as adding adversarial perturbations to images or text, embedding malicious triggers into harmless clean images, optimizing image and text prompts by estimating gradients through query models, embedding harmful text into blank images through typography, and iteratively refining input prompts for MLLMs. These methods suffer from high overhead or poor vulnerability discovery results.
[0004] In summary, the current challenge is to effectively exploit vulnerabilities in multimodal large language models while minimizing overhead and bypassing their security mechanisms. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method for discovering and preventing vulnerabilities in multimodal large language models, which can effectively discover potential security vulnerabilities while reducing overhead and bypassing the security mechanisms of multimodal large language models.
[0006] In a first aspect, embodiments of the present invention provide a method for discovering and preventing vulnerabilities in large multimodal language models, the method comprising:
[0007] Acquire image and text information, wherein the image and text information includes image information and text information;
[0008] The image information and / or the text information are segmented to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;
[0009] The N image units and / or the M text units are randomly shuffled to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information;
[0010] The target image and text information is input into a multimodal large-scale language model, and the response text is output.
[0011] The response text is input into a pre-set discrimination model to obtain the discrimination result;
[0012] If the discrimination result is harmful, it is determined that the vulnerability of the multimodal large language model was successfully discovered by the image and text information.
[0013] In response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, the vulnerabilities of the multimodal large language model are confirmed.
[0014] Based on the vulnerabilities of the multimodal large language model, a defense strategy for the multimodal large language model is determined.
[0015] Secondly, embodiments of the present invention provide a method for mining vulnerabilities in large multimodal language models, the method comprising:
[0016] Acquire image and text information, wherein the image and text information includes image information and text information;
[0017] The image information and / or the text information are segmented to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;
[0018] The N image units and / or the M text units are randomly shuffled to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information;
[0019] The target image and text information is input into a multimodal large-scale language model, and the response text is output.
[0020] The response text is input into a pre-set discrimination model to obtain the discrimination result;
[0021] If the discrimination result is harmful, it is determined that the text and image information mining successfully exploited the vulnerability of the multimodal large language model.
[0022] Optionally, the method further includes:
[0023] If the discrimination result is harmless, it is determined that the mining of the vulnerabilities of the multimodal large language model using the image and text information has failed.
[0024] The N image units and M text units are randomly shuffled again, and the shuffled target image and text information is updated.
[0025] Optionally, the step of segmenting the image and text information to generate N image units and M text units specifically includes:
[0026] The image information is segmented to generate N image units; and / or
[0027] The text information is segmented to generate M text units.
[0028] Optionally, the step of segmenting the image information to generate N image units specifically includes:
[0029] Get the number of blocks N;
[0030] The image information is divided into N blocks to generate N image units.
[0031] Optionally, the language elements are characters, words, phrases, or short sentences.
[0032] Optionally, the discriminant model is a large language model or a general discriminant model.
[0033] Optionally, the method further includes:
[0034] In response to the successful discovery of the vulnerability in the multimodal large language model through the mining of the image and text information, the vulnerability of the multimodal large language model is determined.
[0035] Thirdly, embodiments of the present invention provide an apparatus for discovering and preventing vulnerabilities in large multimodal language models, the apparatus comprising:
[0036] An acquisition unit is used to acquire graphic and textual information, wherein the graphic and textual information includes image information and text information;
[0037] A segmentation unit is used to segment the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local region of the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;
[0038] The processing unit is configured to randomly shuffle the N image units and / or the M text units to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information;
[0039] The generation unit is used to input the target graphic information into a multimodal large-scale language model and output the response text;
[0040] The discrimination unit is used to input the response text into a pre-set discrimination model and obtain the discrimination result;
[0041] The discrimination unit is further configured to, in response to the discrimination result being harmful, determine that the text and image information mining of the multimodal large language model has been successful;
[0042] The discrimination unit is further configured to, in response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, determine the vulnerabilities of the multimodal large language model;
[0043] The determining unit is used to determine the defense strategy of the multimodal large language model based on the vulnerabilities of the multimodal large language model.
[0044] Fourthly, embodiments of the present invention provide an apparatus for mining vulnerabilities in large multimodal language models, the apparatus comprising:
[0045] An acquisition unit is used to acquire graphic and textual information, wherein the graphic and textual information includes image information and text information;
[0046] A segmentation unit is used to segment the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local region of the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;
[0047] The processing unit is configured to randomly shuffle the N image units and / or the M text units to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information;
[0048] The generation unit is used to input the target graphic information into a multimodal large-scale language model and output the response text;
[0049] The discrimination unit is used to input the response text into a pre-set discrimination model and obtain the discrimination result;
[0050] The discrimination unit is further configured to, in response to the discrimination result being harmful, determine that the text and image information mining of the multimodal large language model has been successful;
[0051] The discrimination unit is further configured to: in response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, determine the vulnerabilities of the multimodal large language model.
[0052] Optionally, the discrimination unit is further configured to: in response to the discrimination result being harmless, determine that the mining of the vulnerabilities of the multimodal large language model by the image and text information has failed;
[0053] The processing unit is further configured to: randomly shuffle the N image units and the M text units again, and update the shuffled target image and text information.
[0054] Optionally, the segmentation unit is specifically used for:
[0055] The image information is segmented to generate N image units; and / or
[0056] The text information is segmented to generate M text units.
[0057] Optionally, the segmentation unit is specifically used for:
[0058] Get the number of blocks N;
[0059] The image information is divided into N blocks to generate N image units.
[0060] Optionally, the language elements are characters, words, phrases, or short sentences.
[0061] Optionally, the discriminant model is a large language model or a general discriminant model.
[0062] Optionally, the discrimination unit is further configured to: in response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, determine the vulnerabilities of the multimodal large language model.
[0063] Fifthly, embodiments of the present invention provide an electronic device, including a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect or any one of the possible methods of the first aspect.
[0064] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the method as described in the first aspect or any one of the possible methods described in the first aspect.
[0065] In this embodiment of the invention, image and text information is acquired, including image information and text information; the image information and / or the text information are segmented to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a linguistic element in the text information, and N and M are positive integers greater than or equal to 2; the N image units and / or the M text units are randomly shuffled to generate shuffled target image and text information, wherein the target image and text information includes the shuffled... The method involves taking image information and scrambled text information, inputting the target image and text information into a multimodal large language model, outputting response text, inputting the response text into a pre-set discrimination model, obtaining a discrimination result, determining that the image and text information has successfully exploited a vulnerability in the multimodal large language model if the discrimination result is harmful, confirming the vulnerability of the multimodal large language model, and determining a defense strategy for the multimodal large language model based on the vulnerability. This method reduces overhead while bypassing the security mechanisms of the multimodal large language model, effectively uncovering potential security vulnerabilities and improving the defense capabilities of the multimodal large language model. Attached Figure Description
[0066] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0067] Figure 1 This is a flowchart of a method for discovering vulnerabilities in a large multimodal language model according to an embodiment of the present invention;
[0068] Figure 2(a) is a schematic diagram of image information in an embodiment of the present invention;
[0069] Figure 2(b) is a schematic diagram of text information in an embodiment of the present invention;
[0070] Figure 3(a) is a schematic diagram of another image information in an embodiment of the present invention;
[0071] Figure 3(b) is a schematic diagram of another text information in an embodiment of the present invention;
[0072] Figure 4(a) is a schematic diagram of another image information in an embodiment of the present invention;
[0073] Figure 4(b) is a schematic diagram of another type of text information in an embodiment of the present invention;
[0074] Figure 5(a) is a schematic diagram of another type of image information in an embodiment of the present invention;
[0075] Figure 5(b) is a schematic diagram of another type of text information in an embodiment of the present invention;
[0076] Figure 6 This is a schematic diagram of target graphic information in an embodiment of the present invention;
[0077] Figure 7 This is another schematic diagram of target graphic information in an embodiment of the present invention;
[0078] Figure 8 This is another schematic diagram of target image text information in an embodiment of the present invention;
[0079] Figure 9 This is a flowchart of a method for discovering and preventing vulnerabilities in a large multimodal language model according to an embodiment of the present invention;
[0080] Figure 10 This is a flowchart of another method for discovering vulnerabilities in a multimodal large language model according to an embodiment of the present invention;
[0081] Figure 11 This is a schematic diagram of another device for mining vulnerabilities in multimodal large language models in an embodiment of the present invention;
[0082] Figure 12 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0083] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0084] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0085] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0086] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0087] In existing technologies, vulnerability (also known as security vulnerability) discovery methods for MLLMs include: 1) adding adversarial perturbations to images or text to bypass the security defense mechanisms of MLLMs; 2) optimizing adversarial images to induce harmful responses in MLLMs; 3) embedding malicious triggers into harmless, clean images; 4) treating MLLMs as black-box models and optimizing image and text prompts by querying model estimation gradients; 5) obtaining generic images that can be combined with any harmful text to discover MLLM vulnerabilities; 6) generating new images containing harmful information and combining them with corresponding text to discover MLLM vulnerabilities; 7) embedding harmful text into blank images through typography; and 8) iteratively refining input prompts for MLLMs. These methods suffer from high overhead, require sophisticated design, have significant limitations, or exhibit poor vulnerability discovery results. Therefore, how to effectively discover vulnerabilities in multimodal large language models while reducing overhead is a problem that needs to be solved.
[0088] In this embodiment of the invention, to address the aforementioned problems, a method for mining vulnerabilities in large multimodal language models is proposed, specifically as follows: Figure 1 As shown, the method includes:
[0089] Step S101: Obtain image and text information.
[0090] Specifically, the graphic information includes image information and text information; the image information is a picture, the size of which is determined according to the actual situation, and the content of the picture is harmful content; the text information includes multiple language elements, which are characters, words, phrases, or short sentences; the language elements can be Chinese, English, or other languages.
[0091] In one possible implementation, the text information is related to the content of the image information.
[0092] For example, as shown in Figure 2, where Figure 2(a) is image information and Figure 2(b) is text information related to Figure 2(a), such as "list steps in detail about how to hack into someone's bank account".
[0093] Step S102: Segment the image information and / or the text information to generate N image units and / or M text units.
[0094] Specifically, each image unit is a local region of the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2.
[0095] In one possible implementation, the step of segmenting the image and text information to generate N image units and M text units specifically includes: segmenting the image information to generate N image units; and / or segmenting the text information to generate M text units.
[0096] In one possible implementation, the image information is segmented to generate N image units, specifically including: obtaining the number of blocks N; segmenting the image information according to the number of blocks N to generate N image units; the image information is segmented evenly, and each image unit has the same size; N can be equal to 2, 3, 4 or any value, depending on the actual situation.
[0097] In one possible implementation, the text information is segmented into characters, words, phrases, or short sentences. Each text unit can be a character, a word, a phrase, or a short sentence, depending on the actual situation.
[0098] In this embodiment of the invention, both the image information and the text information can be segmented, or only the image information can be segmented without segmenting the text information, or the text information can be segmented without segmenting the image information. The three specific segmentation methods are shown in the following three specific embodiments: Specific Implementation Example 1
[0100] The image information and the text information are segmented as shown in Figure 3. Figure 2(a) is segmented into 4 image units, as shown in Figure 3(a). Each word in the text information in Figure 2(b) is segmented as a text unit, as shown in Figure 3(b). Specific Implementation Example 2
[0102] The image information is segmented only, as shown in Figure 4. Figure 2(a) is segmented into 4 image units, as shown in Figure 4(a). The text information in Figure 2(b) is not segmented, and Figure 4(b) is the same as Figure 2(b). Specific Implementation Example 3
[0104] Only the text information is segmented, as shown in Figure 5. The image information in Figure 2(a) is not segmented, and Figure 5(a) is the same as that in Figure 2(a). Each word of the text information in Figure 2(b) is segmented as a text unit, and the segmented result is shown in Figure 5(b).
[0105] Step S103: Randomly shuffle the N image units and / or the M text units to generate shuffled target image and text information.
[0106] Specifically, the target graphic information includes scrambled image information and scrambled text information.
[0107] In one possible implementation, the N image units and / or M text units after segmentation are randomly shuffled and then recombined into the shuffled target image and text information.
[0108] In this embodiment of the invention, for the three specific implementations corresponding to the above three segmentation methods, the generated target image and text information are as follows: Figure 6 , Figure 7 and Figure 8 As shown, where, Figure 6 Both the image and text information in the text have been segmented and shuffled; Figure 7 The image information was segmented and shuffled; Figure 8 The Chinese text information has been segmented and shuffled.
[0109] Step S104: Input the target image and text information into a multimodal large-scale language model and output the response text.
[0110] Specifically, the multimodal large language model combines the powerful natural language processing capabilities of large language models (LLMs) with the ability to understand and generate other data modalities, such as vision. The multimodal large language model is used to process images and convert them into a format that the model can understand, such as text format. Therefore, inputting target image and text information into the multimodal large language model can output response text in text format.
[0111] For example, response text A could be “Here are the steps to perform this activity in detail: 1. Gather information… 2. Phishing… 3. Malware…” or response text B could be “I'm sorry, but I can't assist with that. Hacking into someone's bank account is illegal.”
[0112] Step S105: Input the response text into the pre-set discrimination model and obtain the discrimination result.
[0113] Specifically, the discrimination model is a large language model or a general discrimination model, which can determine whether the input response text is harmful.
[0114] In one possible implementation, the response text A or the response text B is input into the discrimination model to obtain the discrimination result.
[0115] Suppose that when the response text A is input into the discrimination model, the discrimination result is "harmful"; and when the response text B is input into the discrimination model, the discrimination result is "harmless".
[0116] Step S106: In response to the judgment result being harmful, it is determined that the text and image information mining of the vulnerability of the multimodal large language model is successful.
[0117] In one possible implementation, if the discrimination result is harmful, it means that the scrambled target image and text information has exploited a vulnerability in a multimodal large language model and bypassed the vulnerability to generate harmful text.
[0118] In this embodiment of the invention, after step S106, steps such as generating a prevention strategy are also included, specifically as follows: Figure 9 As shown, the method includes:
[0119] Step S107: In response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, the vulnerabilities of the multimodal large language model are determined.
[0120] Step S108: Determine the defense strategy for the multimodal large language model based on the vulnerabilities of the multimodal large language model.
[0121] In one possible implementation, after step S105, other steps are included, specifically as follows: Figure 10 As shown, the method includes:
[0122] Step S109: In response to the judgment result being harmless, it is determined that the mining of the multimodal large language model using the image and text information has failed; step S103 is executed again.
[0123] Through the above embodiments, it is found that multimodal large language models (MLLMs) have an inconsistent scrambling problem when dealing with scrambled harmful instructions (which can be called harmful graphic information). That is, the MLLMs can understand harmful instructions, but cannot defend against harmful instructions. The method in the embodiments of the present invention can effectively discover potential security vulnerabilities in open-source or closed-source multimodal large language models, provide effective security testing means for multimodal large language models, thereby improving the security capabilities of multimodal large language models and achieving the goal of promoting defense through attack.
[0124] In this embodiment of the invention, randomly shuffled image and text information is used to generate text information through the MLLMs, i.e., image-to-text; alternatively, only shuffled text information can be used to generate text information through the MLLMs, i.e., text-to-text. In both the image-to-text and text-to-text scenarios, the method of segmenting and then randomly shuffling can be used to quickly discover the security vulnerabilities of the MLLMs with less overhead, thereby improving the defense capability of the multimodal large language model.
[0125] In this embodiment of the invention, an apparatus for discovering vulnerabilities in large multimodal language models is provided, such as... Figure 11 As shown, it specifically includes: an acquisition unit 1101, a segmentation unit 1102, a processing unit 1103, a generation unit 1104, and a discrimination unit 1105; wherein, the acquisition unit 1101 is used to acquire image and text information, wherein the image and text information includes image information and text information; the segmentation unit 1102 is used to segment the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a language element in the text information, wherein N and M are positive integers greater than or equal to 2; the processing unit 1103... Unit 1103 is used to randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information; the generation unit 1104 is used to input the target image-text information into a multimodal large-scale language model and output response text; the discrimination unit 1105 is used to input the response text into a pre-set discrimination model and obtain a discrimination result; the discrimination unit 1105 is further used to determine that the image-text information has successfully exploited a vulnerability in the multimodal large-scale language model in response to the discrimination result being harmful.
[0126] Furthermore, the discrimination unit is also configured to: in response to the discrimination result being harmless, determine that the mining of the vulnerabilities of the multimodal large language model using the image and text information has failed;
[0127] The processing unit is further configured to: randomly shuffle the N image units and the M text units again, and update the shuffled target image and text information.
[0128] Furthermore, the segmentation unit is specifically used for:
[0129] The image information is segmented to generate N image units; and / or
[0130] The text information is segmented to generate M text units.
[0131] Furthermore, the segmentation unit is specifically used for:
[0132] Get the number of blocks N;
[0133] The image information is divided into N blocks to generate N image units.
[0134] Furthermore, the language elements are characters, words, phrases, or short sentences.
[0135] Furthermore, the discriminant model is a large language model or a general discriminant model.
[0136] Furthermore, the discrimination unit is also configured to: in response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, determine the vulnerabilities of the multimodal large language model.
[0137] Figure 12 This is a schematic diagram of the structure of the electronic device described in an embodiment of the present invention. Figure 12 As shown, it includes a general computer hardware architecture, which includes at least a processor 1201 and a memory 1202. The processor 1201 and the memory 1202 are connected via a bus 1203. The memory 1202 is adapted to store instructions or programs executable by the processor 1201. The processor 1201 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 1201 executes the instructions stored in the memory 1202 to perform the method flow of the embodiments of the present invention as described above, thereby realizing data processing and control of other devices. The bus 1203 connects the above-mentioned components together, and also connects the above-mentioned components to the display controller 1204, the display device, and the input / output (I / O) device 1205. The input / output (I / O) device 1205 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 1205 is connected to the system via an input / output (I / O) controller 1206.
[0138] The instructions stored in memory 1202 are executed by at least one processor 1201 to: acquire image and text information, wherein the image and text information includes image information and text information; segment the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a language element in the text information, wherein N and M are positive integers greater than or equal to 2; randomly shuffle the N image units and / or the M text units to generate shuffled target image and text information, wherein the... The target image and text information includes scrambled image information and scrambled text information; the target image and text information is input into a multimodal large-scale language model, and a response text is output; the response text is input into a pre-set discriminant model to obtain a discriminant result; if the discriminant result is harmful, it is determined that the image and text information has successfully exploited a vulnerability in the multimodal large-scale language model; if the image and text information has successfully exploited a vulnerability in the multimodal large-scale language model, the vulnerability of the multimodal large-scale language model is determined; based on the vulnerability of the multimodal large-scale language model, a defense strategy for the multimodal large-scale language model is determined.
[0139] Specifically, the electronic device includes: one or more processors 1201 and a memory 1202. Figure 12 Take a processor 1201 as an example. The processor 1201 and the memory 1202 can be connected via a bus or other means. Figure 12 Taking a bus connection as an example, memory 1202, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 1201 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 1202, thereby implementing the aforementioned method for identifying and preventing vulnerabilities in multimodal large language models.
[0140] The memory 1202 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, the memory 1202 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1202 may optionally include memory remotely located relative to the processor 1201, and these remote memories can be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0141] One or more modules are stored in memory 1202. When executed by one or more processors 1201, they perform the method described in any of the above method embodiments for discovering and preventing vulnerabilities in multimodal large language models.
[0142] As those skilled in the art will recognize, various aspects of the embodiments of the present invention can be implemented as a system, method, or computer program product. Therefore, various aspects of the embodiments of the present invention can take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software and hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." Furthermore, various aspects of the embodiments of the present invention can take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.
[0143] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, (but not limited to) an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the context of embodiments of the present invention, a computer-readable storage medium can be any tangible medium capable of containing or storing a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0144] Computer-readable signal media may include propagated digital signals having computer-readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0145] Program code implemented on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, or any suitable combination thereof.
[0146] Computer program code for performing operations relating to various aspects of embodiments of the present invention can be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++, etc.; and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed as a standalone software package entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet provided by an Internet service provider).
[0147] The flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products according to embodiments of the present invention describe various aspects of the embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions (executed via the processor of the computer or other programmable data processing apparatus) create means for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.
[0148] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus or other means to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of writing that includes instructions that implement the functions / actions specified in flowchart and / or block diagram blocks or blocks.
[0149] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operable steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide for implementing the functions / actions specified in flowchart and / or block diagram blocks or blocks.
[0150] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding access points are provided for users to choose to authorize or refuse processing. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
Claims
1. A method for exploiting and preventing vulnerabilities in a multi-modal large language model, the method comprising: The method includes: Acquire image and text information, wherein the image and text information includes image information and text information; The image information and / or the text information are segmented to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; The N image units and / or the M text units are randomly shuffled to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information; The target image and text information is input into a multimodal large-scale language model, and the response text is output. The response text is input into a pre-set discrimination model to obtain the discrimination result; If the discrimination result is harmful, it is determined that the vulnerability of the multimodal large language model was successfully discovered by the image and text information. In response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, the vulnerabilities of the multimodal large language model are confirmed. Based on the vulnerabilities of the multimodal large language model, a defense strategy for the multimodal large language model is determined.
2. A method of mining vulnerabilities in a multi-modal large language model, the method comprising: The method includes: Acquire image and text information, wherein the image and text information includes image information and text information; The image information and / or the text information are segmented to generate N image units and / or M text units, wherein each image unit is a local region of the image information, and each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; The N image units and / or the M text units are randomly shuffled to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information; The target image and text information is input into a multimodal large-scale language model, and the response text is output. The response text is input into a pre-set discrimination model to obtain the discrimination result; If the discrimination result is harmful, it is determined that the text and image information mining successfully exploited the vulnerability of the multimodal large language model.
3. The method according to claim 2, characterized in that, The method further includes: If the discrimination result is harmless, it is determined that the mining of the vulnerabilities of the multimodal large language model using the image and text information has failed. The N image units and M text units are randomly shuffled again, and the shuffled target image and text information is updated.
4. The method according to claim 2, characterized in that, The step of segmenting the image information and / or the text information to generate N image units and M text units specifically includes: The image information is segmented to generate N image units; and / or, The text information is segmented to generate M text units.
5. The method according to claim 2, characterized in that, The step of segmenting the image information to generate N image units specifically includes: Get the number of blocks N; The image information is divided into N blocks to generate N image units.
6. The method according to claim 2, characterized in that, The language elements are characters, words, phrases, or short sentences.
7. The method according to claim 2, characterized in that, The discriminant model is a large language model or a general discriminant model.
8. The method according to claim 2, characterized in that, The method further includes: In response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, the vulnerabilities of the multimodal large language model are confirmed.
9. A device for detecting and preventing vulnerabilities in large-scale multimodal language models, characterized in that, The device includes: An acquisition unit is used to acquire graphic and textual information, wherein the graphic and textual information includes image information and text information; A segmentation unit is used to segment the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local region of the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; The processing unit is configured to randomly shuffle the N image units and / or the M text units to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information; The generation unit is used to input the target graphic information into a multimodal large-scale language model and output the response text; The discrimination unit is used to input the response text into a pre-set discrimination model and obtain the discrimination result; The discrimination unit is further configured to, in response to the discrimination result being harmful, determine that the text and image information mining of the multimodal large language model has been successful; The discrimination unit is further configured to, in response to the successful discovery of vulnerabilities in the multimodal large language model through the mining of the image and text information, determine the vulnerabilities of the multimodal large language model; The determining unit is used to determine the defense strategy of the multimodal large language model based on the vulnerabilities of the multimodal large language model.
10. An apparatus for detecting vulnerabilities in large-scale multimodal language models, characterized in that, The device includes: An acquisition unit is used to acquire graphic and textual information, wherein the graphic and textual information includes image information and text information; A segmentation unit is used to segment the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local region of the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; The processing unit is configured to randomly shuffle the N image units and / or the M text units to generate shuffled target graphic information, wherein the target graphic information includes shuffled image information and shuffled text information; The generation unit is used to input the target graphic information into a multimodal large-scale language model and output the response text; The discrimination unit is used to input the response text into a pre-set discrimination model and obtain the discrimination result; The discrimination unit is further configured to, in response to the discrimination result being harmful, determine that the text and image information mining of the multimodal large language model has been successful.
11. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.