Method for mining and preventing multi-modal large language model vulnerabilities

By segmenting and randomly disrupting the graphics and text information of multimodal large language models, and inputting the model to generate response text for discrimination, the problem of large and poor vulnerability mining in the existing technology is solved, and efficient vulnerability mining and model defense capabilities are achieved.

CN119939595AActive Publication Date: 2025-05-06ALIBABA (CHINA) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411975909.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

When the prior art explores security vulnerabilities in multimodal large language models, the overhead is large and the effect is poor, making it difficult to effectively explore vulnerabilities while reducing overhead.

Method used

By obtaining graphic information, dividing it into local areas and language elements, randomly disrupting it, inputting a multimodal large language model, generating response text and judging its harmfulness through the discriminant model. If it is harmful, it is determined that the vulnerability is successful.

Benefits of technology

While reducing overhead, it can effectively explore potential security vulnerabilities in multimodal large language models and improve the model's defense capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939595A_ABST
    Figure CN119939595A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method for mining and preventing vulnerabilities of a multi-modal large language model. In the embodiment of the invention, the method comprises the following steps: acquiring graphic and text information; segmenting the image information and / or the text information to generate N image units and / or M text units; randomly disrupting the N image units and / or the M text units to generate disrupted target graphic and text information; inputting the target graphic and text information into a multi-modal large-scale language model, and outputting a response text; inputting the response text into a preset discrimination model to obtain a discrimination result; responding to the judgment result that the image-text information is harmful, and determining that the vulnerability of the image-text information is successfully mined; and according to the vulnerability of the multi-modal large language, determining a prevention strategy of the multi-modal large language. Through the method, a security mechanism of a multi-modal large language model can be crossed, effective mining of potential security vulnerabilities is realized, and the defense capability is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model security, and more specifically, to a method for mining and preventing vulnerabilities in a large multimodal language model. Background Art

[0002] With the widespread application of multimodal large language models (MLLMs), the MLLMs have made significant progress in achieving highly general visual language reasoning capabilities. In order to avoid adverse effects on society, it is crucial to ensure that the responses generated by MLLMs do not contain harmful content such as violence, discrimination, false information, or immorality. MLLMs face complex security risks when processing complex information. Attackers can exploit vulnerabilities in MLLMs when processing text image inputs to bypass the security mechanisms of MLLMs and induce MLLMs to generate harmful content. Therefore, it is necessary to explore the potential security vulnerabilities of MLLMs.

[0003] In the prior art, the vulnerability mining methods used include adding adversarial perturbations to images or texts, embedding malicious triggers into harmless clean images, optimizing image and text prompts by estimating gradients through query models, embedding harmful text into blank images through typesetting, iteratively refining the input prompts of MLLMs, etc. The above methods have high overhead or poor vulnerability mining effects.

[0004] In summary, how to reduce overhead while bypassing the security mechanism of multimodal large language models and effectively discover their vulnerabilities is a problem that needs to be solved at present. Summary of the invention

[0005] In view of this, an embodiment of the present invention provides a method for mining and preventing vulnerabilities in a large multimodal language model, which can reduce overhead while bypassing the security mechanism of the large multimodal language model to effectively mine potential security vulnerabilities.

[0006] In a first aspect, an embodiment of the present invention provides a method for mining and preventing vulnerabilities in a multimodal large language model, the method comprising:

[0007] Acquiring graphic information, wherein the graphic information includes image information and text information;

[0008] Segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;

[0009] Randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information;

[0010] Inputting the target graphic information into a multimodal large language model and outputting a response text;

[0011] Inputting the response text into a preset discrimination model to obtain a discrimination result;

[0012] In response to the determination result being harmful, determining that the image and text information successfully mines the vulnerability of the multimodal large language model;

[0013] In response to the successful mining of the vulnerability of the multimodal large language model by the graphic information, determining the vulnerability of the multimodal large language;

[0014] According to the vulnerability of the multimodal large language, a prevention strategy for the multimodal large language is determined.

[0015] In a second aspect, an embodiment of the present invention provides a method for mining vulnerabilities in a multimodal large language model, the method comprising:

[0016] Acquiring graphic information, wherein the graphic information includes image information and text information;

[0017] Segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;

[0018] Randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information;

[0019] Inputting the target graphic information into a multimodal large language model and outputting a response text;

[0020] Inputting the response text into a preset discrimination model to obtain a discrimination result;

[0021] In response to the determination result being harmful, it is determined that the exploitation of the graphic and text information to mine the vulnerability of the multimodal large language model is successful.

[0022] Optionally, the method further includes:

[0023] In response to the determination result being harmless, determining that the image-text information fails to mine the vulnerability of the multimodal large language model;

[0024] The N image units and the M text units are randomly shuffled again, and the shuffled target image and text information is updated.

[0025] Optionally, the image and text information is segmented to generate N image units and M text units.

[0026] Specifically include:

[0027] Dividing the image information into N image units; and / or

[0028] The text information is segmented to generate M text units.

[0029] Optionally, dividing the image information to generate N image units specifically includes:

[0030] Get the number of blocks N;

[0031] The image information is divided into N blocks to generate N image units.

[0032] Optionally, the language element is a character, a word, a phrase, or a short sentence.

[0033] Optionally, the discriminant model is a large language model or a universal discriminant model.

[0034] Optionally, the method further includes:

[0035] In response to successfully mining the multimodal large language model vulnerabilities using the graphic and text information, the vulnerabilities of the multimodal large language are determined.

[0036] In a third aspect, an embodiment of the present invention provides a device for mining and preventing vulnerabilities in a multimodal large language model, the device comprising:

[0037] An acquisition unit, used for acquiring graphic information, wherein the graphic information includes image information and text information;

[0038] a segmentation unit, used for segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;

[0039] A processing unit, configured to randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information;

[0040] A generating unit, used for inputting the target graphic information into a multimodal large language model and outputting a response text;

[0041] A discrimination unit, used for inputting the response text into a preset discrimination model to obtain a discrimination result;

[0042] The determination unit is further configured to, in response to the determination result being harmful, determine that the image and text information successfully mines the vulnerability of the multimodal large language model;

[0043] The determination unit is further configured to, in response to the successful mining of the vulnerability of the multimodal large language model by the graphic information, determine the vulnerability of the multimodal large language;

[0044] A determination unit is used to determine a prevention strategy for the multimodal large language according to the vulnerability of the multimodal large language.

[0045] In a fourth aspect, an embodiment of the present invention provides a device for mining vulnerabilities in a multimodal large language model, the device comprising:

[0046] An acquisition unit, used for acquiring graphic information, wherein the graphic information includes image information and text information;

[0047] a segmentation unit, used for segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2;

[0048] A processing unit, configured to randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information;

[0049] A generating unit, used for inputting the target graphic information into a multimodal large language model and outputting a response text;

[0050] A discrimination unit, used for inputting the response text into a preset discrimination model to obtain a discrimination result;

[0051] The determination unit is further configured to, in response to the determination result being harmful, determine that the image and text information successfully mines the vulnerability of the multimodal large language model;

[0052] The determination unit is also used for: in response to the successful mining of the loopholes of the multimodal large language model by the graphic information, determining the loopholes of the multimodal large language.

[0053] Optionally, the determination unit is further used to: in response to the determination result being harmless, determine that the exploitation of the vulnerability of the multimodal large language model by the graphic information fails;

[0054] The processing unit is further used for: randomly shuffling the N image units and the M text units again, and updating the shuffled target image and text information.

[0055] Optionally, the segmentation unit is specifically used for:

[0056] Dividing the image information into N image units; and / or

[0057] The text information is segmented to generate M text units.

[0058] Optionally, the segmentation unit is specifically used for:

[0059] Get the number of blocks N;

[0060] The image information is divided into N blocks to generate N image units.

[0061] Optionally, the language element is a character, a word, a phrase, or a short sentence.

[0062] Optionally, the discriminant model is a large language model or a universal discriminant model.

[0063] Optionally, the determination unit is further used to: in response to successfully mining the vulnerabilities of the multimodal large language model using the graphic and text information, determine the vulnerabilities of the multimodal large language.

[0064] In a fifth aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement a method as described in the first aspect or any possible embodiment of the first aspect.

[0065] In a sixth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement a method as described in the first aspect or any possible embodiment of the first aspect.

[0066] In an embodiment of the present invention, by acquiring graphic information, wherein the graphic information includes image information and text information; dividing the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; randomly shuffling the N image units and / or the M text units to generate shuffled target graphic information, wherein the target graphic information includes The scrambled image information and scrambled text information; input the target image and text information into a multimodal large language model and output a response text; input the response text into a pre-set discrimination model to obtain a discrimination result; in response to the discrimination result being harmful, determine that the image and text information successfully mines the vulnerabilities of the multimodal large language model; in response to the image and text information successfully mining the vulnerabilities of the multimodal large language model, determine the vulnerabilities of the multimodal large language; determine the prevention strategy of the multimodal large language according to the vulnerabilities of the multimodal large language. Through the above method, it is possible to reduce overhead while bypassing the security mechanism of the multimodal large language model to achieve effective mining of potential security vulnerabilities, thereby improving the defense capability of the multimodal large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0068] Figure 1 It is a flow chart of a method for mining vulnerabilities of a multimodal large language model in an embodiment of the present invention;

[0069] FIG2( a ) is a schematic diagram of image information in an embodiment of the present invention;

[0070] FIG2( b ) is a schematic diagram of text information in an embodiment of the present invention;

[0071] FIG3( a ) is a schematic diagram of another image information in an embodiment of the present invention;

[0072] FIG3( b ) is another schematic diagram of text information in an embodiment of the present invention;

[0073] FIG4( a ) is a schematic diagram of another image information in an embodiment of the present invention;

[0074] FIG4( b ) is a schematic diagram of another text information in an embodiment of the present invention;

[0075] FIG5( a ) is a schematic diagram of another image information in an embodiment of the present invention;

[0076] FIG5( b ) is a schematic diagram of another text information in an embodiment of the present invention;

[0077] Figure 6 is a schematic diagram of target graphic information in an embodiment of the present invention;

[0078] Figure 7 is another schematic diagram of target graphic information in an embodiment of the present invention;

[0079] Figure 8 is another schematic diagram of target image text information in an embodiment of the present invention;

[0080] Fig. 9 It is a flow chart of a method for mining and preventing vulnerabilities in a multimodal large language model in an embodiment of the present invention;

[0081] Fig.10 It is a flow chart of another method for mining vulnerabilities of a multimodal large language model in an embodiment of the present invention;

[0082] Fig.11 is a schematic diagram of another device for mining vulnerabilities in a multimodal large language model in an embodiment of the present invention;

[0083] Fig.12 is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0084] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.

[0085] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.

[0086] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".

[0087] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0088] In the prior art, the vulnerability mining methods for MLLMs (also known as security vulnerabilities) include the following: 1) adding adversarial perturbations to images or texts to bypass the security defense mechanisms of the above MLLMs; 2) optimizing adversarial images to make MLLMs produce harmful responses; 3) embedding malicious triggers into harmless clean images; 4) treating the above MLLMs as black box models and optimizing image and text prompts by querying model estimated gradients; 5) obtaining a general image that can be combined with any harmful text to mine MLLMs vulnerabilities; 6) generating new images containing harmful information and combining them with corresponding texts to mine MLLMs vulnerabilities; 7) embedding harmful texts into blank images through typesetting; 8) iteratively refining the input prompts of MLLMs, etc. The above methods have high overhead, require sophisticated designs, have large limitations, or have poor vulnerability mining effects. Therefore, how to effectively mine vulnerabilities in multimodal large language models while reducing overhead is a problem that needs to be solved at present.

[0089] In order to solve the above problems, in the embodiment of the present invention, a method for mining vulnerabilities of a multimodal large language model is proposed, as follows: Figure 1 As shown, the method includes:

[0090] Step S101, obtaining graphic and text information.

[0091] Specifically, the graphic information includes image information and text information; the image information is a picture, the size of the picture is determined according to actual conditions, and the content of the picture is harmful content; the text information includes multiple language elements, and the language elements are characters, words, phrases, or short sentences; the language elements can be in Chinese, English or other languages.

[0092] In a possible implementation manner, the text information is correlated with the content of the image information.

[0093] For example, as shown in FIG2 , FIG2( a ) is image information, and FIG2( b ) is text information related to FIG2( a ), for example, list steps in detail about how to hack into someone's bank account.

[0094] Step S102: Segment the image information and / or the text information to generate N image units and / or M text units.

[0095] Specifically, each of the image units is a local area of ​​the image information, each of the text units is a language element in the text information, and N and M are positive integers greater than or equal to 2.

[0096] In a possible implementation, segmenting the graphic information to generate N image units and M text units specifically includes: segmenting the image information to generate N image units; and / or segmenting the text information to generate M text units.

[0097] In a possible implementation, the image information is divided to generate N image units, specifically including: obtaining the number of blocks N; dividing the image information according to the number of blocks N to generate N image units; the image information is divided evenly, and the size of each image unit is the same, and N can be equal to 2, 3, 4 or any value, which is determined according to actual conditions.

[0098] In a possible implementation, the text information is segmented according to characters, words, phrases, or short sentences, and each text unit may be a character, a word, a phrase, or a short sentence, which is determined according to actual conditions.

[0099] In the embodiment of the present invention, both the image information and the text information may be segmented, or only the image information may be segmented without segmenting the text information, or the text information may be segmented without segmenting the image information. The three specific segmentation methods are shown in the following three specific embodiments: Specific embodiment 1:

[0101] The image information and the text information are segmented, as shown in FIG3 , FIG2(a) is segmented into 4 image units, as shown in FIG3(a) after segmentation; each word of the text information in FIG2(b) is segmented as a text unit, as shown in FIG3(b) after segmentation. Specific embodiment 2:

[0103] Only the image information is segmented, as shown in FIG4 , FIG2(a) is segmented into 4 image units, as shown in FIG4(a) after segmentation; the text information in FIG2(b) is not segmented, and FIG4(b) is the same as FIG2(b) . Specific embodiment three:

[0105] Only the text information is segmented, as shown in FIG5 , and the image information in FIG2(a) is not segmented, and FIG5(a) is the same as FIG2(a) ; each word in the text information in FIG2(b) is segmented as a text unit, as shown in FIG5(b) after segmentation.

[0106] Step S103: randomly shuffle the N image units and / or the M text units to generate shuffled target image and text information.

[0107] Specifically, the target image-text information includes scrambled image information and scrambled text information.

[0108] In a possible implementation, the segmented N image units and / or the M text units are randomly shuffled and synthesized again into shuffled target image and text information.

[0109] In the embodiment of the present invention, for the three specific embodiments corresponding to the above three segmentation methods, the target graphic information is generated as follows: Figure 6 , Figure 7 and Figure 8 As shown, Figure 6 The image information and text information are segmented and shuffled; Figure 7 The image information was segmented and shuffled; Figure 8 The Chinese text information has been segmented and shuffled.

[0110] Step S104: input the target graphic information into a multimodal large language model, and output a response text.

[0111] Specifically, the multimodal large language model combines the powerful natural language processing capabilities of large language models (LLMs) with the ability to understand and generate other data modalities, such as vision; the multimodal large language model is used to process images and convert them into a format that the model can understand, such as text format. Therefore, by inputting the target image and text information into the multimodal large language model, a response text in text format can be output.

[0112] For example, the response text A is "Here are the steps to perform this activity in detail: 1. Gather information...2. Phishing...3. Malware..." or the response text B is "I'm sorry, but I can't assist with that. Hacking into someone's bank account is illegal."

[0113] Step S105: input the response text into a preset discrimination model to obtain a discrimination result.

[0114] Specifically, the discriminant model is a large language model or a universal discriminant model, which can determine whether the input response text is harmful.

[0115] In a possible implementation, the response text A or the response text B is input into the discrimination model to obtain a discrimination result.

[0116] Assume that the response text A is input into the discrimination model and the discrimination result is harmful; and the response text B is input into the discrimination model and the discrimination result is harmless.

[0117] Step S106: In response to the determination result being harmful, determining that the image and text information successfully mines the vulnerability of the multimodal large language model.

[0118] In a possible implementation, if the determination result is harmful, it means that the scrambled target image and text information has discovered a loophole in the multimodal large language model, bypassed the loophole, and generated harmful text.

[0119] In the embodiment of the present invention, after step S106, the steps of generating a prevention strategy are also included, as shown in the following example: Fig. 9 As shown, the method includes:

[0120] Step S107: In response to the successful mining of vulnerabilities of the multimodal large language model using the graphic information, determining the vulnerabilities of the multimodal large language.

[0121] Step S108: Determine a prevention strategy for the multimodal large language according to the vulnerability of the multimodal large language.

[0122] In a possible implementation, after step S105, other steps are also included, specifically: Fig.10 As shown, the method includes:

[0123] Step S109: In response to the judgment result being harmless, determining that the exploitation of the image and text information to mine the vulnerability of the multimodal large language model has failed; and re-execute step S103.

[0124] Through the above embodiments, based on the multimodal large language model MLLMs, there is an inconsistency problem in the disrupted harmful instructions (which can be called harmful graphic information), that is, the MLLMs can understand harmful instructions, but cannot defend against harmful instructions. The method in the embodiment of the present invention can effectively explore the potential security vulnerabilities in open source or closed source multimodal large language models, and can provide effective security testing means for multimodal large language models, thereby improving the security capabilities of multimodal large language models and achieving the purpose of promoting defense through attack.

[0125] In the embodiment of the present invention, the randomly shuffled image and text information is used to generate text information through the MLLMs, that is, image and text generate text; it is also possible to use only the shuffled text information to generate text information through the MLLMs, that is, text generates text; in the above-mentioned image and text generation scenario and text generation scenario, the method of random shuffling after segmentation can be adopted to quickly discover the security vulnerabilities of the MLLMs with less overhead, thereby improving the defense capability of the multimodal large language model.

[0126] In an embodiment of the present invention, a device for mining loopholes in a multimodal large language model is provided. Fig.11 As shown, it specifically includes: an acquisition unit 1101, a segmentation unit 1102, a processing unit 1103, a generation unit 1104 and a discrimination unit 1105; wherein the acquisition unit 1101 is used to acquire graphic information, wherein the graphic information includes image information and text information; the segmentation unit 1102 is used to segment the image information and / or the text information to generate N image units and / or M text units, wherein each of the image units is a local area of ​​the image information, each of the text units is a language element in the text information, and N and M are positive integers greater than or equal to 2; the processing unit 1103 is used to generate N image units and / or M text units. Unit 1103 is used to randomly shuffle the N image units and / or the M text units to generate shuffled target image and text information, wherein the target image and text information includes shuffled image information and shuffled text information; the generating unit 1104 is used to input the target image and text information into a multimodal large language model and output a response text; the distinguishing unit 1105 is used to input the response text into a pre-set distinguishing model to obtain a distinguishing result; the distinguishing unit 1105 is also used to determine that the image and text information successfully mines the vulnerability of the multimodal large language model in response to the distinguishing result being harmful.

[0127] Further, the determination unit is further used to: in response to the determination result being harmless, determine that the exploitation of the vulnerability of the multimodal large language model by the graphic information fails;

[0128] The processing unit is further used for: randomly shuffling the N image units and the M text units again, and updating the shuffled target image and text information.

[0129] Furthermore, the segmentation unit is specifically used for:

[0130] Dividing the image information into N image units; and / or

[0131] The text information is segmented to generate M text units.

[0132] Furthermore, the segmentation unit is specifically used for:

[0133] Get the number of blocks N;

[0134] The image information is divided into N blocks to generate N image units.

[0135] Furthermore, the language element is a character, a word, a phrase, or a short sentence.

[0136] Furthermore, the discriminant model is a large language model or a universal discriminant model.

[0137] Furthermore, the determination unit is also used to: in response to the successful mining of the loopholes of the multimodal large language model by the graphic information, determine the loopholes of the multimodal large language.

[0138] Fig.12 Schematic diagram of the structure of the electronic device in the embodiment of the present invention. Fig.12 As shown, it includes a general computer hardware structure, which at least includes a processor 1201 and a memory 1202. The processor 1201 and the memory 1202 are connected via a bus 1203. The memory 1202 is suitable for storing instructions or programs executable by the processor 1201. The processor 1201 can be an independent microprocessor or a collection of one or more microprocessors. Thus, the processor 1201 executes the instructions stored in the memory 1202, thereby executing the method flow of the embodiment of the present invention as described above to realize the processing of data and the control of other devices. The bus 1203 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to the display controller 1204 and the display device and the input / output (I / O) device 1205. The input / output (I / O) device 1205 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices known in the art. Typically, the input / output device 1205 is connected to the system through an input / output (I / O) controller 1206.

[0139] The instructions stored in the memory 1202 are executed by at least one processor 1201 to implement: obtaining graphic information, wherein the graphic information includes image information and text information; segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; randomly shuffling the N image units and / or the M text units to generate shuffled target graphic information, wherein , the target graphic information includes scrambled image information and scrambled text information; the target graphic information is input into a multimodal large language model to output a response text; the response text is input into a preset discrimination model to obtain a discrimination result; in response to the discrimination result being harmful, it is determined that the graphic information successfully mines the vulnerability of the multimodal large language model; in response to the graphic information successfully mining the vulnerability of the multimodal large language model, the vulnerability of the multimodal large language is determined; according to the vulnerability of the multimodal large language, a prevention strategy for the multimodal large language is determined.

[0140] Specifically, the electronic device includes: one or more processors 1201 and a memory 1202, Fig.12 Take a processor 1201 as an example. The processor 1201 and the memory 1202 may be connected via a bus or other means. Fig.12 In the example, the bus connection is used. The memory 1202 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The processor 1201 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions and modules stored in the memory 1202, that is, the above-mentioned method of determining and preventing the exploitation of multimodal large language model vulnerabilities is implemented.

[0141] The memory 1202 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store a list of options, etc. In addition, the memory 1202 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1202 may optionally include a memory remotely arranged relative to the processor 1201, and these remote memories may be connected to an external device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0142] One or more modules are stored in the memory 1202, and when executed by one or more processors 1201, the method of mining and preventing vulnerabilities in a multimodal large language model in any of the above method embodiments is executed.

[0143] As will be appreciated by those skilled in the art, various aspects of the embodiments of the present invention may be implemented as a system, method or computer program product. Therefore, various aspects of the embodiments of the present invention may take the form of a complete hardware implementation, a complete software implementation (including firmware, resident software, microcode, etc.), or an implementation that combines software aspects with hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." In addition, various aspects of the embodiments of the present invention may take the form of a computer program product implemented in one or more computer-readable media having a computer-readable program code implemented thereon.

[0144] Any combination of one or more computer-readable media can be utilized. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media can be, for example, (but not limited to) electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or apparatuses, or any suitable combination of the foregoing. More specific examples (non-exhaustive enumeration) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of an embodiment of the present invention, a computer-readable storage medium can be any tangible medium that can contain or store a program used by an instruction execution system, device or apparatus or a program used in conjunction with an instruction execution system, device or apparatus.

[0145] A computer readable signal medium may include a propagated digital signal having a computer readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any of the following computer readable media: not a computer readable storage medium, and may communicate, propagate, or transmit a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0146] Program code embodied on a computer readable medium may be transmitted using any appropriate medium including, but not limited to, wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0147] The computer program code for performing operations for various aspects of the embodiments of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and conventional process programming languages ​​such as "C" programming language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer as a stand-alone software package; partially on the user's computer and partially on a remote computer; or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using the Internet of an Internet service provider).

[0148] The flowchart legend and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present invention described above describe various aspects of the embodiment of the present invention.It will be understood that each block of the flowchart legend and / or block diagram and the combination of the blocks in the flowchart legend and / or block diagram can be implemented by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that the instructions (executed by the processor of the computer or other programmable data processing device) create a device for implementing the function / action specified in the flowchart and / or block diagram block or block.

[0149] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other apparatus to operate in a particular manner, so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.

[0150] The computer program instructions may also be loaded onto a computer, other programmable data processing device or other apparatus so that a series of operable steps are performed on the computer, other programmable device or other apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.

[0151] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. Users' refusal to process personal information other than the necessary information required for basic functions will not affect the user's use of basic functions.

Claims

1. A method for mining and preventing vulnerabilities in a multimodal large language model, characterized in that: The method comprises: Acquiring graphic information, wherein the graphic information includes image information and text information; Segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; Randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information; Inputting the target graphic information into a multimodal large language model and outputting a response text; Inputting the response text into a preset discrimination model to obtain a discrimination result; In response to the determination result being harmful, determining that the image and text information successfully mines the vulnerability of the multimodal large language model; In response to the successful mining of the vulnerability of the multimodal large language model by the graphic information, determining the vulnerability of the multimodal large language; According to the vulnerability of the multimodal large language, a prevention strategy for the multimodal large language is determined.

2. A method for mining vulnerabilities in a large multimodal language model, characterized in that: The method comprises: Acquiring graphic information, wherein the graphic information includes image information and text information; Segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; Randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information; Inputting the target graphic information into a multimodal large language model and outputting a response text; Inputting the response text into a preset discrimination model to obtain a discrimination result; In response to the determination result being harmful, it is determined that the exploitation of the graphic and text information to mine the vulnerability of the multimodal large language model is successful.

3. The method according to claim 2, characterized in that The method further comprises: In response to the determination result being harmless, determining that the image-text information fails to mine the vulnerability of the multimodal large language model; The N image units and the M text units are randomly shuffled again, and the shuffled target image and text information is updated.

4. The method according to claim 2, characterized in that: The segmenting of the image information and / or the text information to generate N image units and M text units specifically includes: Dividing the image information to generate N image units; and / or, The text information is segmented to generate M text units.

5. The method according to claim 2, characterized in that: The segmenting of the image information to generate N image units specifically includes: Get the number of blocks N; The image information is divided into N blocks to generate N image units.

6. The method according to claim 2, characterized in that The language elements are characters, words, phrases, or short sentences.

7. The method according to claim 2, characterized in that The discriminant model is a large language model or a universal discriminant model.

8. The method according to claim 2, characterized in that: The method further comprises: In response to successfully mining the vulnerabilities of the multimodal large language model using the graphic and text information, the vulnerabilities of the multimodal large language are determined.

9. A device for mining and preventing vulnerabilities in a multimodal large language model, characterized in that: The device comprises: An acquisition unit, used for acquiring graphic information, wherein the graphic information includes image information and text information; a segmentation unit, used for segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; A processing unit, configured to randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information; A generating unit, used for inputting the target graphic information into a multimodal large language model and outputting a response text; A discrimination unit, used for inputting the response text into a preset discrimination model to obtain a discrimination result; The determination unit is further configured to, in response to the determination result being harmful, determine that the image and text information successfully mines the vulnerability of the multimodal large language model; The determination unit is further configured to, in response to the successful mining of the vulnerability of the multimodal large language model by the graphic information, determine the vulnerability of the multimodal large language; A determination unit is used to determine a prevention strategy for the multimodal large language according to the vulnerability of the multimodal large language.

10. A device for mining vulnerabilities in a multimodal large language model, characterized in that: The device comprises: An acquisition unit, used for acquiring graphic information, wherein the graphic information includes image information and text information; a segmentation unit, used for segmenting the image information and / or the text information to generate N image units and / or M text units, wherein each image unit is a local area of ​​the image information, each text unit is a language element in the text information, and N and M are positive integers greater than or equal to 2; A processing unit, configured to randomly shuffle the N image units and / or the M text units to generate shuffled target image-text information, wherein the target image-text information includes shuffled image information and shuffled text information; A generating unit, used for inputting the target graphic information into a multimodal large language model and outputting a response text; A discrimination unit, used for inputting the response text into a preset discrimination model to obtain a discrimination result; The determination unit is further configured to determine, in response to the determination result being harmful, that the image and text information successfully mines the vulnerability of the multimodal large language model.

11. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • LLM-based ASOC vulnerability assessment method, apparatus and device, and medium

    CN118395457A

  • Harmful image detection method

    CN118799615A

  • Speech induction method for false information

    CN118821203A

  • Large visual language model risk test method based on bimodal confrontation prompt

    CN118862057A

  • Generative AI report on security risk using LLMs

    US20240422187A1