Capability information determination method and device of multi-modal large language model, server and storage medium
By generating adversarial images and comparing bounding boxes, the visual localization capability of a multimodal large language model is evaluated, which solves the problem of lack of capability information determination in existing technologies and improves the model's resistance to attacks in visual localization tasks.
Patent Information
- Application Number
- CN202410553361.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2025-11-07
AI Technical Summary
Existing multimodal large language models lack effective methods for determining capability information when performing visual localization tasks, making it impossible to assess their robustness against adversarial attacks, especially in visual localization tasks other than image captioning or question answering.
By acquiring adversarial images of the initial image, an attack is launched against a multimodal large language model using these adversarial images. The output is a target image labeled with bounding boxes. The model's capabilities are evaluated by comparing the bounding boxes of the target image and the initial image. Specific methods include adjusting embedded features or predicted bounding boxes to generate adversarial images, and calculating the intersection-over-union ratio (IoU) or distance of the bounding boxes to assess the model's robustness.
This study effectively evaluates the adversarial robustness of multimodal large language models in visual localization tasks, identifies the degree to which the model output is affected by the input, and improves the model's resistance to attacks in visual localization tasks.
Smart Images

Figure CN120913012A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a capability information determination method and device of a multi-modal large language model, a server and a storage medium. BACKGROUND
[0002] The multi-modal large language model is a powerful artificial intelligence model that can process various modalities of content such as text, pictures, videos or voice input and output corresponding information. Currently, the multi-modal large language model faces the risk of adversarial attacks, that is, during the process of inputting prompt text and pictures, an attacker modifies the pictures and inputs the prompt text and the modified pictures into the multi-modal large language model to make the multi-modal large language model output incorrect information.
[0003] To cope with the above-mentioned risk of adversarial attacks, the current method makes the multi-modal large language model output corresponding text according to the input modified pictures and prompt text. Based on the text and a preset reference text, the adversarial robustness of the multi-modal large language model, i.e., the capability information of the multi-modal large language model, is determined. Through the above-mentioned acquisition of the capability information, the capability of the multi-modal large language model to cope with the risk of adversarial attacks can be evaluated.
[0004] The above-mentioned method is mainly aimed at tasks in which the output content is text such as generating subtitles for pictures or asking questions according to pictures. However, with the development of the multi-modal large language model, the multi-modal large language model has been able to perform a visual positioning task, i.e., outputting a picture with positioning information of a preset object according to input prompt text and pictures. Therefore, the current method is not applicable to determining the capability information of the multi-modal large language model whose output content is pictures, and there is a lack of a capability information determination method of the multi-modal large language model to evaluate the capability of the multi-modal large language model for performing the visual positioning task. SUMMARY
[0005] The embodiment of the present application provides a capability information determination method, device, server and storage medium of a multi-modal large language model to evaluate the capability of the multi-modal large language model for performing a visual positioning task. The technical solution is as follows:
[0006] In a first aspect, a method for determining capability information of a multi-modal large language model is provided. The method includes: obtaining a prompt text and an initial image, the prompt text being used for positioning a preset object; obtaining an adversarial image of the initial image based on the initial image, the adversarial image being different from the initial image; inputting the prompt text and the adversarial image into a first multi-modal large language model to output a target image, the target image being labeled with a bounding box, the bounding box being used to indicate a position of the preset object in the target image; and determining capability information of the first multi-modal large language model based on a reference bounding box of the initial image and the bounding box of the target image, the capability information indicating a degree to which an output of the first multi-modal large language model is affected by an input.
[0007] In the method for determining capability information of a multi-modal large language model, the adversarial image is used to attack the first multi-modal large language model to make the first multi-modal large language model output the target image. Since the first multi-modal large language model uses the bounding box of the target image to indicate a position of the preset object predicted by the first multi-modal large language model to implement a visual positioning task, and the reference bounding box of the initial image indicates an actual position of the preset object, the capability of the multi-modal large language model for performing the visual positioning task can be evaluated by comparing the bounding box of the target image and the reference bounding box of the initial image.
[0008] In some embodiments, the obtaining, based on the initial image, the adversarial image of the initial image includes: adjusting, according to an embedding feature of the initial image, an embedding feature of a preset adversarial image to increase a distance between the embedding feature of the preset adversarial image and the embedding feature of the initial image; and generating the adversarial image based on the adjusted embedding feature.
[0009] In some embodiments, the adjusting, according to the embedding feature of the initial image, the embedding feature of the preset adversarial image includes: obtaining a distance between the embedding feature of the initial image and an embedding feature of the preset adversarial image; and adjusting the embedding feature of the preset adversarial image based on the distance.
[0010] In some embodiments, the obtaining, based on the initial image, the adversarial image of the initial image includes: obtaining, by a second multi-modal large language model, a first predicted bounding box of the initial image and a first loss value, the first loss value being used to represent a difference between the first predicted bounding box of the initial image and a reference bounding box of the initial image; adjusting the first predicted bounding box of the initial image based on the first loss value to make a difference between the adjusted first predicted bounding box and the reference bounding box larger; and generating the adversarial image based on the adjusted first predicted bounding box.
[0011] In some embodiments, the obtaining the adversarial image based on the initial image comprises: obtaining, by the second multi-modal large language model, a second predicted bounding box of the initial image and a second loss value, the second loss value being used to represent a difference between the second predicted bounding box of the initial image and a target bounding box of the initial image, the target bounding box indicating an object other than the preset object in the initial image; adjusting the second predicted bounding box of the initial image based on the second loss value, so that a difference between the adjusted second predicted bounding box and the target bounding box is smaller; and generating the adversarial image based on the adjusted second predicted bounding box.
[0012] In some embodiments, the obtaining the target bounding box comprises: obtaining a bounding box of the kth preset object indicated by the prompt text, k being an integer greater than 1; and determining the bounding box of the kth preset object as a target bounding box of the (k-1)th preset object indicated by the prompt text.
[0013] In some embodiments, the determining the capability information of the first multi-modal large language model based on the reference bounding box of the initial image and the bounding box of the target image comprises: calculating an intersection-over-union between the bounding box of the target image and the reference bounding box of the initial image, the intersection-over-union being a ratio of an area of an intersection between the bounding box of the target image and the reference bounding box to an area of a union between the bounding box of the target image and the reference bounding box; if the intersection-over-union is greater than or equal to an intersection-over-union threshold, an evaluation index of the capability information is incremented by 0; and if the intersection-over-union is less than the intersection-over-union threshold, the evaluation index of the capability information is incremented by 1.
[0014] In a second aspect, a device for determining capability information of a multi-modal large language model is provided. The device comprises: an image obtaining module configured to obtain a prompt text and an initial image, the prompt text being used to locate a preset object; an adversarial image obtaining module configured to obtain an adversarial image of the initial image based on the initial image, the adversarial image being different from the initial image; a target image obtaining module configured to input the prompt text and the adversarial image into a first multi-modal large language model, and output a target image, the target image being labeled with a bounding box, the bounding box being used to indicate a position of the preset object in the target image; and a capability information determining module configured to determine capability information of the first multi-modal large language model based on a reference bounding box of the initial image and the bounding box of the target image, the capability information indicating a degree to which an output of the first multi-modal large language model is affected by an input.
[0015] In some embodiments, the adversarial image obtaining module comprises: a first adjusting unit configured to adjust an embedding feature of a preset adversarial image based on an embedding feature of the initial image, so that a distance between the embedding feature of the preset adversarial image and the embedding feature of the initial image is increased; and a first adversarial image generating unit configured to generate the adversarial image based on the adjusted embedding feature.
[0016] In some embodiments, the first adjusting unit is configured to: obtain a distance between the embedding feature of the initial picture and the embedding feature of the preset adversarial picture; and adjust the embedding feature of the preset adversarial picture based on the distance.
[0017] In some embodiments, the adversarial picture obtaining module is configured to: obtain, by the second multi-modal large language model, a first predicted bounding box of the initial picture and a first loss value, the first loss value being used to represent a difference between the first predicted bounding box of the initial picture and a reference bounding box of the initial picture; adjust the first predicted bounding box of the initial picture based on the first loss value, so that a difference between the adjusted first predicted bounding box and the reference bounding box becomes larger; and generate the adversarial picture based on the adjusted first predicted bounding box.
[0018] In some embodiments, the adversarial picture obtaining module is configured to: obtain, by the second multi-modal large language model, a second predicted bounding box of the initial picture and a second loss value, the second loss value being used to represent a difference between the second predicted bounding box of the initial picture and a target bounding box of the picture, the target bounding box indicating an object other than the preset object in the initial picture; adjust the second predicted bounding box of the initial picture based on the second loss value, so that a difference between the adjusted second predicted bounding box and the target bounding box becomes smaller; and generate the adversarial picture based on the adjusted second predicted bounding box.
[0019] In some embodiments, the adversarial picture obtaining module is configured to obtain the target bounding box, and the obtaining process comprises: obtaining a bounding box of a kth preset object indicated by the prompt text, k being an integer greater than 1; and determining the bounding box of the kth preset object as a target bounding box of a (k-1)th preset object indicated by the prompt text.
[0020] In some embodiments, the capability information determining module is configured to: calculate an intersection-over-union between the bounding box of the target picture and the reference bounding box of the initial picture, the intersection-over-union being a ratio of an area of an intersection between the bounding box of the target picture and the reference bounding box to an area of a union between the bounding box of the target picture and the reference bounding box; if the intersection-over-union is greater than or equal to an intersection-over-union threshold, the evaluation index of the capability information is incremented by 0; and if the intersection-over-union is less than the intersection-over-union threshold, the evaluation index of the capability information is incremented by 1.
[0021] In a third aspect, a server is provided, which includes a processor and a memory, the memory being configured to store at least one piece of computer program, the at least one piece of computer program being loaded and executed by the processor to implement the operations performed by the capability information determining method of the multi-modal large language model provided in the above first aspect or various optional implementation manners of the first aspect.
[0022] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores at least one piece of computer program. The at least one piece of computer program is loaded and executed by a processor to implement the operations performed by the capability information determination method of the multi-modal large language model provided in the first aspect or various optional implementation manners of the first aspect.
[0023] In a fifth aspect, a computer program product or computer program is provided, and the computer program product or computer program includes computer program code stored in a computer-readable storage medium. The computer program code is read by a processor of a server from the computer-readable storage medium, and the processor executes the computer program code to enable the server to perform the capability information determination method of the multi-modal large language model provided in the first aspect or various optional implementation manners of the first aspect.
[0024] On the basis of the implementation manners of the above aspects provided by the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0026] Figure 1 is an implementation environment schematic diagram of a capability information determination method of a multi-modal large language model provided by an embodiment of the present application;
[0027] Figure 2 is a flowchart of a capability information determination method of a multi-modal large language model provided by an embodiment of the present application;
[0028] Figure 3 is a flowchart of another capability information determination method of a multi-modal large language model provided by an embodiment of the present application;
[0029] Figure 4 is a flowchart of another capability information determination method of a multi-modal large language model provided by an embodiment of the present application;
[0030] Figure 5 is an attack result schematic diagram of a targetless adversarial attack provided by an embodiment of the present application;
[0031] Figure 6 is a flowchart of another capability information determination method of a multi-modal large language model provided by an embodiment of the present application;
[0032] Figure 7 is an attack result schematic diagram of an exclusive target adversarial attack provided by an embodiment of the present application;
[0033] Figure 8 is another capability information determination method flowchart of a multi-modal large language model provided by an embodiment of the present application;
[0034] Figure 9 is an attack result schematic diagram of a permutation target adversarial attack provided by an embodiment of the present application;
[0035] Figure 10 is a structural block diagram of a capability information determination apparatus of a multi-modal large language model provided by an embodiment of the present application;
[0036] Figure 11 is a structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0038] In the present application, the terms "first", "second", etc. are used to distinguish the same or similar items with basically the same function and purpose, and it should be understood that there is no logical or time sequence relationship between "first", "second", "nth", and the number and execution order are not limited.
[0039] In the present application, the term "at least one" means one or more, and the term "multiple" means two or more.
[0040] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the pictures involved in the present application are obtained under full authorization.
[0041] The capability information determination method of a multi-modal large language model provided by an embodiment of the present application involves computer vision and natural language processing in the field of artificial intelligence, for example, related technologies based on computer vision are used to extract features of a picture to obtain embedding features of the picture, and for another example, related technologies based on natural language processing are used to process prompt text.
[0042] AI (Artificial Intelligence) is the theory, method, technology and application system that use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0043] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0044] CV (Computer Vision) is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify and measure targets and other machine vision, and further to do image processing, so that the computer processing becomes more suitable for human eye observation or image transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, trying to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology brings important changes to the development of computer vision technology. Swin-transformer, ViT (Vision Transformer), V-MOE, MAE (Masked Autoencoders), and other pre-training models in the field of vision can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (Three-Dimensional) technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.
[0045] NLP (Nature Language processing, natural language processing) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing involves natural language, i.e. the language used in daily life, and is closely related to linguistic research, as well as computer science and mathematics. The pre-training model, an important technology in artificial intelligence model training, is developed from the LLM (Large Language Model) in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph, etc.
[0046] Figure 1 is an implementation environment schematic diagram of a method for determining the capability information of a multi-modal large language model provided by an embodiment of the present application, see Figure 1 The implementation environment includes a terminal 101 and a server 102. Figure 1 The number of terminals and servers shown is only an example. For example, the number of terminals and servers can both be multiple, and the embodiments of the present application do not limit this. The implementation environment is used to call a multi-modal large language model to perform a visual positioning task, i.e. to obtain the position of a preset object in a picture in the prompt text according to the prompt text and the picture. The multi-modal large language model for performing the visual positioning task is widely used in various fields of life, such as intelligent customer service, intelligent marketing, role playing, advertising copywriting, script creation data analysis, content analysis, automatic driving field, robot field and face payment field. For example, in the automatic driving field, by calling the multi-modal large language model to perform the visual positioning task, the objects in the environment around the vehicle can be positioned to determine the driving plan. In the robot field, the multi-modal large language model can be applied to service robots such as food delivery robots or sweeping robots, so that the robot can interact with the surrounding environment through visual positioning. In the face payment field, by calling the multi-modal large language model to perform the visual positioning task, the face can be accurately positioned to realize the function of paying by face.
[0047] In some embodiments, the terminal 101 receives a picture and corresponding prompt text, and sends the picture and the corresponding prompt text to the server 102. The server 102 invokes the multi-modal large language model to process the picture and the corresponding prompt text, and outputs the position of the preset object in the picture to the terminal 101. The multi-modal large language model capability information determination method provided by the embodiments of the present application can be executed by the server 102 based on the picture, the prompt text, and the position of the preset object in the picture output by the multi-modal large language model, or can be executed by the terminal 101 based on the received picture, the prompt text, and the position of the preset object in the picture obtained by interacting with the server 102.
[0048] The server can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0049] Figure 2 is a flow chart of a multi-modal large language model capability information determination method provided by the embodiments of the present application, as shown in Figure 2 The method includes the following steps:
[0050] 201, the server obtains a prompt text and an initial picture, and the prompt text is used to locate a preset object.
[0051] The initial picture contains at least one object, which can be a person, an animal, or an object, etc., which is not limited in the embodiments of the present application. The prompt text contains a description of at least one preset object, for example, a certain prompt text is “a blue and yellow bus”. The preset object can be an object in the initial picture, or can not be an object in the initial picture.
[0052] The above-mentioned locating the preset object refers to obtaining information representing a position, which is the position of the preset object in the picture. The position information can be represented in the form of a bounding box. The bounding box can be represented in the form of a geometric box, or in the form of text, which is called bounding box text. The bounding box text is used to record the position of the bounding box in the picture, such as the coordinates of the bounding box in the picture, the coordinates of the preset point in the bounding box, or the number of the bounding box, etc.
[0053] The embodiments of the present application are described by taking the identification of the preset object in the picture through the bounding box as an example. In some embodiments, the preset object in the picture can be identified in other manners, for example, the preset object in the picture is identified by using an arrow, and the like, which is not limited in the embodiments of the present application.
[0054] 202. The server obtains, based on the initial picture, an adversarial picture of the initial picture, the adversarial picture being different from the initial picture.
[0055] The purpose of the adversarial picture is to mislead the multimodal large language model, so that the multimodal large language model outputs an incorrect picture according to the adversarial picture, the incorrect picture being a picture labeled with an incorrect bounding box, the incorrect bounding box indicating an incorrect position of the preset object in the adversarial picture. The adversarial picture being different from the initial picture means that there are some different pixel points between the adversarial picture and the initial picture, wherein the difference between the pixel value of any pixel point in the adversarial picture and the pixel value of the pixel point in the initial picture is within a preset modification range. For example, the pixel points of the initial picture are 237, 28, 148, 69 and 178, and the preset modification range is 16, and the corresponding pixel points in the adversarial picture can be 229, 41, 148, 58 and 163.
[0056] In the embodiments of the present application, the server adds invisible tiny adversarial perturbations on the basis of the initial picture or the preset adversarial picture to obtain the adversarial picture. The preset adversarial picture is similar to the initial picture, and can be obtained by preprocessing the initial picture. The preprocessing includes modifying the pixel values of the initial picture or cropping the initial picture. The adding of the invisible tiny adversarial perturbations means that the pixel values of the initial picture are modified according to a preset step within a preset modification range. The modification can be performed only once or multiple times through iteration. The preset modification range is the maximum modification range of any pixel value of the initial picture, referred to as the perturbation amplitude. The multiple pixel values can be all the pixel values of the initial picture or all the pixel values of a preset region in the initial picture, which is not limited in the embodiments of the present application. For example, the perturbation amplitude is 16, indicating that the absolute value of the difference between any pixel value of the adversarial picture and the corresponding pixel value of the initial picture is less than 16. The preset step is the amplitude of the modification of any pixel value of the initial picture each time. For example, the preset step is 1, indicating that the amplitude of the modification of any pixel value of the initial picture each time is 1.
[0057] 203. The server inputs the prompt text and the adversarial picture into the first multimodal large language model to output a target picture, the target picture being labeled with a bounding box, the bounding box being used to indicate the position of the preset object in the target picture.
[0058] The first multi-modal large language model is a multi-modal large language model capable of performing a visual positioning task, such as MiniGPT-v2. Taking the first multi-modal large language model as MiniGPT-v2 as an example, the model architecture of the first multi-modal large language model includes a visual backbone model, a linear projection layer, and a large language model. The visual backbone model is connected to the linear projection layer, and the linear projection layer is connected to the large language model. The visual backbone model is used to convert an input picture into a plurality of visual tokens, each visual token being a string. The linear projection layer is used to project the plurality of visual tokens from the visual backbone model to the large language model, and the large language model is used to process the received plurality of visual tokens and output a processing result. The training process of the first multi-modal large language model includes three stages. The first stage is pre-training, that is, the first multi-modal large language model is trained using a weakly labeled dataset and a fine-grained dataset. In the training process, the weakly labeled dataset is given a higher sampling rate, so that the first multi-modal large language model obtains more diverse knowledge. The second stage is multi-task training, that is, the first multi-modal large language model is trained only using the fine-grained dataset to improve the performance of the first multi-modal large language model in processing different tasks. The third stage is multi-modal instruction fine-tuning, that is, the first multi-modal large language model is adjusted using the fine-grained dataset and the instruction dataset. In the training process, the fine-grained dataset is given a lower sampling rate, and the instruction dataset is given a higher sampling rate, so as to improve the dialog ability of the first multi-modal large language model.
[0059] In the process of obtaining the target picture, the server calls an encoder in the first multi-modal large language model to process the adversarial picture to obtain embedding features of the adversarial picture, and performs word segmentation on the prompt text to obtain a plurality of words. The encoder is called to process the plurality of words to obtain embedding features of the prompt text. The first multi-modal large language model calculates the embedding features of the prompt text and the embedding features of the adversarial picture to obtain a plurality of candidate pictures and a probability of each candidate picture. Each candidate picture is labeled with a bounding box. The first multi-modal large language model determines the candidate picture with the highest probability as the target picture and outputs the target picture. The probability of each candidate picture indicates the possibility of the first multi-modal large language model outputting the candidate picture.
[0060] In the above process, if the preset object is an object in the initial picture, the first multi-modal large language model can output the target picture based on the initial picture and the prompt text. If the preset object is not an object in the initial picture, the first multi-modal large language model outputs a positioning failure information indicating that the first multi-modal large language model fails to locate the preset object in the initial picture, and the visual positioning task fails.
[0061] 204、The server determines the capability information of the first multi-modal large language model based on the reference bounding box of the initial picture and the bounding box of the target picture, and the capability information indicates the degree to which the output of the first multi-modal large language model is affected by the input.
[0062] The reference bounding box is used to indicate the correct position of the preset object in the initial picture. The adversarial robustness of the multi-modal large language model is used to reflect the capability information of the multi-modal large language model, and the evaluation index of the adversarial robustness is a numerical value, for example, the evaluation index is 85 in a certain capability information acquisition process. The stronger the adversarial robustness, the smaller the degree to which the output of the multi-modal large language model is affected by the input. The weaker the adversarial robustness, the greater the degree to which the output of the multi-modal large language model is affected by the input. The greater the degree to which the output of the multi-modal large language model is affected by the input, the more likely it is that the multi-modal large language model will output an incorrect picture according to the adversarial picture of the input.
[0063] In the embodiments of the present application, the step 204 includes: for each of the preset number of initial pictures, the server calculates the IoU (intersection over union) between the bounding box of the target picture corresponding to the initial picture and the reference bounding box of the initial picture. The IoU is the ratio of the area of the intersection between the bounding box of the target picture and the reference bounding box to the area of the union between the bounding box of the target picture and the reference bounding box. If the IoU is greater than or equal to the intersection over union threshold, the evaluation index of the adversarial robustness is added by 0, and if the IoU is less than the intersection over union threshold, the evaluation index of the adversarial robustness is added by 1. The greater the evaluation index, the weaker the adversarial robustness, and the smaller the evaluation index, the stronger the adversarial robustness. For example, the intersection over union threshold is 50%, and the server uses 100 initial pictures to obtain the capability information of the first multi-modal large language model, of which 67 initial pictures correspond to an IoU less than 50%. The evaluation index corresponding to the first multi-modal large language model is 67.
[0064] In some embodiments, the server calculates a distance between the reference bounding box and the bounding box of the target picture, and takes the distance as an evaluation index to evaluate the adversarial robustness of the multi-modal large language model. The greater the distance, the weaker the adversarial robustness, and the smaller the distance, the stronger the adversarial robustness. Alternatively, for each of the preset number of initial pictures, the server calculates a distance between the reference bounding box of the initial picture and the bounding box of the target picture corresponding to the initial picture, and takes a sum of distances corresponding to the preset number of initial pictures or an average of distances corresponding to the preset number of initial pictures as an evaluation index. The greater the evaluation index, the weaker the adversarial robustness, and the smaller the evaluation index, the stronger the adversarial robustness. The distance can be a distance between a preset point in the reference bounding box and a preset point in the bounding box of the target picture, or a text distance between a bounding box text of the reference bounding box and a bounding box text of the target picture. The text distance between the two bounding box texts can be a distance between numerical values represented by the two bounding box texts, or a distance between embedding features of the two bounding box texts. The distance can be an Euclidean distance, and the embodiments of the present application are not limited in this regard.
[0065] In the method for determining the capability information of the multi-modal large language model provided by the embodiments of the present application, the first multi-modal large language model is attacked by using the adversarial picture to make the first multi-modal large language model output the target picture. Since the first multi-modal large language model uses the bounding box of the target picture to indicate the position of the preset object predicted by the first multi-modal large language model to implement the visual positioning task, and the reference bounding box of the initial picture indicates the actual position of the preset object, the capability of the multi-modal large language model for performing the visual positioning task can be evaluated by comparing the bounding box of the target picture and the reference bounding box of the initial picture.
[0066] The above content is combined with Figure 2 The method for evaluating the capability information of the multi-modal large language model provided by the embodiments of the present application is briefly described, and the following content is combined with Figures 3 to 9 The above method is specifically described. The following content relates to a plurality of different adversarial attack methods, and the capability information determination process of the multi-modal large language model under each adversarial attack method is described in detail. The plurality of different adversarial attack methods are divided into targetless adversarial attack and targeted adversarial attack. The targetless adversarial attack is divided into image embedding attack and text bounding box attack, and the targeted adversarial attack is divided into exclusive target adversarial attack and replacement target adversarial attack.
[0067] Wherein, the purpose of the no-target adversarial attack is to make the position predicted by the multimodal large language model deviate from the correct position of the preset object. For example, the position of the object in the picture is identified in the form of a bounding box. If the no-target adversarial attack is successful, the distance between the bounding box in the picture output by the multimodal large language model and the reference bounding box of the picture is greater than or equal to the distance threshold. If the no-target adversarial attack fails, the distance between the bounding box in the picture output by the multimodal large language model and the reference bounding box of the picture is less than the distance threshold. The following description of Figure 3 and Figure 4 respectively relates to the description of the image embedding attack and the text bounding box attack.
[0068] The purpose of the targeted adversarial attack is to make the position predicted by the multimodal large language model be the position of the preset error object in the picture, which is an object other than the preset object in the picture. For example, if the targeted adversarial attack is successful, the distance between the bounding box in the picture output by the multimodal large language model and the bounding box of the preset error object is less than the distance threshold. Wherein, the preset error object of the exclusive target adversarial attack is any object in the picture other than the preset object, and the preset error object of the replacement target adversarial attack is another preset object indicated by the prompt text, for example, the next preset object indicated by the prompt text. The purpose of the replacement target adversarial attack is to disturb the bounding boxes of all preset objects in the picture, so that the predicted bounding box of any preset object is the correct bounding box of another preset object. The following description of Figure 6 and Figure 7 respectively relates to the description of the exclusive target adversarial attack and the replacement target adversarial attack.
[0069] Figure 3 is another process flow chart of a method for determining the capability information of a multimodal large language model provided by an embodiment of the present application, which is used to describe the process of determining the capability information of the multimodal large language model under the image embedding attack. In this process, the process of obtaining the adversarial picture includes multiple iterations, and each iteration adjusts the adversarial picture after the last iteration until the iteration ends. As shown in Figure 3 , taking the server executing the above-mentioned method for determining the capability information of the multimodal large language model as an example, the method includes the following steps.
[0070] 301. The server obtains a prompt text and an initial picture, and the prompt text is used to locate a preset object.
[0071] This step 301 is the same as the above-mentioned step 201, and the embodiments of the present application will not be repeated here.
[0072] The following steps 302 to 308 are the process of obtaining the adversarial picture by the server through multiple iterations.
[0073] 302. In the first iteration process, the server obtains the distance between the embedding features of the initial picture and the preset adversarial picture through the multi-modal large language model.
[0074] The multi-modal large language model applied in step 302 and the first multi-modal large language model or the second multi-modal large language model involved in the embodiments of the present application can be the same model or different models.
[0075] In the embodiments of the present application, the server inputs the initial picture into the multi-modal large language model, calls the encoder in the multi-modal large language model, processes the initial picture, and outputs the embedding features of the initial picture. Similarly, the embedding features of the preset adversarial picture can be obtained. The server calculates the distance between the embedding features of the initial picture and the preset adversarial picture.
[0076] The embodiments of the present application are described by taking the multi-modal large language model outputting the embedding features of the initial picture and the embedding features of the preset adversarial picture and the server calculating the distance between the two embedding features as an example. In some embodiments, the multi-modal large language model calculates the distance between the embedding features of the initial picture and the preset adversarial picture after obtaining the embedding features of the initial picture and the embedding features of the preset adversarial picture, and outputs the distance.
[0077] 303. The server adjusts the embedding features of the preset adversarial picture based on the distance to increase the distance between the embedding features of the preset adversarial picture and the embedding features of the initial picture.
[0078] In the embodiments of the present application, the function used to calculate the distance is called a distance function. The server calculates the gradient of the embedding features of the preset adversarial picture according to the distance function, the embedding features of the initial picture, and the embedding features of the preset adversarial picture. The gradient is a vector, and each element in the gradient corresponds to an element in the embedding features of the preset adversarial picture, indicating the adjustment direction of the element in the embedding features of the preset adversarial picture. If the element in the gradient is greater than zero, the corresponding element in the embedding features of the preset adversarial picture is increased. If the element in the gradient is less than zero, the corresponding element in the embedding features of the preset adversarial picture is decreased. If the element in the gradient is equal to zero, the corresponding element in the embedding features of the preset adversarial picture remains unchanged. The server adjusts the embedding features of the preset adversarial picture according to the gradient and a preset step, thereby increasing the distance between the embedding features of the preset adversarial picture and the embedding features of the initial picture.
[0079] For example, the embedding features of a certain preset adversarial picture are (3, 5, 7), the gradient of the embedding features of the preset adversarial picture is (2, 0, -1), and the preset step is 1. The adjusted embedding features of the preset adversarial picture are (4, 5, 6).
[0080] The steps 302 to 303 are a possible implementation of the server adjusting the embedding feature of the preset adversarial image according to the embedding feature of the initial image. In this possible implementation, the server adjusts the embedding feature of the preset adversarial image according to the preset step length, which can effectively control the modification range of the preset adversarial image, so that the final adversarial image meets the preset modification range and the multimodal large language model cannot recognize the adversarial image.
[0081] 304. The server generates the adversarial image after the first iteration based on the adjusted embedding feature, and the adversarial image is different from the initial image.
[0082] In the embodiment, the server performs back propagation calculation according to the adjusted embedding feature based on the structure and parameters of the encoder in the multimodal large language model to obtain the adversarial image after the first iteration. The multimodal large language model is the same model as the multimodal large language model used in step 302.
[0083] 305. If the first iteration does not meet the iteration end condition, the server performs the next iteration. In the ith iteration, the server obtains the distance between the embedding feature of the initial image and the embedding feature of the adversarial image after the (i-1)th iteration. If the first iteration meets the iteration end condition, the server stops the iteration, determines the adversarial image after the first iteration as the adversarial image of the initial image, and i is an integer greater than 1.
[0084] The iteration end condition includes at least one of the number of completed iterations reaching a preset iteration number and the adversarial image after iteration reaching a preset modification range.
[0085] In some embodiments, the iteration end condition is that the adversarial image after each iteration reaches a preset modification range. For example, the server obtains the absolute value of the difference between each pixel value of the initial image and each pixel value of the adversarial image after each iteration, with respect to the preset modification range for all pixel values of the initial image. If the adversarial image after each iteration reaches the preset modification range, the iteration end condition is met, the server stops iteration, modifies the pixel value after each iteration corresponding to the absolute value that exceeds the preset modification range, so that the adjustment of the pixel value is controlled within the preset modification range, and the modified adversarial image after each iteration is determined as the adversarial image of the initial image. For example, if the preset modification range is 16, a pixel value in the initial image is 260, and the pixel value after each iteration is 280, the server modifies the pixel value after each iteration to 276. The adversarial image after each iteration reaching the preset modification range can be that any absolute value is greater than or equal to the perturbation amplitude, the absolute value of the first number (which can be obtained by the number of all pixels in the initial image and the first proportion) is greater than or equal to the perturbation amplitude, or the average of multiple absolute values is greater than or equal to the absolute value threshold. If the adversarial image after each iteration reaches the preset modification range and the number of absolute values greater than the perturbation amplitude is greater than or equal to the number threshold, the iteration end condition is met, the server stops iteration, adjusts the preset adversarial image or the preset step, and reiterates. If the adversarial image after each iteration does not reach the preset modification range, the iteration end condition is not met, and the server performs the next iteration.
[0086] For example, the server obtains the absolute value of the difference between each pixel value of the initial image in the preset region and each pixel value of the adversarial image in the preset region after each iteration, with respect to the preset modification range for all pixel values of the initial image in the preset region. Whether the iteration end condition is met is determined according to the absolute value. The determination process is the same as that for all pixel values of the initial image, which will not be repeated here.
[0087] In some embodiments, the iteration end condition is that the number of completed iterations reaches a preset iteration number. After each iteration, if the number of completed iterations is greater than or equal to the preset iteration number, the iteration end condition is met, the server stops the iteration, if the adversarial image after the iteration does not reach the preset modification range, the server determines the adversarial image after the iteration as the adversarial image of the initial image; if the adversarial image after the iteration reaches the preset modification range, the server modifies the pixel value after the iteration corresponding to the absolute value of the preset modification range that exceeds the preset modification range, so that the adjustment of the pixel value is controlled within the preset modification range, and the modified adversarial image after the iteration is determined as the adversarial image of the initial image; if the adversarial image after the iteration reaches the preset modification range and the number of absolute values greater than the perturbation amplitude is greater than or equal to a number threshold, the server adjusts the preset adversarial image or the preset step size and reiterates. After each iteration, if the number of completed iterations is less than the preset iteration number, the iteration end condition is not met, and the server performs the next iteration.
[0088] In some embodiments, the iteration end condition includes that the number of completed iterations reaches a preset iteration number and that the adversarial image after the iteration reaches a preset modification range. After each iteration, if the adversarial image after the iteration reaches the preset modification range, the server stops the iteration, if the adversarial image after the iteration does not reach the preset modification range, the server determines whether the number of completed iterations is greater than or equal to the preset iteration number, if the number of completed iterations is greater than or equal to the preset iteration number, the server stops the iteration, if the number of completed iterations is less than the preset iteration number, the server starts the next iteration.
[0089] 306、The server adjusts the embedding feature of the adversarial image after the i-1th iteration based on the distance, so that the distance between the embedding feature of the adversarial image after the i-1th iteration and the embedding feature of the initial image increases.
[0090] 307、The server generates the adversarial image after the ith iteration based on the adjusted embedding feature.
[0091] The steps 305 to 307 are the same as the steps 302 to 304, and the embodiments of the present application will not be repeated here.
[0092] 308、If the ith iteration does not meet the iteration end condition, the server performs the next iteration, if the ith iteration meets the iteration end condition, the server determines the adversarial image after the ith iteration as the adversarial image of the initial image.
[0093] The step 308 is the same as the step 305, and the embodiments of the present application will not be repeated here.
[0094] The steps 302 to 308 are a possible implementation of the server obtaining the adversarial image of the initial image based on the initial image. In this possible implementation, the server obtains the final adversarial image by gradually increasing the distance between the embedding features of the initial image and the embedding features of the adversarial image through multiple iterations, which is equivalent to obtaining the adversarial image by destroying the embedding features of the initial image, and can make the attack of the adversarial image on the multimodal large language model more effective.
[0095] 309. The server inputs the prompt text and the adversarial image into the first multimodal large language model to output a target image, and the target image is labeled with a bounding box, and the bounding box is used to indicate the position of the preset object in the target image.
[0096] 310. The server determines the capability information of the first multimodal large language model based on the reference bounding box of the initial image and the bounding box of the target image, and the capability information indicates the degree to which the output of the first multimodal large language model is affected by the input.
[0097] The steps 309 to 310 are the same as the steps 203 to 204, and will not be described here.
[0098] In the method for determining the capability information of the multimodal large language model provided by the embodiments of the present application, the adversarial image corresponding to the image embedding attack is used to attack the first multimodal large language model to make the first multimodal large language model output a target image. Since the first multimodal large language model uses the bounding box of the target image to indicate the position of the preset object predicted by the first multimodal large language model to implement the visual positioning task, and the reference bounding box of the initial image indicates the actual position of the preset object, the capability of the multimodal large language model for performing the visual positioning task can be evaluated by comparing the bounding box of the target image and the reference bounding box of the initial image. Since the adversarial image corresponding to the image embedding attack is obtained by destroying the embedding features of the initial image, and the embedding features of the initial image have a greater impact on the performance of the visual positioning task, the attack of the multimodal large language model using this adversarial image is more effective, and thus the capability information of the multimodal large language model for performing the visual positioning task can be determined more accurately.
[0099] The above content is combined with Figure 3 The capability information determination process of the multimodal large language model under the image embedding attack is described in detail, Figure 4 is another method flowchart for determining the capability information of a multimodal large language model provided by the embodiments of the present application, which is used to describe the capability information determination process of the multimodal large language model under the text bounding box attack. In this process, the process of obtaining the adversarial image also includes multiple iterations. For example, Figure 4As shown, taking the method for determining the capability information of the server for executing the above multi-modal large language model as an example, the method comprises the following steps.
[0100] 401. The server obtains prompt text and an initial picture, and the prompt text is used to locate a preset object.
[0101] The step 401 is the same as the step 201 described above, and thus will not be described herein.
[0102] The following steps 402 to 408 are the process of obtaining the adversarial picture through multiple iterations by the server.
[0103] 402. In the first iteration process, the server obtains a first predicted bounding box of the initial picture and a first loss value through a second multi-modal large language model, and the first loss value is used to represent the difference between the first predicted bounding box of the initial picture and a reference bounding box of the initial picture.
[0104] The first predicted bounding box of the initial picture is a bounding box obtained by the second multi-modal large language model according to the initial picture, and is used to indicate the position of the preset object predicted by the second multi-modal large language model. The greater the first loss value, the greater the distance between the first predicted bounding box of the initial picture and the reference bounding box of the initial picture, and the greater the difference between the first predicted bounding box of the initial picture and the reference bounding box of the initial picture. The smaller the first loss value, the smaller the distance between the first predicted bounding box of the initial picture and the reference bounding box of the initial picture, and the smaller the difference between the first predicted bounding box of the initial picture and the reference bounding box of the initial picture. The model architecture and training process of the second multi-modal large language model are the same as those of the first multi-modal large language model, and thus will not be described herein.
[0105] In the embodiment of the application, the server inputs the initial picture and the prompt text into the second multi-modal large language model, calls the encoder in the second multi-modal large language model to process the initial picture to obtain the embedding features of the initial picture, performs word segmentation on the prompt text to obtain a plurality of words, and calls the encoder to process the plurality of words to obtain the embedding features of the prompt text. The server calls the second multi-modal large language model to calculate the embedding features of the prompt text and the embedding features of the initial picture to obtain a plurality of candidate first predicted bounding boxes, calculates the first loss value of each candidate first predicted bounding box according to a preset loss function, and the first loss value reflects the distance between each candidate first predicted bounding box and the reference bounding box. The second multi-modal large language model determines the candidate first predicted bounding box with the smallest first loss value as the first predicted bounding box, and outputs the first predicted bounding box and the first loss value of the first predicted bounding box.
[0106] In some embodiments, the preset loss function is a distance function, and the process of calculating the first loss value of each candidate first prediction bounding box according to the distance function is the same as the process of calculating the distance between the reference bounding box and the bounding box of the target picture in the above step 204. The embodiments of the present application will not be repeated here.
[0107] The embodiments of the present application are described by taking the second multi-modal large language model outputting the first loss value as an example. In some embodiments, the second multi-modal large language model converts the first loss value of each candidate first prediction bounding box into the probability of each candidate first prediction bounding box, determines the candidate first prediction bounding box with the maximum probability as the first prediction bounding box, and outputs the first prediction bounding box and the probability of the first prediction bounding box. The probability is also called the logarithmic likelihood of the autoregressive loss.
[0108] The embodiments of the present application are described by taking one prompt text indicating one preset object and one initial picture corresponding to one first prediction bounding box as an example. In some embodiments, one prompt text indicates multiple preset objects, and one initial picture corresponds to multiple first prediction bounding boxes. Correspondingly, each first prediction bounding box corresponds to one reference bounding box. The server calls the second multi-modal large language model to calculate the embedding features of the prompt text and the embedding features of the initial picture to obtain multiple candidate first prediction bounding box combinations. Each candidate first prediction bounding box combination includes multiple candidate first prediction bounding boxes. For each candidate first prediction bounding box combination, the first loss value of each candidate first prediction bounding box in the candidate first prediction bounding box combination is calculated according to the preset loss function. The second multi-modal large language model adds the first loss value of each candidate first prediction bounding box in the candidate first prediction bounding box combination to obtain the first loss value of the candidate first prediction bounding box combination. The second multi-modal large language model determines the candidate first prediction bounding box combination with the minimum first loss value as the first prediction bounding box combination, and outputs the first prediction bounding box combination and the first loss value of the first prediction bounding box combination.
[0109] 403、The server adjusts the first prediction bounding box of the initial picture based on the first loss value, so that the difference between the adjusted first prediction bounding box and the reference bounding box becomes larger.
[0110] In the embodiments of the present application, the server obtains a vector corresponding to the first predicted bounding box, which can be an embedding feature of the bounding box text of the first predicted bounding box. The embodiments of the present application do not limit this, and the vector corresponding to the reference bounding box is the same. The server calculates the gradient of the first predicted bounding box according to the preset loss function, the vector corresponding to the first predicted bounding box and the vector corresponding to the reference bounding box. The gradient is a vector, and each element in the gradient corresponds to an element in the vector corresponding to the first predicted bounding box, indicating the adjustment direction of the element in the vector corresponding to the first predicted bounding box. If the element in the gradient is greater than zero, the corresponding element in the vector is increased. If the element in the gradient is less than zero, the corresponding element in the vector is decreased. If the element in the gradient is equal to zero, the corresponding element in the vector remains unchanged. The server adjusts the vector according to the gradient and the preset step, thereby increasing the distance between the first predicted bounding box and the reference bounding box.
[0111] 404. The server generates an adversarial picture after the first iteration based on the adjusted first predicted bounding box. The adversarial picture is different from the initial picture.
[0112] In the embodiments of the present application, the server performs back propagation calculation based on the structure and parameters of the second multi-modal large language model according to the bounding box text of the adjusted first predicted bounding box or the vector corresponding to the adjusted first predicted bounding box, to obtain the adversarial picture after the first iteration.
[0113] 405. If the first iteration does not satisfy the iteration end condition, the server performs the next iteration. In the jth iteration, the server obtains the first predicted bounding box and the first loss value of the adversarial picture after the j-1th iteration through the second multi-modal large language model. If the first iteration satisfies the iteration end condition, the server stops iteration and determines the adversarial picture after the first iteration as the adversarial picture of the initial picture. j is an integer greater than 1.
[0114] 406. The server adjusts the first predicted bounding box of the adversarial picture after the j-1th iteration based on the first loss value, so that the difference between the adjusted first predicted bounding box and the reference bounding box becomes larger.
[0115] 407. The server generates an adversarial picture after the jth iteration based on the adjusted first predicted bounding box.
[0116] The steps 405 to 407 described above are the same as the steps 402 to 404 described above, and the embodiments of the present application will not be described here.
[0117] 408. If the jth iteration does not satisfy the iteration end condition, the server performs the next iteration. If the jth iteration satisfies the iteration end condition, the server determines the adversarial picture after the jth iteration as the adversarial picture of the initial picture.
[0118] The step 408 is the same as the step 308, and thus details are not repeated here.
[0119] The steps 402 to 408 are a possible implementation of the server obtaining the adversarial image of the initial image based on the initial image. In this possible implementation, the server obtains the final adversarial image by increasing the distance between the first predicted bounding box and the reference bounding box through multiple iterations, so that the second multi-modal large language model outputs the first predicted bounding box deviating from the reference bounding box according to the adversarial image.
[0120] 409. The server inputs the prompt text and the adversarial image into the first multi-modal large language model to output a target image, and the target image is labeled with a bounding box indicating the position of the preset object in the target image.
[0121] 410. The server determines the capability information of the first multi-modal large language model based on the reference bounding box of the initial image and the bounding box of the target image, and the capability information indicates the degree to which the output of the first multi-modal large language model is affected by the input.
[0122] The steps 409 to 410 are the same as the steps 203 to 204, and thus details are not repeated here.
[0123] In the method for determining the capability information of the multi-modal large language model provided by the embodiments of the present application, the corresponding adversarial image of the text bounding box attack is used to attack the first multi-modal large language model, so that the first multi-modal large language model outputs a target image. Since the first multi-modal large language model uses the bounding box of the target image to indicate the position of the preset object predicted by the first multi-modal large language model to implement the visual positioning task, and the reference bounding box of the initial image indicates the actual position of the preset object, the capability of the multi-modal large language model for performing the visual positioning task can be evaluated by comparing the bounding box of the target image and the reference bounding box of the initial image. Since the corresponding adversarial image of the text bounding box attack is obtained by increasing the distance between the first predicted bounding box and the reference bounding box, the first predicted bounding box deviates from the reference bounding box when the multi-modal large language model is attacked by the adversarial image, and the capability information of the multi-modal large language model when facing the text bounding box attack can be effectively determined according to the adversarial image.
[0124] The above Figure 3 and Figure 4 respectively describe the capability information determination process of the multi-modal large language model under the image embedding attack in the target-free adversarial attack and the text bounding box attack, Figure 5 is a schematic diagram of an attack result of a target-free adversarial attack provided by the embodiments of the present application, wherein the dashed box is a bounding box. As Figure 5As shown, the prompt text is "black triangle", and the multimodal large language model outputs a picture in which the bounding box is a black triangle, and the multimodal large language model outputs a picture in which the bounding box is empty of objects.
[0125] The above Figures 3 to 5 The process of evaluating the ability information of the multimodal large language model according to the adversarial pictures corresponding to the targetless adversarial attacks is described in detail below. Figures 6 to 9 The process of evaluating the ability information of the multimodal large language model according to the adversarial pictures corresponding to the targetless adversarial attacks is described in detail below.
[0126] Figure 6 is another process flowchart of a method for determining the ability information of a multimodal large language model provided by the embodiments of the present application, which is used to describe the process of determining the ability information of a multimodal large language model under exclusive target adversarial attacks. In this process, the process of obtaining adversarial pictures also includes multiple iterations. As shown in Figure 6 The method includes the following steps.
[0127] 601. The server obtains a prompt text and an initial picture, and the prompt text is used to locate a preset object.
[0128] This step 601 is the same as the above step 201, and the embodiments of the present application will not be repeated here.
[0129] The following steps 602 to 608 are the process of obtaining adversarial pictures by the server through multiple iterations.
[0130] 602. In the first iteration process, the server obtains a second predicted bounding box and a second loss value of the initial picture through a second multimodal large language model, and the second loss value is used to represent the difference between the second predicted bounding box of the initial picture and a target bounding box of the initial picture. The target bounding box indicates an object other than the preset object in the initial picture.
[0131] The second prediction bounding box of the initial picture is a bounding box predicted by the second multi-modal large language model according to the initial picture, and is used to indicate the position of the preset object predicted by the second multi-modal large language model. The target bounding box indicates the preset error object in the initial picture. For example, the initial picture contains a person, a table and a chair, the preset object is the person in the initial picture, and the target bounding box indicates the table in the initial picture. The embodiments of the present application are described by taking one prompt text indicating one preset object and one initial picture corresponding to one second prediction bounding box as an example. In some embodiments, one prompt text indicates multiple preset objects, and one initial picture corresponds to multiple second prediction bounding boxes. Correspondingly, each second prediction bounding box corresponds to a target bounding box. The target bounding boxes corresponding to different second prediction bounding boxes can be the same or different.
[0132] The step 602 is the same as the step 402 described above, and the embodiments of the present application will not be described here again.
[0133] 603. The server adjusts the second prediction bounding box of the initial picture based on the second loss value, so that the difference between the adjusted second prediction bounding box and the target bounding box is smaller.
[0134] In the embodiments of the present application, the server obtains a vector corresponding to the second prediction bounding box. The vector can be an embedding feature of the bounding box text of the second prediction bounding box, and the embodiments of the present application do not limit this. The vector corresponding to the target bounding box is the same. The server calculates the gradient of the second prediction bounding box according to the preset loss function, the vector corresponding to the second prediction bounding box and the vector corresponding to the target bounding box. The gradient is a vector, and each element in the gradient corresponds to an element in the vector corresponding to the second prediction bounding box, indicating the adjustment direction of the element in the vector corresponding to the second prediction bounding box. If the element in the gradient is greater than zero, the corresponding element in the vector is reduced. If the element in the gradient is less than zero, the corresponding element in the vector is increased. If the element in the gradient is equal to zero, the corresponding element in the vector remains unchanged. The server adjusts the vector according to the gradient and the preset step length, thereby reducing the distance between the second prediction bounding box and the target bounding box.
[0135] 604. The server generates an adversarial picture after the first iteration based on the adjusted second prediction bounding box. The adversarial picture is different from the initial picture.
[0136] The step 604 described above is the same as the step 404 described above, and the embodiments of the present application will not be described here again.
[0137] 605、If the first iteration does not satisfy the iteration end condition, the server performs the next iteration. In the mth iteration, the server obtains, by the second multi-modal large language model, a second predicted bounding box and a second loss value of the adversarial picture after the (m-1)th iteration. If the first iteration satisfies the iteration end condition, the server stops the iteration, determines the adversarial picture after the first iteration as the adversarial picture of the initial picture, and m is an integer greater than 1.
[0138] 606、The server adjusts the second predicted bounding box of the adversarial picture after the (m-1)th iteration based on the second loss value, so that the difference between the adjusted second predicted bounding box and the target bounding box becomes smaller.
[0139] 607、The server generates the adversarial picture after the mth iteration based on the adjusted second predicted bounding box.
[0140] The steps 605-607 described above are the same as the steps 602-604 described above, and the embodiments of the present application will not be repeated here.
[0141] 608、If the mth iteration does not satisfy the iteration end condition, the server performs the next iteration. If the mth iteration satisfies the iteration end condition, the server determines the adversarial picture after the mth iteration as the adversarial picture of the initial picture.
[0142] The step 608 described above is the same as the step 308 described above, and the embodiments of the present application will not be repeated here.
[0143] The steps 602-608 described above are a possible implementation of the server obtaining the adversarial picture of the initial picture based on the initial picture. In this possible implementation, the server narrows the distance between the second predicted bounding box and the target bounding box through multiple iterations to obtain the final adversarial picture, which can enable the second multi-modal large language model to output the first predicted bounding box deviating from the reference bounding box and close to the target bounding box according to the adversarial picture.
[0144] 609、The server inputs the prompt text and the adversarial picture into the first multi-modal large language model to output a target picture. The target picture is labeled with a bounding box, and the bounding box is used to indicate the position of the preset object in the target picture.
[0145] 610、The server determines the capability information of the first multi-modal large language model based on the reference bounding box of the initial picture and the bounding box of the target picture. The capability information indicates the degree to which the output of the first multi-modal large language model is affected by the input.
[0146] The steps 609-610 described above are the same as the steps 203-204 described above, and the embodiments of the present application will not be repeated here.
[0147] The method for determining the capability information of the multi-modal large language model provided in the embodiments of the present application includes: performing an exclusive target adversarial attack on the first multi-modal large language model to make the first multi-modal large language model output a target picture. Since the first multi-modal large language model uses a bounding box of the target picture to indicate the position of a preset object predicted by the first multi-modal large language model to implement a visual positioning task, and a reference bounding box of the initial picture indicates the actual position of the preset object, the capability of the multi-modal large language model for performing the visual positioning task can be evaluated by comparing the bounding box of the target picture with the reference bounding box of the initial picture. Since the adversarial picture corresponding to the exclusive target adversarial attack is obtained by reducing the distance between the second predicted bounding box and the target bounding box, the first predicted bounding box obtained by performing the attack on the multi-modal large language model using the adversarial picture deviates from the reference bounding box and is close to the specified target bounding box, and the capability information of the multi-modal large language model when facing the exclusive target adversarial attack can be effectively determined according to the adversarial picture.
[0148] The above Figure 6 The capability information determination process of the multi-modal large language model under the exclusive target adversarial attack in the targeted adversarial attack is described, Figure 7 is a schematic diagram of an attack result of an exclusive target adversarial attack provided in the embodiments of the present application, wherein the dashed box is a bounding box. As Figure 7 indicated, the prompt text is "the black rectangle on the rightmost side", the target bounding box indicates the black triangle on the leftmost side, in the picture output by the multi-modal large language model according to the initial picture, the black rectangle on the rightmost side is in the bounding box, and in the picture output by the multi-modal large language model according to the adversarial picture, the black triangle on the leftmost side is in the bounding box.
[0149] The above content is combined with Figure 6 The capability information determination process of the multi-modal large language model under the exclusive target adversarial attack is described in detail, Figure 8 is a flowchart of another method for determining the capability information of a multi-modal large language model provided in the embodiments of the present application, which is used to describe the capability information determination process of the multi-modal large language model under the replacement target adversarial attack. In this process, the process of obtaining the adversarial picture also includes multiple iterations. As Figure 8 indicated, taking the server performing the above method for determining the capability information of the multi-modal large language model as an example, the method includes the following steps.
[0150] 801. The server obtains a prompt text and an initial picture, and the prompt text is used to locate a preset object.
[0151] This step 801 is the same as the above step 201, and the embodiments of the present application will not be described here.
[0152] 802、The server obtains a target bounding box of the initial picture based on the prompt text.
[0153] In some embodiments, one prompt text corresponds to one preset object, the server obtains a bounding box of the preset object indicated by the kth prompt text, and takes the bounding box as a bounding box of the kth preset object, where k is an integer greater than 1. The server determines the preset object indicated by the (k-1)th prompt text as the (k-1)th preset object, and determines the bounding box of the kth preset object as a target bounding box of the (k-1)th preset object. For example, the initial picture contains a person, a table and a chair, the target bounding box corresponding to the person is the bounding box corresponding to the table, the target bounding box corresponding to the table is the bounding box corresponding to the chair, and the target bounding box corresponding to the chair is the bounding box corresponding to the person.
[0154] In other embodiments, one prompt text corresponds to multiple preset objects, and the server obtains a bounding box of the kth preset object indicated by the prompt text, and determines the bounding box of the kth preset object as a target bounding box of the (k-1)th preset object indicated by the prompt text.
[0155] The embodiments of the present application are described by taking the determination of the bounding box of the kth preset object as the target bounding box of the (k-1)th preset object as an example. In some embodiments, the server determines the bounding box of the (k-1)th preset object as the target bounding box of the kth preset object, which is not limited in the embodiments of the present application.
[0156] 803、In the first iteration process, the server obtains a second predicted bounding box of the initial picture and a second loss value through the second multi-modal large language model, the second loss value is used to represent the difference between the second predicted bounding box of the initial picture and a target bounding box of the initial picture, and the target bounding box indicates an object other than the preset object in the initial picture.
[0157] 804、The server adjusts the second predicted bounding box of the initial picture based on the second loss value, so that the difference between the adjusted second predicted bounding box and the target bounding box becomes smaller.
[0158] 805、The server generates an adversarial picture after the first iteration based on the adjusted second predicted bounding box, and the adversarial picture is different from the initial picture.
[0159] 806、If the first iteration does not satisfy the iteration end condition, the server performs the next iteration, and in the nth iteration process, the server obtains a second predicted bounding box of the adversarial picture after the (n-1)th iteration and a second loss value through the second multi-modal large language model, and if the first iteration satisfies the iteration end condition, the server stops iteration, determines the adversarial picture after the first iteration as the adversarial picture of the initial picture, and n is an integer greater than 1.
[0160] 807、The server adjusts the second predicted bounding box of the adversarial image after the n-1th iteration based on the second loss value, so that the difference between the adjusted second predicted bounding box and the target bounding box becomes smaller.
[0161] 808、The server generates an adversarial image after the n th iteration based on the adjusted second predicted bounding box.
[0162] 809、If the n th iteration does not satisfy the iteration end condition, the server performs the next iteration, and if the n th iteration satisfies the iteration end condition, the server determines the adversarial image after the n th iteration as the adversarial image of the initial image.
[0163] 810、The server inputs the prompt text and the adversarial image into the first multi-modal large language model to output a target image, and the target image is labeled with a bounding box, which is used to indicate the position of the preset object in the target image.
[0164] 811、The server determines the capability information of the first multi-modal large language model based on the reference bounding box of the initial image and the bounding box of the target image, and the capability information indicates the degree to which the output of the first multi-modal large language model is affected by the input.
[0165] The steps 803-811 described above are the same as the steps 602-610 described above, and the embodiments of the present application will not be repeated here.
[0166] In the method for determining the capability information of the multi-modal large language model provided by the embodiments of the present application, the adversarial image corresponding to the replacement target adversarial attack is used to attack the first multi-modal large language model, so that the first multi-modal large language model outputs a target image. Since the first multi-modal large language model uses the bounding box of the target image to indicate the position of the preset object predicted by the first multi-modal large language model to implement the visual positioning task, and the reference bounding box of the initial image indicates the actual position of the preset object, by comparing the bounding box of the target image and the reference bounding box of the initial image, the capability of the multi-modal large language model for performing the visual positioning task can be evaluated. Since the adversarial image corresponding to the replacement target adversarial attack is obtained by reducing the distance between the second predicted bounding box and the target bounding box indicating other preset objects, the first predicted bounding box obtained by attacking the multi-modal large language model using the adversarial image deviates from the reference bounding box and is close to the reference bounding box of other preset objects. Therefore, the capability information of the multi-modal large language model when facing the replacement target adversarial attack can be effectively determined according to the adversarial image.
[0167] The above Figure 8 The capability information determination process of the multi-modal large language model under the replacement target adversarial attack in the targeted adversarial attack is described, Figure 9is a schematic diagram of an attack result of a substitution target adversarial attack provided by an embodiment of the present application, wherein the dashed box is a bounding box. As shown in Figure 9 the first prompt text is "the black rectangle on the rightmost side" and the second prompt text is "the black triangle on the leftmost side". The multimodal large language model outputs the black rectangle on the rightmost side in the bounding box of the first picture and the black triangle on the leftmost side in the bounding box of the second picture in the two pictures output according to the initial picture. The multimodal large language model outputs the black triangle on the leftmost side in the bounding box of the first picture and the black rectangle on the rightmost side in the bounding box of the second picture in the two pictures output according to the two adversarial pictures.
[0168] The four multimodal large language model capability information determination methods provided by the embodiments of the present application all use the gradient descent method to optimize the process of adding a small adversarial disturbance to obtain an adversarial picture for misleading the multimodal large language model. The gradient descent algorithm refers to modifying the picture according to the gradient. The above method of obtaining an adversarial picture and evaluating the capability of a multimodal large language model for performing a visual positioning task according to the adversarial picture can be included in a large model comprehensive safety evaluation system to evaluate the capability of the multimodal large language model for performing the visual positioning task. The large model comprehensive safety evaluation system can perform a comprehensive safety evaluation in advance on a multimodal large language model to be disclosed within a company or other entity. In addition to the gradient descent algorithm, heuristic optimization algorithms or genetic algorithms can also be used to obtain adversarial pictures. For multimodal large language models for processing modal data such as text, video, and speech, the method of obtaining adversarial pictures in the present method can also be applied to obtain adversarial data of various modalities, such as adversarial text, adversarial video, and adversarial speech, to evaluate the capability of the corresponding multimodal large language model according to the adversarial data of various modalities, so as to realize the safety evaluation of multimodal large language models with multiple targets or multiple levels.
[0169] Evaluating a multimodal large language model for performing a visual positioning task according to the present method can discover potential weaknesses of the multimodal large language model in the visual positioning task, thereby improving the multimodal large language model to improve the model performance. In addition, the above method helps to develop standards and methods for evaluating multimodal large language models to ensure the safety and reliability of multimodal large language models, and also helps to improve the understanding of researchers and developers on the safety of multimodal large language models, so as to pay more attention to the safety protection of multimodal large language models in the design and development process of multimodal large language models.
[0170] The 7B version of MiniGPT-v2 is taken as the second multi-modal large language model in the embodiments of the present application. The preset number of iterations is set to 100, the perturbation amplitude is 16, and the preset step size is 1. In the iteration process, the gradient descent method in the above method is used to obtain the adversarial image. The RefCOCO, RefCOCO+, and RefCOCOg data sets are used for experiments, and the IoU between two boundary boxes is used as the evaluation index of the above adversarial image, which is the ratio of the intersection area of the two boundary boxes to the union area of the two boundary boxes. The threshold of IoU is set to 0.5, denoted as IoU@0.5, and if the IoU between two boundary boxes is greater than 0.5, it means that the two boundary boxes are very close. For example, if the IoU between the boundary box of the target image and the reference boundary box is greater than 0.5, it means that the boundary box of the target image and the reference boundary box are very close, and it can be considered that the target image output by the multi-modal large language model is correct. For the adversarial image corresponding to the no-target adversarial attack, the IoU is the IoU between the boundary box of the target image and the reference boundary box. In the case of IoU@0.5, the smaller the IoU, the greater the influence of the adversarial image on the output of the second multi-modal large language model, that is, the more effective the adversarial image is in attacking the second multi-modal large language model. For the targeted adversarial attack, the IoU is the IoU between the boundary box of the target image and the target boundary box. In the case of IoU@0.5, the greater the IoU, the greater the influence of the adversarial image on the output of the second multi-modal large language model, that is, the more effective the adversarial image is in attacking the second multi-modal large language model. Among the adversarial images obtained by the method, the adversarial image corresponding to the image embedding attack in the no-target adversarial attack can reduce the average IoU from 81.15% to 23.24%. And the adversarial image corresponding to the text boundary box attack in the no-target adversarial attack can reduce the average IoU to 33.90%. For the targeted adversarial attack, the exclusive target adversarial attack can increase the average IoU from 0.13% to 61.84%. The replacement target adversarial attack can increase the average IoU from 8.89% to 30.14%.
[0171] Figure 10 is a structural block diagram of a multi-modal large language model capability information determination device provided by the embodiments of the present application. The device is used to execute the steps of the above-mentioned multi-modal large language model capability information determination method, see Figure 10 , the multi-modal large language model capability information determination device comprises:
[0172] The picture acquisition module 1001 is configured to acquire a prompt text and an initial picture, and the prompt text is used to locate a preset object.
[0173] The adversarial image acquisition module 1002 is configured to acquire an adversarial image of the initial picture based on the initial picture, and the adversarial image is different from the initial picture.
[0174] The target picture acquisition module 1003 is configured to input the prompt text and the adversarial picture into the first multi-modal large language model, and output a target picture. The target picture is labeled with a bounding box. The bounding box is used to indicate the position of the preset object in the target picture.
[0175] The capability information determination module 1004 is configured to determine capability information of the first multi-modal large language model based on the reference bounding box of the initial picture and the bounding box of the target picture. The capability information indicates the degree to which the output of the first multi-modal large language model is affected by the input.
[0176] In some embodiments, the adversarial picture acquisition module 1002 includes:
[0177] The first adjustment unit is configured to adjust the embedding feature of the preset adversarial picture according to the embedding feature of the initial picture, so that the distance between the embedding feature of the preset adversarial picture and the embedding feature of the initial picture is increased.
[0178] The first adversarial picture generation unit is configured to generate the adversarial picture based on the adjusted embedding feature.
[0179] In some embodiments, the first adjustment unit is configured to:
[0180] Obtain the distance between the embedding feature of the initial picture and the embedding feature of the preset adversarial picture.
[0181] Adjust the embedding feature of the preset adversarial picture based on the distance.
[0182] In some embodiments, the adversarial picture acquisition module 1002 is configured to:
[0183] Obtain, by the second multi-modal large language model, a first predicted bounding box of the initial picture and a first loss value. The first loss value is used to represent the difference between the first predicted bounding box of the initial picture and the reference bounding box of the initial picture.
[0184] Adjust the first predicted bounding box of the initial picture based on the first loss value, so that the difference between the adjusted first predicted bounding box and the reference bounding box is increased.
[0185] Generate the adversarial picture based on the adjusted first predicted bounding box.
[0186] In some embodiments, the adversarial picture acquisition module 1002 is configured to:
[0187] The second prediction bounding box of the initial picture and a second loss value are obtained through the second multi-modal large language model, the second loss value is used to represent a difference between the second prediction bounding box of the initial picture and a target bounding box of the initial picture, and the target bounding box indicates an object other than the preset object in the initial picture;
[0188] Based on the second loss value, the second prediction bounding box of the initial picture is adjusted, so that a difference between the adjusted second prediction bounding box and the target bounding box is smaller;
[0189] Based on the adjusted second prediction bounding box, the adversarial picture is generated.
[0190] In some embodiments, the adversarial picture obtaining module 1002 is configured to obtain the target bounding box, and the obtaining process includes:
[0191] The bounding box of the kth preset object indicated by the prompt text is obtained, and k is an integer greater than 1;
[0192] The bounding box of the kth preset object is determined as the target bounding box of the k-1th preset object indicated by the prompt text.
[0193] In some embodiments, the capability information determining module 1004 is configured to:
[0194] The intersection over union between the bounding box of the target picture and the reference bounding box of the initial picture is calculated, and the intersection over union is a ratio of an area of an intersection between the bounding box of the target picture and the reference bounding box to an area of a union between the bounding box of the target picture and the reference bounding box;
[0195] If the intersection over union is greater than or equal to an intersection over union threshold, the evaluation index of the capability information is added by 0;
[0196] If the intersection over union is less than the intersection over union threshold, the evaluation index of the capability information is added by 1.
[0197] It should be noted that: the apparatus provided in the above embodiments is only used as an example to illustrate the division of the above functional modules when determining the capability information of the multi-modal large language model, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.
[0198] Figure 11A structural schematic diagram of a server is provided according to an embodiment of the present application. The server 1100 can have great differences due to different configurations or performances, and can include one or more CPUs (Central Processing Units, processors) 1101 and one or more memories 1102. The memory 1102 stores at least one computer program, which is loaded and executed by the processor 1101 to implement the multi-modal large language model capability information determination method provided by each method embodiment described above. Of course, the server can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for realizing the functions of the device, and the like, so as to perform input and output. The server can also include other components for realizing the functions of the device, and the like, which are not described herein.
[0199] The present application also provides a computer readable storage medium having at least one computer program stored therein. The at least one computer program is loaded and executed by a processor of a server to implement the operations performed by the server in the multi-modal large language model capability information determination method of the above embodiments. For example, the computer readable storage medium can be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0200] In some embodiments, the computer program related to the embodiments of the present application can be deployed to execute on one server, or on multiple servers located in one place, or on multiple servers distributed in multiple places and interconnected through a communication network, which can constitute a blockchain system.
[0201] The present application also provides a computer program product, including a computer program stored in a computer readable storage medium. The processor of the server reads the computer program from the computer readable storage medium, and the processor executes the computer program to enable the server to perform the multi-modal large language model capability information determination method provided in the various optional implementation manners described above.
[0202] Those of ordinary skill in the art can understand that all or part of the steps of the above-described embodiments can be completed by hardware, or by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, such as a read-only memory, a magnetic disk or an optical disk.
[0203] The above merely provides the optional embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining capability information of a multi-modal large language model, characterized in that, The method comprises: acquiring prompt text and an initial picture, the prompt text being used for positioning a preset object; based on the initial picture, acquiring an adversarial picture of the initial picture, the adversarial picture being different from the initial picture; inputting the prompt text and the adversarial picture into a first multi-modal large language model to output a target picture, the target picture being labeled with a bounding box, the bounding box being used to indicate the position of the preset object in the target picture; based on a reference bounding box of the initial picture and the bounding box of the target picture, determining the capability information of the first multi-modal large language model, the capability information indicating the degree to which the output of the first multi-modal large language model is affected by the input.
2. The method of claim 1, wherein, The acquiring of the adversarial picture of the initial picture based on the initial picture comprises: adjusting the embedding feature of a preset adversarial picture according to the embedding feature of the initial picture, so as to increase the distance between the embedding feature of the preset adversarial picture and the embedding feature of the initial picture; generating the adversarial picture based on the adjusted embedding feature.
3. The method of claim 2, wherein, The adjusting of the embedding feature of the preset adversarial picture according to the embedding feature of the initial picture comprises: acquiring the distance between the embedding feature of the initial picture and the embedding feature of the preset adversarial picture; adjusting the embedding feature of the preset adversarial picture based on the distance.
4. The method according to any one of claims 1 to 3, characterized in that, The acquiring of the adversarial picture of the initial picture based on the initial picture comprises: acquiring, by a second multi-modal large language model, a first predicted bounding box of the initial picture and a first loss value, the first loss value being used to represent the difference between the first predicted bounding box of the initial picture and a reference bounding box of the initial picture; adjusting the first predicted bounding box of the initial picture based on the first loss value, so that the difference between the adjusted first predicted bounding box and the reference bounding box becomes larger; generating the adversarial picture based on the adjusted first predicted bounding box.
5. The method according to any one of claims 1 to 4, characterized in that, The acquiring of the adversarial picture of the initial picture based on the initial picture comprises: acquiring, by a second multi-modal large language model, a second predicted bounding box of the initial picture and a second loss value, the second loss value being used to represent the difference between the second predicted bounding box of the initial picture and a target bounding box of the initial picture, the target bounding box indicating an object other than the preset object in the initial picture; adjusting the second predicted bounding box of the initial picture based on the second loss value, so that the difference between the adjusted second predicted bounding box and the target bounding box becomes smaller; generating the adversarial picture based on the adjusted second predicted bounding box.
6. The method of claim 5, wherein, The acquisition process of the target bounding box comprises: acquiring the bounding box of the kth preset object indicated by the prompt text, k being an integer greater than 1; determining the bounding box of the kth preset object as the target bounding box of the k-1th preset object indicated by the prompt text.
7. The method according to any one of claims 1 to 6, characterized in that, The determination of the capability information of the first multi-modal large language model based on the reference bounding box of the initial picture and the bounding box of the target picture comprises: calculate an intersection-over-union between the bounding box of the target picture and a reference bounding box of the initial picture, the intersection-over-union being a ratio of an area of an intersection between the bounding box of the target picture and the reference bounding box to an area of a union between the bounding box of the target picture and the reference bounding box; if the intersection-over-union is greater than or equal to an intersection-over-union threshold, the evaluation index of the capability information is incremented by 0; if the intersection-over-union is less than the intersection-over-union threshold, the evaluation index of the capability information is incremented by 1. 8.A capability information determination apparatus of a multi-modal large language model, characterized in that, The apparatus comprises: a picture acquisition module configured to acquire prompt text and an initial picture, the prompt text being used to locate a preset object; an adversarial picture acquisition module configured to acquire an adversarial picture of the initial picture based on the initial picture, the adversarial picture being different from the initial picture; a target picture acquisition module configured to input the prompt text and the adversarial picture into a first multi-modal large language model, and output a target picture, the target picture being labeled with a bounding box, the bounding box being used to indicate a position of the preset object in the target picture; a capability information determination module configured to determine capability information of the first multi-modal large language model based on a reference bounding box of the initial picture and a bounding box of the target picture, the capability information indicating a degree to which an output of the first multi-modal large language model is affected by an input.
9. A server, characterized by The server comprises a processor and a memory, the memory being used to store at least one piece of computer program, the at least one piece of computer program being loaded and executed by the processor to implement the method in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium is used to store at least one piece of computer program, the at least one piece of computer program being used to implement the method in any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method in any one of claims 1 to 7.