Safety evaluation method, device and equipment for multi-modal large model and storage medium
The method addresses privacy and security challenges in LMMs by quantizing and encrypting model weights, performing secure collaborative reasoning, and validating service providers, ensuring secure and efficient deployment in pervasive computing environments.
Patent Information
- Application Number
- CN202510797457.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing multimodal large models face privacy security challenges in a universal computing environment, especially when users may disclose user privacy information without explicit authorization, and malicious attackers can steal model weight parameters for inference attacks. The existing security inference schemes are incompatible in multimodal scenarios and are expensive to calculate.
By quantizing the multimodal large model and approximating nonlinear functions, initialization models are generated and encrypted weights, combining secret sharing and MPC multi-party security calculations, the client builds a hybrid query data set and performs secure collaborative inference. Finally, the client decrypts and reconstructs and verifies the inference results to judge the security of the service provider.
It effectively protects the privacy and security of multimodal large models, reduces model size and computing overhead, ensures the reliability of service providers and the security of client data, and avoids malicious attacks.
Smart Images

Figure CN120321041A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of machine learning, and in particular, to a security evaluation method, device, equipment, and storage medium for multi-modal large models. Background Art
[0002] In recent years, with the rise of the Transformer model architecture and the popularization of the pervasive computing paradigm, large multimodal models (LMMs) have shown great application potential and been widely used in fields such as image-text retrieval, speech recognition, and cross-domain predictive analysis due to their powerful multi-modal data (such as text, images, videos, and audio, etc.) processing and understanding capabilities. However, the powerful functions of LMMs are accompanied by increasingly severe privacy and security challenges. On the one hand, service providers may illegally collect and analyze users' request data without the explicit authorization of users, resulting in the leakage of users' privacy information. On the other hand, malicious attackers may attempt to steal or abuse the weight parameters of LMMs, and then conduct inference attacks on user data, causing deeper privacy violations. Especially in the pervasive computing environment, LMMs inference services are usually deployed on edge devices or in the cloud, further exacerbating the risk of privacy leakage. In addition, existing security inference schemes mainly focus on single-modal scenarios, and directly migrating to multi-modal inference will face problems of computational incompatibility and a large amount of inference overhead.
[0003] Therefore, there is an urgent need for a lightweight and secure security evaluation method for multi-modal large models to address the threats of malicious service providers and ensure the secure and reliable application of LMMs in the pervasive computing environment. Summary of the Invention
[0004] According to the embodiments of the present application, there is provided a security evaluation method, device, equipment, and storage medium for multi-modal large models, which can effectively avoid the threats of malicious service providers.
[0005] In the first aspect of the present application, there is provided a security evaluation method for multi-modal large models. The method includes: The model owner performs quantization processing on the multi-modal large model to obtain a quantized multi-modal large model, and then performs non-linear function approximation processing on the quantized multi-modal large model to obtain an initialized multi-modal large model, and encrypts and sends the weights of the initialized multi-modal large model to different security agents; The client merges and shuffles the public data set and the query data set to construct a mixed query data set, and then performs hierarchical vectorization on the mixed query data set to obtain an inference query data set, and encrypts and sends it to different security agents; Different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multimodal large model respectively, obtain the inference results, and send the inference results to the client; The client decrypts and reconstructs the inference results sent by different security agents, obtains the inference prediction results, and then verifies the inference prediction results according to the true label values in the inference query dataset to determine whether the service provider is secure and available.
[0006] In a possible implementation, the model owner quantizes the multimodal large model, including: Quantize the weights of the multimodal large model to float16.
[0007] In a possible implementation, perform non-linear function approximation processing on the quantized multimodal large model, including: Replace the sigmoid function in the quantized multimodal large model with the Approxsigmoid function; Replace the GeLU function in the quantized multimodal large model with the ApproxGeLU function; Replace the Softmax function in the quantized multimodal large model with the ApproxSoftmax function.
[0008] In a possible implementation, encrypt and send the weights of the initialized multimodal large model to different security agents, including: Perform secret sharing on the weights of the initialized multimodal large model; Split and encrypt the weights of the initialized multimodal large model to obtain encrypted weight blocks; Send the encrypted weight blocks to different security agents respectively.
[0009] In a possible implementation, perform hierarchical vectorization on the mixed query dataset to obtain the inference query dataset, including: According to the modality of the data in the mixed query dataset, use different preprocessing methods to perform independent feature encoding on the data to obtain the inference query dataset.
[0010] In a possible implementation, different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multimodal large model respectively, obtain the inference results, including: The secure collaborative inference adopts the MPC multi-party secure computing mechanism; Different security agents perform secure inference respectively according to the inference query dataset and the weights of the initialized multimodal large model to obtain the initial inference results; Perform privacy encryption calculation according to the initial inference results and Beaver triples to obtain the inference results.
[0011] In a possible implementation, the inference prediction result is verified according to the true label value in the inference query dataset to determine whether the service provider is safe and available, including: Calculating the number of secure data in the inference prediction result that is the same as the true label value in the inference request dataset; If the number of secure data is greater than the verification threshold, the service provider is safe and available; If the number of secure data is less than or equal to the verification threshold, the service provider will not be adopted; The service provider includes different security agents and model owners.
[0012] In a second aspect of the present application, a security evaluation device for a multimodal large model is provided. The device includes: A model preprocessing module, where the model owner quantifies the multimodal large model to obtain a quantized multimodal large model, then performs a non-linear function approximation process on the quantized multimodal large model to obtain an initialized multimodal large model, and encrypts and sends the weights of the initialized multimodal large model to different security agents; A data preprocessing module, where the client merges and shuffles the public dataset and the query dataset to construct a mixed query dataset, and then performs hierarchical vectorization on the mixed query dataset to obtain an inference query dataset, and encrypts and sends it to different security agents; An inference model, where different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multimodal large model respectively to obtain an inference result, and send the inference result to the client; A verification module, where the client decrypts and reconstructs the inference results sent by different security agents to obtain an inference prediction result, and then verifies the inference prediction result according to the true label value in the inference query dataset to determine whether the service provider is safe and available.
[0013] In a third aspect of the present application, an electronic device is provided. The electronic device includes: a memory and a processor, and a computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.
[0014] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present application is implemented.
[0015] The security evaluation method for multimodal large models provided by the embodiments of this application. First, the model owner performs quantization processing on the multimodal large model to obtain a quantized multimodal large model. Then, non-linear function approximation processing is performed on the quantized multimodal large model to obtain an initialized multimodal large model, and the weights of the initialized multimodal large model are encrypted and sent to different security agents. Next, the client combines and shuffles the public dataset and the query dataset to construct a mixed query dataset. Then, hierarchical vectorization is performed on the mixed query dataset to obtain an inference query dataset, which is encrypted and sent to different security agents. Secondly, different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multimodal large model respectively to obtain inference results, and send the inference results to the client. Finally, the client decrypts and reconstructs the inference results sent by different security agents to obtain inference prediction results, and then verifies the inference prediction results according to the true label values in the inference query dataset to determine whether the service provider is safe and available. Through quantization processing, the model is compressed to half its size. Non-linear function approximation processing enables the model to adapt to different security computing requirements. At the same time, merging, shuffling, and hierarchical vectorization of the dataset improve the inference accuracy of the model. Finally, decrypting, reconstructing, and verifying the inference results effectively realizes the identification of unreliable service providers and avoids threats or attacks from unreliable service providers to the client.
[0016] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understandable through the following description. Brief Description of the Drawings
[0017] Combined with the drawings and referring to the following detailed description, the above and other features, advantages, and aspects of the embodiments of this application will become more obvious. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 is a flowchart of the security evaluation method for multimodal large models according to the embodiments of this application; Figure 2 is a flowchart of secret sharing according to the embodiments of this application; Figure 3 is a system architecture diagram related to the method provided by the embodiments of this application; Figure 4 is a block diagram of the security evaluation device for multimodal large models according to the embodiments of this application; Figure 5 is a schematic structural diagram of a terminal device or a server suitable for implementing the embodiments of this application. Detailed Description of the Embodiments
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0019] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0020] Figure 1 The flowchart of the security evaluation method for a multi-modal large model according to an embodiment of the present disclosure is shown as Figure 1 shown.
[0021] S101. The model owner performs quantization processing on the multi-modal large model to obtain a quantized multi-modal large model, and then performs non-linear function approximation processing on the quantized multi-modal large model to obtain an initialized multi-modal large model, and encrypts and sends the weights of the initialized multi-modal large model to different security agents.
[0022] In this embodiment, the quantization processing reduces the size of the multi-modal large model, reduces the storage and transmission costs, and also increases the difficulty of reverse engineering the model to a certain extent. Then, non-linear function approximation processing is performed on the quantized multi-modal large model to further obscure the structure and parameter information of the original model, and provide a basis for subsequent secure privacy computing.
[0023] Optionally, the model owner performs quantization processing on the multi-modal large model, including: Quantize the weights of the multi-modal large model to float16.
[0024] In this embodiment, float16 is a 16-bit floating-point number format. Quantizing the weights of the multi-modal large model to float16 can reduce the memory occupancy by half compared to the original 32-bit floating-point number format float32, thereby reducing the storage and transmission costs.
[0025] Optionally, the non-linear function approximation processing on the quantized multi-modal large model includes: Replace the sigmoid function in the quantized multi-modal large model with the Approxsigmoid function; Replace the GeLU function in the quantized multimodal large model with the ApproxGeLU function; Replace the Softmax function in the quantized multimodal large model with the ApproxSoftmax function.
[0026] Among them, the Sigmoid function, GeLU function, and Softmax function are three indispensable activation functions in multimodal large models. They play a crucial role in the model training process to help the model learn and represent input data more accurately. Among them, the Sigmoid function mainly plays its unique role in gating mechanisms, binary classification tasks, and feature combination. The GeLU function, initially widely used in Transformer models, has become the key to improving model generalization with its excellent performance improvement and ability to combat gradient vanishing, and has gradually become popular in the field of multimodal models. The Softmax function is prominent in multi-classification tasks and attention mechanisms. It can effectively convert the model output into a probability distribution for probability interpretation and class decision-making. The core of the multimodal large model LMMs in this application is the Transformer module. The Transformer module can capture and model the complex relationships between different modal features through the attention mechanism. The Transformer module includes an attention mechanism, a feed-forward neural network, and a normalization layer. Among them, the Softmax function is used to calculate the similarity between query values and key values in the attention mechanism, and the GeLU function or Sigmoid function is used as the activation function in the feed-forward neural network.
[0027] In the present invention, the calculation formula of the sigmoid function in the quantized multimodal large model is: 。
[0028] Among them, the input variable x of the sigmoid function is the output vector of the fully connected layer in the feed-forward neural network. Then, the calculation formula of the Approxsigmoid function after non-linear function approximation is: 。
[0029] Among them, and are preset hyperparameters with values ranging from 1 to 10. The calculation formula of the GeLU function in the quantized multimodal large model is: 。
[0030] Among them, the input variable x of the GeLU function is also the output vector of the fully connected layer in the feed-forward neural network. Then, the calculation formula of the ApproxGeLU function after being approximated by a non-linear function is: 。
[0031] The calculation formula of the Softmax function in the quantized multi-modal large model is: 。
[0032] Among them, the input variable x of the GeLU function is the dot product value between the query value and the key value in the attention mechanism. Then, the calculation formula of the Approxsoftmax function after being approximated by a non-linear function is: 。
[0033] Among them, is the mean of all variables x of the input multi-modal large model within one round of training, is a preset hyperparameter with a value range of 0 to 1. In this embodiment, the approximation processing by the non-linear function can effectively protect the privacy and security of the model, prevent the model from being reverse-engineered or illegally used, and provide a unified approximation calculation function for subsequent secure privacy calculations.
[0034] Optionally, the weights of the initialized multi-modal large model are encrypted and sent to different security agents, including: Performing secret sharing on the weights of the initialized multi-modal large model; Splitting and encrypting the weights of the initialized multi-modal large model to obtain encrypted weight blocks; Sending the encrypted weight blocks to different security agents respectively.
[0035] Among them, Secret Shares is a cryptographic protocol. Its core idea is to split a secret into multiple "shares" and distribute these shares to multiple participants. Only when a sufficient number of participants combine their shares can the original secret be restored, and any set of less than the sufficient number of shares cannot obtain any information about the secret.
[0036] Figure 2 is the flowchart of secret sharing according to the embodiment of the present application, as Figure 2 shown.
[0037] In a possible implementation manner, the weights of the initialized multi-modal large model are split and encrypted into n encrypted weight blocks , and then, the initialized multi-modal large model will first send the encrypted weight blocks to the service provider, and the service provider will send the encrypted weight blocks Distribute to different security agents.
[0038] In this embodiment, the weights for initializing the multi-modal large model are encrypted and sent to different security agents through a secret sharing protocol, ensuring the security and privacy of the weights for initializing the multi-modal large model, and thus ensuring the security of the transmission process.
[0039] S102. The client merges and shuffles the public dataset and the query dataset to construct a mixed query dataset, and then hierarchically vectorizes the mixed query dataset to obtain an inference query dataset, which is encrypted and sent to different security agents.
[0040] In a possible implementation, there is a public dataset , and a query dataset . Then, the mixed query dataset constructed by merging and shuffling the public dataset and the query dataset can be expressed as .
[0041] In this embodiment, merging and shuffling the public dataset and the query dataset to construct a mixed query dataset improves the robustness of model training, increases the difficulty for malicious service providers to infer the client's private data, and in addition, unifies different modal data through hierarchical vectorization for the input and processing of the model.
[0042] Optionally, hierarchically vectorizing the mixed query dataset to obtain an inference query dataset includes: Independently feature-encoding the data in the mixed query dataset using different preprocessing methods according to the modality of the data to obtain an inference query dataset.
[0043] For example, for text-modal data, the preprocessing methods include but are not limited to word segmentation, stop-word removal, and stemming; for image-modal data, the preprocessing methods include but are not limited to scaling, cropping, and flipping; for voice-modal data, the preprocessing methods include but are not limited to silence detection and removal, sound enhancement, and normalization. Then, independently feature-encode the preprocessed text-modal data, image-modal data, and voice-modal data respectively to obtain an inference query dataset. Independent feature encoding includes but is not limited to text feature encoding, image feature encoding, and audio feature encoding, where text feature encoding includes but is not limited to the bag-of-words model, TF-IDF, and Word2Vec, image feature encoding includes but is not limited to convolutional neural networks (CNNs, Convolutional Neural Networks), and audio feature encoding includes but is not limited to extracting Mel Frequency Cepstral Coefficients (MFCCs, Mel Frequency Cepstral Coefficient).
[0044] In this embodiment, by hierarchically vectorizing the mixed query dataset, the key features of different modality data can be accurately captured, improving the accuracy of inference queries.
[0045] S103. Different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multi-modal large model respectively, obtain the inference results, and send the inference results to the client.
[0046] In this embodiment, the inference task is completed without exposing the original data and model parameters, further improving the security of the data and the model.
[0047] Optionally, different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multi-modal large model respectively, obtain the inference results, including: The secure collaborative inference adopts the MPC multi-party secure computing mechanism; Different security agents perform secure inference based on the inference query dataset and the weights of the initialized multi-modal large model respectively to obtain the initial inference results; Perform privacy encryption calculation according to the initial inference results and Beaver triples to obtain the inference results.
[0048] Among them, MPC multi-party secure computing (Multi-Party Computation) is a method that allows multiple participants to jointly calculate a function without revealing their respective input data. The calculation results will only be made public with the consent of all participants, while their respective input data remains confidential. Beaver triples are an auxiliary method used to optimize calculations in MPC multi-party secure computing, consisting of three secretly shared values a, b, c, where a and b are two random numbers, and c = a * b.
[0049] In a possible implementation manner, the security agent constructs a multi-modal large model according to the weights of the initialized multi-modal large model, and then inputs the inference query dataset into the pre-trained multi-modal large model to obtain the initial inference results , where i is the serial number of the security agent. There exists a Beaver triple , and the formula for performing privacy encryption calculation according to the Beaver triple and the initial inference results can be: , Among them, and are auxiliary calculation elements, , , and The calculation formulas are respectively , 。 is the initial inference result of the i-th security agent, and is the weight of the initialized multi-modal large model of the i-th security agent. Finally, the set of inference results of different security agents can be expressed as 。
[0050] In this embodiment, through the MPC multi-party secure computing mechanism, each security agent only processes the local inference query dataset and the weight of the initialized multi-modal large model, and the data does not need to leave its secure environment, minimizing the risk of data leakage.
[0051] S104. The client decrypts and reconstructs the inference results sent by different security agents to obtain the inference prediction results, and then verifies the inference prediction results according to the true label values in the inference query dataset to determine whether the service provider is secure and available.
[0052] In this embodiment, it is possible to determine whether the service provider is secure and available, estimate whether there may be malicious behaviors in the future, and avoid security risks in a timely manner.
[0053] Optionally, verifying the inference prediction results according to the true label values in the inference query dataset to determine whether the service provider is secure and available includes: Calculating the number of secure data in the inference prediction results that is the same as the true label values in the inference request dataset; If the number of secure data is greater than the verification threshold, the service provider is secure and available; If the number of secure data is less than or equal to the verification threshold, the service provider will not be adopted; The service provider includes different security agents and model owners.
[0054] For example, the total data volume in the inference request dataset is 50000. By statistics, the number of secure data in the inference prediction results that is the same as the true label values in the inference request dataset is 35000, and the verification threshold is 40000. Then the number of secure data is less than or equal to the verification threshold, and the service provider is regarded as insecure and unreliable and will not be adopted.
[0055] In this embodiment, by determining whether the service provider is secure and available, the security and reliability of the service provider in multi-modal large model inference are ensured.
[0056] Figure 3 is the system architecture diagram involved in the method provided by the embodiment of the present application, as Figure 3 shown: First, the model owner sequentially performs quantization and approximation of non-linear functions on the multi-modal large model to generate an initial multi-modal large model. Subsequently, the weights of the model are encrypted and securely transmitted to various different security proxy nodes in the service provider. Then, the client merges and shuffles the public dataset and the query dataset to construct a mixed query dataset. The mixed query dataset is transformed into an inference query dataset through hierarchical vectorization and then sent to each security proxy in an encrypted form as well. Next, each security proxy performs secure collaborative inference operations on the encrypted inference query dataset and the weights of the initial multi-modal large model. After the inference is completed, the inference results are sent back to the client. Finally, the client decrypts and reconstructs these inference results to obtain the inference prediction results. Moreover, the client will also verify the inference prediction results based on the true label values contained in the inference query dataset to evaluate the security and availability of the service provider, avoid attacks on the client by malicious service providers, and protect the data security of the client.
[0057] According to the embodiments of the present disclosure, the following technical effects are achieved: 1. Protect the privacy and security of the query dataset of the client, and at the same time keep the weights of the multi-modal large model confidential.
[0058] 2. Optimize unnecessary computational overhead, laying a foundation for the secure inference of multi-modal large models deployed on resource-constrained servers.
[0059] 3. Effectively avoid attacks on the client by untrusted service providers or malicious provision of incorrect inference results.
[0060] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0061] The above is the introduction of the method embodiments. The following further illustrates the solution of this application through device embodiments.
[0062] Figure 4 The block diagram of a security evaluation device for a multi-modal large model according to an embodiment of the present application is shown, as Figure 4 shown and includes: The model preprocessing module 401 quantizes the multi-modal large model by the model owner to obtain a quantized multi-modal large model, and then performs a non-linear function approximation process on the quantized multi-modal large model to obtain an initialized multi-modal large model, and encrypts and sends the weights of the initialized multi-modal large model to different security agents; The data preprocessing module 402 merges and shuffles the public data set and the query data set by the client to construct a mixed query data set, and then performs hierarchical vectorization on the mixed query data set to obtain an inference query data set, and encrypts and sends it to different security agents; The inference model 403 performs secure collaborative inference on the inference query data set and the weights of the initialized multi-modal large model by different security agents to obtain an inference result, and sends the inference result to the client; The verification module 404 decrypts and reconstructs the inference results sent by different security agents by the client to obtain an inference prediction result, and then verifies the inference prediction result according to the true label value in the inference query data set to determine whether the service provider is secure and available.
[0063] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0064] Figure 5 The structure diagram of a terminal device or a server suitable for implementing the embodiments of the present application is shown.
[0065] As Figure 5 shown, the terminal device or the server includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the terminal device or the server are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0066] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. as well as a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 510 as required so that a computer program read therefrom is installed into the storage section 508 as required.
[0067] Specifically, according to an embodiment of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a machine-readable medium, the computer program including program codes for performing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by a central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.
[0068] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0069] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in a block can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0070] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0071] As another aspect, this application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above-mentioned programs are executed by one or more processors, the methods described in this application are implemented.
[0072] The above description is only a preferred embodiment of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the application involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above application concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions described in this application.
Claims
1. A security evaluation method for multi-modal large models, characterized in that Including: The model owner performs quantization processing on the multi-modal large model to obtain a quantized multi-modal large model, then performs non-linear function approximation processing on the quantized multi-modal large model to obtain an initialized multi-modal large model, and encrypts and sends the weights of the initialized multi-modal large model to different security agents; The client merges and shuffles the public dataset and the query dataset to construct a mixed query dataset, and then performs hierarchical vectorization on the mixed query dataset to obtain an inference query dataset, and encrypts and sends it to the different security agents; The different security agents perform secure collaborative inference on the inference query dataset and the weights of the initialized multi-modal large model respectively to obtain inference results, and send the inference results to the client; The client decrypts and reconstructs the inference results sent by the different security agents to obtain an inference prediction result, and then verifies the inference prediction result according to the true label value in the inference query dataset to determine whether the service provider is secure and available.
2. The security evaluation method for multi-modal large models according to claim 1, characterized in that The model owner performs quantization processing on the multi-modal large model, including: Quantizing the weights of the multi-modal large model to float16.
3. The security evaluation method for multi-modal large models according to claim 1, wherein, The non-linear function approximation processing on the quantized multi-modal large model includes: Replacing the sigmoid function in the quantized multi-modal large model with the Approxsigmoid function; Replacing the GeLU function in the quantized multi-modal large model with the ApproxGeLU function; Replacing the Softmax function in the quantized multi-modal large model with the ApproxSoftmax function.
4. The security evaluation method for multi-modal large models according to claim 1, wherein The encrypting and sending the weights of the initialized multi-modal large model to different security agents includes: Performing secret sharing on the weights of the initialized multi-modal large model; Splitting and encrypting the weights of the initialized multi-modal large model to obtain encrypted weight blocks; Sending the encrypted weight blocks to the different security agents respectively.
5. The security evaluation method for multi-modal large models according to claim 1, characterized in that The hierarchical vectorization of the mixed query dataset to obtain an inference query dataset includes: Independently encoding features of the data in the mixed query dataset using different preprocessing methods according to the modality of the data to obtain the inference query dataset.
6. The security evaluation method for multi-modal large models according to claim 1, characterized in that, The different security agents performing secure collaborative inference on the inference query dataset and the weights of the initialized multi-modal large model respectively to obtain inference results includes: The secure collaborative inference uses the MPC multi-party secure computing mechanism; The different security agents respectively perform secure inference according to the inference query dataset and the weights of the initialized multi-modal large model to obtain initial inference results; Performing privacy encryption calculation according to the initial inference results and Beaver triples to obtain the inference results.
7. The security evaluation method for multi-modal large models according to claim 1, wherein The verifying the inference prediction result according to the true label value in the inference query dataset to determine whether the service provider is secure and available includes: Calculating the number of secure data in the inference prediction result that is the same as the true label value in the inference request dataset; If the quantity of the security data is greater than the verification threshold, the service provider is secure and available; If the quantity of the security data is less than or equal to the verification threshold, the service provider will not be adopted; The service providers include different security agents and model owners.
8. A security evaluation device for multi-modal large models, characterized in that, Including: A model preprocessing module, where the model owner quantizes the multi-modal large model to obtain a quantized multi-modal large model, then performs a non-linear function approximation process on the quantized multi-modal large model to obtain an initialized multi-modal large model, and encrypts and sends the weights of the initialized multi-modal large model to different security agents; A data preprocessing module, where the client merges and shuffles the public data set and the query data set to construct a mixed query data set, then performs a hierarchical vectorization on the mixed query data set to obtain an inference query data set, and encrypts and sends it to different security agents; An inference model, where different security agents perform secure collaborative inference on the inference query data set and the weights of the initialized multi-modal large model respectively to obtain an inference result, and send the inference result to the client; A verification module, where the client decrypts and reconstructs the inference results sent by the different security agents to obtain an inference prediction result, and then verifies the inference prediction result according to the true label values in the inference query data set to determine whether the service provider is secure and available.
9. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal supercomputing network for decentralized private data training
CN118714144A
Medical metadata processing and de-identification method and system for large medical model
CN119293643A
Reliable reasoning scheduling method based on edge hybrid expert large model
CN119903923A
Equipment intelligent guarantee system based on off-line large model
CN119919126A
Large model privacy protection reasoning method and system based on secure multi-party computing
CN119995821A