Large model security assessment method, device, equipment and storage medium based on Zero-Shot NAS
Through the visual inference method based on zero-sample NAS, ViT model is selected and trained, combining secret sharing and MPC protocol, the problems of computing complexity and delay in multi-party security computing are solved, and efficient and secure ViT model evaluation and reasoning are achieved.
Patent Information
- Application Number
- CN202510789338.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-13
AI Technical Summary
When the prior art is applied to ViT models with multi-party security computing, there are computational complexity and significant inference delay problems, and requires strong computing resources and professional knowledge, which are difficult to promote and apply.
Through a visual inference method based on zero-sample NAS, multiple search spaces are obtained, the first ViT model is randomly selected, the second ViT model is searched, and the model is selected using downstream data to retrain the model, and collaborative inference is performed through secret sharing sending architecture and weights, combining secret sharing and MPC protocol to protect data privacy.
It effectively reduces the inference delay of multi-party security computing in the ViT model, improves model evaluation efficiency and accuracy, protects data privacy and model security, and reduces computing resources and time costs.
Smart Images

Figure CN120316783B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of machine learning, and in particular to large-scale model security assessment methods, devices, equipment, and storage media based on Zero-Shot NAS. Background Art
[0002] In recent years, secure multi-party computation (MPC) technology has been widely used in the inference process of machine learning models. It enables multi-party joint computation while protecting the data privacy of all parties involved. In particular, the Vision Transformer (ViT), a deep learning model for image classification, has achieved remarkable results by extending the Transformer architecture, previously used primarily for natural language processing tasks, to the field of computer vision. With the assistance of secure MPC, the ViT model significantly improves data privacy and security during inference. However, when applying secure MPC to the ViT model, the numerous nonlinear operations it incorporates lead to communication complexity and significant computational latency. Although some optimized solutions for applying secure MPC to the ViT model exist, these solutions typically require deep expertise and powerful computing resources. These solutions face high barriers to adoption and are difficult to scale and implement in practice, hindering the application of secure MPC to the ViT model.
[0003] Therefore, how to select an efficient and secure ViT model and effectively reduce the inference delay of multi-party secure computing in the application of the ViT model is a problem that needs to be solved. Summary of the Invention
[0004] According to the embodiments of the present application, a visual reasoning solution based on zero-sample NAS is provided, which can realize efficient and secure ViT model evaluation selection and effectively reduce the reasoning delay of multi-party secure computing in ViT model applications.
[0005] In a first aspect of the present application, a visual reasoning method based on zero-sample NAS is provided. The method comprises:
[0006] Obtain a multi-party search space, and randomly select a first ViT model from the multi-party search space;
[0007] Perform a Zero-Shot NAS search on the first ViT model based on its accuracy and inference latency, and select the second ViT model.
[0008] Retrain the second ViT model based on downstream data to obtain a pre-trained ViT model, and send the architecture and weights of the pre-trained ViT model to different servers through secret sharing;
[0009] Obtain the inference dataset sent by the client, and different servers perform collaborative inference based on the MPC protocol, inference dataset, and the architecture and weights of the pre-trained ViT model to obtain the inference results.
[0010] In one possible implementation, the multi-party search space includes a ViT model that has been approximated by a nonlinear function;
[0011] Nonlinear function approximation processing includes approximating the Softmax function and GeLU function in the ViT model.
[0012] In one possible implementation, the accuracy evaluation indicators of the first ViT model include the MSA score and the MLP score;
[0013] The MSA score is calculated based on the nuclear norm of the gradient matrix and the nuclear norm of the weight matrix in the first ViT model;
[0014] The MLP score is calculated based on the Frobenius inner product between the gradient matrix and the weight matrix in the first ViT model.
[0015] In one possible implementation, the inference latency of the first ViT model is calculated based on the sum of the MSA inference latency and the MLP inference latency of each layer in the first ViT model;
[0016] The MSA inference delay is approximately calculated based on the inference delay of the multi-head attention mechanism in the first ViT model;
[0017] The MLP inference latency is approximately calculated based on the inference latency of the GeLU activation function in the first ViT model.
[0018] In one possible implementation, performing a Zero-Shot NAS search on the first ViT model based on the accuracy and inference latency of the first ViT model to select the second ViT model includes:
[0019] Calculate the search evaluation score of the first ViT model based on latency sensitivity, accuracy of the first ViT model, and inference latency;
[0020] The model with the largest search evaluation score from the first ViT model is selected as the second ViT model.
[0021] In one possible implementation, a calculation formula for calculating the search evaluation score of the first ViT model based on delay sensitivity, the accuracy of the first ViT model, and the inference delay is as follows:
[0022] ,
[0023] Among them, TES is the search evaluation score, is the normalization function, is the MSA score of the lth layer in the first ViT model, is the MLP score of the lth layer in the first ViT model, is the inference delay of the first ViT model, is delay sensitivity.
[0024] In one possible implementation, the inference dataset sent by the client is obtained, and different servers perform collaborative inference based on the MPC protocol, the inference dataset, and the architecture and weights of the pre-trained ViT model to obtain the inference results, including:
[0025] The client sends the inference dataset to different servers through secret sharing;
[0026] Different servers perform collaborative calculations based on the MPC protocol, inference dataset, and the architecture and weights of the pre-trained ViT model to obtain different inference result blocks, which are then sent to the client through secret sharing.
[0027] The client reconstructs and decrypts different inference result blocks to obtain the inference results.
[0028] In a second aspect of the present application, a visual reasoning device based on zero-sample NAS is provided. The device comprises:
[0029] A first selection module is used to obtain a multi-party search space and randomly select a first ViT model from the multi-party search space;
[0030] A second selection module is configured to perform a Zero-Shot NAS search on the first ViT model based on the accuracy and inference latency of the first ViT model and select a second ViT model;
[0031] A pre-training module is used to retrain the second ViT model based on downstream data to obtain a pre-trained ViT model, and send the architecture and weights of the pre-trained ViT model to different servers through secret sharing;
[0032] The inference module obtains the inference data set sent by the client, and different servers perform collaborative inference based on the MPC protocol, inference data set, and the architecture and weights of the pre-trained ViT model to obtain the inference results.
[0033] In a third aspect of the present application, an electronic device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the program.
[0034] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present application is implemented.
[0035] The visual reasoning method based on zero-shot NAS provided in the embodiment of the present application obtains a multi-party search space and randomly selects a first ViT model from the multi-party search space. Then, a Zero-Shot NAS search is performed on the first ViT model based on the accuracy and reasoning delay of the first ViT model to select a second ViT model. The second ViT model is then retrained based on downstream data to obtain a pre-trained ViT model, and the architecture and weights of the pre-trained ViT model are sent to different servers through secret sharing. Finally, the reasoning data set sent by the client is obtained, and different servers perform collaborative reasoning based on the MPC protocol, the reasoning data set, the architecture and weights of the pre-trained ViT model to obtain the reasoning results. In other words, through the multi-party search space, a wider range of model architectures are covered, and the architecture of the first ViT model found is optimized. Then, the first ViT model is evaluated using Zero-Shot NAS search, and the second ViT model is selected from the first ViT model. There is no need to fully train each first ViT model, saving computing resources and time for model evaluation. Next, the second ViT model is retrained based on downstream data, making it better adapted to specific tasks and data, thereby improving the model's accuracy and practicality. Finally, through secret sharing and MPC protocols, the secure sharing and inference of the pre-trained model are ensured, protecting data privacy and model security.
[0036] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present application, nor are they intended to limit the scope of the present application. Other features of the present application will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0038] Figure 1 Flowchart of a large model security assessment method based on Zero-Shot NAS according to an embodiment of the present application;
[0039] Figure 2 Schematic diagram of the structure of the ViT model according to an embodiment of the present application;
[0040] Figure 3 A diagram of the system architecture involved in the method provided in the embodiments of this application;
[0041] Figure 4 1 is a block diagram of a large model security assessment device based on Zero-Shot NAS according to an embodiment of the present application;
[0042] Figure 5 A schematic diagram of the structure of a terminal device or server suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0043] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.
[0044] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0045] Figure 1 The flowchart of the large model security evaluation method based on Zero-Shot NAS according to the embodiment of the present disclosure is shown. Figure 1 , the method comprising:
[0046] S101: Obtain a multi-party search space and randomly select a first ViT model from the multi-party search space.
[0047] In this embodiment, compared with a single search space, the multi-party search space increases the diversity of the model architecture, providing a basis for subsequently improving the inference speed of multi-party secure computing in ViT model applications.
[0048] Optionally, the multi-party search space includes a ViT model that has been approximated by a nonlinear function;
[0049] Nonlinear function approximation processing includes approximating the Softmax function and GeLU function in the ViT model.
[0050] The ViT model is a model that applies the Transformer structure to image classification. Figure 2 is a structural diagram of the ViT model according to an embodiment of the present application, as shown in Figure 2 As shown:
[0051] The ViT model consists of an input layer, an embedding layer, an encoder layer, a multi-layer perceptron head, and an output layer. Among them, the input layer is responsible for standardizing and preprocessing the input data to make it conform to the input format of the model. The embedding layer is responsible for encoding the preprocessed data and converting the preprocessed data into a coded sequence. The encoder layer is responsible for feature extraction of the coded sequence data. The number of layers of the encoder layer is greater than one, and the specific number of layers can be customized as needed. The multi-head attention mechanism head is responsible for integrating the feature data from the multi-layer perceptron. The output layer is responsible for inference and prediction based on the integrated feature data to obtain the final inference result value. In addition, the structure of the encoder layer includes a normalization layer, a multi-head self-attention mechanism, a normalization layer, and a multi-layer perceptron in sequence. The normalization layer is responsible for normalization processing. The multi-head self-attention mechanism is the core of the encoder layer, capturing the dependencies in the sequence from different angles. The core calculation of the multi-head self-attention is based on the Softmax function, and the calculation formula is as follows:
[0052] ,
[0053] Among them, Q, K, and V are the query, key, and value values in the multi-head self-attention mechanism. This is the embedding encoding dimension of the key value, preventing excessively large values from causing instability in the multi-head self-attention mechanism. The multilayer perceptron is a fully connected feedforward neural network responsible for performing nonlinear transformations on the features learned by the self-attention mechanism, including but not limited to GeLU function transformations.
[0054] In one possible implementation, methods for approximating the Softmax function in the multi-head self-attention mechanism in the ViT model include, but are not limited to, ReLU approximation, second-order approximation, and ScaleAttn. The ReLU (Rectified Linear Unit) function is a commonly used activation function. ReLUSoftmax uses the ReLU function to approximate the Softmax function, significantly reducing computational cost. The calculation formula for the ReLUSoftmax function is as follows:
[0055] ,
[0056] in, is a preset very small value to avoid the error of dividing by zero. Then, the approximate calculation formula of the Softmax function of the ReLUSoftmax function applied to the multi-head self-attention mechanism can be expressed as:
[0057] .
[0058] The second-order approximation (quadratic approximation) is an approximation method that uses a quadratic function to approximate the local behavior of a given function near a certain point. Because it takes into account the curvature of the function, it is more accurate than the linear approximation (first-order approximation). The calculation formula 2QuadSoftmax using the second-order approximation to approximate the Softmax function is as follows:
[0059] ,
[0060] Among them, c is a preset hyperparameter with a value range of 1 to 10. Furthermore, the approximate calculation formula of the Softmax function applied to the multi-head self-attention mechanism can be expressed as:
[0061] .
[0062] ScaleAttn is an approximation method that removes the activation function and uses only linear operations to implement part of the attention mechanism calculation. It can provide performance comparable to the traditional attention mechanism and reduce computational complexity and inference time. The calculation formula of the ScaleAttn approximation function is as follows:
[0063] .
[0064] Among them, the value of N is equal to the number of multi-head self-attention in the multi-head self-attention mechanism.
[0065] In one possible implementation, methods for approximating the GeLU function in the multilayer perceptron in the ViT model include, but are not limited to, piecewise approximation. Piecewise approximation (PW) achieves approximation accuracy while reducing computational cost by using low-order polynomials for approximation. The calculation formula for piecewise approximation of the GeLU function in the multilayer perceptron in the ViT model is as follows:
[0066] ,
[0067] in, and The calculation formula is as follows:
[0068] ,
[0069]
[0070] In this embodiment, by performing nonlinear function approximation processing on the ViT model, the computational complexity of the model is reduced while the inference speed of the model is improved.
[0071] S102: Perform a Zero-Shot NAS search on the first ViT model based on the accuracy and inference delay of the first ViT model, and select a second ViT model.
[0072] NAS (Neural Architecture Search) is an automated neural network architecture design method that automatically constructs and searches for the optimal network structure in the search space, evaluates its performance, and continuously optimizes the model structure to improve model performance and efficiency. However, NAS requires a large amount of training data and neural network architecture, which in turn requires a large amount of computing resources. Zero-shot NAS, on the other hand, directly predicts the model's performance on the target task based on its own characteristics (typically based on the model structure and initialization weights) without any training or verification on the target task. It then uses this prediction to conduct network architecture search, allowing for rapid evaluation of model performance, saving computing resources, and improving the efficiency of model evaluation.
[0073] In this embodiment, Zero-Shot NAS search does not require actual training of all first ViT models, which significantly reduces computational costs and improves the efficiency of model evaluation.
[0074] Optionally, the accuracy evaluation index of the first ViT model includes an MSA score and an MLP score;
[0075] The MSA score is calculated based on the nuclear norm of the gradient matrix and the nuclear norm of the weight matrix in the first ViT model;
[0076] The MLP score is calculated based on the Frobenius inner product between the gradient matrix and the weight matrix in the first ViT model.
[0077] In one possible implementation, the MSA score is calculated as follows:
[0078] ,
[0079] in, is the MSA score of the l-th layer multi-head self-attention mechanism in the first ViT model, p is the number of weight matrices in the l-th layer multi-head self-attention mechanism in the first ViT model, is the nuclear norm of the gradient matrix in the l-th layer multi-head self-attention mechanism in the first ViT model, is the nuclear norm of the weight matrix in the l-th layer multi-head self-attention mechanism in the first ViT model.
[0080] In one possible implementation, the MLP score is calculated as follows:
[0081] ,
[0082] in, is the MLP score of the l-th layer multilayer perceptron in the first ViT model, q is the number of weight matrices in the l-th layer multilayer perceptron in the first ViT model, is the Frobenius inner product between the gradient matrix and the weight matrix in the first layer of the multilayer perceptron in the first ViT model. The Frobenius inner product, also known as the matrix inner product, is a binary operation on matrices of equal size, and its result is a scalar. The formula for calculating the Frobenius inner product is as follows:
[0083] ,
[0084] Among them, j, k are the number of rows and columns of the gradient matrix and weight matrix.
[0085] In this embodiment, the MSA score and the MLP score are used to evaluate the accuracy of the first ViT model, and the performance of the models is compared from multiple perspectives, providing a basis for the evaluation and selection of the subsequent second ViT model.
[0086] Optionally, the inference latency of the first ViT model is calculated based on the sum of the MSA inference latency and the MLP inference latency of each layer in the first ViT model;
[0087] The MSA inference delay is approximately calculated based on the inference delay of the multi-head attention mechanism in the first ViT model;
[0088] The MLP inference latency is approximately calculated based on the inference latency of the GeLU activation function in the first ViT model.
[0089] In one possible implementation, the MSA inference delay is approximately calculated based on the inference delay of the multi-head attention mechanism in the first ViT model, that is, , where h is the number of heads of the l-th layer multi-head attention mechanism in the first ViT model, and e is the embedding dimension of the multi-head attention mechanism. The MLP inference delay is approximately calculated based on the inference delay of the GeLU activation function in the first ViT model, that is, , where r is the ratio of the multilayer perceptron in the first layer of the ViT model, and e is the embedding dimension of the multilayer perceptron. Then, the inference delay of the first ViT model is The calculation formula is as follows:
[0090] .
[0091] In this embodiment, the inference delay of the first ViT model is calculated by summing the MSA inference delay and the MLP inference delay of each layer in the first ViT model, which reduces the amount of calculation, improves the speed of model evaluation, and facilitates subsequent optimization and improvement of the model.
[0092] Optionally, performing a Zero-ShotNAS search on the first ViT model according to the accuracy and inference latency of the first ViT model to select a second ViT model includes:
[0093] Calculate the search evaluation score of the first ViT model based on latency sensitivity, accuracy of the first ViT model, and inference latency;
[0094] The model with the largest search evaluation score from the first ViT model is selected as the second ViT model.
[0095] In one possible implementation, the second ViT model The selection formula can be expressed as:
[0096] ,
[0097] in, is the i-th first ViT model, and TES is the search evaluation score of the first ViT model.
[0098] In this embodiment, the search evaluation score of the first ViT model is calculated directly based on delay sensitivity, the accuracy of the first ViT model, and the inference delay through Zero-Shot NAS search, without the need for extensive training and verification, which significantly reduces the computational cost and time overhead in the search process.
[0099] Optionally, a calculation formula for calculating the search evaluation score of the first ViT model based on delay sensitivity, the accuracy of the first ViT model, and the inference delay is as follows:
[0100] ,
[0101] Among them, TES is the search evaluation score, is the normalization function, is the MSA score of the lth layer in the first ViT model, is the MLP score of the lth layer in the first ViT model, is the inference delay of the first ViT model, is delay sensitivity.
[0102] Among them, delay sensitivity The value can be customized according to user needs and the need for model inference speed.
[0103] In this embodiment, the search evaluation score of the first ViT model comprehensively considers delay sensitivity, accuracy, and inference delay, ensuring that the selected second ViT model meets the accuracy requirements while not affecting the actual deployment effect due to excessive inference delay.
[0104] S103: Retrain the second ViT model according to the downstream data to obtain a pre-trained ViT model, and send the architecture and weights of the pre-trained ViT model to different servers through secret sharing.
[0105] In this embodiment, the second ViT model is retrained using relevant downstream data to obtain the weights for the model's optimal performance. The architecture and weights are then sent to different servers through secret sharing, effectively protecting the model's data privacy and improving security.
[0106] S104: Obtain the inference data set sent by the client, and different servers perform collaborative inference based on the MPC protocol, the inference data set, the architecture and weights of the pre-trained ViT model to obtain the inference results.
[0107] In this embodiment, the MPC protocol is used to ensure the data privacy of multi-party collaborative computing and improve the security of pre-trained ViT model reasoning.
[0108] Optionally, the inference dataset sent by the client is obtained, and different servers perform collaborative inference based on the MPC protocol, the inference dataset, and the architecture and weights of the pre-trained ViT model to obtain inference results, including:
[0109] The client sends the inference dataset to different servers through secret sharing;
[0110] Different servers perform collaborative calculations based on the MPC protocol, inference dataset, and the architecture and weights of the pre-trained ViT model to obtain different inference result blocks, which are then sent to the client through secret sharing.
[0111] The client reconstructs and decrypts different inference result blocks to obtain the inference results.
[0112] Secret sharing is a cryptographic technique used to split a secret into multiple parts (called shares), which are distributed to different participants. The original secret can only be recovered when a certain number of participants (usually a predefined threshold) combine their shares. This ensures that no single participant can obtain the secret independently, thereby enhancing data security. The MPC protocol (Secure Multi-party Computation) allows multiple participants to collaboratively compute a function without disclosing their private data. This ensures that each participant only obtains the final computation result and cannot access the input data of other participants.
[0113] In this embodiment, by combining secret sharing and the MPC protocol, efficient and secure model reasoning is achieved while protecting data privacy.
[0114] Figure 3 This is a system architecture diagram of the method provided in the embodiment of the present application, such as Figure 3 As shown:
[0115] First, on the model provider side, the first ViT model is randomly selected from the multi-party search space. Then, a Zero-Shot NAS search is performed on the first ViT model based on its accuracy and inference latency, and a second ViT model is selected. The second ViT model is then retrained using downstream data to obtain a pre-trained ViT model. The architecture and weights of the pre-trained ViT model are sent to different servers in the service provider side through secret sharing. At the same time, the client sends the inference dataset to different servers in the service provider side through secret sharing. Finally, different servers perform collaborative inference based on the MPC protocol, inference dataset, architecture and weights of the pre-trained ViT model to obtain different inference result blocks, which are then sent to the client through secret sharing. The client reconstructs and decrypts the different inference result blocks to obtain the final inference result.
[0116] According to the embodiments of the present disclosure, the following technical effects are achieved:
[0117] 1. By incorporating ViT models that have been processed with nonlinear function approximation into the multi-party search space, a more robust and diverse ViT model architecture is provided.
[0118] 2. Using Zero-Shot NAS search effectively saves the time of the first ViT model evaluation, and there is no need to waste time and resources to train each model.
[0119] 3. The second ViT model evaluated has the best architecture and weights, which improves the speed and quality of model inference.
[0120] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0121] The above is an introduction to the method embodiment. The following is a device embodiment to further illustrate the solution described in this application.
[0122] Figure 4 FIG shows a block diagram of a large model security evaluation device based on Zero-Shot NAS according to an embodiment of the present application, as shown in FIG. Figure 4 Shown include:
[0123] A first selection module 401 is configured to obtain a multi-party search space and randomly select a first ViT model from the multi-party search space;
[0124] A second selection module 402 is configured to perform a Zero-Shot NAS search on the first ViT model based on the accuracy and inference latency of the first ViT model, and select a second ViT model;
[0125] A pre-training module 403 is used to retrain the second ViT model according to the downstream data to obtain a pre-trained ViT model, and send the architecture and weights of the pre-trained ViT model to different servers through secret sharing;
[0126] The reasoning module 404 obtains the reasoning data set sent by the client, and different servers perform collaborative reasoning based on the MPC protocol, the reasoning data set, the architecture and weights of the pre-trained ViT model to obtain the reasoning results.
[0127] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0128] Figure 5 A schematic diagram of the structure of a terminal device or server suitable for implementing an embodiment of the present application is shown.
[0129] like Figure 5As shown, the terminal device or server includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the terminal device or server are also stored. The CPU 501, ROM 502 and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0130] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, and the like; an output section 507 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or a modem. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 510 as needed, so that computer programs read therefrom can be installed into the storage section 508 as needed.
[0131] In particular, according to an embodiment of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the system of the present application are executed.
[0132] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the aforementioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0134] The units or modules involved in the embodiments described in this application may be implemented in software or hardware. The units or modules described may also be provided in a processor. The names of these units or modules do not, in certain circumstances, constitute limitations on the units or modules themselves.
[0135] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device. The computer-readable storage medium stores one or more programs, which, when used by one or more processors, execute the method described in the present application.
[0136] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of application involved in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the aforementioned application concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions applied for in this application.
Claims
1. A large model security evaluation method based on Zero-Shot NAS, characterized by: include: Obtain a multi-party search space, and randomly select a first ViT model from the multi-party search space; Performing a Zero-Shot NAS search on the first ViT model based on the accuracy and inference latency of the first ViT model to select a second ViT model; calculating a search evaluation score of the first ViT model based on latency sensitivity, the accuracy and inference latency of the first ViT model; The calculation formula is as follows: , Wherein, TES is the search evaluation score, is the normalization function, is the MSA score of the first layer in the ViT model, is the MLP score of the first layer in the first ViT model, is the inference latency of the first ViT model, is the delay sensitivity; Selecting the model with the largest search evaluation score from the first ViT models as the second ViT model; Retraining the second ViT model based on downstream data to obtain a pre-trained ViT model, and sending the architecture and weights of the pre-trained ViT model to different servers through secret sharing; The inference data set sent by the client is obtained, and the different servers perform collaborative inference according to the MPC protocol, the inference data set, and the architecture and weight of the pre-trained ViT model to obtain the inference result.
2. The large model security evaluation method based on Zero-Shot NAS according to claim 1 is characterized in that: The multi-party search space includes a ViT model that has been approximated by a nonlinear function; The nonlinear function approximation processing includes approximating the Softmax function and the GeLU function in the ViT model.
3. The large model security evaluation method based on Zero-Shot NAS according to claim 1 is characterized in that: The accuracy evaluation indicators of the first ViT model include MSA score and MLP score; The MSA score is calculated based on the nuclear norm of the gradient matrix and the nuclear norm of the weight matrix in the first ViT model; The MLP score is calculated based on the Frobenius inner product between the gradient matrix and the weight matrix in the first ViT model.
4. The large model security evaluation method based on Zero-Shot NAS according to claim 1, characterized in that: The inference latency of the first ViT model is calculated based on the sum of the MSA inference latency and the MLP inference latency of each layer in the first ViT model; The MSA inference delay is approximately calculated based on the inference delay of the multi-head attention mechanism in the first ViT model; The MLP inference delay is approximately calculated based on the inference delay of the GeLU activation function in the first ViT model.
5. The large model security evaluation method based on Zero-Shot NAS according to claim 1, characterized in that: The acquiring of the inference data set sent by the client, and the collaborative inference by the different servers according to the MPC protocol, the inference data set, and the architecture and weight of the pre-trained ViT model to obtain the inference result, include: The client sends the inference data set to the different servers through secret sharing; The different servers perform collaborative calculations according to the MPC protocol, the reasoning dataset, the architecture and weights of the pre-trained ViT model, obtain different reasoning result blocks, and send the different reasoning result blocks to the client through secret sharing; The client reconstructs and decrypts the different inference result blocks to obtain the inference results.
6. A large model security assessment device based on Zero-Shot NAS, characterized in that: include: A first selection module is configured to obtain a multi-party search space and randomly select a first ViT model from the multi-party search space; a second selection module, configured to perform a Zero-Shot NAS search on the first ViT model based on the accuracy and inference latency of the first ViT model to select a second ViT model; and calculate a search evaluation score of the first ViT model based on delay sensitivity, the accuracy and inference latency of the first ViT model; The calculation formula is as follows: , Wherein, TES is the search evaluation score, is the normalization function, is the MSA score of the first layer in the ViT model, is the MLP score of the first layer in the first ViT model, is the inference latency of the first ViT model, is the delay sensitivity; Selecting the model with the largest search evaluation score from the first ViT models as the second ViT model; A pre-training module, configured to retrain the second ViT model according to downstream data to obtain a pre-trained ViT model, and send the architecture and weights of the pre-trained ViT model to different servers through secret sharing; The reasoning module obtains the reasoning data set sent by the client, and the different servers perform collaborative reasoning according to the MPC protocol, the reasoning data set, the architecture and weights of the pre-trained ViT model to obtain the reasoning results.
7. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method and device for evaluating model reasoning ability and storage medium
CN119066381A