Model training method and apparatus, feature extraction method and apparatus, device, and medium
By using a feature extraction model based on MPCFormer and SETR models, one-dimensional sequence processing and self-attention mechanism calibration of images is solved, and the technical challenges of image feature extraction in privacy calculations are achieved, and efficient and accurate feature extraction is achieved.
Patent Information
- Application Number
- PCT/CN2024/142026
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-22
AI Technical Summary
In privacy computing, there are technical challenges in how to quickly and accurately extract high-precision image feature vectors, especially in multi-party security computing scenarios.
A feature extraction model trained by training logic based on the MPCFormer model framework structure and semantic visual Transformer model SETR is obtained by performing one-dimensional sequence processing on the image and using the self-attention mechanism to extract and calibrate more accurate feature vectors.
It realizes the rapid and accurate extraction of high-precision image feature vectors, improves the efficiency and accuracy of multi-party security calculations, reduces calculation time and reduces communication bottlenecks.
Smart Images

Figure CN2024142026_22052025_PF_FP_ABST
Abstract
Description
Model training method, feature extraction method, device, equipment and medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on November 14, 2023, with application number 202311518180.8 and application name "Model training method, feature extraction method, device, equipment and medium", all contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of data processing technology, and in particular to a feature extraction model training method, a convolutional neural network training method, a model training method, a feature extraction method, an apparatus, a device and a medium. Background Art
[0004] Secure Multi-Party Computation (MPC) in privacy computing is a method in which multiple parties collaborate to complete computing goals without a trusted third party, ensuring that each party cannot obtain any input information from other parties except the calculation results, which can effectively protect the user's privacy information.
[0005] If the input information is an image containing the user's private information, to protect the user's privacy, feature extraction can be performed on the image and multi-party secure computation can be performed based on the extracted feature vectors. However, how to quickly and accurately extract high-precision image feature vectors is a technical problem that needs to be solved urgently. Summary of the Invention
[0006] The present application provides a feature extraction model training method, a convolutional neural network training method, a model training method, a feature extraction method, an apparatus, a device and a medium for quickly and accurately extracting high-precision image feature vectors.
[0007] In a first aspect, the present application provides a feature extraction model training method, the method comprising:
[0008] For any first sample image obtained, a first feature vector of the first sample image is obtained based on a trained convolutional neural network; and based on a preset image processing method, the first sample image is processed into a one-dimensional sequence, the one-dimensional sequence is input into a feature extraction model to be trained, and a second feature vector of the first sample image is obtained based on the feature extraction model; based on the second feature vector, the first feature vector is calibrated to obtain a calibrated first reference feature vector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0009] The feature extraction model is trained based on the first benchmark feature vector, the label feature vector carried by the first sample image, and the configured first loss function.
[0010] In a possible implementation, after obtaining the second eigenvector of the first sample image and before calibrating the first eigenvector based on the second eigenvector, the method further includes:
[0011] Perform global average pooling on the second feature vector to obtain global feature vector information corresponding to each dimensional feature in the second feature vector, and based on the probability distribution function, obtain the importance weight coefficient corresponding to each dimensional feature in the global feature vector information; calibrate the second feature vector based on the importance weight coefficient corresponding to each dimensional feature to obtain a calibrated second feature vector; based on the calibrated second feature vector, perform subsequent steps of calibrating the first feature vector based on the second feature vector.
[0012] In a possible implementation, the first loss function includes an importance weight coefficient corresponding to each dimensional feature.
[0013] In a possible implementation, calibrating the first eigenvector based on the second eigenvector to obtain a calibrated first reference eigenvector includes:
[0014] Perform an XOR operation on each dimensional feature included in the second eigenvector and each dimensional feature included in the first eigenvector to obtain a calibrated first reference eigenvector.
[0015] In a possible implementation, after performing an XOR operation on each dimensional feature included in the second feature vector and each dimensional feature included in the first feature vector, and before obtaining the calibrated first reference feature vector, the method further includes:
[0016] The vector obtained after the XOR operation is normalized, and a calibrated first reference eigenvector is determined based on the normalized vector.
[0017] In a possible implementation, the excitation function used by the feature extraction model includes: a 2quad excitation function.
[0018] In a possible implementation, the feature extraction model is a model obtained based on knowledge distillation.
[0019] In a second aspect, the present application provides a convolutional neural network training method, the method comprising:
[0020] For any acquired second sample image, the second sample image is input into a convolutional neural network to be trained, and a third eigenvector of the second sample image is obtained based on the convolutional neural network; and based on a preset image processing method, the second sample image is processed into a one-dimensional sequence, and the one-dimensional sequence is input into a trained feature extraction model, and a fourth eigenvector of the second sample image is obtained based on the feature extraction model; and based on the fourth eigenvector, the third eigenvector is calibrated to obtain a calibrated second reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0021] The convolutional neural network is trained based on the second benchmark feature vector, the label feature vector carried by the second sample image, and the configured second loss function.
[0022] In a possible implementation, calibrating the third eigenvector based on the fourth eigenvector to obtain a calibrated second reference eigenvector includes:
[0023] Perform an XOR operation on each dimensional feature included in the fourth eigenvector and each dimensional feature included in the third eigenvector to obtain a calibrated second reference eigenvector.
[0024] In a possible implementation, after performing an XOR operation on each dimensional feature included in the fourth eigenvector and each dimensional feature included in the first eigenvector, and before obtaining the calibrated second reference eigenvector, the method further includes:
[0025] The vector obtained after the XOR operation is normalized, and a calibrated second reference eigenvector is determined based on the normalized vector.
[0026] In a third aspect, the present application provides a model training method, the method comprising:
[0027] For any third sample image obtained, a fifth eigenvector of the third sample image is obtained based on the convolutional neural network to be trained; and based on a preset image processing method, the third sample image is processed into a one-dimensional sequence, the one-dimensional sequence is input into a feature extraction model to be trained, and a sixth eigenvector of the third sample image is obtained based on the feature extraction model; and based on the sixth eigenvector, the fifth eigenvector is calibrated to obtain a calibrated third reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0028] The feature extraction model is trained based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured first loss function; and the convolutional neural network is trained based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured second loss function.
[0029] In a fourth aspect, the present application provides a feature extraction method, comprising:
[0030] receiving an image to be processed;
[0031] A feature vector of the image to be processed is obtained based on a feature extraction model trained by the method according to any one of the first and third aspects, or a convolutional neural network trained by the method according to any one of the second and third aspects.
[0032] In one possible implementation, the method further includes:
[0033] Based on the feature vector, multi-party secure computation is performed.
[0034] In a fifth aspect, the present application provides a feature extraction model training device, the device comprising:
[0035] A first calibration module is configured to obtain, for any acquired first sample image, a first feature vector of the first sample image based on a trained convolutional neural network; process the first sample image into a one-dimensional sequence based on a preset image processing method, input the one-dimensional sequence into a feature extraction model to be trained, and obtain a second feature vector of the first sample image based on the feature extraction model; calibrate the first feature vector based on the second feature vector to obtain a calibrated first reference feature vector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0036] The first training module is used to train the feature extraction model based on the first benchmark feature vector, the label feature vector carried by the first sample image, and a configured first loss function.
[0037] In a possible implementation manner, the first calibration module is further configured to:
[0038] Perform global average pooling on the second feature vector to obtain global feature vector information corresponding to each dimensional feature in the second feature vector, and based on the probability distribution function, obtain the importance weight coefficient corresponding to each dimensional feature in the global feature vector information; calibrate the second feature vector based on the importance weight coefficient corresponding to each dimensional feature to obtain a calibrated second feature vector; based on the calibrated second feature vector, perform subsequent steps of calibrating the first feature vector based on the second feature vector.
[0039] In a possible implementation, the first loss function includes an importance weight coefficient corresponding to each dimensional feature.
[0040] In a possible implementation manner, the first calibration module is specifically configured to:
[0041] Perform an XOR operation on each dimensional feature included in the second eigenvector and each dimensional feature included in the first eigenvector to obtain a calibrated first reference eigenvector.
[0042] In a possible implementation manner, the first calibration module is further configured to:
[0043] The vector obtained after the XOR operation is normalized, and a calibrated first reference eigenvector is determined based on the normalized vector.
[0044] In a possible implementation, the excitation function used by the feature extraction model includes: a 2quad excitation function.
[0045] In a possible implementation, the feature extraction model is a model obtained based on knowledge distillation.
[0046] In a sixth aspect, the present application provides a convolutional neural network training device, the device comprising:
[0047] A second calibration module is configured to input any acquired second sample image into a convolutional neural network to be trained, and obtain a third eigenvector of the second sample image based on the convolutional neural network; process the second sample image into a one-dimensional sequence based on a preset image processing method, input the one-dimensional sequence into a trained feature extraction model, and obtain a fourth eigenvector of the second sample image based on the feature extraction model; calibrate the third eigenvector based on the fourth eigenvector to obtain a calibrated second reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0048] The second training module is used to train the convolutional neural network based on the second benchmark feature vector, the label feature vector carried by the second sample image, and the configured second loss function.
[0049] In a possible implementation manner, the second calibration module is specifically configured to:
[0050] Perform an XOR operation on each dimensional feature included in the fourth eigenvector and each dimensional feature included in the third eigenvector to obtain a calibrated second reference eigenvector.
[0051] In a possible implementation manner, the second calibration module is further configured to:
[0052] The vector obtained after the XOR operation is normalized, and a calibrated second reference eigenvector is determined based on the normalized vector.
[0053] In a seventh aspect, the present application provides a model training device, comprising:
[0054] a third calibration module configured to obtain, for any acquired third sample image, a fifth eigenvector of the third sample image based on the convolutional neural network to be trained; and, based on a preset image processing method, process the third sample image into a one-dimensional sequence, input the one-dimensional sequence into a feature extraction model to be trained, and obtain, based on the feature extraction model, a sixth eigenvector of the third sample image; and calibrate the fifth eigenvector based on the sixth eigenvector to obtain a calibrated third reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0055] The third training module is used to train the feature extraction model based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured first loss function; and to train the convolutional neural network based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured second loss function.
[0056] In an eighth aspect, the present application provides a feature extraction device, comprising:
[0057] A receiving module, configured to receive an image to be processed;
[0058] An extraction module is used to obtain the feature vector of the image to be processed based on a feature extraction model trained based on the method described in any one of the first and third aspects, or a convolutional neural network trained based on the method described in any one of the second and third aspects.
[0059] In a possible implementation, the device further includes:
[0060] The multi-party secure computing module is used to perform multi-party secure computing based on the feature vector.
[0061] In a ninth aspect, the present application also provides an electronic device, which includes at least a processor and a memory, and the processor is used to implement the steps of any of the above methods when executing a computer program stored in the memory.
[0062] In a tenth aspect, the present application provides a computer-readable storage medium storing a computer program, which implements the steps of any of the above methods when executed by a processor.
[0063] Since in the embodiment of the present application, the first sample image can be processed into a one-dimensional sequence, and the one-dimensional sequence is input into the feature extraction model, so that the feature extraction model can perform feature extraction on the first sample image, and the feature extraction model in the embodiment of the present application is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the SETR model, the feature extraction model can capture more accurate feature vectors such as faces (second feature vectors) based on the self-attention mechanism, and can calibrate the first feature vector obtained by the convolutional neural network based on the second feature vector to obtain a first baseline feature vector that can more accurately express the image features. Based on the first baseline feature vector and the configured first loss function, the feature extraction model is trained, which can improve the accuracy of the trained feature extraction model in determining (extracting) the image feature vector, and achieve the purpose of quickly and accurately extracting high-precision image feature vectors.
[0064] In addition, the feature extraction model in the embodiment of the present application is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the SETR model, so that the feature extraction model can obtain pixel-level accurate feature vectors, thereby improving the accuracy of the trained feature extraction model in determining (extracting) image feature vectors. In addition, in the embodiment of the present application, unimportant features in the first feature vector, etc. can be weakened or removed (ignored) based on the self-attention mechanism, so that fewer and more precise feature vectors can be obtained, thereby improving the accuracy of the obtained feature vectors while reducing computational time and improving efficiency without incurring additional computational overhead, and can also weaken the communication bottleneck of MPC. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the implementation methods in the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0066] FIG1 shows a schematic diagram of a first feature extraction model training process provided by some embodiments;
[0067] FIG2A shows a schematic diagram of an image corresponding to a first feature vector obtained based on a convolutional neural network in related art, provided by some embodiments;
[0068] FIG2B shows a schematic diagram of an image corresponding to a second feature vector obtained based on a feature extraction model provided by some embodiments;
[0069] FIG3 shows a schematic diagram of a second feature extraction model training process provided by some embodiments;
[0070] FIG4 shows a schematic diagram of a third feature extraction model training process provided by some embodiments;
[0071] FIG5 is a schematic diagram showing the communication duration of different excitation functions provided by some embodiments;
[0072] FIG6 shows a schematic diagram of a Transformer model provided by some embodiments;
[0073] FIG7 shows a schematic diagram of a SETR model provided by some embodiments;
[0074] FIG8 shows a schematic diagram of a convolutional neural network training process provided by some embodiments;
[0075] FIG9 shows a schematic diagram of a multi-party secure computation process provided by some embodiments;
[0076] FIG10 shows a schematic diagram of a model training process provided by some embodiments;
[0077] FIG11 shows a schematic diagram of a feature extraction process provided by some embodiments;
[0078] FIG12 shows a schematic diagram of a feature extraction model training device provided by some embodiments;
[0079] FIG13 shows a schematic diagram of a convolutional neural network training device provided by some embodiments;
[0080] FIG14 shows a schematic diagram of a model training device provided by some embodiments;
[0081] FIG15 shows a schematic diagram of a feature extraction device provided by some embodiments;
[0082] FIG16 shows a schematic structural diagram of an electronic device provided in some embodiments. DETAILED DESCRIPTION
[0083] In order to quickly and accurately extract (obtain) high-precision image feature vectors, the present application provides a feature extraction model training method, a convolutional neural network training method, a model training method, a feature extraction method, an apparatus, a device and a medium.
[0084] In order to make the purpose and implementation of this application clearer, the exemplary implementation of this application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only part of the embodiments of this application, not all of the embodiments.
[0085] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0086] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.
[0087] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0088] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
[0090] All implementation methods of the embodiments of this application comply with the relevant provisions of national laws and regulations regarding the acquisition, storage, use, and processing of data.
[0091] Example 1:
[0092] FIG1 shows a schematic diagram of a first feature extraction model training process provided by some embodiments. As shown in FIG1 , the process includes the following steps:
[0093] S101: For any first sample image obtained, based on the trained convolutional neural network, obtain a first feature vector of the first sample image; and based on a preset image processing method, process the first sample image into a one-dimensional sequence, input the one-dimensional sequence into a feature extraction model to be trained, and obtain a second feature vector of the first sample image based on the feature extraction model; based on the second feature vector, calibrate the first feature vector to obtain a calibrated first baseline feature vector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR.
[0094] In one possible implementation, the feature extraction model may be a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the Semantic Visual Transformer (SETR) model. How to obtain the feature extraction model in the embodiment of the present application based on the framework structure of the MPCFormer model and the training logic of the SETR model will be introduced later and will not be repeated here.
[0095] In one possible embodiment, when training a feature extraction model, to improve the accuracy of feature vectors obtained by the feature extraction model, any first sample image can be obtained from the first sample image set and input into a trained convolutional neural network (CNN), such as a deep residual network (ResNet). The convolutional neural network then performs feature extraction on the first sample image to obtain a first feature vector for the first sample image. The first sample image can be an image containing user privacy information, such as a biometric feature such as a face, and this application does not specifically limit the first sample image.
[0096] In a possible implementation, the first sample image can be processed into a one-dimensional sequence based on a preset image processing method, the one-dimensional sequence can be input into a feature extraction model to be trained, and features can be extracted from the first sample image based on the feature extraction model to obtain a second feature vector of the first sample image; based on the second feature vector, the first feature vector can be calibrated to obtain a more accurate feature vector (first baseline feature vector).
[0097] Specifically, considering that the framework structure of the feature extraction model is based on the MPCFormer model, the MPCFormer model in the related art is usually used to process natural language. In order to enable the feature extraction model to process images, in an embodiment of the present application, the first sample image can be processed into a one-dimensional sequence, and the one-dimensional sequence can be input into the feature extraction model. The training logic of the semantic visual Transformer (SEgmentation TRansformer, SETR) model is used to train the feature extraction model, so that the operation logic of the feature extraction model is the same as the operation logic of the SETR model, so that the feature extraction model can not only process images, but also obtain more accurate image features at the pixel level based on the operation logic of the SETR model, and obtain more accurate (few and fine) image features (feature vectors) based on the multi-head self-attention mechanism, etc., thereby improving the speed and accuracy of multi-party secure computing (MPC). Among them, how to process the first sample image into a one-dimensional sequence and how to obtain the feature extraction model in the embodiment of this application based on the framework structure of the MPCFormer model and the training logic of the SETR model will be introduced later. The following is an explanation of the process of calibrating the first eigenvector based on the second eigenvector.
[0098] For ease of understanding, in the embodiment of the present application, the first eigenvector is represented by F, and the second eigenvector is represented by A. Based on the second eigenvector A, when calibrating the first eigenvector, the second eigenvector A can be used to pool the first eigenvector F, that is, each dimensional feature contained in the second eigenvector (Attention Maps) is subjected to an XOR operation (⊙) with each dimensional feature contained in the first eigenvector. For example, the vector v obtained by performing an XOR operation on the kth vector (Attention Map) in the second eigenvector (Attention Maps) is k =∑ i,j F i,j,k ⊙A i,j,k , where “i,j,k” represents the pixel in the i-th row and j-th column of the k-th vector (Attention Map).
[0099] In one possible implementation, when performing an XOR operation on each dimensional feature included in the second feature vector and each dimensional feature included in the first feature vector, if the pixels of the corresponding features in both the second feature vector and the first feature vector are high, the pixels of the feature are considered to be of higher quality, and the pixels of the corresponding feature obtained after the XOR operation are configured to be higher pixels. If the pixels of the features corresponding to either side are poor, the pixels of the feature are considered to be poor, and the pixels of the corresponding feature obtained after the XOR operation are configured to be lower pixels, so that more accurate (fewer but more precise) image features (feature vectors) can be obtained through calibration.
[0100] Optionally, taking the second eigenvector as an example, if the pixels of a certain feature are higher than the average pixels of the features in the second eigenvector, the pixels of the feature can be considered to be higher. Conversely, if the pixels of a certain feature are not higher than the average pixels of the features in the second eigenvector, the pixels of the feature can be considered to be lower. Similarly, if the pixels of a certain feature in the first eigenvector are higher than the average pixels of the features in the first eigenvector, the pixels of the feature can be considered to be higher. Conversely, if the pixels of the feature are not higher than the average pixels of the features in the first eigenvector, the pixels of the feature can be considered to be lower. The above-mentioned specific calculation process of performing the same or operation on each dimensional feature contained in the second eigenvector and each dimensional feature contained in the first eigenvector is only an exemplary description, and this application does not specifically limit the specific calculation process of the same or operation.
[0101] In one possible implementation, after performing an XOR operation on each dimension of the second feature vector and each dimension of the first feature vector, the vectors obtained after the XOR operation can be normalized based on a fully connected layer to normalize them to feature spaces of the same size, and then the normalized vectors are averaged to obtain a calibrated first reference feature vector, i.e., the final calibrated first reference feature vector.
[0102] Refer to Figures 2A and 2B, wherein Figure 2A shows a schematic diagram of an image corresponding to a first feature vector obtained based on a convolutional neural network in related technologies, as provided in some embodiments. Figure 2B shows a schematic diagram of an image corresponding to a second feature vector obtained based on a feature extraction model, as provided in some embodiments. The black areas in Figure 2B are irrelevant feature areas. It can be seen that the feature extraction model can use the self-attention mechanism to capture more specific facial features in the image, while ignoring the remaining irrelevant feature areas in the image. The Attention mechanism allows the features of the image subject part that are truly valuable to be extracted. Specifically for facial features, the effect can be manifested as Attention can filter out facial features, while shielding other image information such as lighting and background in the environment, and representing these other image information with black areas, thereby obtaining fewer, more precise, and more accurate facial features.
[0103] In a possible implementation, taking the first sample image as an image containing a face as an example, the face in the image may have a portion that is blocked or tilted. In order to weaken this portion of the feature, thereby obtaining a more prominent and accurate facial feature, after obtaining the second eigenvector A of the first sample image, before calibrating the first eigenvector F based on the second eigenvector A, the second eigenvector A may also be calibrated. Based on the calibrated second eigenvector A ~ The first eigenvector is calibrated to further improve the accuracy of the first reference eigenvector obtained. The second eigenvector A is calibrated to obtain the calibrated second eigenvector A. ~ The process can be as follows:
[0104] Optionally, the second eigenvector A can be subjected to global average pooling (GAP) to obtain the global eigenvector information corresponding to each dimension of the second eigenvector, and the obtained global eigenvector information can be mapped to the softmax probability distribution function (f ex(·)=softmax(·)), based on the softmax probability distribution function (also known as the normalized exponential function), the more important feature information in the second feature vector A is further amplified, and the importance weight coefficient s corresponding to each dimension of the feature in the global feature vector information is obtained. For example, the importance weight coefficient s can be: Among them, the importance weight coefficient s corresponding to the more important features in the second eigenvector (such as facial features that are not blocked and have no skewed angles) may be larger, while the importance weight coefficient s corresponding to the less important features in the second eigenvector (such as facial features that are blocked or skewed at an angle) may be smaller, thereby weakening the parts of features such as the face that are blocked or skewed at an angle, setting a lower degree of activation (weight coefficient) for the parts in the Attention Maps that are blocked or skewed at an angle, thereby further improving the accuracy of the determined first benchmark eigenvector.
[0105] In a possible implementation, the second eigenvector A can be transformed using a sigmoid function, and then the transformed second eigenvector (such as sigmoid (A k )) and the corresponding importance weight coefficient s k Multiply them to get the calibrated second eigenvector A ~ For example, the calibrated second eigenvector
[0106] Get the calibrated second eigenvector A ~ After that, we can use the calibrated second eigenvector A ~ , calibrate the first eigenvector F to obtain the first reference eigenvector f. ~ The process of calibrating the first eigenvector F is similar to the process of calibrating the first eigenvector F based on the second eigenvector A. For example, the second eigenvector (Attention Maps) A ~ Each dimension of the feature contained in is ANDed with each dimension of the feature contained in the first feature vector (⊙), and the kth vector (Attention Map) obtained after the OR operation is: I will not go into details here.
[0107] S102: Training the feature extraction model based on the first benchmark feature vector, the label feature vector carried by the first sample image, and a configured first loss function.
[0108] In one possible implementation, the first sample image may carry a preconfigured label feature vector, wherein this application does not specifically limit the configuration process of the label feature vector. A feature extraction model may be trained based on the first baseline feature vector f, the label feature vector carried by the first sample image, and the configured first loss function L, thereby obtaining a feature extraction model that can be deployed online for extracting image feature vectors (e.g., feature vectors of facial images).
[0109] Optionally, the first loss function L may include a cross entropy loss function, a weighted differential regularization loss function, and an L2 regularization loss function. In order to extract high-precision image feature vectors, the first loss function may include an importance weight coefficient corresponding to each dimension of the feature. For example, when configuring the cross entropy loss function in the first loss function L, the importance weight coefficient s may be used based on the above-mentioned importance weight coefficient s. k And the cross entropy loss parameter L CE,k , to determine the cross entropy loss function. For example, we can use the importance weight coefficient s corresponding to each dimension feature k The cross entropy loss parameter L of the corresponding dimension CE,k The sum of the products of and determines the cross entropy loss function. For ease of understanding, the cross entropy loss function L can be expressed in formula form. wCE Expressed as: L wCE =∑ k s k ·L CE,k The cross entropy loss function L in the embodiment of this application is used. wCE The loss caused by the obscured or angled unimportant parts of the Attention Maps to the model convergence can be constrained by its own importance weight coefficient s, so that this part of the loss does not participate in the calculation of model convergence, thereby improving the accuracy of the feature vector determined by the trained model (feature extraction model). Among them, the cross entropy loss parameter L can be determined based on existing technology CE,k , I will not go into details here.
[0110] Optionally, when configuring the weighted differential regularized loss function in the first loss function L, the importance weight coefficient s can be used based on the above k And the weighted differential loss term parameter P i,j,k , to determine the weighted differential regularization loss function L wDIV For ease of understanding, the weighted differential regularization loss function L can be expressed in formula form as wDIV Expressed as: L wDIV =1-∑ i,j max k (s k ·P i,j,k), where the weighted differential loss term parameter Weighted differential loss parameter P i,j,k It can be considered as a penalty term for cross-overlapping features, which can be based on the weighted differentiation regularization loss function L wDIV Weaken or remove the overlapping features in the feature vector so that the loss caused by these features does not participate in the calculation of model convergence, thereby improving the accuracy of the feature vector determined by the trained model (feature extraction model). The weighted differential loss term parameter P can be determined based on existing technology. i,j,k , which will not be described here. In addition, the existing technology can also be used to determine the L2 regularization loss function L REG , and I will not go into details.
[0111] Optionally, it can be the cross entropy loss function L in the first loss function wCE , weighted differential regularization loss function L wDIV And the L2 regularization loss function L REG Each of these three sub-loss functions is assigned a corresponding weight, and the loss function L is determined based on the sum of the products of each sub-loss function and the corresponding weight. For example, the cross entropy loss function L wCE The corresponding weight is λ wCE Represents the weighted differential regularization loss function L wDIV The corresponding weight is λ wDIV Indicates that the L2 regularization loss function L REG The corresponding weight is λ REG Indicates that the first loss function L = λ wCE L wCE +λ wDIV L wDIV +λ REG L REG . Among them, the specific training process of training the feature extraction model based on the first loss function can adopt existing technology and will not be repeated here. This application does not specifically limit the weights corresponding to each sub-loss function and can be flexibly set according to needs.
[0112] Optionally, when training the feature extraction model, the accuracy of the feature extraction model's recognition results can be determined based on whether the first baseline feature vector is consistent with the label feature vector carried by the first sample image. In a specific implementation, if they are inconsistent, indicating that the feature extraction model's recognition results are inaccurate, the parameters of the feature extraction model need to be adjusted to gradually bring the first baseline feature vector obtained based on the feature extraction model closer to the label feature vector, thereby training the feature extraction model.
[0113] In a specific implementation, when adjusting the parameters in the feature extraction model, a gradient descent algorithm can be used to back-propagate the gradients of the parameters of the feature extraction model, thereby training the feature extraction model.
[0114] In one possible implementation, the above operation may be performed for each first sample image in the first sample set. When a preset convergence condition is met, the feature extraction model training is determined to be complete. The preset convergence condition may be satisfied by the original feature extraction model correctly identifying more than a set number of first sample images in the first sample set, or by the number of iterations of training the feature extraction model reaching a set maximum number of iterations. This setting may be flexible in specific implementations and is not specifically limited here.
[0115] In one possible implementation, when training the original feature extraction model, the sample images in the sample set can be divided into training sample images and test sample images. The original feature extraction model is first trained based on the training sample images, and then the reliability of the trained feature extraction model is verified based on the test sample images.
[0116] Since in the embodiment of the present application, the first sample image can be processed into a one-dimensional sequence, and the one-dimensional sequence is input into the feature extraction model, so that the feature extraction model can perform feature extraction on the first sample image, and the feature extraction model in the embodiment of the present application is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the SETR model, the feature extraction model can capture more accurate feature vectors such as faces (second feature vectors) based on the self-attention mechanism, and can calibrate the first feature vector obtained by the convolutional neural network based on the second feature vector to obtain a first baseline feature vector that can more accurately express the image features. Based on the first baseline feature vector and the configured first loss function, the feature extraction model is trained, which can improve the accuracy of the trained feature extraction model in determining (extracting) the image feature vector, and achieve the purpose of quickly and accurately extracting high-precision image feature vectors.
[0117] In addition, the feature extraction model in the embodiment of the present application is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the SETR model, so that the feature extraction model can obtain pixel-level accurate feature vectors, thereby improving the accuracy of the trained feature extraction model in determining (extracting) image feature vectors. In addition, in the embodiment of the present application, unimportant features in the first feature vector, etc. can be weakened or removed (ignored) based on the self-attention mechanism, so that fewer and more precise feature vectors can be obtained, thereby improving the accuracy of the obtained feature vectors while reducing computational time and improving efficiency without incurring additional computational overhead, and can also weaken the communication bottleneck of MPC.
[0118] For ease of understanding, the feature extraction model training process provided by this application is explained below through a specific embodiment. Referring to Figure 3, Figure 3 shows a schematic diagram of the second feature extraction model training process provided by some embodiments, which includes the following steps:
[0119] S301: For any acquired first sample image, a first feature vector of the first sample image is obtained based on a trained convolutional neural network. The first sample image is then processed into a one-dimensional sequence based on a preset image processing method. The one-dimensional sequence is input into a feature extraction model to be trained, and a second feature vector of the first sample image is obtained based on the feature extraction model. The feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the SETR model.
[0120] S302: Perform global average pooling on the second eigenvector to obtain global eigenvector information corresponding to each dimensional feature in the second eigenvector, and obtain the importance weight coefficient corresponding to each dimensional feature in the global eigenvector information based on the probability distribution function; calibrate the second eigenvector based on the importance weight coefficient corresponding to each dimensional feature to obtain a calibrated second eigenvector; calibrate the first eigenvector based on the calibrated second eigenvector to obtain a calibrated first baseline eigenvector.
[0121] S303: Training a feature extraction model based on the first benchmark feature vector, the label feature vector carried by the first sample image, and the configured first loss function.
[0122] For ease of understanding, the feature extraction model training process provided by this application is further explained below through a specific embodiment. Referring to FIG4 , FIG4 shows a schematic diagram of a third feature extraction model training process provided by some embodiments, which includes the following steps:
[0123] After obtaining any first sample image, the first sample image can be input into a trained convolutional neural network (such as a CNN neural network such as ResNet), and feature extraction is performed on the first sample image based on the convolutional neural network to obtain a feature map of the first sample image, i.e., a first feature vector (Feature Maps) F. At the same time, the first sample image can be processed into a one-dimensional sequence, wherein the process of processing the first sample image into a one-dimensional sequence can be as follows:
[0124] A 2D first sample image (such as a face image) with a length × width × number of color channels (C) of H × W × 3 (unit: pixel) can be split into The blocks are stretched into a one-dimensional (1D) sequence of length L = HW / 256, and the 1D sequence is used as input embedding (Imput Embedding) and input into the feature extraction model (Transformer) to be trained. Optionally, a position embedding algorithm can be used to encode the original image position information, and for each vector e in the 1D sequence, i Add the corresponding position code p i , and finally get the one-dimensional sequence (input sequence) E that is input to the feature extraction model, where E can be expressed as: {e1+p 1, e2+p2,…,e L +p L}.
[0125] In one possible implementation, a one-dimensional sequence can be input into a feature extraction model to be trained. The framework structure of the feature extraction model can adopt the framework structure of the MPCFormer model, and the operation logic (also called training logic) adopts the operation logic of the SETR model. Based on the encoder (Encoder) in the feature extraction model, the second feature vector (Attention Maps) extracted based on the self-attention mechanism can be obtained. Optionally, each encoder in the feature extraction model can be composed of a multi-head self-attention layer (MSA) and a feedforward neural network (also called a multi-layer perceptron, MLP). Assuming that there are L e Encoder. Among them, the input of the l-th layer of self-attention is composed of Z l-1 The calculated three-dimensional tuple (query, key, value): where query = Z l-1 W Q , key = Z l-1 W k ,value=Z l-1 WV Among them, W Q , W k , W V . is a learnable weight matrix, d is the dimension of the three-dimensional tuple (query, key, value), then self-attention can be expressed as:
[0126] Multi-head (multi-layer) self-attention is obtained by splicing m SAs together, MSA (Z l-1 )=[SA1(Z l-1 );SA2(Z l-1 );…;SA m (Z l-1 )]Wo, where Wo∈R md×C . The output of each Encoder: Z l =MSA(Z l-1 )+MLP(MSA(Z l-1 ))∈R L×C , the total output of all Encoders is {Z 1 ,Z 2 ,…,Z Le}. Among them, converting a two-dimensional image into a one-dimensional sequence can be considered as converting the image x∈R C×H×W Transformed into sequence Z∈R L×C Optionally, when performing feature extraction reasoning based on the feature extraction model, the feature extraction model may be used to implement the extraction process of the second feature vector A (Attention Maps) through a general MPC computing engine.
[0127] Optionally, after obtaining the second feature vector A based on the feature extraction model, the second feature vector A can be adjusted (referred to as recalibration based on Attention Maps in the figure). When calibrating the second feature vector A, the second feature vector can be firstly subjected to global average pooling (global pooling) to obtain the global feature vector information corresponding to each dimension of the second feature vector, and based on the softmax probability distribution function (referred to as the SoftMax normalized exponential function (fex( . ))), obtain the importance weight coefficient s (can also be represented by capital S) corresponding to each dimension of the global feature vector information. In addition, the second feature vector A can be transformed using the sigmoid function, and then the transformed second feature vector (sigmoid (A k )) is multiplied by the corresponding importance weight coefficient s (calibration product) to obtain the calibrated second eigenvector A ~ For example, the calibrated second eigenvector
[0128] Get the calibrated second eigenvector A ~ After that, we can use the calibrated second eigenvector A ~ , calibrate the first eigenvector F to obtain the calibrated first reference eigenvector f. For example, the second eigenvector (Attention Maps) A ~ Each dimension of the feature contained in is ANDed with each dimension of the feature contained in the first feature vector (⊙), and the kth vector (Attention Map) obtained after the OR operation is: The process of calibrating the first feature vector by the above-mentioned XOR operation can also be called an attention-based pooling process.
[0129] Optionally, after obtaining the first baseline feature vector, a feature extraction model may be trained based on the first baseline feature vector and a configured first loss function, etc. The trained feature extraction model may be deployed online to extract feature vectors of images such as facial images.
[0130] Exemplarily, when feature extraction is performed using a trained feature extraction model, upon receiving an image to be processed for feature extraction, the image to be processed can be input into the trained feature extraction model, and a feature vector of the image to be processed can be obtained based on the output result of the feature extraction model.
[0131] In one possible implementation, after obtaining the feature vector of the image to be processed, a multi-party secure computation process can be performed based on the feature vector. For example, a feature extraction model can be encrypted and deployed on a server, with one or more clients encrypting and uploading images and other data. Multi-party collaborative reasoning or extraction of image feature vectors can then be performed to perform a multi-party secure computation process. The process of performing multi-party secure computation based on feature vectors can utilize existing technologies and will not be further described here.
[0132] Example 2:
[0133] To ensure the execution efficiency of multi-party secure computation and reduce the communication latency between different participants in the multi-party secure computation, based on the above embodiment, in the embodiment of the present application, the excitation function used in the feature extraction model includes: a 2quad excitation function. In addition, the feature extraction model can also be a model derived from knowledge distillation. In one possible implementation, the selection of different excitation functions can have different effects on the computational efficiency of the MPC computing engine. Referring to Figure 5, Figure 5 shows a schematic diagram of the communication duration of different excitation functions provided in some embodiments. It can be seen that when the excitation function is an activation function based on a Gaussian Error Linear Unit (GELU) function, the total runtime (referred to as total time in the figure) is 11 seconds, of which the communication duration between participants (referred to as comm time in the figure) is 9.6 seconds. When the excitation function is quad, the total runtime is 0.5 seconds, of which the communication duration between participants is 9.6 seconds. When the excitation function is softmax, the total runtime is 40 seconds, of which the communication duration between participants is 34.1 seconds. When the activation function is 2relu, the total running time is 14.7s, of which the communication time between participants is 12.8s. When , the total running time is 3.2s, of which the communication time between participants is 1.7s. It can be seen that when the excitation function of the feature extraction model includes the 2quad excitation function, the communication time between participants and the total running time can be greatly reduced. Optionally, in addition to the 2quad excitation function, the excitation function of the feature extraction model can also include the quad excitation function (GeLU(x)≈0.125x 2 +0.25x+0.5).
[0134] In one possible implementation, to improve the inference speed and accuracy of the feature extraction model, the feature extraction model can be a small model derived through knowledge distillation. For example, a larger feature extraction model (MPCFormer model) can be derived through knowledge distillation to form a smaller model, and the small model derived through knowledge distillation can be used for feature extraction.
[0135] The following introduces the process of obtaining the feature extraction model based on the framework structure of the MPCFormer model and the training logic of the SETR model.
[0136] The Transformer model is the foundation of the ChatGPT large model. It is a deep neural network model based on the self-attention mechanism that can efficiently process sequence data in parallel. The Transformer model is a neural network model that learns context and thus meaning by tracking relationships in sequence data (such as the words in this sentence). It has achieved great success in language understanding, especially in large models. Before the emergence of Transformer, users had to use large labeled data sets to train neural networks. The production cost of the data sets was high and time-consuming. As a basic model, the Transformer model can replace Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) in many cases. The Transformer model can be a large encoder and decoder module for processing data, implemented based on the encoder-decoder framework. See Figure 6, which shows a schematic diagram of a Transformer model provided by some embodiments. The Transformer model is roughly divided into two parts: encoder and decoder, corresponding to the left and right parts in Figure 6 (left and right shown in the figure, where the left part corresponds to the encoder and the right part corresponds to the decoder).
[0137] Specifically, the input of the Transformer model consists of two parts: input embedding (word embedding conversion) and position-based word embedding (also called positional encoding). Whether it's text or images, the digital representation of the input embedding can be converted into a vector representation, capturing the relationships between text or images in a high-dimensional space. Positional encoding is added after the input embedding to incorporate information with different meanings generated by different input positions into the vector, supplementing the positional information. Both the encoder and decoder contain inputs, and the input structure of the two parts is the same. The difference lies in their usage during inference. The encoder performs inference only once, while the decoder performs recursive inference similar to an RNN, continuously generating prediction results.
[0138] The encoder portion of the Transformer model consists of N stacked encoder layers, each of which consists of two connected sublayers. The first sublayer includes a multi-head self-attention (MSA) layer, a normalization layer, and a residual connection. The second sublayer includes a feedforward fully connected layer, a normalization layer, and a residual connection, and so on. Both sublayers are supplemented with an Add & Normalization layer. Each encoder layer performs a feature extraction process on the input, also known as the encoding process.
[0139] The Transformer model's decoder also stacks N identical layers, but the structure of each layer differs slightly from that of the encoder. In addition to the two sublayers in the encoder, each decoder layer also includes a masked multi-head attention layer (Mask Multi-Head Attention). As shown in the figure, each sublayer also uses an Add & Normalization layer.
[0140] The Transformer model's decoder extracts features from the target, a process known as decoding, based on the encoder's output and the previous prediction. At the output, a linear transformation is performed through a fully connected output layer (Linear), converting the output to a specific dimension. A softmax layer is then used to scale the numbers in the one-dimensional vector to a probability range of 0-1, with the sum of these values equal to 1, resulting in the model's prediction (Output Probabilities).
[0141] Optionally, the Transformer model can be pre-trained on a large dataset to gain general understanding and then fine-tuned on a small downstream dataset to learn task-specific features.
[0142] The reasoning process of a Transformer model can be expressed as a two-party computation (2PC). In 2PC, for example, a user inputs data, and a model provider inputs the Transformer model. They jointly compute an inference result. Throughout the inference process, 2PC ensures that both parties only know information about their own inputs and results. Multi-party computation (MPC) and knowledge distillation techniques can be used to implement Transformer reasoning with multi-party participation. The MPCFormer model uses a trained (or fine-tuned) Transformer model. The MPCFormer model is a Transformer model with low inference latency and high ML performance. It can use a given MPC for partial approximate computation. Knowledge distillation techniques can be used to construct a high-performance approximate Transformer model during computation. During inference, the MPCFormer model leverages the MPC engine to implement private model reasoning.
[0143] The reasoning process based on the MPCFormer model can be described as: a convolution calculation phase and an inference phase. The convolution calculation phase can be described as: S = MPCFormer(T, D, A), where T is the trained Transformer model, D is the input dataset, and A is the MPC-based calculation process. The inference phase can be described as: y = MPCs(X). In the MPC-based Transformer model inference, the GeLU function and the Softmax function are the main sources of communication bottlenecks (high communication complexity). The GeLU function is slow to calculate (time-consuming) because the calculation of the error function requires a high-order Taylor expansion (involving a large number of multiplication operations), and the Softmax function is slow to calculate because the exponential function needs to be evaluated through multiple square iterations, which also involves a large number of multiplication operations.
[0144] The SETR model is introduced below. SETR consists of three main parts: input → transformation → output. Referring to Figure 7 , a schematic diagram of a SETR model provided by some embodiments is shown. Part (a) of Figure 7 primarily depicts the input preprocessing and feature extraction process of the SETR model, part (b) primarily depicts the progressive upsampling process, and part (c) primarily depicts the multi-level feature aggregation process.
[0145] First, the original input image needs to be processed into a format that can be supported by Transformer (SETR). That is, the input image is sliced, and each 2D image slice (patch) is sliced into a one-dimensional (1D) sequence, and the one-dimensional sequence is input into the model as a whole. In order to encode the spatial information of each slice, a specific embedding can be learned for each local position and added to a linear projection function to form the final input sequence. In this way, even though Transformer (SETR) is disordered, the corresponding spatial position information (p) can still be retained.
[0146] The one-dimensional sequence is input into the SETR model. The fully connected layers (linear projection layers) in the SETR model process the one-dimensional sequence to obtain the slice vector (Patch Embedding) and the position information vector (Position Embedding) corresponding to the slice. The SETR model contains several transformer layers. Optionally, any transformer layer mainly consists of two parts: the Multi-head Self-Attention layer (MSA) and the Multilayer Perceptron (MLP). Both parts are connected to the layer normalization process (Layer Normalization, Layer Norm). Among them, Layer Norm is a normalization technology used in deep neural networks. It can normalize the output of each neuron in the network so that the output of each layer in the network has a similar distribution.
[0147] Feature extraction can be performed by inputting a one-dimensional sequence into the Transformer architecture (SETR model). The encoder is used to compress the spatial resolution of the original input image and gradually extract more advanced abstract semantic features. The decoder is used to upsample the high-level features extracted by the encoder to the original input resolution for pixel-level prediction, thereby achieving image segmentation.
[0148] Referring to part (b) of Figure 7, the goal of the Decoder is to generate segmentation results on the original two-dimensional image (H×W). The output Z of the Encoder needs to be reshaped from the two-dimensional HW / 256×C to the three-dimensional feature map H / 16×W / 16×C.
[0149] The decoder can use progressive upsampling (PUP). Specifically, considering that direct one-step upsampling will introduce a lot of noise, the method of alternating convolution and upsampling can be used to set the upsampling multiple to 2 to reduce the noise. Le To get the original image size, a total of 4 upsampling operations are required.
[0150] Refer to part (c) of Figure 7 for the process of multi-level feature fusion (MLA). Multi-level feature fusion is similar to feature pyramid, but the difference is that feature Z L From each layer of Transformer and have the same resolution. Every L e / M layers extract one layer of features {Z m}(m∈{L e / M,2L e / M,…,ML e / M), a total of M paths are extracted. Among them, each layer first converts Z Le The image is reshaped from HW / 256×C to H / 16×W / 16×C, then processed using a three-layer neural network (convolution kernel sizes 1×1, 3×3, and 3×3). The number of channels is halved in the first and third layers, respectively, and then upsampled by 4x bilinear interpolation after the third layer. Starting from the second path, the features of the previous paths are sequentially fused to enhance information interaction between each path. After another 3×3 convolution layer, the feature maps obtained from each path are concatenated in the channel dimension and upsampled by 4x bilinear interpolation to restore the original image size.
[0151] Among them, the process of obtaining the feature extraction model based on the framework structure of the MPCFormer model and the training logic of SETR can be understood as replacing the SETR model framework structure in part (a) of Figure 7 with the MPCFormer model framework structure, and still using the training logic (operation logic) in parts (b) and (c) of Figure 7 to train the MPCFormer model, so as to obtain the feature extraction model.
[0152] Example 3:
[0153] Based on the same technical concept, the present application also provides a convolutional neural network training method. Referring to FIG8 , FIG8 shows a schematic diagram of a convolutional neural network training process provided by some embodiments. The process includes the following steps:
[0154] S801: For any second sample image obtained, the second sample image is input into the convolutional neural network to be trained, and based on the convolutional neural network, the third eigenvector of the second sample image is obtained; and based on a preset image processing method, the second sample image is processed into a one-dimensional sequence, and the one-dimensional sequence is input into the trained feature extraction model, and based on the feature extraction model, the fourth eigenvector of the second sample image is obtained; based on the fourth eigenvector, the third eigenvector is calibrated to obtain a calibrated second reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR.
[0155] In one possible implementation, when training a convolutional neural network (CNN), such as a deep residual network (ResNet), in order to improve the accuracy of the feature vector obtained by the convolutional neural network, any second sample image can be obtained from the second sample image set, and the obtained second sample image can be input into the convolutional neural network to be trained. The second sample image and the first sample image can be the same or different, and this application does not specifically limit this. For example, the second sample image and the first sample image can be images containing occluded faces or faces at an angle, etc.
[0156] Optionally, the convolutional neural network performs feature extraction on the second sample image to obtain a third eigenvector F of the second sample image, wherein for convenience of description, the third eigenvector and the first eigenvector obtained by the convolutional neural network are both represented by F.
[0157] In one possible implementation, the second sample image can be processed into a one-dimensional sequence based on a preset image processing method, and the one-dimensional sequence can be input into a trained feature extraction model. Feature extraction can be performed on the second sample image based on the feature extraction model to obtain a fourth eigenvector of the second sample image. The feature extraction model in the embodiment of the present application is the same as the feature extraction model in the above-described embodiments. Both are models based on the framework structure of the MPCFormer model and the training logic of the SETR model, and are not further described here. For ease of description, the fourth eigenvector and the second eigenvector obtained by the feature extraction model are both represented by A.
[0158] Optionally, in order to improve the accuracy of the feature vector obtained by the convolutional neural network, the third feature vector F can be calibrated based on the fourth feature vector A to obtain a calibrated second reference feature vector f. The process of calibrating the third feature vector F based on the fourth feature vector A is similar to the process of calibrating the first feature vector F based on the second feature vector A in the above embodiment. For example, each dimensional feature contained in the fourth feature vector and each dimensional feature contained in the third feature vector can be subjected to an XOR operation to obtain the calibrated second reference feature vector. It is also possible to perform an XOR operation on each dimensional feature contained in the fourth feature vector and each dimensional feature contained in the third feature vector, normalize the vector obtained after the XOR operation, and determine the calibrated second reference feature vector based on the normalized vector. In addition, the fourth eigenvector can also be globally averaged pooled to obtain the global eigenvector information corresponding to each dimensional feature in the fourth eigenvector, and based on the probability distribution function, the importance weight coefficient corresponding to each dimensional feature in the global eigenvector information is obtained. Based on the importance weight coefficient corresponding to each dimensional feature, the fourth eigenvector is calibrated to obtain the calibrated fourth eigenvector (A~); based on the calibrated fourth eigenvector, the third eigenvector is calibrated, which will not be repeated here.
[0159] S802: Training the convolutional neural network based on the second benchmark feature vector, the label feature vector carried by the second sample image, and the configured second loss function.
[0160] In one possible implementation, the second sample image may carry a preconfigured label feature vector, wherein this application does not specifically limit the configuration process of the label feature vector. A convolutional neural network may be trained based on the second baseline feature vector f, the label feature vector carried by the second sample image, and the configured second loss function, thereby obtaining a convolutional neural network that can be deployed online for extracting image feature vectors (e.g., feature vectors of facial images).
[0161] Optionally, the second loss function may include a cross entropy loss function, a weighted differential regularization loss function, and an L2 regularization loss function. The second loss function may be configured using existing techniques, which will not be described in detail here.
[0162] Optionally, when training the convolutional neural network, the accuracy of the convolutional neural network's recognition result can be determined based on whether the second baseline feature vector is consistent with the label feature vector carried by the second sample image. In a specific implementation, if they are inconsistent, it indicates that the recognition result of the convolutional neural network is inaccurate, and the parameters of the convolutional neural network need to be adjusted to gradually bring the second baseline feature vector obtained based on the convolutional neural network closer to the label feature vector, thereby training the convolutional neural network.
[0163] In a specific implementation, when adjusting the parameters in the convolutional neural network, a gradient descent algorithm can be used to backpropagate the gradients of the parameters of the convolutional neural network, thereby training the convolutional neural network.
[0164] In one possible implementation, the above operation may be performed for each second sample image in the second sample set. When a preset convergence condition is met, the convolutional neural network training is determined to be complete. The preset convergence condition may be satisfied when the second sample images in the second sample set are correctly identified by the original convolutional neural network, and the number of sample images correctly identified exceeds a set number, or the number of iterations of training the convolutional neural network reaches a set maximum number of iterations. This setting may be flexible in specific implementations and is not specifically limited here.
[0165] In one possible implementation, when training the original convolutional neural network, the sample images in the sample set can be divided into training sample images and test sample images. The original convolutional neural network is first trained based on the training sample images, and then the reliability of the trained convolutional neural network is verified based on the test sample images.
[0166] Since in the embodiment of the present application, the second sample image can be input into the feature extraction model, so that the feature extraction model can perform feature extraction on the second sample image, and the feature extraction model in the embodiment of the present application is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the SETR model, the feature extraction model can capture more accurate feature vectors such as faces (fourth feature vectors) based on the self-attention mechanism, and can calibrate the third feature vector obtained by the convolutional neural network based on the fourth feature vector to obtain a second baseline feature vector that can more accurately express the image features. Based on the second baseline feature vector and the configured second loss function, the convolutional neural network is trained, which can improve the accuracy of the trained convolutional neural network in determining (extracting) image feature vectors, and achieve the purpose of quickly and accurately extracting high-precision image feature vectors.
[0167] For example, when using a trained convolutional neural network for feature extraction, upon receiving an image to be processed for feature extraction, the image to be processed can be input into the trained convolutional neural network, and a feature vector of the image to be processed can be obtained based on the output of the convolutional neural network. After obtaining the feature vector of the image to be processed, a multi-party secure computation process can be performed based on the feature vector. For example, referring to FIG9 , FIG9 shows a schematic diagram of a multi-party secure computation process provided by some embodiments. The convolutional neural network can be deployed on a server (server side). After obtaining the feature vector of the image to be processed through the trained convolutional neural network, the feature vector can be secretly sharded and sent to multiple clients (client sides) collaborating on the multi-party secure computation. For example, the feature vector can be divided into secret sharing shard 1 and secret sharing shard 2. Secret sharing shard 1 can be distributed to the client of participant 1 performing the multi-party secure computation, and secret sharing shard 2 can be distributed to the client of participant 2 performing the multi-party secure computation, and so on. Each client side collaborates to extract features based on the convolutional neural network (CNN) in a secret sharing manner to perform the multi-party secure computation. This multi-party secure computation method can prevent the leakage of users' biometric privacy information such as their faces, thereby improving the security of user data. The process of performing multi-party secure computation based on feature vectors can be implemented using existing technologies and will not be further described here.
[0168] Example 4:
[0169] Based on the same technical concept, the present application also provides a model training method. Referring to FIG10 , FIG10 shows a schematic diagram of a model training process provided by some embodiments. The process includes the following steps:
[0170] S1001: For any third sample image obtained, based on the convolutional neural network to be trained, obtain the fifth eigenvector of the third sample image; and based on a preset image processing method, process the third sample image into a one-dimensional sequence, input the one-dimensional sequence into the feature extraction model to be trained, and based on the feature extraction model, obtain the sixth eigenvector of the third sample image; based on the sixth eigenvector, calibrate the fifth eigenvector to obtain a calibrated third baseline eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR.
[0171] In one possible implementation, the convolutional neural network and the feature extraction model can be trained separately and simultaneously. For example, when the convolutional neural network and the feature extraction model are trained, in order to improve the accuracy of the feature vectors obtained by the trained convolutional neural network and the feature extraction model, any third sample image can be obtained from the third sample image set, and the obtained third sample image can be input into the convolutional neural network to be trained. The third sample image can be the same as or different from the first sample image and the second sample image, and this application does not impose any specific restrictions on this. For example, the third sample image can be an image containing an occluded face or a face with an angled angle, etc.
[0172] Optionally, a convolutional neural network performs feature extraction on the second sample image to obtain a fifth eigenvector F of the third sample image, where, for convenience of description, the fifth eigenvector, the third eigenvector and the first eigenvector obtained by the convolutional neural network are all represented by F.
[0173] In one possible implementation, the third sample image can be processed into a one-dimensional sequence based on a preset image processing method, and the one-dimensional sequence can be input into a feature extraction model to be trained. Feature extraction can be performed on the third sample image based on the feature extraction model to obtain a sixth eigenvector of the third sample image. The feature extraction model in the embodiment of the present application is the same as the feature extraction model in the aforementioned embodiments. Both are models based on the framework structure of the MPCFormer model and the training logic of the SETR model, and are not further described here. For ease of description, the sixth eigenvector, the fourth eigenvector, and the second eigenvector obtained by the feature extraction model are all represented by A.
[0174] Optionally, in order to improve the accuracy of the feature vectors obtained by the convolutional neural network and the feature extraction model, the fifth feature vector F can be calibrated based on the sixth feature vector A to obtain a calibrated third reference feature vector f. The process of calibrating the fifth feature vector F based on the sixth feature vector A is similar to the process of calibrating the first feature vector F based on the second feature vector A and the process of calibrating the third feature vector F based on the fourth feature vector A in the above embodiment. For example, each dimensional feature contained in the sixth feature vector can be subjected to an XOR operation with each dimensional feature contained in the fifth feature vector to obtain the calibrated third reference feature vector. Alternatively, each dimensional feature contained in the sixth feature vector can be subjected to an XOR operation with each dimensional feature contained in the fifth feature vector, and the vector obtained after the XOR operation can be normalized. Based on the normalized vector, the calibrated third reference feature vector can be determined. In addition, the sixth eigenvector can also be globally averaged pooled to obtain the global eigenvector information corresponding to each dimensional feature in the sixth eigenvector, and based on the probability distribution function, the importance weight coefficient corresponding to each dimensional feature in the global eigenvector information can be obtained. Based on the importance weight coefficient corresponding to each dimensional feature, the sixth eigenvector is calibrated to obtain the calibrated sixth eigenvector (A~); based on the calibrated sixth eigenvector, the fifth eigenvector is calibrated, etc., which will not be repeated here.
[0175] S1002: Based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured first loss function, the feature extraction model is trained; and based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured second loss function, the convolutional neural network is trained.
[0176] Among them, the process of training the feature extraction model based on the third baseline feature vector, the label feature vector carried by the third sample image and the configured first loss function is similar to the process of training the feature extraction model based on the first baseline feature vector, the label feature vector carried by the first sample image and the configured first loss function in the above embodiment, and will not be repeated here.
[0177] The process of training the convolutional neural network based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured second loss function is similar to the process of training the convolutional neural network based on the second baseline feature vector, the label feature vector carried by the second sample image, and the configured second loss function in the above embodiment, and will not be repeated here.
[0178] Example 5:
[0179] Based on the same technical concept, the present application also provides a feature extraction method. Referring to FIG11 , FIG11 shows a schematic diagram of a feature extraction process provided by some embodiments. The process includes the following steps:
[0180] S1101: Receive an image to be processed.
[0181] S1102: Obtain a feature vector of the image to be processed based on a feature extraction model or convolutional neural network trained by any of the above methods.
[0182] In one possible implementation, the method further includes:
[0183] Based on the feature vector, multi-party secure computation is performed.
[0184] Example 6:
[0185] Based on the same technical concept, the present application provides a feature extraction model training device. Referring to FIG12 , FIG12 shows a schematic diagram of a feature extraction model training device provided by some embodiments. The device includes:
[0186] The first calibration module 1201 is configured to obtain, for any acquired first sample image, a first feature vector of the first sample image based on a trained convolutional neural network; process the first sample image into a one-dimensional sequence based on a preset image processing method, input the one-dimensional sequence into a feature extraction model to be trained, and obtain a second feature vector of the first sample image based on the feature extraction model; calibrate the first feature vector based on the second feature vector to obtain a calibrated first reference feature vector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0187] The first training module 1202 is used to train the feature extraction model based on the first benchmark feature vector, the label feature vector carried by the first sample image, and the configured first loss function.
[0188] In a possible implementation manner, the first calibration module 1201 is further configured to:
[0189] Perform global average pooling on the second feature vector to obtain global feature vector information corresponding to each dimensional feature in the second feature vector, and based on the probability distribution function, obtain the importance weight coefficient corresponding to each dimensional feature in the global feature vector information; calibrate the second feature vector based on the importance weight coefficient corresponding to each dimensional feature to obtain a calibrated second feature vector; based on the calibrated second feature vector, perform subsequent steps of calibrating the first feature vector based on the second feature vector.
[0190] In a possible implementation, the first loss function includes an importance weight coefficient corresponding to each dimensional feature.
[0191] In a possible implementation manner, the first calibration module 1201 is specifically configured to:
[0192] Perform an XOR operation on each dimensional feature included in the second eigenvector and each dimensional feature included in the first eigenvector to obtain a calibrated first reference eigenvector.
[0193] In a possible implementation manner, the first calibration module 1201 is further configured to:
[0194] The vector obtained after the XOR operation is normalized, and a calibrated first reference eigenvector is determined based on the normalized vector.
[0195] In a possible implementation, the excitation function used by the feature extraction model includes: a 2quad excitation function.
[0196] In a possible implementation, the feature extraction model is a model obtained based on knowledge distillation.
[0197] Example 7:
[0198] Based on the same technical concept, the present application provides a convolutional neural network training device. Referring to FIG13 , FIG13 shows a schematic diagram of a convolutional neural network training device provided in some embodiments, the device comprising:
[0199] The second calibration module 1301 is configured to input any acquired second sample image into a convolutional neural network to be trained, and obtain a third eigenvector of the second sample image based on the convolutional neural network; process the second sample image into a one-dimensional sequence based on a preset image processing method, input the one-dimensional sequence into a trained feature extraction model, and obtain a fourth eigenvector of the second sample image based on the feature extraction model; calibrate the third eigenvector based on the fourth eigenvector to obtain a calibrated second reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0200] The second training module 1302 is used to train the convolutional neural network based on the second benchmark feature vector, the label feature vector carried by the second sample image, and the configured second loss function.
[0201] In a possible implementation manner, the second calibration module 1301 is specifically configured to:
[0202] Perform an XOR operation on each dimensional feature included in the fourth eigenvector and each dimensional feature included in the third eigenvector to obtain a calibrated second reference eigenvector.
[0203] In a possible implementation manner, the second calibration module 1301 is further configured to:
[0204] The vector obtained after the XOR operation is normalized, and a calibrated second reference eigenvector is determined based on the normalized vector.
[0205] Example 8:
[0206] Based on the same technical concept, the present application provides a model training device. Referring to FIG14 , FIG14 shows a schematic diagram of a model training device provided by some embodiments, the device comprising:
[0207] The third calibration module 1401 is configured to obtain, for any acquired third sample image, a fifth eigenvector of the third sample image based on the convolutional neural network to be trained; process the third sample image into a one-dimensional sequence based on a preset image processing method, input the one-dimensional sequence into a feature extraction model to be trained, and obtain a sixth eigenvector of the third sample image based on the feature extraction model; and calibrate the fifth eigenvector based on the sixth eigenvector to obtain a calibrated third reference eigenvector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR;
[0208] The third training module 1402 is used to train the feature extraction model based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured first loss function; and to train the convolutional neural network based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured second loss function.
[0209] Example 9:
[0210] Based on the same technical concept, the present application provides a feature extraction device. Referring to FIG15 , FIG15 shows a schematic diagram of a feature extraction device provided by some embodiments, the device comprising:
[0211] Receiving module 1501, configured to receive an image to be processed;
[0212] The extraction module 1502 is used to obtain the feature vector of the image to be processed based on the feature extraction model trained by the method described in any one of the first and third aspects, or the convolutional neural network trained by the method described in any one of the second and third aspects.
[0213] In a possible implementation, the device further includes:
[0214] The multi-party secure computing module is used to perform multi-party secure computing based on the feature vector.
[0215] Example 10:
[0216] Based on the same technical concept, the present application also provides an electronic device. FIG16 shows a schematic structural diagram of an electronic device provided in some embodiments. As shown in FIG16 , the electronic device includes: a processor 1601, a communication interface 1602, a memory 1603, and a communication bus 1604. The processor 1601, the communication interface 1602, and the memory 1603 communicate with each other via the communication bus 1604.
[0217] The memory 1603 stores a computer program. When the program is executed by the processor 1601, the processor 1601 executes the steps in any of the above method embodiments, which will not be described in detail here.
[0218] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0219] The communication interface 1602 is used for communication between the electronic device and other devices.
[0220] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk memory. Alternatively, the memory may be at least one storage device located away from the processor.
[0221] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it can also be a digital signal processing processor (DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.
[0222] Example 11:
[0223] Based on the same technical concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program that can be executed by an electronic device. When the program runs on the electronic device, the electronic device implements the steps in any of the above method embodiments when executing, which will not be repeated here.
[0224] Based on the same technical concept, the present application provides a computer program product, which includes: computer program code, which, when executed on a computer, enables the computer to implement the method described in any of the above method embodiments applied to an electronic device.
[0225] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any combination thereof, and may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions that, when loaded and executed on a computer, fully or partially generate the processes or functions described in the embodiments of the present application.
[0226] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0227] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flow charts and / or one or more boxes in the block diagram.
[0228] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0229] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0230] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A feature extraction model training method, the method comprising: For any first sample image obtained, based on the trained convolutional neural network, obtain a first feature vector of the first sample image; Based on a preset image processing method, the first sample image is processed into a one-dimensional sequence, the one-dimensional sequence is input into a feature extraction model to be trained, and based on the feature extraction model, a second feature vector of the first sample image is obtained; based on the second feature vector, the first feature vector is calibrated to obtain a calibrated first reference feature vector, wherein the feature extraction model is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR; The feature extraction model is trained based on the first benchmark feature vector, the label feature vector carried by the first sample image, and the configured first loss function.
2. The method according to claim 1, wherein: After obtaining the second eigenvector of the first sample image and before calibrating the first eigenvector based on the second eigenvector, the method further includes: Perform global average pooling on the second feature vector to obtain global feature vector information corresponding to each dimensional feature in the second feature vector, and based on a probability distribution function, obtain the importance weight coefficient corresponding to each dimensional feature in the global feature vector information; based on the importance weight coefficient corresponding to each dimensional feature, calibrate the second feature vector to obtain a calibrated second feature vector; based on the calibrated second feature vector, perform subsequent steps of calibrating the first feature vector based on the second feature vector.
3. The method according to claim 2, wherein: The first loss function includes the importance weight coefficient corresponding to each dimensional feature.
4. The method according to claim 1, wherein: The step of calibrating the first feature vector based on the second feature vector to obtain a calibrated first reference feature vector includes: An XOR operation is performed on each dimensional feature included in the second feature vector and each dimensional feature included in the first feature vector to obtain a calibrated first reference feature vector.
5. The method according to claim 4, wherein: After performing an XOR operation on each dimensional feature included in the second feature vector and each dimensional feature included in the first feature vector, and before obtaining the calibrated first reference feature vector, the method further includes: The vector obtained after the XOR operation is normalized, and a calibrated first reference feature vector is determined based on the normalized vector.
6. The method according to claim 1, wherein: The excitation function used by the feature extraction model includes: 2quad excitation function.
7. The method according to claim 1, wherein: The feature extraction model is a model obtained based on knowledge distillation.
8. A convolutional neural network training method, the method comprising: For any second sample image obtained, the second sample image is input into a convolutional neural network to be trained, and based on the convolutional neural network, a third feature vector of the second sample image is obtained; and based on a preset image processing method, the second sample image is processed into a one-dimensional sequence, and the one-dimensional sequence is input into a trained feature extraction model, and based on the feature extraction model, a fourth feature vector of the second sample image is obtained; based on the fourth feature vector, the third feature vector is calibrated to obtain a calibrated second reference feature vector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR; The convolutional neural network is trained based on the second benchmark feature vector, the label feature vector carried by the second sample image, and the configured second loss function.
9. The method according to claim 8, wherein: The step of calibrating the third eigenvector based on the fourth eigenvector to obtain a calibrated second reference eigenvector includes: Perform an XOR operation on each dimensional feature included in the fourth eigenvector and each dimensional feature included in the third eigenvector to obtain a calibrated second reference eigenvector.
10. The method according to claim 9, wherein: After performing an XOR operation on each dimensional feature included in the fourth eigenvector and each dimensional feature included in the first eigenvector, and before obtaining the calibrated second reference eigenvector, the method further includes: The vector obtained after the XOR operation is normalized, and a calibrated second reference feature vector is determined based on the normalized vector.
11. A model training method, the method comprising: For any third sample image obtained, based on the convolutional neural network to be trained, obtain a fifth eigenvector of the third sample image; Based on a preset image processing method, the third sample image is processed into a one-dimensional sequence, the one-dimensional sequence is input into a feature extraction model to be trained, and based on the feature extraction model, a sixth feature vector of the third sample image is obtained; based on the sixth feature vector, the fifth feature vector is calibrated to obtain a calibrated third reference feature vector, wherein the feature extraction model is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR; The feature extraction model is trained based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured first loss function; and the convolutional neural network is trained based on the third baseline feature vector, the label feature vector carried by the third sample image, and the configured second loss function.
12. A feature extraction method, the method comprising: receiving an image to be processed; A feature vector of the image to be processed is obtained based on a feature extraction model trained by the method according to any one of claims 1-7 and 11, or a convolutional neural network trained by the method according to any one of claims 8-11.
13. The method according to claim 12, further comprising: Based on the feature vector, multi-party secure computation is performed.
14. A feature extraction model training device, the device comprising: A first calibration module, configured to obtain, for any first sample image acquired, a first feature vector of the first sample image based on a trained convolutional neural network; Based on a preset image processing method, the first sample image is processed into a one-dimensional sequence, the one-dimensional sequence is input into a feature extraction model to be trained, and based on the feature extraction model, a second feature vector of the first sample image is obtained; based on the second feature vector, the first feature vector is calibrated to obtain a calibrated first reference feature vector, wherein the feature extraction model is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR; The first training module is used to train the feature extraction model based on the first benchmark feature vector, the label feature vector carried by the first sample image, and a configured first loss function.
15. A convolutional neural network training device, the device comprising: A second calibration module is used for inputting any acquired second sample image into a convolutional neural network to be trained, and obtaining a third feature vector of the second sample image based on the convolutional neural network; and processing the second sample image into a one-dimensional sequence based on a preset image processing method, and inputting the one-dimensional sequence into a trained feature extraction model, and obtaining a fourth feature vector of the second sample image based on the feature extraction model; and calibrating the third feature vector based on the fourth feature vector to obtain a calibrated second reference feature vector, wherein the feature extraction model is a model trained based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR; The second training module is used to train the convolutional neural network based on the second benchmark feature vector, the label feature vector carried by the second sample image and the configured second loss function.
16. A model training device, comprising: A third calibration module, configured to obtain, for any acquired third sample image, a fifth eigenvector of the third sample image based on the convolutional neural network to be trained; Based on a preset image processing method, the third sample image is processed into a one-dimensional sequence, the one-dimensional sequence is input into a feature extraction model to be trained, and based on the feature extraction model, a sixth feature vector of the third sample image is obtained; based on the sixth feature vector, the fifth feature vector is calibrated to obtain a calibrated third reference feature vector, wherein the feature extraction model is a model obtained by training based on the framework structure of the MPCFormer model and the training logic of the semantic visual Transformer model SETR; The third training module is used to train the feature extraction model based on the third baseline feature vector, the label feature vector carried by the third sample image and the configured first loss function; and to train the convolutional neural network based on the third baseline feature vector, the label feature vector carried by the third sample image and the configured second loss function.
17. A feature extraction device, comprising: A receiving module, used for receiving an image to be processed; An acquisition module is used to obtain the feature vector of the image to be processed based on a feature extraction model trained by the method described in any one of claims 1-7 and 11, or a convolutional neural network trained by the method described in any one of claims 8-11.
18. An electronic device, comprising at least a processor and a memory, wherein the processor is configured to implement the steps of any one of the methods of claims 1 to 13 when executing a computer program stored in the memory.
19. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Information processing method and device, equipment, storage medium and program product
CN114154650A
Biological feature extraction method and device for multi-party secure computing system
CN114511705A
Feature extraction model training method and device, equipment, medium and program product
CN114861829A
Biological feature extraction method and device
CN115439903A
Model training method and device, computer equipment and readable storage medium
CN116452922A