An eye area feature extraction method, device, equipment and medium
Through the training periophthalmic feature extraction network, the residual module, convolutional component and transformer structure are used to extract and fuse the periophthalmic image, which solves the problem of insufficient accuracy of periophthalmic features and achieves higher identity recognition accuracy.
Patent Information
- Application Number
- CN202210542426.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-18
AI Technical Summary
How to efficiently extract stable and unique periophthalmic features for identity recognition, the accuracy of a single local or global feature in the prior art is insufficient.
A feature extraction network trained by multiple periophthalmic image samples is used to extract and fuse the periophthalmic image through residual modules, convolutional components and transformer structures, including multiple processing and fusion of local features and global features.
It improves the accuracy and reliability of periophthalmic features, avoids the problem of insufficient accuracy of a single feature, and enhances the accuracy of identity recognition.
Smart Images

Figure CN114913591B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of periorbital feature extraction, and in particular, to a method, apparatus, device, and medium for periorbital feature extraction. Background Art
[0002] In the past two decades, due to the advantages of biometrics such as universality, uniqueness, stability, and convenience, and being difficult to forge and imitate, it has made great contributions in many fields. Among numerous biometrics, periorbital features have very great advantages due to their stability, uniqueness, and non-invasiveness within a certain period of time, and have broad market prospects and scientific research value.
[0003] Therefore, how to extract periorbital features has become an important issue. Summary of the Invention
[0004] In order to extract periorbital features, this application provides a method, apparatus, device, and medium for periorbital feature extraction.
[0005] In a first aspect, this application provides a method for periorbital feature extraction, adopting the following technical solution:
[0006] Obtain a to-be-detected periorbital image;
[0007] Input the to-be-detected periorbital image into a periorbital feature extraction network to obtain periorbital features, where the periorbital feature extraction network is obtained by training based on multiple periorbital image samples.
[0008] By adopting the above technical solution, the to-be-detected periorbital image is subjected to periorbital feature extraction through a periorbital feature extraction network trained by multiple periorbital image samples, so that periorbital features can be obtained.
[0009] In a possible implementation manner, the periorbital features are used for identity recognition.
[0010] Through the above technical solution, identity recognition can be performed using the extracted periorbital features.
[0011] In a possible implementation manner, the step of inputting the to-be-detected periorbital image into a periorbital feature extraction network to obtain periorbital features includes:
[0012] Extract features from the to-be-detected periorbital image to obtain local features;
[0013] Perform feature processing based on the local features to obtain global features;
[0014] Fuse the local features and the global features to obtain periorbital features.
[0015] By adopting the above technical solution, the periorbital features are obtained by fusing local features and global features, avoiding the situation of insufficient accuracy of single local features or global features, and improving the reliability of periorbital feature extraction.
[0016] In a possible implementation manner, the extracting local features from the to-be-detected periorbital image includes:
[0017] Based on the to-be-detected periorbital image, local features are obtained through a feature extraction module, where the feature extraction module includes: multiple groups of residual modules and multiple groups of convolution components.
[0018] By adopting the above technical solution, local features are extracted from the to-be-detected periorbital image through multiple groups of residual modules and multiple groups of convolution operations, improving the accuracy of local feature extraction.
[0019] In a possible implementation manner, the obtaining local features based on the to-be-detected periorbital image through a feature extraction module includes:
[0020] Input the to-be-detected periorbital image into the first residual module to obtain the first local feature;
[0021] Input the first local feature into the second residual module to obtain the second local feature, and input the second local feature into the third residual module to obtain the third local feature;
[0022] Fuse the second local feature and the third local feature to obtain the initial first local fusion feature;
[0023] Convolve the initial first local fusion feature to obtain the first local fusion feature;
[0024] Input the first local fusion feature into the fourth residual module to obtain the fourth local feature;
[0025] Fuse the first local feature and the fourth local feature to obtain the second local fusion feature;
[0026] Pass the second local fusion feature through multiple groups of convolution components in sequence to obtain the local features.
[0027] By adopting the above technical solution, four groups of residual modules and three groups of convolutional components are used to extract local features. Compared with the case where there are three groups of residual modules and three groups of convolutional components, since the second local feature and the third local feature are fused, and the fused initial first local fused feature is convolved and then input into the fourth residual module to obtain the second local fused feature, and the local feature obtained after passing through three groups of convolutional components, the obtained local feature is more accurate than the local feature obtained by passing through three groups of residual modules and three groups of convolutional components, and the calculation speed is faster compared with five or more groups of residual modules and multiple groups of convolutional components.
[0028] In a possible implementation manner, the feature processing based on the local feature to obtain the global feature includes:
[0029] Based on the local feature, a first target feature is obtained through a first convolutional component;
[0030] Based on the first target feature, a second target feature is obtained through a second convolutional component;
[0031] Based on the first target feature, a third target feature is obtained through multiple transformer structures;
[0032] Based on the second target feature and the third target feature, feature fusion is performed to obtain the global feature.
[0033] By adopting the above technical solution, the third target feature is obtained according to the local feature through the first convolutional component, the second convolutional component, and multiple transformer structures in sequence; then the third target feature and the second target feature are fused to obtain the global feature. Since it is obtained by feature processing of the local feature, the accuracy of the global feature is improved.
[0034] In a possible implementation manner, the obtaining of the third target feature based on the first target feature through multiple transformer structures includes:
[0035] Based on the first target feature, a first feature map, a second feature map, and a third feature map are obtained;
[0036] The matrix product calculation is performed on the first feature map and the second feature map to obtain the feature map product;
[0037] The feature map product is scaled to obtain the weight of the feature map product;
[0038] The weight of the feature map product is pooled to obtain the pooled weight;
[0039] The pooled weight is normalized to obtain the normalized weight;
[0040] Performing addition calculation based on the position information of each feature in the third feature map to obtain a feature sum;
[0041] Performing matrix calculation based on the feature sum and the normalization weight to obtain a calculation result;
[0042] Pooling the first target feature to obtain a pooled first target feature, and fusing the pooled first target feature with the calculation result to obtain an initial third target feature;
[0043] Using the initial third target feature as the first target feature and passing it through the next transformer structure until the third target feature is obtained.
[0044] By adopting the above technical solution, the third target feature is obtained through the fusion of the first target feature and the calculation result, increasing the hierarchy of the third target feature.
[0045] In a possible implementation manner, the fusing the local feature and the global feature to obtain a periorbital feature includes:
[0046] Fusing the local feature and the global feature to obtain an initial fusion feature;
[0047] Sequentially inputting the initial fusion feature into multiple groups of residual modules to obtain a target fusion feature;
[0048] Sequentially inputting the target fusion feature into multiple groups of fully connected layers to obtain the periorbital feature.
[0049] By adopting the above technical solution, inputting the initial fusion feature into multiple groups of residual modules and multiple groups of fully connected layers improves the accuracy of the obtained periorbital feature.
[0050] In a possible implementation manner, the training process of the periorbital feature extraction network includes:
[0051] Obtaining a plurality of periorbital image samples, where the plurality of periorbital image samples include a plurality of periorbital image sample pictures and their respective corresponding labels;
[0052] Obtaining a periorbital feature extraction network to be trained;
[0053] Training the periorbital feature extraction network to be trained according to the plurality of periorbital image samples to obtain the periorbital feature extraction network.
[0054] By adopting the above technical solution, training the periorbital feature extraction network to be trained based on periorbital image samples, and finally obtaining a periorbital feature extraction network that meets the requirements for extracting periorbital features from periorbital images to be detected.
[0055] In a possible implementation manner, the tag includes an identity information tag. Training the to-be-trained periorbital feature extraction network based on the multiple periorbital image samples to obtain the periorbital feature extraction network includes:
[0056] Inputting the multiple periorbital image sample pictures into the to-be-trained model to obtain the predicted identity information corresponding to each sample picture. Among them, the to-be-trained model includes: a to-be-trained periorbital feature extraction network, a to-be-trained fully connected layer, and a to-be-trained classifier. Among them, the to-be-trained periorbital feature extraction network is used to extract the periorbital features of the sample pictures;
[0057] Determining a loss value according to the predicted identity information corresponding to each periorbital image sample picture and its corresponding identity information tag by using a preset loss function;
[0058] Iteratively training the to-be-trained model according to the loss value and the multiple periorbital image sample pictures until the loss value reaches a preset loss threshold to obtain a final trained model. The final trained model includes: a periorbital feature extraction network, a fully connected layer, and a classifier;
[0059] Extracting the periorbital feature extraction network from the final trained model.
[0060] By adopting the above technical solution, the accuracy of the periorbital feature extraction network is improved by training the to-be-trained periorbital feature extraction network.
[0061] In a second aspect, the present application provides a periorbital feature extraction device, adopting the following technical solution:
[0062] A periorbital feature extraction device includes
[0063] An acquisition module: used to acquire the to-be-detected periorbital image;
[0064] A obtaining module: used to input the to-be-detected periorbital image into the periorbital feature extraction network to obtain the periorbital features. Among them, the periorbital feature extraction network is obtained by training based on multiple periorbital image samples.
[0065] By adopting the above technical solution, the to-be-detected periorbital image is subjected to periorbital feature extraction through the periorbital feature extraction network trained by multiple periorbital image samples, so that periorbital features can be obtained.
[0066] In a third aspect, the present application provides an electronic device, adopting the following technical solution:
[0067] At least one processor;
[0068] A memory;
[0069] At least one application program, wherein the at least one application program is stored in a memory and configured to be executed by at least one processor, and the at least one application program is configured to: execute the method shown in any possible implementation manner of the first aspect.
[0070] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution:
[0071] A computer-readable storage medium stores at least one instruction, at least one segment of program, a code set or an instruction set, and the at least one instruction, at least one segment of program, the code set or the instruction set is loaded and executed by a processor to implement the method shown in any possible implementation manner of the first aspect.
[0072] In summary, the present application includes at least one of the following beneficial technical effects:
[0073] 1. Extract the periorbital features of the periorbital image to be detected through a periorbital feature extraction network trained by a plurality of periorbital image samples, so as to obtain the periorbital features. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a schematic flowchart of a periorbital feature extraction method provided by an embodiment of the present application;
[0075] Figure 2 is a schematic structural diagram of a periorbital feature extraction network provided by an embodiment of the present application;
[0076] Figure 3 is a schematic structural diagram of a residual module provided by an embodiment of the present application;
[0077] Figure 4 is a schematic structural diagram of a feature extraction module with three groups of residual modules and three groups of convolutional components provided by an embodiment of the present application;
[0078] Figure 5 is a schematic structural diagram of a feature extraction module with four groups of residual modules and three groups of convolutional components provided by an embodiment of the present application;
[0079] Figure 6 is a schematic structural diagram of obtaining a global feature based on a local feature provided by an embodiment of the present application;
[0080] Figure 7 is a schematic diagram of a transformer structure provided by an embodiment of the present application;
[0081] Figure 8 is a schematic structural diagram of a feature fusion module provided by an embodiment of the present application;
[0082] Figure 9 It is a schematic structural diagram of an eye feature extraction device provided by an embodiment of the present application;
[0083] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0084] The following will further describe the present application in detail with reference to the Figure 1 - appended Figure 10 drawings.
[0085] Those skilled in the art can make modifications to this embodiment without creative contributions according to needs after reading this specification, but as long as they are within the scope of the embodiments of the present application, they are protected by the patent law.
[0086] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts fall within the scope of protection of the embodiments of the present application.
[0087] In addition, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.
[0088] Specifically, an embodiment of the present application provides an eye feature extraction method, which is executed by an electronic device. The electronic device can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not make limitations here.
[0089] Combined with Figure 1 , Figure 1 It is a schematic flowchart of an eye feature extraction method provided by an embodiment of the present application. The method includes step S100 and step S101, where:
[0090] Step S100: Obtain an eye image to be detected.
[0091] Among them, the periorbital area is the area around the eyes. The periorbital area ranges from the inner canthus to the temple on the outside, from the upper orbital margin to the lower orbital margin on the upper and lower sides, and the middle part of this area.
[0092] The embodiments of the present application do not limit the application scenarios, which can be the entrance of a laboratory, a bank window, the entrance of a community, a factory, etc., where periorbital feature extraction can be performed. For example, when the application scenario is the entrance of a laboratory, a camera device is placed in advance at the entrance of the laboratory, and the camera device can at least capture the periorbital image of a person. When a periorbital feature extraction request is detected, the periorbital image to be detected is obtained.
[0093] Furthermore, a monitoring program is pre-integrated in the electronic device. The monitoring program is used to monitor the triggering behavior of the periorbital feature extraction request. Once it monitors that the periorbital feature extraction request is triggered, it obtains the facial image of the person and preprocesses the facial image of the person to obtain the periorbital image to be detected. Among them, the preprocessing process includes: screening the periorbital area from the facial image of the person.
[0094] Step S101: Input the periorbital image to be detected into the periorbital feature extraction network to obtain periorbital features. Among them, the periorbital feature extraction network is obtained after being trained based on multiple periorbital image samples.
[0095] Among them, the embodiments of the present application do not limit the application of periorbital features. Users can apply periorbital features according to actual needs. Among them, periorbital features can be used as the basis for personnel identity recognition. For example, personnel information is determined through features such as the height from the upper eyelid to the lower eyelid and the width of the canthus in the periorbital features.
[0096] Among them, the periorbital feature extraction network is a model for extracting periorbital features. By training the periorbital feature extraction network, a pre-trained periorbital feature extraction network is obtained. After inputting the periorbital image to be detected into the pre-trained periorbital feature extraction network, periorbital features can be directly obtained.
[0097] The embodiments of the present application do not limit the network structure of the periorbital feature extraction network, as long as it can achieve the purpose of the embodiments of the present application.
[0098] Specifically, the embodiments of the present application provide a method for extracting periorbital features. The periorbital image to be detected is subjected to periorbital feature extraction through a periorbital feature extraction network trained by multiple periorbital image samples, so that periorbital features can be obtained.
[0099] Furthermore, the periorbital features are used for identity recognition.
[0100] Among them, the embodiments of the present application do not limit the application scenarios, which can specifically be scenarios such as the entrance of a community, the entrance of a laboratory, the entrance of a campus, and the entrance and exit of a highway inspection, and can use periorbital features for identity recognition.
[0101] The embodiments of the present application do not limit the method of using periorbital features for identity recognition. Preferably, the method of using periorbital features for identity recognition can be: determining the identity information corresponding to the periorbital features from the corresponding relationship according to the periorbital features, where the corresponding relationship is the corresponding relationship between each periorbital feature and the identity information corresponding to each periorbital feature. Specifically, the corresponding relationship between each periorbital feature and the identity information is pre-stored in the electronic device. After the periorbital features are extracted, based on the periorbital features, through the corresponding relationship between each periorbital feature and the identity information corresponding to the periorbital feature, the identity information corresponding to the periorbital features is determined. The periorbital features in the corresponding relationship can be obtained by any one of computer vision, deep learning algorithms, and manual annotation.
[0102] Further, in order to improve the accuracy of determining the identity information corresponding to the periorbital features from the corresponding relationship, it further includes: regularly updating the periorbital features in the corresponding relationship. By regularly updating the periorbital features in the corresponding relationship, the accuracy of identity recognition is improved.
[0103] Further, in the embodiments of the present application, in order to improve the accuracy of the periorbital features, inputting the to-be-detected periorbital image into the periorbital feature extraction network to obtain the periorbital features includes: step S20 (not shown in the drawings), step S21 (not shown in the drawings), step S22 (not shown in the drawings), where:
[0104] Step S20, performing feature extraction on the to-be-detected periorbital image to obtain local features.
[0105] Among them, the local features are the feature of a certain area of the to-be-detected periorbital image, including the features of any area such as from the upper eyelid to the lower eyelid, the corners of the eyes, etc.
[0106] In this step, the to-be-detected periorbital image can be divided into regions to obtain multiple regions, and feature extraction is performed on each region to obtain the local features corresponding to each region. Each obtained local feature has strong semantic information.
[0107] Step S21, performing feature processing based on the local features to obtain global features.
[0108] In this step, the local features can be processed to obtain global features. The global features have high resolution and more detailed information.
[0109] Specifically, an achievable method for extracting global features is as follows: Global features are obtained by fusing local features with multiple groups of convolutional components and multiple Transformer structures.
[0110] Step S22: Fuse the local features with the global features to obtain periorbital features.
[0111] The periorbital features obtained by fusing the local features with the global features have higher accuracy compared to single local features or global features.
[0112] Specifically, the embodiments of the present application provide a method for obtaining periorbital features. By fusing local features and global features to obtain periorbital features, the situation of insufficient accuracy of single local features or global features is avoided, and the accuracy of periorbital features is improved.
[0113] Further, please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a feature extraction network provided by the embodiments of the present application. Among them, the periorbital feature extraction network includes: a feature extraction module, a Transformer, and a feature fusion module. The feature extraction module is used to extract features from the periorbital image to be detected to obtain local features; the Transformer module is used to process the local features to obtain global features; the feature fusion module is used to fuse the local features with the global features to obtain periorbital features.
[0114] Further, in the embodiments of the present application, in order to improve the accuracy of local feature extraction, step S20 of extracting features from the periorbital image to be detected to obtain local features includes:
[0115] Based on the periorbital image to be detected, local features are obtained through the feature extraction module, where the feature extraction module includes: multiple groups of residual modules and multiple groups of convolutional components.
[0116] The convolutional component includes a convolutional layer, an instance normalization layer, and an activation layer. The role of the convolutional layer is to obtain the weight and digital matrix of the input features. Instance normalization is a way to normalize the digital matrix of the features after convolution. In order to make features of different scales have similar value ranges, the weights of the features after convolution are determined through the normalized data, which can improve the accuracy and calculation speed of the residual module. The activation function corresponding to the activation layer adds non-linear factors to better solve the problem of feature classification.
[0117] Among them, the feature extraction module includes multiple groups of residual modules connected in series in sequence, and multiple groups of convolutional components connected in sequence to the last group of residual modules. Herein, the number of residual modules, the structure of the residual modules, the number of convolutional components, and the structure of the convolutional components are not limited in the embodiments of the present application. Users can set them according to actual needs, as long as the purpose of the embodiments of the present application can be achieved.
[0118] Specifically, the residual module is further elaborated. As Figure 3 shown, Figure 3 is a schematic structural diagram of a residual module provided by the embodiments of the present application. Each group of residual modules includes a left branch and a right branch. The left branch includes a convolutional layer, an instance normalization layer, and an activation layer. The right branch includes a convolution operation. Specifically, the eye periocular image to be detected is input into the left branch to obtain a first left branch feature, the eye periocular image to be detected is input into the right branch to obtain a first right branch feature, and the first left branch feature and the first right branch feature are summed to obtain the feature of the first residual module; then the feature of the first residual module is input into the left branch to obtain a second left branch feature, the feature of the first residual module is input into the right branch to obtain a second right branch feature, and the second left branch feature and the second right branch feature are summed to obtain the feature of the second residual module. The convolution operation is to perform convolution on the eye periocular image to be detected to obtain the convolved feature.
[0119] Among them, when the eye periocular image to be detected is input into the first residual module, the weight of the first left branch feature is obtained through the left branch of the first residual module, the first right branch feature is obtained through the right branch of the first residual module, and the weight of the first left branch feature is summed with the first right branch feature to obtain a first local feature, so that the obtained first local feature is more accurate.
[0120] Since in the eye periocular feature extraction network, the convolution function used in the convolutional layer is generally a linear function, no matter how the network structure is built, the output is a linear combination of the input. Therefore, the activation function corresponding to the activation layer can be introduced to add non-linear factors to better solve the problem of feature classification. Herein, the embodiments of the present application do not limit the activation function corresponding to the activation layer, and it can be any one of the sigmod (S-shaped growth curve) function, the tanh (hyperbolic tangent) function, the ReLU (Rectified Linear Units, rectified linear) function, the ELU (Exponential Linear Unit, exponential linear) function, and the PreLU (Parameteric Rectified Linear Unit, parametric rectified linear) function.
[0121] Specifically, based on the eye periocular image to be detected, local features are obtained through the feature extraction module, which specifically may include:
[0122] When the number of residual modules is 3, the eye periocular image to be detected is input into the first residual module to obtain the first local feature; the first local feature is input into the second residual module to obtain the second local feature; the second local feature is input into the third residual module to obtain the third local feature; the third local feature is fused with the first local feature to obtain the local fusion feature; the local fusion feature is sequentially passed through multiple groups of convolutional components to obtain the local feature; among them, through the feature extraction of multiple groups of residual modules for the eye periocular image to be detected, the obtained local feature has more semantic information and better hierarchy, but after the feature extraction of multiple groups of residual modules, the spatial and detail information is less. By fusing the first local feature with the third local feature, the obtained local fusion feature has the advantages of the first local feature and the third local feature, thereby making the extracted eye periocular feature more accurate.
[0123] When the number of residual modules is greater than 3, the eye periocular image to be detected is input into the first residual module to obtain the first local feature; the first local feature is input into the second residual module to obtain the second local feature; the second local feature is input into the third residual module to obtain the third local feature; the second local feature is fused with the third local feature to obtain the first local fusion feature; after the first local fusion feature is convolved, it is input into the fourth residual module to obtain the fourth local feature;
[0124] If there is no next residual module, the fourth local feature is fused with the first local feature and sequentially input into multiple convolutional components to obtain the local feature;
[0125] If there is a next residual module, the fourth local feature is fused with the first local fusion feature to obtain the second local fusion feature; after the second local fusion feature is convolved, it is input into the next residual module until there is no next residual module to obtain the local feature.
[0126] For example, when the number of residual modules is 3 and the number of convolutional components is 3, as Figure 4 shown, Figure 4 is a schematic structural diagram of a feature extraction module with three groups of residual modules and three groups of convolutional components provided by an embodiment of the present application. The specific process of obtaining the local feature includes: inputting the eye periocular image to be detected into the first residual module to obtain the first local feature; inputting the first local feature into the second residual module to obtain the second local feature; inputting the second local feature into the third residual module to obtain the third local feature; fusing the first local feature with the third local feature to obtain the first fusion feature; sequentially inputting the first fused local feature into three groups of convolutional components to obtain the local feature.
[0127] Again, when there are 4 residual modules, in an embodiment of the present application, based on the eye periocular image to be detected through the feature extraction module, the obtained local feature includes:
[0128] Input the eye contour image to be detected into the first residual module to obtain the first local feature;
[0129] Input the first local feature into the second residual module to obtain the second local feature, and input the second local feature into the third residual module to obtain the third local feature;
[0130] Fuse the second local feature and the third local feature to obtain the initial first local fusion feature;
[0131] Convolve the initial first local fusion feature to obtain the first local fusion feature;
[0132] Input the first local fusion feature into the fourth residual module to obtain the fourth local feature;
[0133] Fuse the first local feature and the fourth local feature to obtain the second local fusion feature;
[0134] Pass the second local fusion feature through multiple groups of convolution components in sequence to obtain the local feature.
[0135] Please refer to Figure 5 , Figure 5 which is a schematic structure of a feature extraction module with four groups of residual modules and three groups of convolution components provided by an embodiment of the present application.
[0136] Taking four groups of residual modules and three groups of convolution components as an example, an embodiment of the present application explains obtaining the local feature from the eye contour image to be detected through the feature extraction module. Compared with the case where there are three groups of residual modules, in the embodiment of the present application, the second local feature and the third local feature are fused, the fused initial first local fusion feature is convolved, and then input into the fourth residual module to obtain the second local fusion feature, and finally the local feature obtained after passing through three groups of convolution components. The obtained local feature is more accurate than the local feature obtained through three groups of residual modules. Each time passing through a group of residual modules, the obtained local feature has a higher level and more semantic information.
[0137] Of course, the number of groups of residual modules can also be 5 groups or 6 groups, which will not be elaborated in the embodiment of the present application.
[0138] Specifically, in the embodiment of the present application, local features are extracted from the eye contour image to be detected through multiple groups of residual modules and multiple groups of convolution operations, improving the accuracy of local feature extraction.
[0139] Furthermore, in combination with Figure 6 , Figure 6It is a schematic structural diagram of a global feature obtained based on local features provided by an embodiment of the present application. This structure at least includes a first convolutional component, a second convolutional component connected to the output of the first convolutional component, multiple Transformer structures, and a fusion module connected to the second convolutional component and multiple Transformer structures. Correspondingly, in the embodiment of the present application, in order to improve the accuracy of the global feature, feature processing is performed based on local features to obtain the global feature, including step S31 (not shown in the drawings), step S32 (not shown in the drawings), step S33 (not shown in the drawings), step S34 (not shown in the drawings), where:
[0140] Step S31: Based on the local feature, through the first convolutional component, obtain the first target feature.
[0141] Among them, the convolutional component includes a convolutional layer, an instance normalization layer, and an activation layer. Input the local feature into the convolutional component, and through convolution, instance normalization, and activation operations, obtain the first target feature.
[0142] Step S32: Based on the first target feature, through the second convolutional component, obtain the second target feature.
[0143] Among them, pass the first target feature through the second convolutional component for convolution, instance normalization, and activation operations to obtain the second target feature.
[0144] Step S33: Based on the first target feature, through multiple Transformer structures, obtain the third target feature.
[0145] Among them, input the first target feature into multiple Transformer structures to obtain the third target feature. Among them, the number of Transformer structures can be user-defined, and can be 2, or 3, or 4.
[0146] Furthermore, before inputting the first target feature into multiple Transformer structures, it also includes: performing the first reshape on the first target feature, and sequentially inputting the first target feature after the first reshape into multiple Transformer structures to obtain the initial third target feature; performing the second reshape on the third target feature to obtain the third target feature.
[0147] Among them, the first reshape is to transform the first target feature into a feature that conforms to the dimension of the Transformer structure. After passing through multiple Transformer structures, perform the second reshape on the feature after passing through multiple Transformer structures, and transform the feature after passing through multiple Transformer structures into a feature with the original dimension as the third target feature.
[0148] Step S34: Based on the second target feature and the third target feature, perform feature fusion to obtain the global feature.
[0149] Among them, fuse the second target feature and the third target feature to obtain the global feature.
[0150] Furthermore, performing feature fusion based on the second target feature and the third target feature to obtain the global feature may include: fusing the second target feature and the third target feature to obtain the initial global feature; after processing the initial global feature through the convolutional component, obtaining the final global feature.
[0151] Specifically, the embodiments of the present application do not limit the way of fusing the second target feature and the third target feature, which can be any one of the concat method and the add method.
[0152] Specifically, the embodiments of the present application propose a method for obtaining the global feature based on the local feature. The global feature is obtained through the local feature, the convolutional component, and the transformer structure, improving the accuracy of the global feature.
[0153] Furthermore, combining Figure 7 , Figure 7 is a schematic diagram of a transformer structure provided by the embodiments of the present application. In the embodiments of the present application, obtaining the third target feature based on the first target feature through multiple transformer structures includes steps S41 (not shown in the drawings), step S42 (not shown in the drawings), step S43 (not shown in the drawings), step S44 (not shown in the drawings), step S45 (not shown in the drawings), step S46 (not shown in the drawings), step S47 (not shown in the drawings), step S48 (not shown in the drawings), step S49 (not shown in the drawings), where:
[0154] Step S41: Based on the first target feature, obtain the first feature map, the second feature map, and the third feature map.
[0155] Among them, input the first target feature into the fully connected layer to obtain the feature map, and copy the feature map as the first feature map, the second feature map, and the third feature map. Among them, in Figure 7 , the first feature map is K, the second feature map is Q, and the third feature map is V. Among them, the fully connected layer weights the first target feature.
[0156] The Transformer structure adopts the self-attention mechanism. The first feature map, the second feature map, and the third feature map can be the same, so the feature map is copied three times.
[0157] Step S42: Perform matrix multiplication on the first feature map and the second feature map to obtain the product of the feature maps.
[0158] Step S43: Normalize the product of the feature maps to obtain the weights of the product of the feature maps.
[0159] Among them, data normalization scales the values of the product of the feature maps to between 0 and 1, which is more convenient for algorithm calculation and learning.
[0160] Step S44: Pool the weights of the product of the feature maps to obtain the pooled weights.
[0161] Among them, the role of pooling is to downsample the weights of the product of the feature maps, that is, to remove redundant information from the weights of the product of the feature maps and compress the features. Using the pooled weights for calculation simplifies the network complexity and reduces the computational amount.
[0162] Step S45: Normalize the pooled weights to obtain the normalized weights.
[0163] Among them, the normalization process is softmax normalization. Softmax normalization uses the softmax function for normalization. Specifically, the softmax function can convert all products of feature maps into probabilities (between 0 and 1), and the sum of all probability values is equal to 1. By using the softmax function, the accuracy of global features is improved.
[0164] Step S46: Perform addition calculation based on the position information of each feature in the third feature map to obtain the sum of the features.
[0165] Step S47: Perform matrix multiplication based on the sum of the features and the normalized weights to obtain the calculation result.
[0166] Step S48: Pool the first target feature to obtain the pooled first target feature, and fuse it with the calculation result to obtain the initial third target feature.
[0167] Since the first target feature is obtained by inputting local features into the convolutional component, it will cause feature redundancy. Therefore, the first target feature can be pooled to remove feature redundancy.
[0168] Fuse the pooled first target feature with the calculation result to obtain the third target feature, which increases the hierarchy of the third target feature.
[0169] Step S49: Use the initial third target feature as the first target feature and pass it through the next transformer structure until the third target feature is obtained.
[0170] Among them, steps S41 to S48 are the processing process of the first target feature of a transformer structure. After inputting the first target feature into the transformer structure to obtain the initial third target feature, the initial third target feature is used as the new first target feature and input into the transformer structure. Specifically, multiple transformer structures are connected in series, the output feature of the previous transformer structure is used as the input feature of the next transformer structure, and the output feature of the last transformer structure is determined as the third target feature.
[0171] The embodiments of the present application do not limit the number of transformer structures, and users can customize the settings according to actual needs.
[0172] Specifically, the third target feature is obtained by fusing the first target feature and the output result of the first target feature input into the transformer structure, which increases the hierarchy of the third target feature.
[0173] Furthermore, in a feasible way of fusing local features and global features to obtain periorbital features, the local features and global features can be directly fused to obtain the fused features, and the fused features are used as the periorbital features.
[0174] Furthermore, the feature fusion module at least includes multiple residual modules and multiple fully connected layers. Figure 8 , Figure 8 is a schematic diagram of a feature fusion module provided by the embodiments of the present application. Figure 8 The feature fusion module is composed of three groups of residual modules and two groups of fully connected layers. Of course, it can also be four groups of residual modules and three groups of fully connected layers, or five groups of residual modules and two groups of fully connected layers. The embodiments of the present application do not limit the number of residual modules and fully connected layers, and users can choose according to actual needs.
[0175] In the embodiments of the present application, in order to improve the accuracy of periorbital features, local features and global features are fused to obtain periorbital features, including step S51 (not shown in the drawings), step S52 (not shown in the drawings), step S53 (not shown in the drawings), where:
[0176] Step S51: Fuse the local features and global features to obtain the initial fused features.
[0177] Feature fusion is an important means to improve the accuracy of periorbital features. Global features have higher resolution and contain more positional and detailed information. However, due to fewer convolutions, their hierarchy is lower and there is more noise. Local features have stronger semantic information and higher hierarchy, but their resolution is very low and their ability to perceive details is poor.
[0178] Specifically, the way to fuse local features with global features can be any one of the concat method and the add method. The embodiments of the present application do not limit the fusion method, and users can choose according to actual needs.
[0179] Step S52: Sequentially input the initial fusion features into multiple groups of residual modules to obtain the target fusion features.
[0180] The embodiments of the present application do not limit the number of residual modules, and users can customize the setting according to actual needs. Preferably, the number of residual modules is three groups. Sequentially input the initial fusion features into the three groups of residual modules to obtain the fusion features.
[0181] Step S53: Input the target fusion features into multiple groups of fully connected layers to obtain periorbital features.
[0182] Among them, the fully connected layer outputs the matrix of the fusion features as a value and maps it back to the periorbital image to be detected, integrates the target fusion features together, and outputs a value, which is the periorbital feature.
[0183] Specifically, the embodiments of the present application provide a method for fusing local features and global features. Compared with the traditional method of directly fusing local features and global features, the present application inputs the initial fusion features into multiple groups of residual modules and multiple groups of fully connected layers, improving the accuracy of periorbital features.
[0184] Further, in the embodiments of the present application, a model training method is provided, including step S61 (not shown in the drawings), step S62 (not shown in the drawings), and step S63 (not shown in the drawings), where:
[0185] Step S61: Obtain multiple periorbital image samples, where the multiple periorbital image samples include multiple periorbital image sample pictures and their respective corresponding labels. Specifically, capture multiple periorbital image sample pictures corresponding to multiple identity information through a photographing device, where each identity information corresponds to at least multiple periorbital image sample pictures. Preferably, each identity information corresponds to at least 10 periorbital image sample pictures. A user can utilize computer vision technology or the annotation tool of an annotation platform to obtain the periorbital features in each periorbital image sample picture. The embodiments of the present application do not limit the manner of obtaining the periorbital features in each periorbital image sample picture. Among them, the periorbital features include any one or more of features such as the height from the upper eyelid to the lower eyelid, the width of the eye corners, the number of crow's feet, and eyelashes.
[0186] Further, before obtaining multiple periorbital image samples, it may further include: obtaining an initial periorbital image sample; performing translation, rotation, and random erasing on the initial periorbital image sample to obtain an extended initial periorbital image sample; using the initial periorbital image sample and the extended initial periorbital image sample as periorbital image samples. By obtaining more periorbital image samples, the accuracy of the periorbital feature extraction network obtained by subsequent training based on the periorbital image samples can be higher.
[0187] Step S62: Obtain a periorbital feature extraction network to be trained.
[0188] Among them, the periorbital feature extraction network to be trained is used to extract periorbital features.
[0189] Step S63: Train the periorbital feature extraction network to be trained according to multiple periorbital image samples to obtain a periorbital feature extraction network.
[0190] Specifically, by training the periorbital feature extraction network to be trained based on the periorbital image samples, a periorbital feature extraction network that meets the requirements is finally obtained, which is used to extract periorbital features from the periorbital images to be detected.
[0191] For example, when the label is the periorbital feature label corresponding to the periorbital image sample picture, input the periorbital image sample picture into the periorbital feature extraction network to be trained to obtain the periorbital features corresponding to the predicted periorbital image sample picture. Based on the periorbital features corresponding to the predicted periorbital image sample picture and the periorbital feature label corresponding to the periorbital features corresponding to the periorbital image sample picture, use a loss function to obtain a loss value. Based on the loss value and the loss value threshold, perform iterative training on the periorbital feature extraction network to be trained until the loss value reaches the loss value threshold. At this time, the corresponding periorbital feature extraction network that has completed training is the periorbital feature extraction network.
[0192] When the label is an identity information label, multiple periorbital image sample pictures are input into the model to be trained, and the identity information corresponding to each predicted sample picture is obtained; the model to be trained is iteratively trained according to the identity information corresponding to each predicted periorbital image sample picture and its corresponding identity information label to obtain a final trained model, where the final trained model includes: a periorbital feature extraction network, a fully connected layer, and a classifier; the periorbital feature extraction network is extracted from the final trained model.
[0193] Further, in the embodiment of the present application, the label includes an identity information label, and the periorbital feature extraction network to be trained is trained according to multiple periorbital image samples, including steps S71 (not shown in the drawings), step S72 (not shown in the drawings), step S73 (not shown in the drawings), step S74 (not shown in the drawings), where:
[0194] Step S71, input multiple periorbital image sample pictures into the model to be trained, and obtain the identity information corresponding to each predicted sample picture; the model to be trained includes: a periorbital feature extraction network to be trained, a fully connected layer to be trained, and a classifier to be trained, where the periorbital feature extraction network to be trained is used to extract the periorbital features of the sample pictures.
[0195] Further, it may also include: constructing a model to be trained according to the periorbital feature extraction network to be trained, the fully connected layer to be trained, and the classifier to be trained, where the fully connected layer to be trained and the classifier to be trained are sequentially connected to the periorbital feature extraction network to be trained.
[0196] Wherein, in the embodiment of the present application, the identity information includes but is not limited to any one of the job number, name, and ID (Identity document).
[0197] The model to be trained includes: a periorbital feature extraction network to be trained, a fully connected layer to be trained, and a classifier to be trained. Among them, the periorbital features are extracted by using the periorbital feature extraction network to be trained, and then the extracted periorbital features are input into the fully connected layer to be trained for weighted processing of the periorbital features, and the feature values obtained after weighting are classified by the classifier to be trained to obtain the identity information corresponding to the predicted periorbital image sample picture. The classifier to be trained can obtain the probability values corresponding to different identity information according to the feature values obtained after weighting the extracted periorbital features through the fully connected layer, and determine the identity information corresponding to the maximum probability value as the identity information corresponding to the predicted periorbital image sample picture.
[0198] Step S72, determine the loss value according to the identity information corresponding to each predicted periorbital image sample picture and its corresponding identity information label by using a preset loss function.
[0199] In the embodiments of the present application, multiple periorbital image sample pictures are input into the model to be trained, and the identity information corresponding to each periorbital image sample picture is obtained. According to the identity information labels corresponding to each periorbital image sample picture and the identity information corresponding to the predicted periorbital image sample pictures, a loss value is determined using a preset loss function. The smaller the loss value, the smaller the gap between the identity information labels corresponding to each periorbital image sample picture and the identity information corresponding to the predicted periorbital image sample pictures, and the more accurate the detection result. The embodiments of the present application do not limit the loss function, and users can select according to actual needs, which can be any one of the focal loss function, cross-entropy loss function, absolute value loss function, square loss function, and triplet loss function. Preferably, the loss function is the triplet loss function. The triplet loss function can better measure the difference value of details compared with other loss functions and is more suitable for the periorbital feature extraction network.
[0200] Step S73: Iteratively train the model to be trained according to the loss value and multiple periorbital image sample pictures until the loss value reaches a preset loss threshold, and obtain a final trained model, where the final trained model includes: a periorbital feature extraction network, a fully connected layer, and a classifier.
[0201] Specifically, the user can preset a preset loss threshold, and the specific numerical value can be determined according to empirical values, user settings, or machine self-selection. After the model to be trained is trained for a preset number of times, the loss value is calculated using the preset loss function according to the identity information corresponding to the predicted periorbital image sample pictures and the identity information labels corresponding to each periorbital image sample picture. When the loss value reaches the preset loss threshold, it is determined that the training of the model to be trained is completed.
[0202] Step S74: Extract the periorbital feature extraction network from the final trained model.
[0203] When the loss value reaches the preset loss threshold, it is determined that the model to be trained has been trained well and meets the actual application conditions. Since the model to be trained is obtained by connecting a fully connected layer and a classifier to the periorbital feature extraction network to be trained, after the training is completed, the pre-trained periorbital feature extraction network is obtained by removing the fully connected layer and the classifier.
[0204] Specifically, the embodiments of the present application provide a training process for the periorbital feature extraction network, which improves the accuracy of the periorbital feature extraction network by training the periorbital feature extraction network to be trained.
[0205] The above embodiments introduce a periorbital feature extraction method from the perspective of the method flow. The following embodiments introduce a periorbital feature extraction device from the perspective of virtual modules or virtual units. For details, please refer to the following embodiments.
[0206] Please refer toFigure 9 , Figure 9 is a schematic structural diagram of an eye - surrounding feature extraction device provided by an embodiment of the present application, including:
[0207] An acquisition module 210: configured to acquire an eye - surrounding image to be detected;
[0208] A obtaining module 220: configured to input the eye - surrounding image to be detected into an eye - surrounding feature extraction network to obtain eye - surrounding features, where the eye - surrounding feature extraction network is obtained by training based on multiple eye - surrounding image samples.
[0209] In a possible implementation manner of the embodiment of the present application, the eye - surrounding features are used for identity recognition.
[0210] In a possible implementation manner of the embodiment of the present application, when the obtaining module 220 executes the operation of inputting the eye - surrounding image to be detected into the eye - surrounding feature extraction network to obtain eye - surrounding features, it is specifically configured to:
[0211] Extract features from the eye - surrounding image to be detected to obtain local features;
[0212] Perform feature processing based on the local features to obtain global features;
[0213] Fuse the local features and the global features to obtain eye - surrounding features.
[0214] In a possible implementation manner of the embodiment of the present application, the eye - surrounding features are used for identity recognition.
[0215] In a possible implementation manner of the embodiment of the present application, when the obtaining module 220 executes the operation of extracting features from the eye - surrounding image to be detected to obtain local features, it is specifically configured to:
[0216] Based on the eye - surrounding image to be detected, obtain local features through a feature extraction module, where the feature extraction module includes: multiple groups of residual modules and multiple groups of convolutional components.
[0217] In a possible implementation manner of the embodiment of the present application, when the obtaining module 220 executes the operation of obtaining local features based on the eye - surrounding image to be detected through the feature extraction module, it is specifically configured to:
[0218] Input the eye - surrounding image to be detected into the first residual module to obtain a first local feature;
[0219] Input the first local feature into the second residual module to obtain a second local feature, and input the second local feature into the third residual module to obtain a third local feature;
[0220] Fuse the second local feature and the third local feature to obtain an initial first local fusion feature;
[0221] Convolve the initial first local fusion feature to obtain the first local fusion feature;
[0222] Input the first local fusion feature into the fourth residual module to obtain the fourth local feature;
[0223] Fuse the first local feature and the fourth local feature to obtain the second local fusion feature;
[0224] Pass the second local fusion feature through multiple groups of convolutional components in sequence to obtain the local feature.
[0225] A possible implementation manner of the embodiment of the present application. When the obtaining module 220 performs feature processing based on the local feature to obtain the global feature, it is specifically used for:
[0226] Based on the local feature, pass through the first convolutional component to obtain the first target feature;
[0227] Based on the first target feature, pass through the second convolutional component to obtain the second target feature;
[0228] Based on the first target feature, pass through multiple transformer structures to obtain the third target feature;
[0229] Based on the second target feature and the third target feature, perform feature fusion to obtain the global feature.
[0230] A possible implementation manner of the embodiment of the present application. When the obtaining module 220 performs passing through multiple transformer structures based on the first target feature to obtain the third target feature, it is specifically used for:
[0231] Based on the first target feature, obtain the first feature map, the second feature map, and the third feature map;
[0232] Perform matrix multiplication calculation on the first feature map and the second feature map to obtain the feature map product;
[0233] Scale the feature map product to obtain the weight of the feature map product;
[0234] Pool the weight of the feature map product to obtain the pooled weight;
[0235] Normalize the pooled weight to obtain the normalized weight;
[0236] Perform addition calculation based on the position information of each feature in the third feature map to obtain the feature sum;
[0237] Perform matrix calculation based on the feature sum and the normalized weight to obtain the calculation result;
[0238] Pool the first target feature to obtain the pooled first target feature, and fuse the pooled first target feature with the calculation result to obtain the initial third target feature;
[0239] Use the initial third target feature as the first target feature and pass it through the next transformer structure until the third target feature is obtained.
[0240] A possible implementation of the embodiment of the present application. When the obtaining module 220 performs feature fusion of the local feature and the global feature to obtain the periorbital feature, it is specifically used for:
[0241] Fuse the local feature and the global feature to obtain the initial fusion feature;
[0242] Sequentially input the initial fusion feature into multiple groups of residual modules to obtain the target fusion feature;
[0243] Sequentially input the target fusion feature into multiple groups of fully connected layers to obtain the periorbital feature.
[0244] A possible way of the embodiment of the present application further includes:
[0245] A periorbital feature extraction network training module, which is used for: obtaining multiple periorbital image samples, where the multiple periorbital image samples include multiple periorbital image sample pictures and their respective corresponding labels;
[0246] Obtain the periorbital feature extraction network to be trained;
[0247] Train the periorbital feature extraction network to be trained according to the multiple periorbital image samples to obtain the periorbital feature extraction network.
[0248] A possible way of the embodiment of the present application. The label includes an identity information label. When the periorbital feature extraction network training module performs training on the periorbital feature extraction network to be trained according to the multiple periorbital image samples to obtain the periorbital feature extraction network, it is specifically used for:
[0249] Input the multiple periorbital image sample pictures into the model to be trained to obtain the predicted identity information corresponding to each sample picture. The model to be trained includes: the periorbital feature extraction network to be trained, the fully connected layer to be trained, and the classifier to be trained, where the periorbital feature extraction network to be trained is used to extract the periorbital feature of the sample picture;
[0250] Determine the loss value according to the predicted identity information corresponding to each periorbital image sample picture and its respective corresponding identity information label by using a preset loss function;
[0251] Iteratively train the model to be trained according to the loss value and multiple periorbital image samples until the loss value reaches a preset loss threshold to obtain a final trained model, where the final trained model includes: a periorbital feature extraction network, a fully connected layer, and a classifier;
[0252] Extract the periorbital feature extraction network from the final trained model.
[0253] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of a periorbital feature extraction device 200 described above can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here.
[0254] An electronic device is provided in an embodiment of the present application, such as Figure 10 shown. Figure 10 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Figure 10 The electronic device 300 shown includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as connected through a bus 302. Optionally, the electronic device 300 may further include a transceiver 304. It should be noted that in actual applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation to the embodiments of the present application.
[0255] The processor 301 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 301 may also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0256] The bus 302 may include a path for transmitting information between the above components. The bus 302 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation,Figure 10 It is represented only by a thick line, but it does not mean that there is only one bus or one type of bus.
[0257] The memory 303 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0258] The memory 303 is used to store the application program code for implementing the solution of this application, and is controlled by the processor 301 to execute. The processor 301 is used to execute the application program code stored in the memory 303 to implement the content shown in the foregoing method embodiments.
[0259] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 10 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0260] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments. That is, the periorbital image to be detected is subjected to periorbital feature extraction through a periorbital feature extraction network trained by multiple periorbital image samples, so as to obtain periorbital features.
[0261] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0262] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An eye perimeter feature extraction method, characterized in that, Including: Obtain the periorbital image to be detected; Input the periorbital image to be detected into the first residual module to obtain the first local feature; Input the first local feature into the second residual module to obtain the second local feature, and input the second local feature into the third residual module to obtain the third local feature; Fuse the second local feature and the third local feature to obtain the initial first local fusion feature; Convolve the initial first local fusion feature to obtain the first local fusion feature; Input the first local fusion feature into the fourth residual module to obtain the fourth local feature; Fuse the first local feature and the fourth local feature to obtain the second local fusion feature; Pass the second local fusion feature through multiple groups of convolution components in sequence to obtain the local feature; Perform feature processing based on the local feature to obtain the global feature; Fuse the local feature and the global feature to obtain the periorbital feature.
2. The eye region feature extraction method according to claim 1, wherein The periorbital feature is used for identity recognition.
3. The eye feature extraction method according to claim 1, wherein The performing feature processing based on the local feature to obtain the global feature includes: Based on the local feature, obtain the first target feature through the first convolution component; Based on the first target feature, obtain the second target feature through the second convolution component; Based on the first target feature, obtain the third target feature through multiple transformer structures; Fuse the second target feature and the third target feature to obtain the global feature.
4. The eye region feature extraction method according to claim 3, wherein The obtaining the third target feature based on the first target feature through multiple transformer structures includes: Based on the first target feature, obtain the first feature map, the second feature map, and the third feature map; Perform matrix multiplication calculation on the first feature map and the second feature map to obtain the feature map product; Scale the feature map product to obtain the weight of the feature map product; Pool the weight of the feature map product to obtain the pooled weight; Normalize the pooled weight to obtain the normalized weight; Perform addition calculation based on the position information of each feature in the third feature map to obtain the feature sum; Perform matrix calculation based on the feature sum and the normalized weight to obtain the calculation result; Pool the first target feature to obtain the pooled first target feature, and fuse the pooled first target feature and the calculation result to obtain the initial third target feature; Take the initial third target feature as the first target feature and pass it through the next transformer structure until the third target feature is obtained.
5. The eye region feature extraction method according to claim 1, wherein The fusing the local feature and the global feature to obtain the periorbital feature includes: Fuse the local feature and the global feature to obtain the initial fusion feature; Input the initial fusion feature into multiple groups of residual modules in sequence to obtain the target fusion feature; Input the target fusion feature into multiple groups of fully connected layers in sequence to obtain the periorbital feature.
6. The method for extracting periorbital features according to claim 1, wherein The training process of the periorbital feature extraction network includes: Obtain multiple periorbital image samples, where the multiple periorbital image samples include multiple periorbital image sample pictures and their respective corresponding labels; Obtain the eye region feature extraction network to be trained; Train the eye region feature extraction network to be trained according to the multiple eye region image samples to obtain the eye region feature extraction network.
7. The eye area feature extraction method according to claim 6, wherein The label includes an identity information label. The training of the eye region feature extraction network to be trained according to the multiple eye region image samples to obtain the eye region feature extraction network includes: Input the multiple eye region image samples into the model to be trained, and obtain the identity information corresponding to each predicted sample image. Among them, the model to be trained includes: an eye region feature extraction network to be trained, a fully connected layer to be trained, and a classifier to be trained. Among them, the eye region feature extraction network to be trained is used to extract the eye region features of the sample image; Determine the loss value according to the predicted identity information corresponding to each sample image and its corresponding identity information label using a preset loss function; Iteratively train the model to be trained according to the loss value and the multiple eye region image samples until the loss value reaches a preset loss threshold to obtain the final trained model. The final trained model includes: an eye region feature extraction network, a fully connected layer, and a classifier; Extract the eye region feature extraction network from the final trained model.
8. An eye perimeter feature extraction device, characterized in that, Include: An acquisition module: used to acquire the eye region image to be detected; A obtaining module: used to input the eye region image to be detected into the first residual module to obtain a first local feature; input the first local feature into the second residual module to obtain a second local feature, and input the second local feature into the third residual module to obtain a third local feature; fuse the second local feature and the third local feature to obtain an initial first local fusion feature; perform convolution on the initial first local fusion feature to obtain a first local fusion feature; input the first local fusion feature into the fourth residual module to obtain a fourth local feature; fuse the first local feature and the fourth local feature to obtain a second local fusion feature; pass the second local fusion feature through multiple sets of convolution components in sequence to obtain a local feature; perform feature processing on the local feature to obtain a global feature; fuse the local feature and the global feature to obtain an eye region feature.
9. An electronic device, characterized in that, Include: At least one processor; A memory; At least one application program, where at least one application program is stored in the memory and is configured to be executed by at least one processor. The at least one application program is configured to: execute the eye region feature extraction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the eye region feature extraction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image enhancement model training method, device, electronic device, and readable storage medium
CN109102483A
Recognition method for face wearing mask
CN114220143A