Image Processing Method, Apparatus, Device, Storage Medium and Program

By acquiring and analyzing the local features and contextual relationships of the user's face, using the multi-head self-attention mechanism and preset model to calculate the global contextual feature similarity, the problem of low face recognition accuracy in occlusion is solved, and higher recognition accuracy is achieved.

CN114519886BActive Publication Date: 2025-07-11WEBANK (CHINA)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210153526.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-07-11
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

In the prior art, when the user's face is blocked, the accuracy of face recognition is low and the user's identity cannot be accurately identified.

Method used

By acquiring the user's first face image collected by the electronic device and the user's second face image pre-stored, the context relationship between multiple local features is determined, and the cosine similarity of the global context features is calculated using the multi-head self-attention mechanism and the preset model to determine the matching result of the two face images.

Benefits of technology

Even if the user's face is blocked, the user's face can still be accurately identified through global context features, improving the accuracy of face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519886B_ABST
    Figure CN114519886B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method, apparatus, device, storage medium and program. The method includes: obtaining a first face image of a user collected by an electronic device and a second face image of the user stored in advance; obtaining a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image; determining a first context relationship among the plurality of first local features and a second context relationship among the plurality of second local features; and determining a matching result between the first face image and the second face image according to the plurality of first local features, the first context relationship, the plurality of second local features and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different. The accuracy of face recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image processing method, apparatus, device, storage medium and program. Background Art

[0002] Face recognition technology can be widely applied to technical fields such as finance, security, and education that require identity verification. When there are obstacles on the user's face (such as sunglasses, masks, scarves, etc.), the accuracy of the server's face recognition of the user is particularly important.

[0003] Currently, when the user's face is blocked, the server can verify the user's identity based on the local features of the face image. For example, when the user wears a mask, the server can obtain the user's face image and determine the local features (such as eyes, nose bridge, etc.) that are not blocked by the mask in the face image, and then verify the user's identity through the local features. However, when the user's face is blocked, the local features cannot fully reflect the user's face features, which causes the server to be unable to accurately recognize the user's face, resulting in low accuracy of face recognition. Summary of the Invention

[0004] The main purpose of the present invention is to provide an image processing method, apparatus, device, storage medium and program, aiming to solve the technical problem of low accuracy of face recognition in the prior art.

[0005] To achieve the above object, the present invention provides an image processing method, which includes:

[0006] Obtain a first face image of a user collected by an electronic device and a second face image of the user stored in advance;

[0007] Obtain a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image;

[0008] Determine a first context relationship between the plurality of first local features and a second context relationship between the plurality of second local features;

[0009] Determine a matching result of the first face image and the second face image according to the plurality of first local features, the first context relationship, the plurality of second local features and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different.

[0010] In a possible implementation manner, determining a matching result between the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship includes:

[0011] Determining a first global context feature according to the multiple first local features and the first context relationship;

[0012] Determining a second global context feature according to the multiple second local features and the second context relationship;

[0013] Determining the matching result according to the first global context feature and the second global context feature.

[0014] In a possible implementation manner, determining the matching result according to the first global context feature and the second global context feature includes:

[0015] Obtaining a cosine similarity between the first global context feature and the second global context feature;

[0016] If the cosine similarity is greater than or equal to a preset threshold, the matching result indicates that the first face image and the second face image are the same; if the cosine similarity is less than the preset threshold, the matching result indicates that the first face image and the second face image are different.

[0017] In a possible implementation manner, determining a first global context feature according to the multiple first local features and the first context relationship includes:

[0018] Obtaining a first preset weight;

[0019] Multiplying the multiple first local features by the first preset weight to obtain a first sub-feature;

[0020] Multiplying the first context relationship and the first sub-feature to obtain a first global context feature.

[0021] In a possible implementation manner, determining a second global context feature according to the multiple second local features and the second context relationship includes:

[0022] Multiplying the multiple second local features by the first preset weight to obtain a second sub-feature;

[0023] Multiplying the second context relationship and the second sub-feature to obtain a second global context feature.

[0024] In a possible implementation, determining a first context relationship between the multiple first local features includes:

[0025] Obtain a second preset weight, a third preset weight, and first position encodings corresponding to the multiple first local features;

[0026] Multiply the multiple first local features by the second preset weight to obtain third sub-features, and multiply the multiple first local features by the third preset weight to obtain fourth sub-features;

[0027] Multiply the third sub-features by the first position encodings to obtain first target encodings;

[0028] Multiply the third sub-features by the fourth sub-features to obtain the first context encoding;

[0029] Determine the sum of the first context encoding and the first target encodings as the first context relationship.

[0030] In a possible implementation, determining a second context relationship between the multiple second local features includes:

[0031] Obtain a second preset weight, a third preset weight, and second position encodings corresponding to the multiple second local features;

[0032] Multiply the multiple second local features by the second preset weight to obtain fifth sub-features, and multiply the multiple second local features by the third preset weight to obtain sixth sub-features;

[0033] Multiply the fifth sub-features by the second position encodings to obtain second target encodings;

[0034] Multiply the fifth sub-features by the sixth sub-features to obtain the second context encoding;

[0035] Determine the sum of the second context encoding and the second target encodings as the second context relationship.

[0036] In a possible implementation, obtaining the multiple first local features corresponding to the first face image and the multiple second local features corresponding to the second face image includes:

[0037] Process the first face image and the second face image respectively through a preset model to obtain the multiple first local features and the multiple second local features;

[0038] Determining the first context relationship between the multiple first local features and the second context relationship between the multiple second local features includes:

[0039] Process the first local feature and the second local feature respectively through a preset model to obtain the first context relationship and the second context relationship;

[0040] Determine the matching result of the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features and the second context relationship, including:

[0041] Process the multiple first local features, the first context relationship, the multiple second local features and the second context relationship through a preset model to obtain the matching result of the first face image and the second face image.

[0042] In a possible implementation manner, the preset model includes at least one convolutional layer and at least one self-attention layer, where,

[0043] Process the first face image and the second face image respectively through a preset model to obtain the multiple first local features and the multiple second local features, including:

[0044] Process the first face image and the second face image respectively through the at least one convolutional layer to obtain the multiple first local features and the multiple second local features;

[0045] Process the first local feature and the second local feature respectively through a preset model to obtain the first context relationship and the second context relationship, including:

[0046] Process the first local feature and the second local feature respectively through the at least one self-attention layer to obtain the first context relationship and the second context relationship;

[0047] Process the multiple first local features, the first context relationship, the multiple second local features and the second context relationship through a preset model to obtain the matching result of the first face image and the second face image, including:

[0048] Process the multiple first local features, the first context relationship, the multiple second local features and the second context relationship through the at least one self-attention layer to obtain the matching result of the first face image and the second face image.

[0049] The present invention also provides an image processing device, including a first acquisition module, a second acquisition module, a first determination module, and a second determination module, where:

[0050] The first acquisition module is configured to acquire a first face image of a user collected by an electronic device and a second face image of the user stored in advance;

[0051] The second acquisition module is configured to acquire a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image;

[0052] The first determination module is configured to determine a first context relationship among the plurality of first local features and a second context relationship among the plurality of second local features;

[0053] The second determination module is configured to determine a matching result between the first face image and the second face image according to the plurality of first local features, the first context relationship, the plurality of second local features, and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different.

[0054] The present invention further provides an image processing device, including: a memory, a processor, and an image processing program stored on the memory and executable on the processor. When the image processing program is executed by the processor, the steps of the image processing method described in any one of the foregoing items are implemented.

[0055] The present invention further provides a computer-readable storage medium, on which an image processing program is stored. When the image processing program is executed by a processor, the steps of the image processing method described in any one of the foregoing items are implemented.

[0056] The present invention further provides a computer program product, including a computer program, which implements the image processing method described in the first aspect when executed by a processor.

[0057] In the present invention, a server acquires a first face image of a user collected by an electronic device and a second face image of the user stored in advance, and acquires a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image. The server determines a first context relationship among the plurality of first local features and a second context relationship among the plurality of second local features, and determines a matching result between the first face image and the second face image according to the plurality of first local features, the first context relationship, the plurality of second local features, and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different. In this way, when the user's face is blocked, the server can not only acquire the local features of the face image, but also acquire the context relationship between the local features of the face image, and then combine the local features of the face and the global features indicated by the context relationship to identify the user's face, avoiding the interference caused by the occluder to face recognition and improving the accuracy of face recognition by the server. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0059] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present invention;

[0060] Figure 2 A flowchart of an image processing method provided by an embodiment of the present invention;

[0061] Figure 3 A schematic diagram of a process for obtaining a plurality of first local features provided by an embodiment of the present invention;

[0062] Figure 4 A schematic diagram of a process of another image processing method provided by an embodiment of the present invention;

[0063] Figure 5 A schematic diagram of the structure of a preset model provided by an embodiment of the present invention;

[0064] Figure 6 A schematic diagram of the structure of a self-attention layer provided by an embodiment of the present invention;

[0065] Figure 7 A schematic diagram of obtaining a first global context feature provided by an embodiment of the present invention;

[0066] Figure 8Schematic diagram of a process of an image processing method provided by an embodiment of the present invention;

[0067] Figure 9 Schematic diagram of the structure of an image processing apparatus provided by an embodiment of the present invention;

[0068] Figure 10 Schematic diagram of the structure of an image processing device provided by an embodiment of the present invention.

[0069] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0070] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0071] When the user's face is blocked, the server can verify the user's identity through the local features in the user's face image. For example, when the user wears a mask, the server can collect the user's face image in real time, obtain the eye features in the face image, and obtain the user's eye features in the face image pre-uploaded by the user, and then verify the user's identity through the two eye features. However, the accuracy of face recognition only through the unblocked local features is relatively low, which in turn leads to a relatively low accuracy of user identity verification.

[0072] To solve this technical problem, an embodiment of the present invention provides a solution. The server obtains a first face image of a user collected by an electronic device and a second face image of the user stored in advance, and obtains a first local feature in the first face image and a second local feature of the second face image. Through the multi-head self-attention mechanism, the first context relationship between the first local features and the second context relationship between the second local features are determined, and through multiple first local features and the first context relationship, the first global context feature of the first face image is determined, and through multiple second local features and the second context relationship, the second global context feature of the second face image is determined. Then, through the cosine similarity between the first global context feature and the second global context feature, the matching result between the first face image and the second face image is determined. Since the global context feature of the face image includes local features and global features indicated by the context relationship, even if the user's face is blocked, the user's face can be recognized through the global context feature, thereby improving the accuracy of face recognition.

[0073] Next, in combination with Figure 1 , the application scenarios of the present invention will be described.

[0074] Figure 1 FIG. Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present invention. Please refer to

[0075] Next, in combination with the accompanying drawings, some embodiments of the present invention will be described in detail. Without conflict between the embodiments, the embodiments and the features in the embodiments can be combined with each other.

[0076] Figure 2 FIG. Figure 2 is a schematic flowchart of an image processing method provided by an embodiment of the present invention. Please refer to

[0077] S201. Obtain a first face image of a user collected by an electronic device and a second face image of the user stored in advance.

[0078] In this embodiment, the execution subject of the method may be a server or an image processing device provided in the server. Among them, the image processing device can be implemented by software or by a combination of software and hardware.

[0079] Optionally, the electronic device can be any device with an image acquisition function. For example, the electronic device can be a mobile phone, a computer, a face recognition access control device, a camera device, etc. Optionally, the first face image can be a face image of a user collected by the electronic device. For example, the first face image can be an image of a user's face collected in real time by the electronic device. For example, when the user approaches the image acquisition device (camera) of the electronic device, the electronic device can capture an image or video including the user's face and determine an image in the image or a frame of the video as the first face image.

[0080] Optionally, the face region of the first face image includes an occlusion image. For example, when the electronic device captures the first face image of the user, there is an occluder on the user's face. For example, when the electronic device captures the face image of the user, the user may wear an occluder such as a scarf, a hat, sunglasses or a mask.

[0081] Optionally, the second face image may be an image including the user's complete face. For example, the second face image may be an image such as the user's ID photo or selfie. Optionally, the user may pre-store the second face image in the server. For example, the user may upload an image including the complete face to the server in advance. When verifying the user's identity, the server may obtain the second face image of the user from the database.

[0082] Optionally, when the electronic device captures the first face image of the user, the electronic device may send the first face image and the user's identifier to the server, so that the server obtains the first face image of the user and obtains the second face image of the user according to the user's identifier. For example, in the actual application process, when the user unlocks the mobile phone while wearing a mask, the mobile phone can capture the first face image of the user (the face image of the user including the mask) and send the first face image of the user and the user's identifier (when the user has business requirements such as unlocking and verification, the user's identifier can be obtained) to the server. The server obtains the pre-stored second face image of the user according to the user's identifier, and then verifies the user's identity according to the first face image and the second face image. If the first face image and the second face image match, the mobile phone unlocks successfully. If the first face image and the second face image do not match, the mobile phone unlocks fails.

[0083] S202. Obtain a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image.

[0084] Optionally, the first local feature may be an image feature in the first face image. For example, the first face image is an image with an occluder. The first local feature may be an occluder feature, a feature of the unoccluded region, etc. For example, if the first face image is a face image of the user wearing a mask, the first local feature may be a mask feature, an eye feature, a nose bridge feature, a face shape feature, etc.; if the first face image is a face image of the user wearing sunglasses, the first local feature may be a sunglasses feature, a nose feature, a mouth feature, a face shape feature, etc.

[0085] Optionally, the second local feature may be an image feature in the second face image. For example, the second local feature may be the user's facial features such as the user's eye feature, nose feature, mouth feature, and face shape feature.

[0086] Optionally, the first facial image can be subjected to feature extraction through an image feature extraction algorithm to obtain a plurality of first local features. Optionally, the first facial image can be subjected to feature extraction through a pre-trained machine learning model to obtain a plurality of first local features. For example, the machine learning model can be a convolutional model including a plurality of convolutional layers. After the machine learning model is trained, the first facial image can be input into the machine learning model, and the machine learning model can output a plurality of first local features corresponding to the first facial image.

[0087] Optionally, the second facial image can be subjected to feature extraction through an image feature extraction algorithm or a pre-trained machine learning model to obtain a plurality of second local features. For example, by inputting the second facial image into the algorithm or the pre-trained machine learning model, the algorithm or the machine learning model can output a plurality of second local features corresponding to the second facial image.

[0088] Since the methods for obtaining the first local features and the second local features are the same, therefore, in combination with Figure 3 taking the process of obtaining a plurality of first local features as an example, the process of obtaining a plurality of first local features will be described.

[0089] Figure 3 FIG. is a schematic diagram of a process for obtaining a plurality of first local features provided by an embodiment of the present invention. Please refer to Figure 3 , which includes a convolutional model and a first facial image. Among them, the first facial image includes a sunglass image. When the convolutional model receives the first facial image, it can perform multiple convolutional processes on the first facial image, and then output feature vectors of the sunglasses, hairstyle, ears, mouth, etc. corresponding to the first facial image.

[0090] S203. Determine a first context relationship between a plurality of first local features and a second context relationship between a plurality of second local features.

[0091] Optionally, the first context relationship is used to indicate the degree of correlation between a plurality of first local features. For example, if the first local features include eye features, nose features, and mouth features, then for the eye features, the first context relationship can indicate the degree of correlation between the eye features and the nose features and the mouth features respectively. For the nose features, the first context relationship can indicate the degree of correlation between the nose features and the mouth features.

[0092] Optionally, the second context relationship is used to indicate the degree of correlation between multiple second local features. For example, if the second local features include eye features, ear features, and face shape features, then for the eye features, the second context relationship can indicate the degree of correlation between the eye features and the ear features and the face shape features respectively. For the ear features, the second context relationship can indicate the degree of correlation between the nose features and the face shape features.

[0093] Optionally, the first context relationship between multiple first local features can be determined according to the following feasible implementation methods: Obtain a second preset weight, a third preset weight, and first position encodings corresponding to multiple first local features. The second preset weight and the third preset weight are preset weight parameters. For example, the second preset weight can be a 1*1 convolutional layer, and through model training (such as the self-attention algorithm), etc., the parameters of this convolutional layer are adjusted to obtain the second preset weight. The third preset weight can also be a 1*1 convolutional layer, and through model training (such as the self-attention algorithm), etc., the parameters of this convolutional layer are adjusted to obtain the third preset weight.

[0094] The first position encoding is used to indicate the positions of multiple first local features on the planar feature map. For example, if multiple first local features are 512-dimensional vectors, the first position encoding is used to indicate the positions of each vector when the 512-dimensional vectors are compressed into a plane. For example, the vectors of multiple first local features can be compressed into a feature map of a plane, and the first position encoding includes the positions of the vectors of each local feature in the feature map.

[0095] Multiply multiple first local features by the second preset weight to obtain a third sub-feature. For example, the second preset weight is the parameter of a 1*1 convolutional layer. Input the first local feature into the convolutional layer, and the convolutional layer can output the third sub-feature corresponding to multiple first local features.

[0096] Multiply multiple first local features by the third preset weight to obtain a fourth sub-feature. For example, the third preset weight is the parameter of another 1*1 convolutional layer. Input multiple first local features into this convolutional layer, and the convolutional layer can output the fourth sub-feature corresponding to multiple first local features.

[0097] Multiply the third sub-feature by the first position encoding to obtain a first target encoding. For example, perform a dot product between the vector of the third sub-feature and the first position encoding to obtain the first target encoding.

[0098] Multiply the third sub-feature by the fourth sub-feature to obtain a first context encoding. For example, perform a dot product between the third sub-feature and the fourth sub-feature to obtain the first context encoding.

[0099] Determine the sum of the first context encoding and the first target encoding as the first context relationship. For example, add the first context encoding and the first target encoding to obtain the first context relationship corresponding to the first face image.

[0100] Optionally, the second context relationship between multiple second local features can be determined according to the following feasible implementation methods: Obtain the second preset weight, the third preset weight, and the second position encodings corresponding to the multiple second local features.

[0101] The second position encoding is used to indicate the positions of the multiple second local features on the planar feature map. For example, if the multiple second local features are vectors of 512 dimensions, the second position encoding is used to indicate the positions of each vector when the 512-dimensional vectors are compressed into a plane. For example, the vectors of the multiple second local features can be compressed into a planar feature map, and the second position encoding includes the positions of the vectors of each local feature in the feature map.

[0102] Multiply the multiple second local features by the second preset weight to obtain the fifth sub-feature. For example, the second preset weight is the parameter of a 1*1 convolutional layer. Input the second local feature into the convolutional layer, and the convolutional layer can output the vectors of the fifth sub-features corresponding to the multiple second local features.

[0103] Multiply the multiple second local features by the third preset weight to obtain the sixth sub-feature. For example, the third preset weight is the parameter of another 1*1 convolutional layer. Input the multiple second local features into this convolutional layer, and the convolutional layer can output the sixth sub-features corresponding to the multiple second local features.

[0104] Multiply the fifth sub-feature and the second position encoding to obtain the second target encoding. For example, perform a dot product between the vector of the fifth sub-feature and the second position encoding to obtain the second target encoding.

[0105] Multiply the fifth sub-feature and the sixth sub-feature to obtain the second context encoding. For example, perform a dot product between the fifth sub-feature and the sixth sub-feature to obtain the second context encoding.

[0106] Determine the sum of the second context encoding and the second target encoding as the second context relationship. For example, add the second context encoding and the second target encoding to obtain the second context relationship corresponding to the second face image.

[0107] S204. Determine the matching result of the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship.

[0108] Optionally, the matching result is used to indicate whether the first face image and the second face image are the same or different. For example, if the similarity between the first face image and the second face image is high, the matching result indicates that the first face image and the second face image are the same; if the similarity between the first face image and the second face image is low, the matching result indicates that the first face image and the second face image are different.

[0109] Optionally, the matching result of the first face image and the second face image can be obtained according to the following feasible implementation methods: Determine the first global context feature according to multiple first local features and the first context relationship. The first global context feature includes the local features of the first face image and the global features with context relationships. For example, the server can combine multiple first local features and the first context relationship to achieve the combination of features and obtain the first global context feature.

[0110] Optionally, the first global context feature can be obtained according to the following feasible implementation methods: Obtain the first preset weight. The first preset weight can be the parameter of a 1*1 convolutional layer (which can be obtained through machine learning or can be a preset parameter). Multiply multiple first local features by the first preset weight to obtain the first sub-feature. Multiply the first context relationship by the first sub-feature to obtain the first global context feature. For example, perform a dot product operation on the first context relationship and the first sub-feature to obtain the first global context feature.

[0111] Determine the second global context feature according to multiple second local features and the second context relationship. The second global context feature includes the local features of the second face image and the global features with context relationships. For example, the server can combine multiple second local features and the second context relationship to achieve the combination of features and obtain the first global context feature.

[0112] Optionally, the second global context feature can be determined according to the following feasible implementation methods: Multiply multiple second local features by the first preset weight to obtain the second sub-feature. Multiply the second context relationship by the second sub-feature to obtain the second global context feature. For example, perform a dot product on the second context relationship and the second sub-feature to obtain the second global context feature.

[0113] Determine the matching result according to the first global context feature and the second global context feature. Optionally, the matching result can be determined according to the following feasible implementation methods: Obtain the cosine similarity between the first global context feature and the second global context feature. For example, the first global context feature of the first face image obtained by the server is a 512-dimensional vector, and the second global context feature of the second face image obtained by the server is also a 512-dimensional vector. The cosine value between the two vectors is determined as the cosine similarity between the first global context feature and the second global context feature.

[0114] Optionally, if the cosine similarity is greater than or equal to a preset threshold, the matching result indicates that the first face image and the second face image are the same. If the cosine similarity is less than the preset threshold, the matching result indicates that the first face image and the second face image are different. For example, if the cosine similarity is greater than or equal to the preset threshold, it means that although some features of the first face image are occluded, the similarity between the first face image and the second face image is relatively high, that is, the first face image and the second face image are the same. The server determines that the user indicated by the first face image and the user indicated by the second face image are the same person. If the cosine similarity is less than the preset threshold, it means that the similarity between the first face image and the second face image is relatively low, that is, the first face image and the second face image are different. The server determines that the user indicated by the first face image and the user indicated by the second face image are not the same person.

[0115] An embodiment of the present invention provides an image processing method, which acquires a first face image of a user collected by an electronic device and a second face image of the user stored in advance, acquires a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image, determines a first context relationship between the first local features and a second context relationship between the second local features through a multi-head self-attention mechanism, determines a first global context feature of the first face image through the plurality of first local features and the first context relationship, determines a second global context feature of the second face image through the plurality of second local features and the second context relationship, and further determines a matching result between the first face image and the second face image through the cosine similarity between the first global context feature and the second global context feature. Since the global context feature of the face image includes local features and global features indicated by the context relationship, even if the user's face is occluded, the user's face can be recognized through the global context feature, thereby improving the accuracy of face recognition.

[0116] In Figure 2 On the basis of the shown embodiment, below, in combination with Figure 4 , the process of the above image processing method will be described.

[0117] Figure 4 It is a schematic diagram of the process of another image processing method provided by an embodiment of the present invention. Please refer to Figure 4 , and the method flow includes:

[0118] S401. Obtain the first face image of the user collected by the electronic device and the second face image of the user stored in advance.

[0119] It should be noted that the execution process of step S401 is the same as that of step S201, and the embodiment of the present invention will not elaborate on this again.

[0120] S402. Obtain a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image.

[0121] Optionally, the following feasible implementation manner can be used to obtain a plurality of first local features and a plurality of second local features: Process the first face image and the second face image respectively through a preset model to obtain a plurality of first local features and a plurality of second local features.

[0122] Optionally, the preset model is obtained by training multiple groups of samples, and the preset model includes at least one convolutional layer and at least one self-attention layer. For example, the preset model may include 10 convolutional layers and 4 multi-head self-attention layers.

[0123] Next, in combination with Figure 5 , the structure of the preset model will be described.

[0124] Figure 5 It is a schematic diagram of the structure of a preset model provided by an embodiment of the present invention. Please refer to Figure 5 , the preset model includes a convolutional module and a self-attention module. Among them, each layer structure of the preset model is a residual structure, and the preset model includes 6 residual structures. In the convolutional module, there are 4 residual structures, and each residual structure has a convolutional layer built in. In the self-attention module, there are 2 residual structures, and each residual structure has a multi-head self-attention layer built in.

[0125] Please refer to Figure 5 , when the preset model obtains a face image, through the multi-layer convolutional structure in the convolutional module, obtain the local features of the face image, and input the local features to the self-attention module. The self-attention module can obtain the context information between multiple local features, and process the multiple local features according to the context information, so as to obtain the global context features corresponding to the face image.

[0126] Optionally, when training the preset model, the training sample data consists of unoccluded face data and occluded face data, and the two types of data are randomly distributed. For example, each set of training sample data includes multiple face images and the corresponding class labels for each face image. When training the model, face detection and face alignment can be performed on the multiple face images in advance, and the aligned face images are normalized. Then, the normalized face images are input into the preset model. The preset model outputs the global context features corresponding to the face images (such as a 512-dimensional feature vector). Then, the global context features are normalized, and then the class cosine value of the face image is generated through a fully connected layer. Finally, the class cosine value is input into the loss function to calculate the loss, thereby guiding the learning of the preset model (to maximize the cosine similarity between the global context features of the occluded face images and the unoccluded face images of the same user). When the preset model converges, the preset model training is completed. In this way, when the preset model is used, even if the user's face wears an occluder, the preset model can accurately recognize the user's face.

[0127] Optionally, when training the model, the margin-softmax loss function can be used as the loss function. By using the margin-softmax loss function, the intra-class distance of the model can be made smaller and the inter-class distance can be made larger, thereby improving the discrimination ability of the model. Optionally, the specific representation of the loss function can be shown as the following formula:

[0128]

[0129] where m1, m2, and m3 are the margins in three dimensions; s is the scale; is the angle between the global context feature and the corresponding class center. Optionally, both the global context feature and the class center of the face image need to be normalized.

[0130] Optionally, when using the preset model, multiple first local features corresponding to the first face image and multiple second local features corresponding to the second face image are obtained through the preset model. Specifically: the first face image and the second face image are processed through at least one convolutional layer respectively to obtain multiple first local features and multiple second local features. For example, the first face image is input into the convolutional layer in the preset model, and the convolutional layer performs multi-layer convolutional processing on the first face image to obtain multiple first local features (such as a 512-dimensional feature vector). The second face image is input into the convolutional layer in the preset model, and the convolutional layer performs multi-layer convolutional processing on the second face image to obtain multiple second local features (such as a 512-dimensional feature vector).

[0131] S403. Determine the first context relationship among multiple first local features and the second context relationship among multiple second local features.

[0132] Optionally, the first context relationship and the second context relationship can be determined through the following feasible implementation methods: respectively process the first local features and the second local features through a preset model to obtain the first context relationship and the second context relationship.

[0133] Optionally, obtain the first context relationship and the second context relationship through a preset model. Specifically: process the first local features and the second local features through at least one self-attention layer respectively to obtain the first context relationship and the second context relationship. For example, after processing the first face image through multiple convolutional layers, the convolutional layers transmit multiple first local features to the self-attention layer, and multiple self-attention layers can process the multiple first local features, and then extract the first context relationship among the multiple first local features. It should be noted that the method for determining the second context relationship is the same as that for determining the first context relationship, and the present invention will not elaborate here.

[0134] Optionally, when designing the self-attention layer, a multi-head self-attention layer can be used to increase the number of dimensions of the context relationship obtained by the self-attention layer. For example, the first self-attention layer mainly obtains the context relationship between eyes and eyes, and the second self-attention layer mainly obtains the context relationship between the nose and the mouth, etc. In this way, a comprehensive context relationship can be obtained, thereby improving the accuracy of face recognition.

[0135] Next, in combination with Figure 6 , the structure of the self-attention layer will be described.

[0136] Figure 6 FIG. Figure 6 is a schematic structural diagram of a self-attention layer provided by an embodiment of the present invention. Please refer to

[0137] S404. Determine the first global context feature according to the multiple first local features and the first context relationship.

[0138] Optionally, the first global context feature can be determined according to the following feasible implementation methods: the preset model processes multiple first local features and the first context relationship to obtain the first global context feature corresponding to the first face image. For example, the first global context feature is obtained through at least one self-attention layer for multiple first local features and the first context relationship. For example, multiple self-attention layers are used to process multiple first local features, and then the first context relationship between the multiple first local features is extracted. Based on the first context relationship and the first local features, the first global context feature is output.

[0139] Next, in conjunction with Figure 7 , the process of obtaining the first global context feature through the self-attention layer will be described.

[0140] Figure 7 FIG. Figure 7 is a schematic diagram of obtaining the first global context feature provided by an embodiment of the present invention. Please refer to

[0141] Please refer to Figure 7 . Rh is added to Rw to obtain r. The first local feature is converted to q through Wq, converted to k through Wk, and converted to v through Wv. q and k are dot-multiplied to obtain the first context encoding. q and r are dot-multiplied to obtain the first position encoding. The first position encoding and the first context encoding are added, and the result is input to the softmax layer to obtain the first context relationship. The first context relationship and v are dot-multiplied to obtain the first global context feature corresponding to the first face image.

[0142] S405. Determine the second global context feature according to multiple second local features and the second context relationship.

[0143] It should be noted that the execution process of step S405 can refer to step S404, and the embodiments of the present invention will not elaborate on this.

[0144] S406. Determine the matching result according to the first global context feature and the second global context feature.

[0145] Optionally, the matching result of the first face image and the second face image may be determined according to the cosine similarity between the first global context feature and the second global context feature. Optionally, a preset model may be used to process multiple first local features, the first context relationship, multiple second local features, and the second context relationship to obtain the matching result of the first face image and the second face image. For example, the multiple first local features, the first context relationship, multiple second local features, and the second context relationship are processed through at least one self-attention layer to obtain the matching result of the first face image and the second face image. For example, after the multiple first local features of the first face image are obtained by a multi-layer convolutional layer, the convolutional layer sends the multiple first local features to the self-attention layer. The self-attention layer may extract the first context relationship between the multiple first local features, and combine the multiple first local features and the first context relationship to obtain a fused first global context feature. By the same method, the second global context feature corresponding to the second face image may be obtained, and then the matching result of the first face image and the second face image is obtained through the cosine similarity between the first global context feature and the second global context feature.

[0146] An embodiment of the present invention provides an image processing method, which acquires a first face image of a user collected by an electronic device and a second face image of the user stored in advance, acquires multiple first local features corresponding to the first face image and multiple second local features corresponding to the second face image, determines a first context relationship between the multiple first local features and a second context relationship between the multiple second local features, determines a first global context feature according to the multiple first local features and the first context relationship, determines a second global context feature according to the multiple second local features and the second context relationship, and determines a matching result according to the first global context feature and the second global context feature. Through the training of the preset model and the structure of the self-attention layer, an accurate context relationship can be obtained when the face is occluded. Since the global context feature is determined by the local features and the context relationship, the global context feature includes not only the local features but also the global features of the face image. Therefore, through the global context feature, the matching degree of two face images can be accurately determined, and the accuracy of face recognition when the face is occluded is improved.

[0147] Based on any of the above embodiments, hereinafter, in combination with Figure 8 , the process of the above image processing method will be described.

[0148] Figure 8 FIG. is a schematic diagram of the process of an image processing method provided by an embodiment of the present invention. Please refer to Figure 8, including: a user, an electronic device, and a server. Among them, the electronic device can collect the user's first face image, and the user's first face image includes a sunglasses image. After the electronic device collects the first face image, it can send the first face image to the server. When the server receives the first face image, it can obtain the second face image pre-uploaded by the user in the database, where the second face image is a complete image of the user's face.

[0149] Please refer to Figure 8 , the server inputs the first face image and the second face image into a preset model respectively. When the preset model receives the first face image, it performs 4 convolutional processes on the first face image through a convolutional module to obtain multiple first local features, and inputs the multiple first local features into a self-attention module. The self-attention module extracts the first context relationship between the multiple first local features, and combines the first context relationship to process the multiple first local features to obtain the first global context feature corresponding to the first face image. By the same method, the preset model can obtain the second global context feature of the second face image.

[0150] Please refer to Figure 8 , the preset model determines that the pre-similarity between the first global context feature and the second global context feature is greater than a preset threshold. Therefore, the preset model determines that the first face image and the second face image are the same. In this way, since the global context feature is obtained by combining local features and context relationships, even if there are obstacles on the user's face, the server can still identify the face image based on the global context feature, thereby improving the accuracy of face image recognition.

[0151] Figure 9 It is a schematic structural diagram of an image processing device provided by an embodiment of the present invention. Please refer to Figure 9 , the image processing device 10 includes a first acquisition module 11, a second acquisition module 12, a first determination module 13, and a second determination module 14, where:

[0152] The first acquisition module 11 is used to acquire the user's first face image collected by the electronic device and the user's second face image stored in advance;

[0153] The second acquisition module 12 is used to acquire multiple first local features corresponding to the first face image and multiple second local features corresponding to the second face image;

[0154] The first determination module 13 is used to determine the first context relationship between the multiple first local features and the second context relationship between the multiple second local features;

[0155] The second determination module 14 is configured to determine a matching result between the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different.

[0156] In a possible implementation manner, the second determination module 14 is specifically configured to:

[0157] Determine a first global context feature according to the multiple first local features and the first context relationship;

[0158] Determine a second global context feature according to the multiple second local features and the second context relationship;

[0159] Determine the matching result according to the first global context feature and the second global context feature.

[0160] In a possible implementation manner, the second determination module 14 is specifically configured to:

[0161] Obtain a cosine similarity between the first global context feature and the second global context feature;

[0162] If the cosine similarity is greater than or equal to a preset threshold, the matching result indicates that the first face image and the second face image are the same; if the cosine similarity is less than the preset threshold, the matching result indicates that the first face image and the second face image are different.

[0163] In a possible implementation manner, the second determination module 14 is specifically configured to:

[0164] Obtain a first preset weight;

[0165] Multiply the multiple first local features by the first preset weight to obtain a first sub-feature;

[0166] Multiply the first context relationship by the first sub-feature to obtain a first global context feature.

[0167] In a possible implementation manner, the second determination module 14 is specifically configured to:

[0168] Multiply the multiple second local features by the first preset weight to obtain a second sub-feature;

[0169] Multiply the second context relationship by the second sub-feature to obtain a second global context feature.

[0170] In a possible implementation manner, the first determination module 13 is specifically configured to:

[0171] Obtain a second preset weight, a third preset weight, and first position encodings corresponding to the multiple first local features;

[0172] Multiply the multiple first local features by the second preset weight to obtain third sub-features, and multiply the multiple first local features by the third preset weight to obtain fourth sub-features,

[0173] Multiply the third sub-features by the first position encodings to obtain first target encodings;

[0174] Multiply the third sub-features by the fourth sub-features to obtain the first context encoding;

[0175] Determine the sum of the first context encoding and the first target encodings as the first context relationship.

[0176] In a possible implementation manner, the first determination module 13 is specifically configured to:

[0177] Obtain a second preset weight, a third preset weight, and second position encodings corresponding to the multiple second local features;

[0178] Multiply the multiple second local features by the second preset weight to obtain fifth sub-features, and multiply the multiple second local features by the third preset weight to obtain sixth sub-features;

[0179] Multiply the fifth sub-features by the second position encodings to obtain second target encodings;

[0180] Multiply the fifth sub-features by the sixth sub-features to obtain the second context encoding;

[0181] Determine the sum of the second context encoding and the second target encodings as the second context relationship.

[0182] In a possible implementation manner:

[0183] Obtaining the multiple first local features corresponding to the first face image and the multiple second local features corresponding to the second face image includes:

[0184] Processing the first face image and the second face image respectively through a preset model to obtain the multiple first local features and the multiple second local features;

[0185] Determining a first context relationship between the multiple first local features and a second context relationship between the multiple second local features includes:

[0186] Process the first local feature and the second local feature respectively through a preset model to obtain the first context relationship and the second context relationship;

[0187] Determine the matching result of the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features and the second context relationship, including:

[0188] Process the multiple first local features, the first context relationship, the multiple second local features and the second context relationship through a preset model to obtain the matching result of the first face image and the second face image.

[0189] In a possible implementation manner,

[0190] Process the first face image and the second face image respectively through a preset model to obtain the multiple first local features and the multiple second local features, including:

[0191] Process the first face image and the second face image respectively through the at least one convolutional layer to obtain the multiple first local features and the multiple second local features;

[0192] Process the first local feature and the second local feature respectively through a preset model to obtain the first context relationship and the second context relationship, including:

[0193] Process the first local feature and the second local feature respectively through the at least one self-attention layer to obtain the first context relationship and the second context relationship;

[0194] Process the multiple first local features, the first context relationship, the multiple second local features and the second context relationship through a preset model to obtain the matching result of the first face image and the second face image, including:

[0195] Process the multiple first local features, the first context relationship, the multiple second local features and the second context relationship through the at least one self-attention layer to obtain the matching result of the first face image and the second face image.

[0196] The image processing device provided by the embodiments of the present invention can execute the technical solutions shown in the above method embodiments, and its implementation principles and beneficial effects are similar, and will not be elaborated here.

[0197] The image processing device shown in the embodiments of the present invention may be a chip, a hardware module, a processor, etc. Of course, the image processing device may be in other forms, and the embodiments of the present invention do not specifically limit this.

[0198] Figure 10 It is a schematic structural diagram of an image processing device provided by an embodiment of the present invention. As Figure 10 shown, the device may include: a memory 1001, a processor 1002, and an image processing program stored on the memory 1001 and executable on the processor 1002. When the image processing program is executed by the processor 1002, it implements the steps of the image processing method described in any of the foregoing embodiments.

[0199] Optionally, the memory 1001 can be either independent or integrated with the processor 1002.

[0200] The implementation principle and technical effects of the device provided in this embodiment can be referred to the foregoing embodiments, and will not be elaborated here.

[0201] The embodiments of the present invention also provide a computer-readable storage medium, on which an image processing program is stored. When the image processing program is executed by a processor, it implements the steps of the image processing method described in any of the foregoing embodiments.

[0202] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0203] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in the various embodiments of the present invention.

[0204] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or executed and completed by a combination of hardware and software modules in the processor.

[0205] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0206] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0207] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0208] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.

[0209] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0210] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0211] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An image processing method, characterized in that, Including: Obtaining a first face image of a user collected by an electronic device and a second face image of the user stored in advance; Obtaining a plurality of first local features corresponding to the first face image and a plurality of second local features corresponding to the second face image; Determining a first context relationship among the plurality of first local features and a second context relationship among the plurality of second local features; Determining a matching result between the first face image and the second face image according to the plurality of first local features, the first context relationship, the plurality of second local features, and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different; Determining the first context relationship among the plurality of first local features includes: obtaining a second preset weight, a third preset weight, and first position encodings corresponding to the plurality of first local features; Multiplying the plurality of first local features by the second preset weight to obtain a third sub-feature, and multiplying the plurality of first local features by the third preset weight to obtain a fourth sub-feature; Multiplying the third sub-feature by the first position encoding to obtain a first target encoding; Multiplying the third sub-feature by the fourth sub-feature to obtain a first context encoding; Determining the sum of the first context encoding and the first target encoding as the first context relationship.

2. The method according to claim 1, wherein Determining the matching result between the first face image and the second face image according to the plurality of first local features, the first context relationship, the plurality of second local features, and the second context relationship includes: Determining a first global context feature according to the plurality of first local features and the first context relationship; Determining a second global context feature according to the plurality of second local features and the second context relationship; Determining the matching result according to the first global context feature and the second global context feature.

3. The method according to claim 2, wherein Determining the matching result according to the first global context feature and the second global context feature includes: Obtaining a cosine similarity between the first global context feature and the second global context feature; If the cosine similarity is greater than or equal to a preset threshold, the matching result indicates that the first face image and the second face image are the same; if the cosine similarity is less than the preset threshold, the matching result indicates that the first face image and the second face image are different.

4. The method according to claim 2 or 3, characterized in that Determining the first global context feature according to the plurality of first local features and the first context relationship includes: Obtaining a first preset weight; Multiplying the plurality of first local features by the first preset weight to obtain a first sub-feature; Multiplying the first context relationship by the first sub-feature to obtain a first global context feature.

5. The method according to claim 4, characterized in that, Determining the second global context feature according to the plurality of second local features and the second context relationship includes: Multiplying the plurality of second local features by the first preset weight to obtain a second sub-feature; Multiply the second context relationship and the second sub-feature to obtain a second global context feature.

6. The method according to any one of claims 1-3 and 5, characterized in that, Determine the second context relationship among the multiple second local features, including: Obtain a second preset weight, a third preset weight, and second position encodings corresponding to the multiple second local features; Multiply the multiple second local features by the second preset weight to obtain a fifth sub-feature, and multiply the multiple second local features by the third preset weight to obtain a sixth sub-feature; Multiply the fifth sub-feature and the second position encoding to obtain a second target encoding; Multiply the fifth sub-feature and the sixth sub-feature to obtain the second context encoding; Determine the sum of the second context encoding and the second target encoding as the second context relationship.

7. The method according to any one of claims 1-3, 5, wherein Obtain multiple first local features corresponding to the first face image and multiple second local features corresponding to the second face image, including: Process the first face image and the second face image respectively through a preset model to obtain the multiple first local features and the multiple second local features; Determine the first context relationship among the multiple first local features and the second context relationship among the multiple second local features, including: Process the first local features and the second local features respectively through a preset model to obtain the first context relationship and the second context relationship; Determine the matching result of the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship, including: Process the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship through a preset model to obtain the matching result of the first face image and the second face image.

8. The method according to claim 7, characterized in that, The preset model includes at least one convolutional layer and at least one self-attention layer, wherein Process the first face image and the second face image respectively through a preset model to obtain the multiple first local features and the multiple second local features, including: Process the first face image and the second face image respectively through the at least one convolutional layer to obtain the multiple first local features and the multiple second local features; Process the first local features and the second local features respectively through a preset model to obtain the first context relationship and the second context relationship, including: Process the first local features and the second local features respectively through the at least one self-attention layer to obtain the first context relationship and the second context relationship; Process the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship through a preset model to obtain the matching result of the first face image and the second face image, including: Process the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship through the at least one self-attention layer to obtain a matching result of the first face image and the second face image.

9. An image processing apparatus, characterized in that, Comprising a first acquisition module, a second acquisition module, a first determination module, and a second determination module, where: The first acquisition module is configured to acquire a first face image of a user collected by an electronic device and a second face image of the user stored in advance. The second acquisition module is configured to acquire multiple first local features corresponding to the first face image and multiple second local features corresponding to the second face image. The first determination module is configured to determine a first context relationship between the multiple first local features and a second context relationship between the multiple second local features. The second determination module is configured to determine a matching result of the first face image and the second face image according to the multiple first local features, the first context relationship, the multiple second local features, and the second context relationship, where the matching result is used to indicate whether the first face image and the second face image are the same or different. Specifically, the first determination module is configured to: acquire a second preset weight, a third preset weight, and a first position encoding corresponding to the multiple first local features. Multiply the multiple first local features by the second preset weight to obtain a third sub-feature, and multiply the multiple first local features by the third preset weight to obtain a fourth sub-feature. Multiply the third sub-feature by the first position encoding to obtain a first target encoding. Multiply the third sub-feature by the fourth sub-feature to obtain a first context encoding. Determine the sum of the first context encoding and the first target encoding as the first context relationship.

10. An image processing apparatus, characterized in that, The image processing device includes: a memory, a processor, and an image processing program stored on the memory and executable on the processor. When the image processing program is executed by the processor, the steps of the image processing method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium, characterized in that, An image processing program is stored on the computer-readable storage medium. When the image processing program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Face Recognition Apparatus and Methods

    US20120230545A1