Face relation verification method based on cross-image information interaction

By adding information interaction module after Swin-Transformer block and combining RGB-HSV-Lab color space representation, the problem of insufficient robustness and accuracy of face relationship verification in the prior art is solved, and efficient verification in complex face cases is achieved.

CN119992625APending Publication Date: 2025-05-13SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510116479.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13

Smart Images

  • Figure CN119992625A_ABST
    Figure CN119992625A_ABST
Patent Text Reader

Abstract

The invention provides a face relation verification method based on cross-image information interaction, and belongs to the technical field of face relations, and the method comprises the steps: carrying out the color space transformation of RGB pictures of different face images, and obtaining an RGB-HSV-Lab picture vector; the RGB-HSV-Lab picture vector is input into a Swinin-Transform network fused with an information interaction module, and instance-level feature extraction is guided step by step; the extracted features are fused through a quadratic arithmetical operation and fusion module, and a supervised contrast loss function constraint model is utilized to improve the distinction degree; verifying whether the image pairs have a relationship according to an output result of the classifier; an information interaction module is added behind a Swinin-Transform block to dynamically guide instance-level feature extraction step by step, and face relation verification under different complex face conditions is automatically and efficiently achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of face relationship verification, and relates to a face relationship verification method based on cross-image information interaction. Background Art

[0002] Face relationship verification is a task that uses computer vision technology to determine whether there is a kinship relationship between two face images. Face relationship verification is of great significance in many application scenarios, including automatic organization of family albums, missing persons search, child adoption, genetic research, criminal investigation, and user relationship recommendation in social media. Although face relationship verification has certain similarities with face recognition, it faces more complex challenges. Relative face images have high appearance variability due to differences in age, gender, and genetics. Moreover, in an unconstrained environment, the distance between parent-child images will be increased due to posture, expression, lighting, facial occlusion, etc. The real challenge of face relationship verification is how to effectively learn relative face representations to alleviate the similarity differences between relative images caused by other factors. Current methods enhance the verification effect by multi-scale input, data sample enhancement, potential negative sample information utilization, and consideration of family identity information and differences in the distribution of paired relative faces. The common features of parent and child images are extracted independently, ignoring the differences in the similarity of parent and child images between different relatives. For example, some facial image pairs are similar in the eye or mouth area, but the mouths are very different, while some image pairs have a high appearance similarity in the mouth corner area. As a result, the extracted common features have poor robustness against complex face pairs, and the verification accuracy cannot meet the requirements. Based on this, the present invention proposes a facial relationship verification method based on cross-image information interaction to automatically and efficiently realize facial relationship verification in different complex face situations. Summary of the invention

[0003] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a face relationship verification method based on cross-image information interaction. The design goal of the present invention is to dynamically guide instance-level feature extraction step by step by adding an information interaction module after the Swin-Transformer block, and to provide a more comprehensive color representation view by combining three color spaces as input to improve the verification accuracy.

[0004] To achieve the above objectives, the present invention is implemented through the following technical solutions: Read RGB images of different faces, transform the RGB images to HSV and Lab modes, and obtain three color space representations of the image pairs; Obtain RGB-HSV-Lab picture vectors based on RGB images, normalize them to 0-1 respectively, and concatenate them in the channel dimension to obtain RGB-HSV-Lab picture vectors; The RGB-HSV-Lab image vector obtained based on the RGB image is subjected to feature extraction and fusion, and the RGB-HSV-Lab image vector is input into the Swin-Transformer network integrated with the information interaction module. By performing information interaction in four stages, instance-level feature extraction is guided step by step. Among them, Swin-Transformer consists of a four-stage structure consisting of Blocks. Blocks consist of a layer normalization, window multi-head self-attention, residual connection and re-normalization, shifted window multi-head self-attention, residual connection and re-normalization, feedforward network, residual connection and final normalization. By adding an information interaction layer after Block, an information flow is constructed between different facial features, realizing cross-image information interaction; The features extracted from the RGB-HSV-Lab image vector obtained based on the RGB image are optimized by the contrast loss function to improve the discrimination. The extracted feature pairs are fused through three secondary mathematical operations and fusion modules to fuse the facial features. The supervised contrast loss function is used to constrain the model to maximize the similarity between positive samples and the similarity between negative sample pairs, and minimize the similarity between positive sample pairs and negative sample pairs. The functions are shown in equations (1) and (2):

[0005]

[0006] in, represents the contrast loss function of batch samples; represents the fused features of the face image feature pair; n represents the batch size during network training; Representation sample The contrast loss of yes and The cosine similarity between ; , represents an indicator function, if , the result is 1, otherwise it is 0; The value is 1 if the labels are the same, otherwise it is zero; Indicates taking the logarithm; and means that the previous and subsequent conditions must be met at the same time; e represents the exponential basis.

[0007] The fused feature vector of the extracted face image is input into the Softmax classifier, and the image pair is verified to see if they have a relationship based on the output of the classifier. After the Swin-Transformer extracts the features, the output is passed through the fully connected layer of the two nodes, and then the Softmax function is used for subsequent classification and error measurement during the training process, as shown in Formula 3:

[0008] In the formula, Indicates the number of classification categories (here ), is the jth value of the label, is the value of the jth node of the last fully connected layer, c represents the statistical cross entropy loss, and e represents the exponential basis. In the test phase, the category attribute of the input face image pair is determined according to the node value of the last fully connected layer.

[0009] The beneficial effects of the present invention are: the method cleverly utilizes the similar differences between image pairs, realizes cross-image information interaction through feature splicing and self-attention, extracts instance-level features, and has high robustness.

[0010] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0012] Figure 1 The present invention is a flowchart of a face relationship verification method based on cross-image information interaction. DETAILED DESCRIPTION

[0013] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0014] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in this embodiment have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0015] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0016] In the present invention, terms such as "upper", "lower", "left", "right", "front", "back", "vertical", "horizontal", "side", "bottom", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings. They are relational words determined only for the convenience of describing the structural relationships of the various parts or elements of the present invention, and do not specifically refer to any part or element in the present invention and should not be understood as limitations on the present invention.

[0017] In the present invention, terms such as "fixed connection", "connected", "connection", etc. should be understood in a broad sense, indicating that it can be fixedly connected, integrally connected or detachably connected; it can be directly connected or indirectly connected through an intermediate medium. Relevant scientific research or technical personnel in this field can determine the specific meanings of the above terms in the present invention according to specific circumstances, and they should not be understood as limiting the present invention.

[0018] Embodiment 1, as Figure 1 As shown, this embodiment provides a face relationship verification method based on cross-image information interaction, including: a flowchart of the face relationship verification method based on cross-image information interaction is shown in Figure 1 As shown in the figure, PatchPartition refers to dividing the image into small blocks without overlapping; Linear Embedding refers to linearly mapping three-dimensional graphics to high-dimensional vectors; Stage1, Stage2, Stage3 and Stage4 represent the stages of the backbone network respectively, and the main difference lies in the size of the feature map and the number of network layers; FEIIM represents the feature extraction and information interaction module; Feature Fusion represents the feature fusion module; -dims represents the feature dimension; P represents the parent graph feature; C represents the child graph feature; P2-C2 represents the mathematical operation of the square difference of the feature; (PC)2 represents the square operation of the feature difference; P·C represents the element-by-element product operation of the feature vector; Contrastive Loss represents the contrast loss; CrossEntropy Loss represents the cross entropy loss; PreHead represents the prediction module composed of the fully connected layers; no-kin represents that the two faces are not related; kin represents that the two faces are related; Swin Transformer Block represents Swin Submodule of the Transformer network; AttentionOperation represents the attention operation; Split represents dividing the feature vector into sequences.

[0019] This embodiment obtains the RGB-HSV-Lab picture vector of the face image, inputs it into the Swin-Transformer network that integrates the information interaction module, and guides the instance-level feature extraction step by step. By adding information interaction step by step in the process of feature extraction, information is transferred at different feature scales to achieve dynamic guidance between features. Then, the extracted feature pairs are fused through three secondary mathematical operations and fusion modules, and the supervised contrast loss function is used to constrain the model to maximize the similarity between positive samples and the similarity between negative sample pairs, and minimize the similarity between positive sample pairs and negative sample pairs. Finally, the extracted fused feature vector is input into the Softmax classifier, and the image pairs are verified to see whether they have a relationship based on the output of the classifier. The specific steps are as follows: Step 1. Obtain RGB images of different faces and transform them into HSV and Lab modes: First, transform the RGB images of different face pictures into HSV and Lab modes to obtain three different color space representations. Since the RGB color space uses red, green, and blue light of different intensities to depict colors, it is suitable for device representations such as displays and cameras, but is less effective in describing attributes such as brightness and saturation. The HSV color space decomposes color descriptions into hue, saturation, and value, which is more consistent with human perception. However, hue changes near color boundaries may be unsatisfactory. The Lab color space is device-independent and more suitable for images collected from a variety of devices, although it may not be as effective as HSV in representing color attributes. Concatenating these three color spaces in color channels to form 9 color channels can provide a more comprehensive view of color representation, thereby enhancing feature robustness.

[0020] Step 2, obtain the RGB-HSV-Lab picture vector of the face image for Swin-Transformer feature extraction: perform color space transformation on the RGB picture of the face image, normalize the images of each color space to 0-1, and then splice them in the channel dimension to obtain the RGB-HSV-Lab picture vector of the face image, and input the RGB-HSV-Lab picture vector of the face image into the Swin-Transformer network that integrates the information interaction module. The network guides the instance-level feature extraction step by step through four stages of information interaction. Among them, Swin-Transformer consists of a four-stage structure composed of Blocks, which consists of a layer normalization, window multi-head self-attention, residual connection and re-normalization, shift window multi-head self-attention, residual connection and re-normalization, feedforward network, residual connection and final normalization. By adding an information interaction layer after the Block, the information flow between the face image features is constructed to realize cross-image information interaction. Through adaptive weights, the attention mechanism can focus on the similar areas between image pairs, and the weights mixed with the attention of face image information will adaptively focus on the similar places between image pairs.

[0021] Step 3: Fuse the extracted facial image features and optimize the model through the contrast loss function: Fuse the features of the facial image through three secondary mathematical operations and fusion modules, and use the supervised contrast loss function to constrain the model so that it maximizes the similarity between positive samples and the similarity between negative sample pairs, and minimizes the similarity between positive sample pairs and negative sample pairs. The related functions are shown in equations (4) and (5):

[0022]

[0023] in, represents the contrast loss function of batch samples; represents the fused features of the face image feature pair; n represents the batch size during network training; Representation sample The contrast loss of yes and The cosine similarity between ; , represents an indicator function, if , the result is 1, otherwise it is 0; The value is 1 if the labels are the same, otherwise it is zero; Indicates taking the logarithm; and means that the previous and subsequent conditions must be met at the same time; e represents the exponential basis.

[0024] Step 4: Use the Softmax function to determine whether images of different faces have a relationship: For the relationship verification task, it is essentially a binary classification problem that uses the information in the face image pair to determine whether two faces have a relationship. After obtaining the fusion features of the face image, the output is passed through the fully connected layer of the two nodes, and then the Softmax function is applied for subsequent classification and error measurement during the training process, as shown in formula (5):

[0025] In the formula, Indicates the number of classification categories (here ), is the jth value of the label, is the value of the jth node of the last fully connected layer, c represents the statistical cross entropy loss, and e represents the exponential basis. In the test phase, the category attribute of the input face image pair is determined according to the node value of the last fully connected layer.

[0026] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A face relationship verification method based on cross-image information interaction, characterized in that: Obtain RGB images of different face images; obtain RGB-HSV-Lab image vectors of face images by performing color space transformation on the RGB images of face images; extract features of RGB-HSV-Lab image vectors; fuse features of RGB-HSV-Lab image vectors; use supervised contrast loss function to constrain the model to improve the discrimination of features after face image fusion; use Softmax function to determine whether the images of face images have a relationship.

2. The face relationship verification method based on cross-image information interaction according to claim 1 is characterized in that: The obtaining of the RGB-HSV-Lab picture vector includes: transforming the RGB image into HSV and Lab modes, obtaining three color space representations of the picture pair, normalizing them to 0-1 respectively, and then splicing them in the channel dimension to obtain the RGB-HSV-Lab picture vector.

3. The face relationship verification method based on cross-image information interaction according to claim 1, characterized in that: The feature extraction of the RGB-HSV-Lab image vector is achieved by adding an information interaction module after the Swin-Transformer block to dynamically guide the instance-level feature extraction step by step. After the Swin-Transformer Block extracts the features of the face image, it splices them and introduces a self-attention operation. The weights of the attention mixed with the face image information will adaptively focus on the similarities between the image pairs. By adding the information interaction module after the Block of each stage of the Swin-transformer, cross-image information interaction is performed at different feature scales, thereby achieving dynamic guidance when the features are proposed.

4. The face relationship verification method based on cross-image information interaction according to claim 1, characterized in that: The contrast loss function and related functions are shown in equations (1) and (2): in, represents the contrast loss function of batch samples; represents the fused features of the face image feature pair; n represents the batch size during network training; Representation sample The contrast loss of yes and The cosine similarity between ; , represents an indicator function, if , the result is 1, otherwise it is 0; The value is 1 if the labels are the same, otherwise it is zero; Indicates taking the logarithm; and means that the previous and subsequent conditions must be met at the same time; e represents the exponential basis.

5. The face relationship verification method based on cross-image information interaction according to claim 1, wherein The feature is that the Softmax function is as shown in formula (5): In the formula, Indicates the number of classification categories (here ), is the jth value of the label, is the value of the jth node of the last fully connected layer, c represents the statistical cross entropy loss, and e represents the exponential basis. After obtaining the fusion features of the face image, the output is passed through the fully connected layer of two nodes, and then the Softmax function is used for subsequent classification and measurement of errors in the training process. The category attributes of the input face image pair are judged according to the node values ​​of the last fully connected layer.