Face detection related method and device
By introducing the text description task into the face liveness detection model, sharing the feature extraction network and freezing its parameters, the problem of lack of interpretability of the model is solved, and the interpretability and accuracy of the classification results are maintained.
Patent Information
- Application Number
- CN202510206613.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-09-23
AI Technical Summary
Existing face liveness detection models lack interpretability in identity verification scenarios, and the output only includes the classification results without a description of the classification reasons.
The text description task is introduced into the model training so that the classification task and the text description task share the same feature extraction network. The feature extraction network is frozen after the classification task is optimized to avoid the text description task affecting the classification accuracy.
The interpretability of the model is improved, allowing the model to output true and false categories while providing corresponding classification reason descriptions, keeping the classification accuracy unaffected.
Smart Images

Figure CN120689940A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of face detection technology, and in particular to a face detection related method and device. Background Art
[0002] Facial liveness detection is a method for determining a subject's true physiological characteristics in identity verification scenarios. It's primarily used to verify that the user is truly alive. Currently, anti-counterfeiting algorithms like facial liveness detection have long been treated as independent classification tasks, and the classification output of the algorithm models lacks a clear explanation of the underlying causes. Summary of the Invention
[0003] The present application provides a method and device related to face detection. Based on the classification task of face detection, a text description task is introduced, so that the model has a corresponding classification reason description, thereby improving the interpretability of the model.
[0004] In a first aspect, the present application provides a model training method, wherein the model includes a first feature extraction network, a first classification network, and a first text description network, and the method includes:
[0005] Training the first feature extraction network and the first classification network based on the first training data to obtain a second feature extraction network and a second classification network;
[0006] Training the second feature extraction network and the first text description network based on the second training data to obtain a second text description network;
[0007] Among them, the first training data includes facial images and true and false categories of facial images, the second training data includes facial images and true and false category descriptions, the second feature extraction network is used to extract image features of the input facial images, the second classification network is used to determine the true and false categories of facial images based on the image features, and the second text description network is used to determine the true and false category descriptions of facial images based on the image features.
[0008] It can be seen that in this application, the aforementioned model introduces a text description task on the basis of the classification task, so that the model can obtain a text description of the reasons for true and false classification based on the text description task, thereby improving the interpretability of the model. At the same time, the classification task and the text description task share the same feature extraction network. At this time, the classification task is performed first, so that the feature extraction network performs parameter optimization and adjustment in the training corresponding to the classification task. When the text description task is subsequently performed, the feature extraction network does not perform parameter optimization and adjustment. Based on the fact that the model is dominated by the classification task, the accuracy of the feature extraction network when performing the classification task is guaranteed, and the parameter adjustment and optimization of the feature extraction network during the training corresponding to the text description task is avoided, which will affect the accuracy of the classification task.
[0009] In a second aspect, the present application provides a face detection method, the method comprising:
[0010] Obtain target facial image;
[0011] Input the target facial image into the model and obtain the detection results, which include true or false categories and text descriptions;
[0012] The model is obtained by training using the model training method described in the first aspect.
[0013] For the same reasons mentioned above, since the model training includes a text description task in addition to the classification task, and the classification and text description tasks share a feature extraction network, during the model training process, the model training for the classification task is performed first, and the parameters of the feature extraction network are adjusted and optimized at this time. However, the parameters of the feature extraction network are not adjusted and optimized during the subsequent model training for the text description task. This allows the detection of target facial images based on the aforementioned model to maintain the accuracy of the output true and false categories while also outputting the corresponding explanations of the true and false categories, making the model output results interpretable.
[0014] In a third aspect, the present application provides a model training device, wherein the model includes a first feature extraction network, a first classification network, and a first text description network. The device includes:
[0015] A first processing unit is configured to train the first feature extraction network and the first classification network based on the first training data to obtain a second feature extraction network and a second classification network;
[0016] A second processing unit is configured to train the second feature extraction network and the first text description network based on the second training data to obtain a second text description network;
[0017] Among them, the first training data includes facial images and categories of facial images, the second training data includes facial images and text data, the categories of facial images include real facial image categories and false facial image categories, the text data includes true and false category descriptions of facial images, the second feature extraction network is used to extract image features of the input facial image, the second classification network is used to determine the true and false category of the image based on the image features, and the second text description network is used to determine the true and false category description of the image based on the image features.
[0018] The model training device also includes:
[0019] an acquisition unit, configured to acquire a target facial image;
[0020] The third processing unit is used to input the target facial image into the model to obtain a detection result, where the detection result includes a true or false category and a true or false category description.
[0021] In a fourth aspect, the present application provides an electronic device comprising a processor, a memory, and a communication interface. The processor, memory, and communication interface are interconnected and perform communication with each other. The memory stores executable program code, the communication interface is used for wireless communication, and the processor is used to retrieve the executable program code stored in the memory to execute, for example, some or all of the steps described in any method of the first aspect and / or the second aspect.
[0022] In a fifth aspect, the present application provides a computer-readable storage medium, which stores electronic data. When the electronic data is executed by a processor, it is used to execute the electronic data to implement some or all of the steps described in the first aspect and / or second aspect of the present application.
[0023] In a sixth aspect, the present application provides a computer program product, including a computer program, which is operable to cause a computer to execute some or all of the steps described in the first and / or second aspects of the present application. The computer program product may be a software installation package. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] Figure 1 A schematic diagram of the structure of a face detection system provided in an embodiment of the present application;
[0026] Figure 2 A flowchart of a model training method provided in an embodiment of the present application;
[0027] Figure 3 A flowchart of another model training method provided in an embodiment of the present application;
[0028] Figure 4 A schematic diagram of the structure of a classification model provided in an embodiment of the present application;
[0029] Figure 5 A structural example diagram of a model provided in an embodiment of the present application;
[0030] Figure 6A schematic structural diagram of an adapter provided in an embodiment of the present application;
[0031] Figure 7 A flowchart of a face detection method provided in an embodiment of the present application;
[0032] Figure 8 An application flow chart of a face detection method provided in an embodiment of the present application;
[0033] Figure 9 A block diagram of the functional units of a model training device provided in an embodiment of the present application;
[0034] Figure 10 A block diagram of the functional units of another detection device provided in an embodiment of the present application;
[0035] Figure 11 This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0037] The terms "first," "second," and so on, in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps is not limited to the listed steps but may optionally include steps not listed, or may optionally include other steps inherent to the process, method, product, or apparatus.
[0038] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0039] Currently, liveness detection, a key component of facial detection technology, is a crucial component in identity verification scenarios. It verifies that the user is truly alive. Conventional methods treat liveness detection as a separate classification task, training a model. The trained model then detects the user's face image to determine if the user is alive. However, this results in the model output only including the classification result, without any explanation of the reasons for the classification. This makes the model itself, acting as a black box, lacking interpretability.
[0040] Based on this, the present application provides a method and device related to face detection. For the model training part, in addition to the classification task, a text description task is also introduced, which can enable the model to have a corresponding classification reason description, thereby improving the interpretability of the model. At the same time, since the text description task and the classification task share an image feature extraction network for image feature extraction, in order to avoid the text description task affecting the parameters of the image feature extraction network and thus affecting the overall classification task, the classification task is trained first. After the classification task training is completed, the image feature extraction network is frozen, and the text description task training is performed, thereby avoiding affecting the classification accuracy of the model.
[0041] The following is a detailed description of the embodiments of the present application.
[0042] First, through Figure 1 The system architecture of this application is described in detail.
[0043] See also Figure 1 , Figure 1 A structural diagram of a face detection system provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the face detection system 100 includes an image acquisition device 101 and a detection device 102 .
[0044] The image acquisition device 101 is used to acquire a user's facial image. Optionally, the image acquisition device 101 can be a camera, a scanner, or other devices that can acquire facial images.
[0045] Detection device 102 is used to detect facial images captured by image acquisition device 101, and obtain a true or false classification of the facial image and a true or false classification description. Optionally, image acquisition device 101 can be integrated into detection device 102, which can be a device with computing capabilities, such as a computer. In this application, detection device 102 can also be used for model training and to call the trained model for facial detection.
[0046] Specifically, the detection device 102 performs a first model training on the first feature extraction network and the first classification network to obtain a second feature extraction network and a second classification network, and the second feature extraction network and the second classification network are combined to obtain a classification model; then a second model training is performed on the first text description network to obtain a second text description network, and the second text description network and the second feature extraction network are combined to obtain a text description model, and the text description model and the classification model constitute a model.
[0047] In the application process of the model, the image acquisition device 101 acquires the target facial image of the user and sends it to the detection device 102. The detection device 102 inputs the target facial image into the model to obtain the true or false category corresponding to the target facial image and the true or false category description.
[0048] In this application, the model introduces a text description task alongside the classification task, allowing the model to obtain a textual description of the reasons for true or false classification based on the text description task, thereby improving the model's interpretability. Furthermore, the classification and text description tasks share the same feature extraction network. In this case, the classification task is performed first, allowing the feature extraction network to undergo parameter optimization during training for the classification task. When the text description task is subsequently performed, the feature extraction network does not undergo parameter optimization.
[0049] Since the model is dominated by classification tasks, the accuracy of the feature extraction network is guaranteed when performing classification tasks, and the parameter adjustment and optimization of the feature extraction network during the training corresponding to the text description task is avoided, which will affect the accuracy of the classification task.
[0050] Based on this, an embodiment of the present application provides a model training method, which includes a first feature extraction network, a first classification network, and a first text description network. The embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0051] Example 1: Please refer to Figure 2 , Figure 2 A flow chart of a model training method provided in an embodiment of the present application is provided, and the method is applied to the above-mentioned detection equipment, such as Figure 2 As shown, the method includes the following steps:
[0052] In step S201 , the detection device trains a first feature extraction network and a first classification network based on first training data to obtain a second feature extraction network and a second classification network.
[0053] The models to be trained include a first feature extraction network, a first classification network, and a first text description network. The first training data includes facial images and true or false categories of facial images. The second feature extraction network is used to extract image features of the input facial images, and the second classification network is used to determine the true or false category of the facial images based on the image features. The first classification network and the first feature extraction network can be designed based on a convolutional neural network (CNN) architecture or a Transformer architecture.
[0054] At the same time, the second feature extraction network is designed to extract key features from the input data (such as images). These features are usually used for subsequent classification or regression tasks. In this embodiment, the second feature extraction network can select a basic network based on large-scale image-text pre-training, such as the image encoding branch of CLIP, so that the optimal network initialization weights can be obtained, which is beneficial to downstream classification tasks.
[0055] It is understandable that during the training of the first feature extraction network and the first classification network based on the first training data, the weights and parameters of the first feature extraction network and the first classification network need to be adjusted so that the model can learn to extract useful features from the input and perform accurate classification. Furthermore, a large amount of training data should be used during model training. In this embodiment, the first training data is only used to indicate the type of data and does not limit the amount of training data.
[0056] In step S202 , the detection device trains the second feature extraction network and the first text description network based on the second training data to obtain a second text description network.
[0057] The second training data includes facial images and true or false category descriptions, and the second text description network is used to determine the true or false category descriptions of the facial images based on image features.
[0058] It can be understood that when the second feature extraction network and the first text description network are trained based on the second training data, the facial image included in the second training data is input into the second feature extraction network, and after the second feature extraction network extracts the image features of the facial image, the image features are input into the first text description network, and the first text description network outputs the true and false category description.
[0059] At the same time, when the second feature extraction network and the first text description network are trained based on the second training data, the parameters in the second feature extraction network will not be adjusted, and only the parameters in the first text description network will be adjusted.
[0060] After training, the model uses the input port of the second feature extraction network as the model input port, the features output by the second feature extraction network as the input of the second classification network and the second text description network respectively, and the output ports of the second classification network and the second text description network as the output ports of the model.
[0061] As you can understand, based on the aforementioned training tasks, a model can be built that simultaneously performs feature extraction, classification, and text description. Through a two-stage training process, the model gradually optimizes its feature extraction and classification capabilities, and provides explanations for the classification results through a text description network. This multi-task learning approach helps improve the robustness and interpretability of face liveness detection systems.
[0062] Example 2: Since the fake facial image category includes multiple subcategories, the model training process is described in detail below:
[0063] See also Figure 3 , Figure 3 A flow chart of another model training method provided in an embodiment of the present application, which is also applied to the above-mentioned detection equipment, such as Figure 3 As shown, the method includes the following steps:
[0064] In step S301, the detection device trains a first feature extraction network and a first classification network based on first training data and an asymmetric triplet loss function to obtain a second feature extraction network and a second classification network.
[0065] Among them, the false facial image category includes multiple subcategories, the facial images corresponding to the real facial image category are positive samples in the asymmetric triplet loss function, and the facial images corresponding to the false facial image category are negative samples in the asymmetric triplet loss function. The facial images corresponding to the multiple subcategories have different preset distances from the anchor points in the asymmetric triplet loss function.
[0066] During the training of the first feature extraction network and the first classification network, after the facial image is input into the first feature extraction network, the first feature extraction network extracts image features from the facial image and inputs the image features into the first classification network, and the first classification network performs true and false classification based on the image features.
[0067] The output of the first classification network is determined based on the categories corresponding to the facial images in the training data, i.e., the categories corresponding to the facial images in the training data are used as training targets. When the categories corresponding to the facial images in the training data are used as training targets, the parameters of the first classification network and the first feature extraction network can be optimized based on a cross-entropy loss function, and the classification accuracy can be continuously learned to achieve a preset target.
[0068] The real face image category corresponds to real living subjects, while the fake face image category corresponds to fake living subjects. The fake face image category also includes multiple subcategories, each of which can be used to represent the false manifestation of the fake face. For example, the subcategories can include presentation attack face images, deep fake images, and adversarial attack images. Presentation attack face images can represent reproductions of prosthetic faces such as electronic screens, printed photos, masks, head covers, and head covers; deep fake images can represent videos or images generated using deep synthesis algorithms; and adversarial attack images can represent physical adversarial examples or digital adversarial examples generated by adversarial attack algorithms.
[0069] At the same time, in addition to model optimization based on the cross entropy loss function, model optimization can also be performed based on the asymmetric triplet loss. The triplet loss function is a loss function used to learn deep neural network embeddings. Its goal is to make samples from the same category closer to each other and samples from different categories farther away from each other in the embedding space.
[0070] The triplet loss function requires three samples to calculate the loss: an anchor, a positive sample, and a negative sample. An anchor sample is the sample of interest, or a reference sample for comparison. Positive samples share the same class label as the anchor sample, meaning they belong to the target class. Negative samples have different class labels from the anchor sample, meaning they do not belong to the target class. This is achieved by minimizing the distance between an anchor point and the positive sample, while maximizing the distance between the anchor point and the negative sample.
[0071] Since the fake facial image category in this application includes multiple subcategories, and the common triplet loss function only divides positive samples and negative samples, this embodiment uses an asymmetric triplet loss function (Asymmetric TripletLoss). Its asymmetry is reflected in the use of different preset distances (margins) to constrain the positive samples more closely and the negative samples can be more relaxed.
[0072] Specifically, the preset distance between positive samples is set relatively small, the preset distance between negative samples of the same category is set relatively large, and the preset distance between negative samples of different categories can be set to the maximum, to adapt to the data distribution differences of different samples and avoid jitter during the training process. This not only keeps positive and negative samples away from each other, but also keeps negative samples of different categories relatively far apart.
[0073] The asymmetric triple formula is:
[0074]
[0075] Among them, α is used to represent the preset distance, that is, margin. The training data can be divided into multiple small parts. They are used to represent the feature maps of the anchor point, positive sample, and negative sample of the i-th part respectively.
[0076] In this application, the cross entropy loss function and the asymmetric triplet loss function can be used in combination to simultaneously optimize the classification accuracy and the similarity measure between samples. In general, in addition to directly constructing the loss function of the entire model based on the two loss functions, you can also use the cross entropy loss for preliminary classification training to ensure that the model can correctly distinguish samples of different categories. Then, after the classification accuracy reaches a certain level, the asymmetric triplet loss function is introduced for further training to optimize the distribution of samples in the embedding space. By combining these two loss functions, the model can learn a more refined similarity measure between samples while maintaining a high classification accuracy.
[0077] In this application, an asymmetric triplet loss is used for model optimization, with different preset distances used to constrain the distances between positive samples to be closer and the distance between negative samples to be looser. This allows the feature extraction network and the classification network to learn more discriminative feature representations between real facial images and various types of fake facial images.
[0078] For example, see Figure 4 , Figure 4 A structural diagram of a classification model provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, after the facial image is input into the second feature extraction network (image encoder), image features are obtained, and the image features are input into the second classification network (classifier) to obtain corresponding true and false categories. The true and false categories include real facial images and false facial images, and false facial images can also include presentation attack images, deep fake images, and adversarial sample attack images.
[0079] Step S302: The detection device trains the second feature extraction network and the first text description network based on the facial image in the second training data and the first description data to obtain a third text description network.
[0080] Step S303: The detection device trains the second feature extraction network and the third text description network based on the facial image and the second description data in the second training data to obtain a second text description network.
[0081] The true and false category descriptions in the second training data include first description data and second description data, and the true and false category description features included in the first description data are less than the true and false category description features included in the second description data.
[0082] In order to avoid directly using more detailed true and false category descriptions as training targets to train the first text description network model, which will make it difficult for the first text description network to effectively learn true and false category descriptions. This embodiment divides the current training into two stages. First, the first description data with fewer true and false category description features is used as a training target to train the first text description network. After the training iteration is completed until the loss drops to a preset target, a third text description network is obtained. Subsequently, the second description data with relatively more true and false category description features is used as a training target to train the third text description network. Similarly, after the training iteration is completed until the loss drops to a preset target, a second text description network is obtained.
[0083] The aforementioned true or false category description features can be reflected through data granularity. Data granularity is used to indicate the level of detail of the data. The first description data corresponds to the first data granularity, and the second description data corresponds to the second data granularity. The first data granularity is smaller than the second data granularity.
[0084] In a feasible example, the first data granularity and the second data granularity may also correspond to coarse granularity and fine granularity, respectively. A coarse-grained text description may be a relatively general and holistic description of the image content, which may focus on the main objects, scenes, or activities in the image without involving too much detailed information. This description method is usually used to quickly understand the basic content or theme of the image. A fine-grained text description may be a more detailed and accurate description of the image content, which may not only point out the main objects in the image, but also include detailed information such as the attributes, positional relationships, and actions of these objects.
[0085] In other words, the first descriptive data can be a relatively simple text description generated based on the facial images in the training data and the authenticity classification of the facial images. For example, this is a real face image, this is a presentation attack image, this is a silicone mask, and this is an image generated by a deepfake algorithm. The second descriptive data can be a more complex text description generated based on the facial images in the training data and the authenticity classification of the facial images. For example, this is a real face image in an indoor environment, and this is an image displayed on an electronic screen. The first and second descriptive data can be generated based on other dedicated large language models.
[0086] It is understood that during the aforementioned model training process, the weights and parameters of the first and third text description networks need to be adjusted so that the model can learn to output more detailed description information. Furthermore, a large amount of training data should be used during model training. In this embodiment, the second training data is also used only to indicate the type of data and does not limit the amount of training data.
[0087] In this application, the first description data has fewer features and may represent basic descriptive information, while the second description data has more features and may contain more detailed descriptive information. After the model is trained using the first description data with relatively fewer true and false category description features, the model is then trained using the second description data with relatively more true and false category description features. This can avoid the difficulty of achieving effective convergence when the model directly learns text data with more true and false category description features.
[0088] Furthermore, the third text description network combines the first description data with the output of the second feature extraction network, making the generated text description more comprehensive and accurate. This gradually deepening learning approach helps the model better understand the complexity and diversity of data, thereby achieving better performance in learning text data with a large number of true and false category description features.
[0089] The following is a detailed description of the structure of the second feature extraction network and the second text description network:
[0090] Optionally, in a feasible embodiment, the second feature extraction network includes multiple layers, the second text description network includes an adapter and a large language model, and the second text description network is used to determine the true or false category description of the image based on the image features, including: the adapter is used to determine the text features based on multiple first image features output by multiple layers of the second feature extraction network; the large language model is used to determine the true or false category description based on the text features.
[0091] In this process, both during the training phase of the first text description network and the subsequent inference phase of the second text description network, the input facial image first passes through the second feature extraction network, which then outputs multiple first image features from different layers of the second feature extraction network. For example, a feature extraction network typically consists of multiple convolutional layers, pooling layers, and activation layers, each of which is responsible for extracting features at a different level.
[0092] Then, multiple first image features are input into an adapter, which includes a structure for converting features from different sources into structures suitable for further processing. In this embodiment, the adapter aggregates multiple first image features into text features, and inputs the text features into a large language model (LLM), which outputs a true or false category description corresponding to the facial image based on the text features. Similarly, the second text description network can use a cross-entropy loss function for model optimization when training the second model. For example, a large language model can refer to a complex natural language processing model that can understand and generate natural language text.
[0093] In this application, the second text description network includes an adapter and a large language model. This allows the adapter to convert image features into text features, allowing the large language model to directly output true or false classification descriptions based on the text features. Furthermore, because first image features at different levels represent different levels of abstraction in the image, utilizing these multi-layered first image features can capture a broader and richer range of semantic content in the image, allowing the large language model to generate more accurate and detailed true or false classification descriptions.
[0094] The use of adapters enables efficient conversion of image features at different levels into the features required for text descriptions, improving the efficiency and quality of feature conversion. Large language models can generate detailed true and false category descriptions based on the extracted text features, providing users with a deeper understanding of classification decisions.
[0095] In summary, this example combines deep learning and natural language processing techniques to create a model that can extract features from facial images and generate corresponding true and false classification descriptions. This approach helps improve the performance of face liveness detection systems and enhances user trust in system decisions.
[0096] For example, see Figure 5 , Figure 5 This is a structural example diagram of a model provided in an embodiment of the present application, such as Figure 5 As shown, after the facial image is input into the second feature extraction network, the image features are obtained. At this time, the image features have two input ports, one is input to the second classification network, and the other is input to the adapter.
[0097] After inputting the image features into the second classification network, corresponding true and false categories are obtained. The true and false categories include real facial images and false facial images. The false facial images can also include presentation attack images, deep fake images, and adversarial attack images. The image features for the adapter include the first image features of multiple layers in the second feature extraction network. The adapter derives text features based on these first image features. After inputting the text features into the large language model, the adapter outputs a true and false category description.
[0098] Based on this, the specific structure of the adapter is described in detail below:
[0099] Furthermore, in a feasible embodiment, the adapter includes a gated feature fusion network, and the adapter is used to determine text features based on multiple first image features output by multiple levels of the second feature extraction network, including: the gated feature fusion network performs weight scaling on the multiple first image features according to the weights corresponding to the multiple first image features to obtain multiple second image features; the gated feature fusion network performs feature splicing on the multiple second image features to obtain third image features; and the gated feature fusion network performs feature conversion based on the third image features to obtain text features.
[0100] The gated feature fusion network dynamically adjusts the importance of features through gating units, allowing the model to process input data more flexibly. In this application, the adapter needs to process multiple first image features output by different layers of the second feature extraction network. Through downsampling, normalization, and calculations of gating units, the large language model can more effectively utilize useful information in the image and ignore irrelevant or redundant information.
[0101] Specifically, after multiple first image features are input into the adapter, since the multiple first image features are output from different levels of the second feature extraction network, the scales of the multiple first image features will be different. Therefore, for the multiple first image features, they need to be downsampled at different scales to unify them to the same dimension, and then feature normalization is performed. Then, different feature weight coefficients are obtained according to the gating unit, and finally, they are multiplied with the normalized features to achieve dynamic scaling of the weights of different features to obtain multiple second image features. Through weight scaling and feature splicing, the network can integrate feature information from different levels to generate a richer and more comprehensive image feature representation. After obtaining the multiple second image features, the multiple second image features are spliced in the channel dimension to obtain the third image feature, and then the third image feature is feature converted to obtain text features, so as to adapt to the large language model. The feature conversion step ensures that the generated text features match the requirements of the text description model, thereby improving the performance of the text generation model.
[0102] For example, see Figure 6 , Figure 6 A schematic diagram of the structure of an adapter provided in an embodiment of the present application is shown in FIG. Figure 6As shown in the figure, the adapter is the part in the dotted box. Multiple first image features are input into the multi-channel input end respectively. After down sampling, batch normalization, and then nonlinear transformation through activation function, they are multiplied with the feature weight coefficient obtained by the gate unit. Finally, the features of the multi-channel output are aggregated (concatation) and input into the transformation module (transformer block) for feature conversion, and finally the text features are output.
[0103] In this application, a gated feature fusion network is used to convert multiple first image features into text features. This network effectively weights the contributions of different image features, thereby optimizing the feature fusion process. This allows adaptive determination of which layer of features in the second feature extraction network is most important during model training, enabling the large language model to more effectively utilize useful information in the image, thereby improving the accuracy of the true and false category descriptions output by the large language model.
[0104] The following is a summary of the model training process shown in this application:
[0105] First, conduct the first phase of training.
[0106] First training data is obtained, where the first training data includes facial images and true and false categories of the facial images.
[0107] It is understood that the first training data herein is merely an example of the types of data included. The specific training process should be conducted using a large amount of first training data, or the first training data should include multiple facial images and the true and false classification of each facial image. The true and false classification of facial images includes a real facial image classification and a false facial image classification. The false facial image classification can also include multiple categories. In other words, the number of output categories of the classification model obtained after training corresponds to the categories included in the training data.
[0108] The first feature extraction network and the first classification network are trained based on the first training data, with the true and false classification of the facial images in the first training data serving as the training target. During model training, a loss function corresponding to model training is constructed based on the cross-entropy loss function and the asymmetric triplet loss function, and the parameters of the first feature extraction network and the first classification network are adjusted accordingly. When the classification accuracy reaches the preset target, the second feature extraction network and the second classification network are obtained.
[0109] Then, the second phase of training was carried out.
[0110] Obtain second training data, which includes a facial image and first and second description data corresponding to the facial image. The first description data represents a coarse-grained text description, and the second description data represents a fine-grained text description. Both the first and second description data describe the true and false categories corresponding to the facial image. Similarly, the second training data is merely an enumeration of the data types included; training should be conducted using a large amount of second training data.
[0111] At the same time, the second feature extraction network and the first text description network jointly participate in the current training. The facial image is input into the second feature extraction network, which then outputs multiple first image features at different levels to the adapter in the first text description network. The adapter then determines the text features based on these multiple first image features and outputs them to the large language model, which then outputs the corresponding true or false category descriptions. However, during the entire current training process, the parameters of the second feature extraction network need to be frozen, meaning that no parameter adjustments are made to the second feature extraction network during the entire training process.
[0112] During the current training, training is first performed based on the facial images and the first description data corresponding to the facial images included in the second training data. Specifically, the first description data is used as the training target for training, and a cross-entropy loss function is used to supervise the training of the first text description network. After the training iterations reach a predetermined loss target, a third text description network is obtained.
[0113] Subsequently, training is performed based on the facial images and the second description data corresponding to the facial images included in the second training data. That is, the second description data is used as the training target, and the third text description network is supervised and trained using the cross-entropy loss function. When the training is iterated until the loss reaches the preset target, the second text description network is obtained.
[0114] Finally, the trained model includes a second feature extraction network, a second classification network, and a second text description network. The second feature extraction network serves as a single input, and the second classification network and the second text description network serve as two outputs respectively.
[0115] It is understandable that for facial images in the second training data, the facial images in the first training data can be referenced, and corresponding first description data and second description data can be generated in batches based on the true and false categories corresponding to the facial images in the first training data. Furthermore, when training the first text description network, the first text description network can also incorporate the output of the second classification network, that is, reference the true and false categories output by the second classification network to generate true and false category descriptions. This avoids the output true and false category descriptions being inconsistent with the true and false categories.
[0116] At the same time, in the subsequent model application reasoning process, real-time parameter adjustment updates or batch updates can be performed based on real-time feedback results.
[0117] As can be seen, in this application, the aforementioned model introduces a text description task on top of the classification task, allowing the model to obtain a text description of the reasons for true or false classification based on the text description task, thereby improving the model's interpretability. At the same time, the classification task and the text description task share the same feature extraction network. In this case, the classification task is performed first, allowing the feature extraction network to undergo parameter optimization and adjustment during the training corresponding to the classification task. When the text description task is subsequently performed, the feature extraction network does not undergo parameter optimization and adjustment.
[0118] Since the model is dominated by classification tasks, the accuracy of the feature extraction network is guaranteed when performing classification tasks, and the parameter adjustment and optimization of the feature extraction network during the training corresponding to the text description task is avoided, which will affect the accuracy of the classification task.
[0119] Example 3: After completing the aforementioned model training, the application of the model is described in detail below:
[0120] See also Figure 7 , Figure 7 A flow chart of a face detection method provided in an embodiment of the present application is shown, and the face detection method is applied to the above detection device, such as Figure 7 As shown, the method includes:
[0121] Step S701: The detection device obtains a target facial image.
[0122] The detection device may acquire the target facial image through an image acquisition device, which captures the target facial image of the user and then sends the target facial image to the detection device.
[0123] In step S702, the detection device inputs the target facial image into the model to obtain a detection result.
[0124] The detection results include true and false categories and true and false category descriptions. Based on the aforementioned embodiments, the model shown in this application is mainly used for facial anti-counterfeiting detection. It can be used in combination with other facial recognition technologies and detection technologies. For example, when acquiring a user's facial image, the facial posture, distance, and ambient lighting can be detected through the facial quality verification module, and the user can be prompted to make adjustments immediately; at the same time, the front-end can also use the front-end liveness detection model to intercept simple presentation attacks, such as electronic screens and printed photos. After all the front-end passes, the facial detection shown in this application is performed based on the acquired facial image.
[0125] For example, see Figure 8 , Figure 8 This is an application flow chart of a face detection method provided in an embodiment of the present application, such as Figure 8 As shown, it includes: the user first performs an identity verification process to determine whether it is the identity corresponding to the account; after the verification is completed, the front-end liveness detection is performed, which is mainly a simple presentation attack interception; after the detection is passed, the facial image is extracted; then, facial detection is performed based on the extracted facial image (the facial detection shown in the aforementioned embodiment) to determine whether it is forged, and at the same time output a detailed description of the corresponding classification reason; if it is a forged facial image, it is refused to enter the subsequent process, and if it is not a forged facial image, it enters the subsequent process.
[0126] In a feasible embodiment, obtaining a target facial image includes: obtaining multiple first facial images, where the multiple first facial images are images corresponding to different orientations of the face of the target object; performing facial area recognition on the multiple first facial images respectively, eliminating non-facial areas in the multiple first facial images, and obtaining the target facial image.
[0127] The target subject's face can be captured from different angles, either by asking the user to move their head or through a multi-camera setup. Images from different orientations might include frontal, side, and oblique views. The goal is to capture multiple perspectives of the face to obtain richer geometric structure information, which aids in subsequent liveness detection and recognition.
[0128] Computer vision algorithms are then used to analyze each first facial image to locate and extract facial features (such as eyes, nose, and mouth). Specifically, deep learning models, such as convolutional neural networks (CNNs), can be used. After determining the facial location and boundaries, non-facial areas in the image are removed. This not only improves the efficiency of subsequent facial detection but also reduces the influence of background or other interference factors. The image obtained after removing non-facial areas is the target facial image, which contains only facial-related data.
[0129] As can be seen, in the embodiments of the present application, the aforementioned model introduces a text description task on top of the classification task, allowing the model to obtain a text description of the reasons for true or false classification based on the text description task, thereby improving the model's interpretability. Furthermore, the classification task and the text description task share the same feature extraction network. In this case, the classification task is performed first, allowing the feature extraction network to undergo parameter optimization during training for the classification task. When the text description task is subsequently performed, the feature extraction network does not undergo parameter optimization.
[0130] Since the model is dominated by classification tasks, the accuracy of the feature extraction network is guaranteed when performing classification tasks, and the parameter adjustment and optimization of the feature extraction network during the training corresponding to the text description task is avoided, which will affect the accuracy of the classification task.
[0131] In accordance with the above-mentioned embodiment, please refer to Figure 9 , Figure 9 This is a functional unit block diagram of a model training device provided in an embodiment of the present application. The model training device is the above-mentioned detection device or a part of the detection device, such as Figure 9 As shown, the model training device 90 includes:
[0132] A first processing unit 901 is configured to train the first feature extraction network and the first classification network based on the first training data to obtain a second feature extraction network and a second classification network;
[0133] The second processing unit 902 is configured to train the second feature extraction network and the first text description network based on the second training data to obtain a second text description network.
[0134] Among them, the first training data includes facial images and true and false categories of facial images, the second training data includes facial images and true and false category descriptions, the second feature extraction network is used to extract image features of the input facial images, the second classification network is used to determine the true and false categories of facial images based on the image features, and the second text description network is used to determine the true and false category descriptions of facial images based on the image features.
[0135] In a feasible embodiment, the true and false categories of facial images include a true facial image category and a false facial image category, and the false facial image category includes multiple subcategories. In the aspect of training the first feature extraction network and the first classification network based on the first training data to obtain the second feature extraction network and the second classification network, the first processing unit 901 is specifically configured to:
[0136] Training the first feature extraction network and the first classification network based on the first training data and the asymmetric triplet loss function to obtain a second feature extraction network and a second classification network;
[0137] Among them, the facial images corresponding to the real facial image category are positive samples in the asymmetric triplet loss function, the facial images corresponding to the false facial image category are negative samples in the asymmetric triplet loss function, and the facial images corresponding to the multiple subcategories have different preset distances from the anchor points in the asymmetric triplet loss function.
[0138] In a feasible embodiment, the true and false category descriptions in the second training data include first description data and second description data, the true and false category description features included in the first description data are fewer than the true and false category description features included in the second description data, and in training the second feature extraction network and the first text description network based on the second training data to obtain the second text description network, the second processing unit 902 is specifically configured to:
[0139] Based on the facial images in the second training data and the first description data, the second feature extraction network and the first text description network are trained to obtain a third text description network;
[0140] Based on the facial images and the second description data in the second training data, the second feature extraction network and the third text description network are trained to obtain a second text description network.
[0141] In one feasible embodiment, the second feature extraction network includes multiple layers, the second text description network includes an adapter and a large language model, and the second text description network is used to determine the true or false category description of the image based on the image features, including:
[0142] The adapter is used to determine text features based on multiple first image features output by multiple layers of the second feature extraction network; the large language model is used to determine true and false category descriptions based on the text features.
[0143] In one feasible embodiment, the adapter includes a gated feature fusion network, and the adapter is configured to determine text features based on a plurality of first image features outputted by a plurality of layers of the second feature extraction network, including:
[0144] The gated feature fusion network performs weight scaling on the multiple first image features according to the weights corresponding to the multiple first image features to obtain multiple second image features;
[0145] The gated feature fusion network performs feature concatenation on multiple second image features to obtain a third image feature;
[0146] The gated feature fusion network performs feature conversion based on the third image feature to obtain text features.
[0147] In addition, the model training device also includes:
[0148] An acquisition unit 903 is used to acquire a target facial image;
[0149] The third processing unit 904 is configured to input the target facial image into the model to obtain a detection result, wherein the detection result includes a true or false category and a true or false category description;
[0150] In a feasible embodiment, in terms of acquiring the target facial image, the acquiring unit 903 is specifically configured to:
[0151] Acquire a plurality of first facial images, wherein the plurality of first facial images are images corresponding to different orientations of the face of the target object;
[0152] Facial region recognition is performed on each of the plurality of first facial images, and non-facial regions in the plurality of first facial images are eliminated to obtain a target facial image.
[0153] It can be understood that since the model training method embodiment and the subsequent facial detection method embodiment and the model training device embodiment are different presentation forms of the same technical concept, the content of the model training method embodiment part in this application should be synchronously adapted to the model training device embodiment part and will not be repeated here.
[0154] Meanwhile, the face detection method may also be performed by other independent devices, for example, a face detection device may be used to perform the steps corresponding to the aforementioned face detection method, and the face detection device also includes the aforementioned acquisition unit and third processing unit.
[0155] In the case of integrated units, such as Figure 10 As shown, Figure 10 This is a block diagram of the functional units of another detection device provided in an embodiment of the present application. Figure 10 In the embodiment, the detection device 102 includes: a processing module 1012 and a communication module 1011. The processing module 1012 is used to control and manage the actions of the detection device 102, for example, the steps of the first processing unit 901, the second processing unit 902, the acquisition unit 903 and the third processing unit 904, and / or other processes for executing the technology described herein. The communication module 1011 is used to support the interaction between the detection device 102 and other devices. Figure 10 As shown, the detection device 102 may further include a storage module 1013 , and the storage module 1013 is used to store program codes and data of the detection device 102 .
[0156] The processing module 1012 may be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The communication module 1011 may be a transceiver, an RF circuit, or a communication interface, and the like. The storage module 1013 may be a memory.
[0157] Among them, all relevant contents of each scenario involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here. Figure 2 The model training method shown and Figure 6 The face detection method shown.
[0158] The above embodiments can be implemented in whole or in part through software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0159] Computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. A computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0160] Figure 11 This is a structural block diagram of an electronic device provided in an embodiment of the present application. Figure 11 As shown, the electronic device 1100 may include one or more of the following components: a processor 1101, a memory 1102 and a communication interface 1103. The processor 1101, the memory 1102 and the communication interface 1103 are interconnected and perform communication with each other. The memory 1102 may store one or more computer programs, and the one or more computer programs may be configured to implement the methods described in the above embodiments when executed by one or more processors 1101.
[0161] The processor 1101 may include one or more processing cores. The processor 1101 utilizes various interfaces and circuits to connect various components within the entire electronic device 1100. The processor 1101 executes various functions of the electronic device 1100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1102 and calling data stored in the memory 1102.
[0162] Optionally, the processor 1101 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA).
[0163] The processor 1101 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), a modem, etc. It is understood that the modem may not be integrated into the processor 1101 and may be implemented separately through a communication chip.
[0164] The memory 1102 may include a random access memory (RAM) or a read-only memory (ROM). The memory 1102 may be used to store instructions, programs, codes, code sets, or instruction sets.
[0165] Optionally, the memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created by the electronic device 1100 during use.
[0166] It is understandable that the electronic device 1100 may include more or fewer structural elements than those in the above structural block diagram, for example, a power module, physical buttons, a WiFi (Wireless Fidelity) module, a speaker, a Bluetooth module, a sensor, etc., which are not limited here.
[0167] The electronic device 1100 may be a part of the face detection system 100 or a device independent of the face detection system 100 .
[0168] An embodiment of the present application provides a computer-readable storage medium, wherein program data is stored in the computer-readable storage medium. When the program data is executed by a processor, the program data is used to perform some or all steps of any face detection-related method described in the above method embodiments.
[0169] The present application also provides a computer program product, including a computer program, which is operable to cause a computer to execute some or all of the steps of any face detection method described in the above method embodiments. The computer program product can be a software installation package.
[0170] It should be noted that for the aforementioned method embodiments of any of the face detection-related methods, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0171] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps. The fact that certain measures are recited in different dependent claims does not mean that these measures cannot be combined to produce good results.
[0172] Those skilled in the art will appreciate that all or part of the steps in the various methods of any of the above-mentioned face detection related method embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0173] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of a face detection-related method and device of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, based on the idea of a face detection-related method and device of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
[0174] The present application is described with reference to the flowcharts and / or block diagrams of the methods, hardware products, and computer program products of the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0175] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0177] It can be understood that any product that is controlled or configured to execute the processing method of the flowchart described in the method embodiment of a face detection related method of the present application, such as the terminal and computer program product in the above flowchart, falls within the scope of the related products described in the present application.
[0178] Obviously, those skilled in the art may make various modifications and variations to the face detection method and apparatus provided herein without departing from the spirit and scope of the present application. Thus, if such modifications and variations fall within the scope of the present claims and their equivalents, the present application is intended to encompass such modifications and variations.
Claims
1. A model training method, characterized in that: The model includes a first feature extraction network, a first classification network, and a first text description network, and the method includes: Training the first feature extraction network and the first classification network based on the first training data to obtain a second feature extraction network and a second classification network; Training a second feature extraction network and the first text description network based on second training data to obtain a second text description network; The first training data includes facial images and true or false categories of the facial images, the second training data includes facial images and true or false category descriptions, the second feature extraction network is used to extract image features of the input facial images, the second classification network is used to determine the true or false category of the facial images based on the image features, and the second text description network is used to determine the true or false category description of the facial images based on the image features.
2. The method according to claim 1, characterized in that The true and false categories of the facial images include a true facial image category and a false facial image category, the false facial image category includes multiple subcategories, and the first feature extraction network and the first classification network are trained based on the first training data to obtain a second feature extraction network and a second classification network, including: Training the first feature extraction network and the first classification network based on the first training data and an asymmetric triplet loss function to obtain the second feature extraction network and the second classification network; Among them, the facial image corresponding to the real facial image category is the positive sample in the asymmetric triplet loss function, the facial image corresponding to the false facial image category is the negative sample in the asymmetric triplet loss function, and the facial images corresponding to the multiple subcategories are different from the preset distances of the anchor points in the asymmetric triplet loss function.
3. The method according to claim 1, characterized in that The true and false category descriptions in the second training data include first description data and second description data, the true and false category description features included in the first description data are fewer than the true and false category description features included in the second description data, and the training of the second feature extraction network and the first text description network based on the second training data to obtain the second text description network includes: Training the second feature extraction network and the first text description network based on the facial image in the second training data and the first description data to obtain a third text description network; Based on the facial image and the second description data in the second training data, the second feature extraction network and the third text description network are trained to obtain the second text description network.
4. The method according to claim 1, wherein The second feature extraction network includes multiple layers, the second text description network includes an adapter and a large language model, and the second text description network is used to determine the true or false category description of the image based on the image features, including: The adapter is configured to determine text features based on a plurality of first image features respectively output by a plurality of levels of the second feature extraction network; The large language model is used to determine the true or false category description according to the text features.
5. The method according to claim 4, characterized in that The adapter includes a gated feature fusion network, and the adapter is used to determine text features based on multiple first image features respectively output by multiple levels of the second feature extraction network, including: The gated feature fusion network performs weight scaling on the multiple first image features according to the weights respectively corresponding to the multiple first image features to obtain multiple second image features; The gated feature fusion network performs feature splicing on the plurality of second image features to obtain a third image feature; The gated feature fusion network performs feature conversion based on the third image feature to obtain the text feature.
6. A face detection method, characterized in that: The method comprises: Obtain target facial image; Inputting the target facial image into the model to obtain a detection result, wherein the detection result includes a true or false category and a true or false category description; Wherein, the model is obtained by training through any one of the model training methods according to claims 1 to 5.
7. The method according to claim 6, characterized in that The acquiring of the target facial image comprises: Acquire a plurality of first facial images, wherein the plurality of first facial images are images corresponding to different orientations of the face of the target object; Facial region recognition is performed on each of the plurality of first facial images, and non-facial regions in the plurality of first facial images are eliminated to obtain a target facial image.
8. A model training device, characterized in that: The model includes a first feature extraction network, a first classification network, and a first text description network, and the device includes: A first processing unit is configured to train the first feature extraction network and the first classification network based on the first training data to obtain a second feature extraction network and a second classification network; A second processing unit is configured to train the second feature extraction network and the first text description network based on the second training data to obtain a second text description network; The first training data includes facial images and categories of the facial images, the second training data includes facial images and text data, the categories of the facial images include real facial image categories and false facial image categories, the text data includes true and false category descriptions of the facial images, the second feature extraction network is used to extract image features of the input facial image, the second classification network is used to determine the true and false category of the image based on the image features, and the second text description network is used to determine the true and false category description of the image based on the image features.
9. An electronic device, characterized in that: The device comprises: A processor, a memory, and a communication interface, wherein the processor, the memory, and the communication interface are interconnected and perform communication work among each other; The memory stores executable program code, and the communication interface is used for wireless communication; The processor is used to call the executable program code stored in the memory and execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program for electronic data exchange is stored, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 7.