Model updating method, device, apparatus, storage medium, and program product

By performing feature fusion and loss value determination on the training set of the face recognition model and adjusting the central features of the target cluster, the problem of slow recognition speed of the updated face recognition model was solved, and rapid recognition was achieved.

CN116978087BActive Publication Date: 2026-03-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, re-extracting features from face images in the face database using an updated face recognition model is time-consuming, resulting in slow recognition speed.

Method used

By obtaining the training set of the updated face recognition model, the facial features of the training samples are extracted and fused with the central features of the target cluster. The central features of the target cluster are adjusted, and the compatibility loss value and classification loss value are determined. Then, the face recognition model is trained to obtain the updated target face recognition model.

Benefits of technology

It achieves compatibility between facial features and features in the facial database, reduces the time for feature re-extraction, and improves recognition speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116978087B_ABST
    Figure CN116978087B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a model updating method and device, equipment, a storage medium and a program product. In the embodiments of the present application, the center feature of a target cluster corresponding to the face feature of a training sample and the label of the training sample is fused to obtain a fused feature of the training sample. The center feature of the target cluster is adjusted according to the fused feature and the face feature to obtain an adjusted center feature of the target cluster. The face feature and the adjusted center feature of the target cluster are fused to obtain an adjusted fused feature of the training sample. A compatibility loss value between the fused feature and the adjusted fused feature is determined according to the adjusted fused feature, and a classification loss value of a face recognition model is determined. The face recognition model is trained according to the classification loss value and the compatibility loss value to obtain an updated target face recognition model. The embodiments of the present application can improve the compatibility of the target face recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese invention patent application (application number 202211353988.0; application date November 1, 2022; invention title "Model update method, apparatus, device, storage medium and program product"). Technical Field

[0002] This application relates to the field of image processing technology, specifically to a model update method, apparatus, device, storage medium, and program product. Background Technology

[0003] With the development of science and technology, neural network models are being applied more and more widely. For example, neural network models are being applied to the field of face recognition, which involves extracting facial features from face images using face recognition models, matching the facial features with features in a face database, and obtaining the recognition result of the face image.

[0004] In the process of recognizing faces using a face recognition model, the model is updated based on the face image. However, the facial features extracted by the updated model are incompatible with the features in the face database. This necessitates re-extracting features from the face image in the database using the updated model, which is time-consuming and results in slow recognition speed. Summary of the Invention

[0005] This application provides a model update method, apparatus, device, storage medium, and program product, which can solve the technical problem that it takes a long time and the recognition speed is slow when re-extracting features of face images from the face database through the updated face recognition model.

[0006] This application provides a model update method, including:

[0007] Obtain a training set for updating the face recognition model, wherein the training set includes at least one labeled training sample, which is a face image historically recognized by the face recognition model.

[0008] Feature extraction is performed on the above training samples to obtain the facial features corresponding to the above training samples. The facial features and the central features of the target clusters corresponding to the above labels are then fused to obtain the fused features of the above training samples.

[0009] Based on the above fusion features and the above facial features, the central features of the above target cluster are adjusted to obtain the adjusted central features of the above target cluster.

[0010] The adjusted center features of the above facial features and the above target clusters are fused to obtain the adjusted fused features of the above training samples;

[0011] Based on the adjusted fusion features, the compatibility loss value between the fusion features and the adjusted fusion features is determined, as well as the classification loss value of the face recognition model is determined.

[0012] Based on the above classification loss value and the above compatibility loss value, the above face recognition model is trained to obtain the updated target face recognition model.

[0013] Accordingly, embodiments of this application provide a model update apparatus, including:

[0014] The acquisition module is used to acquire the training set of the updated face recognition model. The training set includes at least one labeled training sample, which is a face image historically recognized by the face recognition model.

[0015] The extraction module is used to extract features from the training samples to obtain the facial features corresponding to the training samples, and to fuse the facial features and the central features of the target clusters corresponding to the labels to obtain the fused features of the training samples.

[0016] The adjustment module is used to adjust the central features of the target cluster based on the above-mentioned fusion features and the above-mentioned facial features, so as to obtain the adjusted central features of the target cluster.

[0017] The fusion module is used to fuse the above facial features and the adjusted center features of the above target clusters to obtain the adjusted fused features of the above training samples.

[0018] The determination module determines the compatibility loss value between the above-mentioned fusion features and the above-mentioned adjusted fusion features, and determines the classification loss value of the above-mentioned face recognition model, based on the above-mentioned adjusted fusion features.

[0019] The training model is performed based on the classification loss value and the compatibility loss value mentioned above to train the face recognition model and obtain the updated target face recognition model.

[0020] Optionally, the adjustment module is specifically used for execution:

[0021] Based on the above fusion features, the predicted labels of the above training samples are determined;

[0022] Based on the facial features and the predicted labels, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0023] Optionally, the extraction module is specifically used to perform:

[0024] The feature extraction layer in the face recognition model described above extracts features from the training samples to obtain the initial facial features corresponding to the training samples.

[0025] The initial facial features are mapped using the feature mapping layer in the aforementioned face recognition model to obtain the facial features corresponding to the training samples.

[0026] Optionally, the adjustment module is specifically used for execution:

[0027] Obtain the first direction information corresponding to the above facial features, and obtain the second direction information of the above initial facial features;

[0028] Based on the first direction information, the second direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0029] Optionally, the adjustment module is specifically used for execution:

[0030] Obtain third-party directional information of the central features of the aforementioned target cluster;

[0031] Based on the first direction information, the second direction information, the third direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0032] Optionally, the fusion module is specifically used to perform:

[0033] The adjusted fusion features and the initial facial features are subjected to a first fusion process to obtain the target fusion features of the training samples.

[0034] Based on the aforementioned target fusion features, determine the compatibility loss value between the aforementioned fusion features and the aforementioned adjusted fusion features.

[0035] Optionally, the model update device also includes:

[0036] The mapping module is used to obtain the initial center features of the target clusters corresponding to the above labels;

[0037] The initial central features are mapped through the central mapping layer of the face recognition model to obtain the central features of the target cluster corresponding to the above labels.

[0038] Optionally, the acquisition module is specifically used for execution:

[0039] Obtain the face images previously recognized by the above face recognition model;

[0040] Select image clusters to be labeled from the image clusters corresponding to the above facial images, and label the above facial images according to the above image clusters to be labeled to obtain labeled facial images.

[0041] Based on the labeled face images above, determine the training set for updating the face recognition model.

[0042] Optionally, the acquisition module is specifically used for execution:

[0043] Obtain historical facial images recognized by the above facial recognition model;

[0044] Extract the key points of the above historical facial images, and determine the quality score corresponding to the above historical facial images based on the key points.

[0045] The historical face images that meet the preset score threshold are used as the face images that the face recognition model has historically recognized.

[0046] Optionally, the acquisition module is specifically used for execution:

[0047] Obtain multiple quality models of the aforementioned historical facial images;

[0048] Based on the aforementioned quality model and key points, multiple initial quality scores for the aforementioned historical facial images are determined.

[0049] Based on the initial quality score, the quality score corresponding to the above historical facial images is determined.

[0050] Optionally, the aforementioned historical facial images have multiple quality evaluation dimensions, each corresponding to a dimensional quality model. Accordingly, the acquisition module is specifically used to perform:

[0051] Based on the aforementioned dimensional quality model and the key points mentioned above, the initial quality scores for each quality evaluation dimension of the aforementioned historical facial images are determined.

[0052] The quality scores corresponding to the historical facial images are obtained by weighting each of the initial quality scores mentioned above.

[0053] Optionally, the model update device also includes:

[0054] The training module is used to perform:

[0055] Obtain the first training set of the dimensional quality model to be trained, wherein the first training set includes multiple first training samples;

[0056] Extract the key points of the first training sample and, based on the key points, determine the pairwise loss value and / or anchor loss value of the quality model of the dimension to be trained.

[0057] Based on the above pairwise loss values ​​and / or anchor loss values, the above-mentioned dimensional quality model is trained to obtain the dimensional quality model.

[0058] Optionally, the training module is specifically used to perform:

[0059] Obtain the anchor labels corresponding to the first training sample mentioned above. The anchor labels represent the level of the quality evaluation dimension mentioned above. The first training set contains at least three types of anchor labels.

[0060] Based on the above key points, the predicted score of the first training sample is determined.

[0061] Based on the above anchor labels, the above occlusion prediction scores, and the score ranges corresponding to the above anchor labels, the anchor loss values ​​of the above-mentioned training dimension quality models are determined.

[0062] Optionally, the multiple quality models of the aforementioned historical facial images include a first comprehensive quality model and a second comprehensive quality model. Accordingly, the acquisition module is specifically used to perform:

[0063] Based on the aforementioned key points and using the first comprehensive quality model, the first initial quality score corresponding to the aforementioned historical facial image is determined.

[0064] Based on the aforementioned key points and using the second comprehensive quality model, the second initial quality score corresponding to the aforementioned historical facial images is determined.

[0065] From the first initial quality score and the second initial quality score, the quality score corresponding to the above historical face image is selected.

[0066] Optionally, the acquisition module is specifically used for execution:

[0067] Get the preset filtering threshold;

[0068] If both the first initial quality score and the second initial quality score are less than or equal to the preset screening threshold, then the first initial quality score will be used as the quality score corresponding to the historical face image.

[0069] If both the first initial quality score and the second initial quality score are greater than the preset screening threshold, then the second initial quality score will be used as the quality score corresponding to the historical face image.

[0070] Optionally, the acquisition module is specifically used for execution:

[0071] Based on the facial features in the image clusters corresponding to the aforementioned facial images, determine the correlation between the image clusters corresponding to the aforementioned facial images;

[0072] Based on the above correlation, image clusters to be labeled are selected from the image clusters corresponding to the above facial images.

[0073] Optionally, the acquisition module is specifically used for execution:

[0074] A data structure is constructed based on the image clusters corresponding to the aforementioned facial images and the aforementioned correlation. The data structure includes multiple nodes and multiple edges. Each node represents an image cluster corresponding to the aforementioned facial image, and the weight of the edge represents the correlation between the image clusters corresponding to the aforementioned facial images.

[0075] Based on the edge weights in the above data structure, the data structure is adjusted in a circular manner to obtain a data structure that does not contain circular connections.

[0076] Node groups are selected from the data structure that does not contain circular connections. These node groups correspond to the image clusters to be labeled.

[0077] Optionally, the acquisition module is specifically used for execution:

[0078] Extract multiple modal information corresponding to the face images in the above image clusters to be labeled;

[0079] Display the above-mentioned image clusters to be labeled and the above-mentioned modal information;

[0080] The system receives annotation information from the user, based on the aforementioned modal information, to annotate the facial images in the aforementioned image cluster to be annotated, and obtains the annotated facial images.

[0081] Furthermore, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the model update method provided in embodiments of this application.

[0082] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute any of the model update methods provided in embodiments of this application.

[0083] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the model update methods provided in this application.

[0084] In this embodiment, a training set for updating the face recognition model is obtained. The training set includes at least one labeled training sample, which is a face image historically recognized by the face recognition model. Feature extraction is performed on the training sample to obtain the corresponding face features. The face features and the center features of the target cluster corresponding to the label are fused to obtain the fused features of the training sample. This allows the center features of the target cluster to be adjusted based on the fused features and the face features, resulting in the adjusted center features of the target cluster. This allows the fusion of the face features and the adjusted center features of the target cluster to obtain the training sample. After adjusting the fused features, the compatibility loss value between the fused features and the adjusted fused features can be determined, as well as the classification loss value of the face recognition model. Based on the classification loss value and the compatibility loss value, the face recognition model is trained to obtain the updated target face recognition model. This ensures that the face features obtained through the target face recognition model are compatible with those obtained through the face recognition model. Consequently, it eliminates the need to re-extract features from the face images in the face database using the target face recognition model, reducing the time spent on feature re-extraction and improving recognition speed. Attached Figure Description

[0085] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0086] Figure 1 This is a schematic diagram of a scenario illustrating the model update process provided in an embodiment of this application;

[0087] Figure 2 This is a schematic diagram of the data closed loop provided in the embodiments of this application;

[0088] Figure 3 This is a flowchart illustrating the model update method provided in an embodiment of this application;

[0089] Figure 4 This is a schematic diagram of a face image provided in an embodiment of this application;

[0090] Figure 5 This is a schematic diagram of the first training set provided in an embodiment of this application;

[0091] Figure 6 This is a schematic diagram of the anchor tag provided in an embodiment of this application;

[0092] Figure 7 This is a schematic diagram of the dimensional quality model provided in an embodiment of this application;

[0093] Figure 8 This is a schematic diagram illustrating the effects of models trained with various loss values ​​according to embodiments of this application;

[0094] Figure 9 This is a schematic diagram illustrating the quality evaluation effect of the unsupervised model provided in the embodiments of this application;

[0095] Figure 10 This is a schematic diagram illustrating the principle of the integrated quality model for model disturbances provided in the embodiments of this application;

[0096] Figure 11 This is a schematic diagram illustrating the principle of the similarity distribution distance integrated quality model provided in the embodiments of this application;

[0097] Figure 12 This is a schematic diagram illustrating the principle of the adaptive feature modulus comprehensive quality model provided in the embodiments of this application;

[0098] Figure 13 This is a schematic diagram illustrating the effects of the first and second integrated quality models provided in the embodiments of this application;

[0099] Figure 14 This is a schematic diagram of the data structure provided in the embodiments of this application;

[0100] Figure 15 This is a schematic diagram of the active learning method provided in an embodiment of this application;

[0101] Figure 16 This is a schematic diagram of the query set provided in an embodiment of this application;

[0102] Figure 17 This is a schematic diagram of another active learning method provided in an embodiment of this application;

[0103] Figure 18 This is a schematic diagram of the labeled interface provided in the embodiments of this application;

[0104] Figure 19 This is a schematic diagram of the new face recognition model and the face recognition model provided in the embodiments of this application;

[0105] Figure 20 This is a schematic diagram illustrating the compatibility provided by the embodiments of this application;

[0106] Figure 21 This is a flowchart illustrating the training set construction method provided in an embodiment of this application;

[0107] Figure 22 This is a flowchart illustrating another model update method provided in an embodiment of this application;

[0108] Figure 23This is a schematic diagram of the structure of the model update device provided in the embodiments of this application;

[0109] Figure 24 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0110] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0111] This application provides a model update method, apparatus, device, storage medium, and program product. The device can be an electronic device, the storage medium can be a computer-readable storage medium, and the program product can be a computer program product. The model update apparatus can be integrated into an electronic device, which can be a server, a terminal, or other similar device.

[0112] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.

[0113] Furthermore, multiple servers can form a blockchain, with the servers being nodes on the blockchain.

[0114] The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited herein.

[0115] For example, such as Figure 1 As shown, the terminal can acquire a facial image and send it to the server. The server uses a facial recognition model to recognize the facial image, obtains the recognition result, and returns the result to the terminal.

[0116] After obtaining a face image, the server can use it as a training sample to construct and update the training set for the face recognition model. Then, the server extracts features from the training samples using the face recognition model to obtain the corresponding facial features. It then fuses the facial features with the central features of the target clusters corresponding to the labels of the training samples to obtain fused features for the training samples. Based on the fused features and facial features, the central features of the target clusters are adjusted to obtain adjusted central features. Next, the server fuses the facial features with the adjusted central features of the target clusters again to obtain adjusted fused features for the training samples. Finally, based on the adjusted fused features, the server determines the compatibility loss value between the fused features and the adjusted fused features, as well as the classification loss value for the face recognition model. Based on the classification loss value and the compatibility loss value, the server trains the face recognition model to obtain the updated target face recognition model.

[0117] In this application's embodiments, "multiple" refers to two or more. The terms "first" and "second," etc., in this application's embodiments are used for distinguishing descriptions and should not be construed as implying relative importance.

[0118] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0119] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0120] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0121] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0122] The solutions provided in this application relate to technologies such as machine learning and computer vision in artificial intelligence, and are specifically illustrated through the following embodiments. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0123] The following reference Figure 2 The following describes the usage scenarios of the embodiments of this application. In the process of recognizing a face image using a face recognition model, the face recognition model is updated based on the face image to obtain an updated face recognition model. Then, the updated face recognition model is used for face recognition. The face image and the face recognition model promote each other, forming a data-model closed-loop feedback system. This process is called data closure.

[0124] For example, when the face recognition model is a face recognition model, the data loop can be as follows: Figure 2 As shown, keypoint detection is first performed on the recognized face images to obtain facial keypoints. Then, quality filtering is performed based on the facial keypoints to obtain filtered face images. A training set is constructed based on the filtered face images. Facial features are extracted from the training set, and the face recognition model is updated based on the facial features to obtain the updated face recognition model.

[0125] Therefore, the method for updating the face recognition model includes the process of constructing a training set for updating the face recognition model and the process of updating the face recognition model.

[0126] In this embodiment, the description will be from the perspective of the model update device. In order to facilitate the explanation of the model update method of this application, the following will describe the model update device integrated into the server in detail, that is, the server will be used as the execution subject in the detailed description.

[0127] Please see Figure 3 , Figure 3 This is a schematic flowchart of a model update method provided in an embodiment of this application. The model update method may include:

[0128] Please see Figure 3 , Figure 3 This is a schematic flowchart of a model update method provided in an embodiment of this application. The model update method may include:

[0129] S301. Obtain the training set for updating the face recognition model. The training set includes at least one labeled training sample, which is a face image recognized by the face recognition model in the past.

[0130] A facial image refers to an image of the front of a head that contains identity information. The type of facial image can be set according to the actual situation. For example, the facial image can be an animal's facial image or a plant's facial image. The animal's facial image can be a human face image or a pet's facial image, etc. This embodiment does not limit this.

[0131] The server can obtain an updated training set for the face recognition model upon receiving a request. Alternatively, the server can obtain an updated training set for the face recognition model when it detects that the number of face images historically recognized by the face recognition model meets a preset condition. Or, the server can periodically obtain and update the training set for the face recognition model.

[0132] The server can either directly use all historical face images previously recognized by the face recognition model as the face images previously recognized by the face recognition model, or it can select a preset number (less than the total number of historical face images) of historical face images from the historical face images previously recognized by the face recognition model as the face images previously recognized by the face recognition model.

[0133] When the quality of historical facial images is low, it becomes impossible to determine the identity information of those images, leading to errors in the recognition results. For example, ... Figure 4As shown, low-quality historical face images cannot be recognized. Therefore, when the server selects a preset number of historical face images from those previously recognized by the face recognition model, it can select a preset number of high-quality historical face images from those previously recognized by the face recognition model.

[0134] Optionally, a preset number of high-quality historical face images can be manually selected from the historical face images recognized by the face recognition model, and then the high-quality historical face images can be used as face images. Alternatively, the server can automatically select a preset number of historical face images with high quality scores from the historical face images recognized by the face recognition model.

[0135] If the server automatically selects a predetermined number of historical face images with high quality scores from the face recognition model's historical face images as face images, then the training set for updating the face recognition model can include:

[0136] Obtain historical facial images recognized by the facial recognition model;

[0137] Extract key points from historical facial images and determine the corresponding quality scores based on the key points;

[0138] Historical face images that meet the preset score threshold and have the corresponding quality scores are used as the face images recognized by the face recognition model in the past.

[0139] In this embodiment, the quality score corresponding to the historical face image is determined based on the key points of the historical face image. Then, the historical face image corresponding to the quality score that meets the preset score threshold is used as the face image recognized by the face recognition model in the past. This realizes the automatic filtering out of historical face images with poor quality, which not only improves the efficiency of updating the face recognition model based on the training set, but also improves the filtering speed.

[0140] The method for determining the quality score of historical facial images based on key points can be selected according to the actual situation. For example, the sharpness of the historical facial image can be determined by using key points, and then the corresponding quality score can be determined based on the sharpness. Alternatively, the deflection angle of the face in the historical facial image can be determined by using key points, and then the corresponding quality score can be determined based on the deflection angle. Alternatively, a quality model can also be used to determine the quality score of the historical facial image based on key points. This embodiment does not impose any limitations on this method.

[0141] When determining the quality score of historical facial images based on key points using a quality model, the accuracy of the obtained quality score is low if only one quality model is used to determine the quality score of historical facial images based on key points.

[0142] To improve the accuracy of quality scores, in some embodiments, the quality score corresponding to historical facial images is determined based on key points, including:

[0143] Obtain multiple quality models of historical facial images;

[0144] Based on key points, multiple initial quality scores are determined for historical facial images using a quality model.

[0145] Based on the initial quality score, determine the quality score corresponding to the historical face image.

[0146] This can be achieved by weighting multiple initial quality scores to obtain the quality score corresponding to the historical face image. Alternatively, an initial quality score can be selected from multiple initial quality scores as the quality score corresponding to the historical face image; this embodiment does not impose any limitations on this approach.

[0147] In this embodiment, multiple initial quality scores of historical face images are obtained through multiple quality models, and then the quality score of historical face images is determined based on the multiple initial quality scores, thereby improving the quality score of historical face images.

[0148] In other embodiments, the historical facial images have multiple quality evaluation dimensions, each corresponding to a dimensional quality model. That is, the quality model can be a dimensional quality model. Based on the key points, multiple initial quality scores are determined for the historical facial images using these quality models. Based on these initial quality scores, the corresponding quality score for the historical facial images is then determined, including:

[0149] Using a dimensional quality model, initial quality scores for each quality evaluation dimension of historical facial images are determined based on key points.

[0150] The initial quality scores are weighted to obtain the quality scores corresponding to the historical face images.

[0151] The quality assessment dimension refers to the dimension in which the factors affecting the quality of historical facial images are located. The quality assessment dimension can be selected according to the actual situation. For example, the quality assessment dimension can be the sharpness of the historical facial image, the angle of the face in the historical facial image, the degree of occlusion of the face in the historical facial image, the brightness of the historical facial image, or the color of the historical facial image, etc. That is, the dimensional quality model (each dimensional quality model can also be called a multi-expert model) can be a sharpness quality model, an angle quality model, an occlusion quality model, a brightness quality model, or a color quality model, etc. The initial quality score can be a sharpness quality score, an angle quality score, an occlusion quality score, a brightness quality score, and a color quality score, etc. This embodiment does not limit it.

[0152] Weighting can refer to directly adding the initial quality scores together, or it can refer to multiplying the initial quality score by the weight corresponding to the initial quality score to obtain the adjusted initial quality score, and then adding the adjusted initial quality scores together. This embodiment does not limit the specifics.

[0153] In this embodiment, each quality evaluation dimension corresponds to a dimensional quality model. Then, through the dimensional quality model, the initial quality score of the historical face image for each quality evaluation dimension is obtained. Finally, the initial quality scores are weighted to obtain the quality score corresponding to the historical face image, thereby realizing the acquisition of quality scores from different quality evaluation dimensions and improving the accuracy of quality scores.

[0154] When the dimensional quality model corresponding to a quality evaluation dimension does not need to be obtained through training, for example, when the quality evaluation dimension is the deflection angle of the face in a historical face image, the dimensional quality model corresponding to the quality evaluation dimension can be an angle quality model. The process of determining the quality score of the deflection angle of the face in a historical face image based on key points using the angle quality model can be as follows:

[0155] The Perspective-n-Points (PNP) method calculates the rotation matrix based on key points and reference coordinates, and then determines the deflection angle of the face in historical face images based on the rotation matrix, without the need to train an angle quality model.

[0156] When the dimensional quality model corresponding to a quality evaluation dimension is obtained through training, for example, when the quality evaluation dimension is the sharpness or occlusion of a historical face image, the dimensional quality model corresponding to the quality evaluation dimension can be a sharpness quality model or an occlusion quality model. The sharpness quality model or the occlusion quality model can be obtained through training, which can be done through supervised training or unsupervised training.

[0157] If the dimensionality quality model is obtained through supervised training, then before determining the quality score of historical facial images based on key points using the dimensionality quality model, the following steps are also included:

[0158] Obtain the first training set of the dimensional quality model to be trained. The first training set includes multiple first training samples.

[0159] Extract the key points of the first training sample, and determine the pairwise loss value and / or anchor loss value of the quality model of the dimension to be trained based on the key points of the sample.

[0160] The dimensionality quality model is trained based on the pairwise loss value and / or anchor loss value to obtain the dimensionality quality model.

[0161] One approach is to use metric learning to determine the pairwise loss value of the dimensional quality model to be trained based on the key points of the samples. The metric learning method can be selected according to the actual situation. For example, the classic contrastive loss method, InfoNCE loss, Margin Ranking Loss (if the pairwise loss value is determined based on the key points of the samples using Margin Ranking Loss, then the training of the dimensional quality model to be trained is transformed into a ranking problem) or N-pair Loss method can be selected as the metric learning method. This embodiment does not limit the method.

[0162] When determining the pairwise loss value of the quality model for the dimension to be trained based on sample key points using the N-pair Loss method, the process of determining the pairwise loss value of the quality model for the dimension to be trained based on sample key points can be as follows:

[0163] From the first sample cluster corresponding to the first training sample, positive training samples corresponding to the first training sample are selected, and from the second sample cluster, multiple negative training samples corresponding to the first training sample are selected. The second sample cluster is the sample cluster other than the first sample cluster in the sample cluster corresponding to the first training set.

[0164] Based on the key points of the first training sample and the key points of the positive training sample, a first distance between the first training sample and the positive training sample is determined, and based on the key points of the first training sample and the key points of the negative training sample, a second distance between the first training sample and the negative training sample is determined.

[0165] Based on the first and second distances, determine the pairwise loss value of the dimensional quality model to be trained.

[0166] Unlike gender, where the presence or absence of a mask can be clearly determined, some quality assessment dimensions (such as clarity and occlusion) reflect a relative relationship. That is, given a pair of images, we judge which image is more occluded or clearer. Therefore, in this example, we convert point-wise labels to pair-wise labels and then use a contrastive loss function for training, so that better images will be assigned higher scores during model training.

[0167] Alternatively, the process of determining the anchor loss value of the dimensional quality model to be trained based on the sample key points can be as follows:

[0168] Obtain the anchor label corresponding to the first training sample, and obtain the score range corresponding to the anchor label;

[0169] Based on the key points of the sample, determine the predicted score of the first training sample;

[0170] Based on the anchor label, the predicted score, and the score range corresponding to the anchor label, determine the anchor loss value of the dimensional quality model to be trained.

[0171] Anchor tags represent true or false categories, and can include both. For example, a true category can be 1 and a false category can be 0, meaning the anchor tag can be 0 or 1.

[0172] Anchor labels represent true or false categories, meaning that when training the quality model for a given dimension, the classification result of the model for the first training sample is either yes or no. However, some quality assessment dimensions have multiple levels. For example, regarding the degree of occlusion and clarity of a face, the degree of occlusion can be varied (strong occlusion, slight occlusion, mild occlusion, no occlusion, etc.), and the degree of clarity can also vary. Therefore, if the anchor labels only include two types when training the quality model for a given dimension, the accuracy of the score obtained by the quality model will be low.

[0173] To further improve the accuracy of the dimensionality quality model, in some embodiments, the anchor loss value of the dimensionality quality model to be trained is determined based on key points, including:

[0174] Obtain the anchor label corresponding to the first training sample. The anchor label represents the level of the quality evaluation dimension. There are at least three anchor labels in the first training set.

[0175] Based on the key points, determine the predicted score of the first training sample;

[0176] Based on the anchor label, occlusion prediction score, and the score range corresponding to the anchor label, determine the anchor loss value of the dimensional quality model to be trained.

[0177] The level of a quality assessment dimension indicates its severity. A lower level indicates a more severe quality assessment. For example, if the quality assessment dimension is occlusion, a lower level indicates more severe occlusion. Similarly, if the quality assessment dimension is sharpness, a lower level indicates less sharpness.

[0178] Optionally, the level of a quality evaluation dimension can be quantified based on the number of visible key points. For example, the higher the number of visible key points, the higher the level of the quality evaluation dimension.

[0179] The first training set may include first training samples at different levels of the quality evaluation dimension, and then the quality model of the dimension to be trained is trained to perform at least three classifications for each first training sample.

[0180] For example, when the quality model to be trained is the occlusion quality model to be trained and the training of the quality model to be trained performs three classifications on each first training sample, the first training set may include strongly occluded first training samples, unoccluded first training samples, and slightly occluded first training samples. In this case, there are three levels of occlusion in the first training set, namely, strong occlusion, slightly occlusion, and no occlusion.

[0181] When the first training set includes strongly occluded first training samples, unoccluded first training samples, and slightly occluded first training samples, the first training set can be as follows: Figure 5 As shown.

[0182] If the first training set has three levels for the quality evaluation dimension, it means that the first training set has at least three anchor labels. When the first training set has three anchor labels and the quality model to be trained is the occlusion quality model to be trained, the three anchor labels can be: Figure 6 As shown in the figure, the horizontal axis represents the score range corresponding to the anchor label. The three anchor labels are 0, 1, and 2, where 0 represents strong occlusion, 1 represents slight occlusion, and 2 represents no occlusion.

[0183] The process of determining the anchor loss value of the quality model for the dimension to be trained based on the anchor label, the predicted score, and the score range corresponding to the anchor label can be as follows:

[0184] Based on the anchor label and predicted score, determine the anchor loss value to be adjusted;

[0185] Determine the correct classification score corresponding to the anchor label based on the score range corresponding to the anchor label;

[0186] The anchor loss value is adjusted based on the difference between the predicted score and the correct classification score corresponding to the anchor label, thus obtaining the anchor loss value of the dimensional quality model to be trained.

[0187] In this embodiment, anchor labels represent the level of the quality evaluation dimension. The first training set contains at least three anchor labels. Then, the quality model of the dimension to be trained is trained based on the first training set, thereby further improving the quality of the face images obtained through the quality model of the dimension, so that the accuracy and efficiency of subsequent annotation of face images can be improved.

[0188] This embodiment can train the quality model of the dimension to be trained based on pairwise loss values. Alternatively, this embodiment can also train the quality model of the dimension to be trained based on anchor loss values. Or, this embodiment can simultaneously train the quality model of the dimension to be trained based on both pairwise loss values ​​and anchor loss values. In this case, when the multiple quality models are respectively an angle quality model, an occlusion quality model, and a sharpness quality model, the multiple quality models in this embodiment can be as follows: Figure 7 As shown. The labels are positive samples and positive samples, positive samples and negative samples.

[0189] When training a dimensional quality model based on pairwise loss values, the resulting dimensional quality model performs better in quality evaluation because the pairwise loss values ​​include the distance between positive sample pairs and the distance between positive and negative samples. This is because training the dimensional quality model based on pairwise loss values ​​results in better performance than training it based on binary classification loss values.

[0190] When training a dimensional quality model based on anchor loss values, unlike binary classification loss values ​​which constrain the categories, anchor loss values ​​can determine the difficulty of predicting a sample based on the difference between the true class prediction probability and the false class prediction probability. Then, the loss value is dynamically re-estimated based on the difficulty of predicting the sample, thereby constraining the score interval (which can be the interval where the true class prediction probability is located), alleviating overfitting, and improving the performance of the dimensional quality model in quality evaluation.

[0191] When the dimensional quality model to be trained is trained simultaneously based on pairwise loss and anchor loss, not only can the distance between positive sample pairs and the distance between positive and negative samples be constrained, but also the score interval can be constrained, which can further improve the performance of the dimensional quality model in quality evaluation.

[0192] The following reference Figure 8 The paper explains the performance of dimensionality quality models trained using various loss values. Figure 8 The horizontal axis in the graph represents the filtering ratio of the dimensionality quality model. The filtering ratio can be defined as the ratio between the number of historical face images filtered out by the dimensionality quality model and the number of historical face images recognized by the face recognition model. Figure 8The vertical axis of 801 represents the True Positive Rate (TPR). Figure 8 The vertical axis of 802 represents the correct filtering rate, which is the ratio of the historical face images that were correctly filtered out of the historical face images filtered out by the dimensionality quality model to the historical face images that were filtered out by the dimensionality quality model.

[0193] from Figure 8 It can be seen that training the dimensional quality model based on binary classification loss results in severe overfitting, leading to low recall and poor filtering results. Training it based on anchor loss alleviates overfitting and improves both recall and filtering results. Adding pairwise loss to the anchor loss further smooths the predicted scores, enhancing the model's performance in quality assessment.

[0194] In this embodiment, initial quality scores are obtained through dimensional quality models of each quality evaluation dimension. Then, a weighted average of these initial quality scores is performed to obtain the quality score corresponding to the historical face image. However, the quality score of the historical face image is not a simple linear summation of the initial quality scores. For example, it is difficult to definitively determine which image—a clear profile or a blurry frontal view—is of better quality. Therefore, it is necessary to accurately measure multiple quality evaluation dimensions to obtain a more accurate quality score.

[0195] To obtain a more accurate quality score, in some embodiments, multiple quality models of historical face images include a first comprehensive quality model and a second comprehensive quality model, which determine the quality score corresponding to the historical face image based on key points, including:

[0196] Based on key points, the first initial quality score corresponding to the historical face image is determined using the first comprehensive quality model.

[0197] The second comprehensive quality model is used to determine the second initial quality score corresponding to the historical face image based on the key points.

[0198] The quality scores corresponding to historical face images are selected from the first initial quality score and the second initial quality score.

[0199] In this embodiment, the first comprehensive quality model and the second comprehensive quality model can simultaneously determine the first initial quality score and the second initial quality score of the historical face image based on each quality evaluation dimension of the historical face image, making the first initial quality score and the second initial quality score more accurate.

[0200] Furthermore, the training method for the dimensional quality models of the aforementioned quality evaluation dimensions is supervised training. Supervised training means that the training of the dimensional quality models to be trained is affected by the annotation accuracy and speed of the first training samples. Therefore, in order to obtain quality scores more accurately, in some embodiments, the first comprehensive quality model and the second comprehensive quality model can be an unsupervised first comprehensive quality model and an unsupervised second comprehensive quality model.

[0201] The process involves selecting quality scores corresponding to historical face images from the first and second initial quality scores, including:

[0202] Get the preset filtering threshold;

[0203] If both the first initial quality score and the second initial quality score are less than or equal to the preset screening threshold, then the first initial quality score will be used as the quality score corresponding to the historical face image.

[0204] If both the first initial quality score and the second initial quality score are greater than the preset screening threshold, then the second initial quality score will be used as the quality score corresponding to the historical face image.

[0205] The first initial quality score and the second initial quality score are both less than or equal to the preset screening threshold. This can be understood as the first initial quality score and the second initial quality score being less than the preset screening threshold, or it can be understood as the first initial quality score and the second initial quality score being equal to the preset screening threshold, or it can be understood as the first initial quality score being less than the preset screening threshold and the second initial quality score being equal to the preset screening threshold, or it can be understood as the first initial quality score being equal to the preset screening threshold and the second initial quality score being less than the preset screening threshold.

[0206] In practical applications, it was found that the comprehensive quality model performs differently on facial images of varying qualities. For example, ... Figure 9As shown, the Unsupervised Estimation of Face Image Quality Based on Stochastic Embedding Robustness (SER-FIQ) and the Similarity Distribution Distance for Face Image Quality Assessment perform better on low-quality face images, meaning that the quality assessment scores obtained from these two models are more accurate. Conversely, the Universal Representation for Face Recognition and Quality Assessment performs better on high-quality face images, indicating that the quality assessment scores obtained from this model are more accurate.

[0207] The principle of the model perturbation integrated quality model can be summarized as follows: Figure 10 As shown, the perturbation-comprehensive quality model includes multiple parallel random subnetworks. During training, neurons in the random subnetworks are randomly dropped out to obtain multiple different features corresponding to the second training sample of the perturbation-comprehensive quality model. The variance of the multiple different features of the second training sample is calculated. The larger the variance, the lower the quality of the second training sample; the smaller the variance, the higher the quality of the second training sample.

[0208] The similarity distribution distance comprehensive quality model, also known as the intra-class similarity distribution and inter-class similarity distribution model, can be based on the following principle: Figure 11 As shown, the third sample cluster where the second training sample is located is obtained through the face recognition model. Then, the intra-class distribution of the similarity between the second training sample and the face images in the third sample cluster is determined, as well as the inter-class distribution of the similarity between the second training sample and the face images in the fourth sample cluster (the sample cluster where the second training sample is not located). Then, the Wasserstein distance between the intra-class distribution and the inter-class distribution is determined through the intra-class distribution and the inter-class distribution. The Wasserstein distance is used as a pseudo-quality score to train the similarity distribution distance comprehensive quality model.

[0209] The principle of the adaptive feature magnitude comprehensive quality model can be summarized as follows: Figure 12 As shown ( Figure 11In the diagram, w' and w represent two class centers, B' represents the boundary of class center w', B ​​represents the boundary of class center w, m represents the angular interval between the two classes, and 1, 2, and 3 represent three face images of different qualities. During training, the feature modulus of the face image is determined, and an adaptive feature modulus comprehensive quality model is trained based on the feature modulus. The larger the feature modulus, the higher the quality of the face image; the smaller the feature modulus, the lower the quality of the face image.

[0210] Therefore, in this embodiment, a comprehensive quality model that can obtain more accurate quality scores for low-quality face images is used as the first comprehensive quality model, and a comprehensive quality model that can obtain more accurate quality scores for high-quality face images is used as the second comprehensive quality model. Then, the initial quality scores of historical face images of different qualities are obtained through the first and second comprehensive quality models, thereby improving the accuracy of the quality scores corresponding to historical face images. However, when determining the first initial quality score corresponding to a historical face image through the first comprehensive quality model and the second initial quality score corresponding to a historical face image through the second comprehensive quality model, the server cannot determine the quality level of the historical face image.

[0211] Therefore, in this embodiment, a preset screening threshold is obtained. If both the first initial quality score and the second initial quality score are less than or equal to the preset screening threshold, it indicates that the historical face image is a low-quality face image. In this case, the first initial quality score is used as the quality score corresponding to the historical face image. If both the first initial quality score and the second initial quality score are greater than the preset screening threshold, it indicates that the historical face image is a high-quality face image. In this case, the second initial quality score is used as the quality score corresponding to the historical face image. This allows for the use of different comprehensive quality models to obtain initial quality scores for historical face images of different qualities, thereby improving the accuracy of the quality scores corresponding to the historical face images.

[0212] It should be understood that in this embodiment, when training the first comprehensive quality model and the second comprehensive quality model, face recognition can be performed on historical face images based on the face recognition model to obtain the recognition result. Then, based on the recognition result and the first sample quality score of the first comprehensive quality model to be trained, the first target loss value is determined. Then, the first comprehensive quality model to be trained is trained based on the first target loss value to obtain the first comprehensive quality model.

[0213] The training process of the second comprehensive quality model can refer to the training process of the first comprehensive quality model; this example does not limit it here.

[0214] The following reference Figure 13 The effectiveness of the integrated quality model will be explained. Figure 13 The vertical axis of 1301 represents the recall rate. Figure 13 The vertical axis of 1302 represents the correct filtration rate, from... Figure 13 As can be seen, the first and second comprehensive quality models can better solve the coupling problem between various quality evaluation dimensions, making the recall and correct filtering rates of the first and second comprehensive quality models better than those of the quality evaluation dimension quality models.

[0215] In other embodiments, obtaining the training set for updating the face recognition model includes:

[0216] Obtain facial images previously recognized by the facial recognition model;

[0217] Select image clusters to be labeled from the image clusters corresponding to the face images, and label the face images according to the image clusters to be labeled to obtain labeled face images;

[0218] Based on labeled facial images, determine the training set for updating the facial recognition model.

[0219] After obtaining the face images recognized in the past by the face recognition model, the server clusters the face images to obtain the image clusters corresponding to the face images. Then, it selects the image clusters to be labeled from the image clusters corresponding to the face images so that the face images can be labeled according to the image clusters to be labeled, and obtain labeled face images. Finally, based on the labeled face images, the training set for updating the face recognition model is determined.

[0220] In this embodiment, labeling facial images based on the image clusters to be labeled can be understood as determining whether facial images within the image clusters belong to the same object. Optionally, the facial images of the image clusters to be labeled can be displayed so that the user can judge the facial images within the image clusters, or the server can automatically determine whether the facial images within the image clusters belong to the same object. This embodiment is not limited to this.

[0221] In other embodiments, the process of selecting the image clusters to be labeled from the image clusters corresponding to the face images can be as follows:

[0222] The first image cluster to be labeled is selected from the candidate image clusters. Then, a first number of second image clusters to be labeled are selected from the candidate image clusters and combined with the first image cluster to form an image cluster group to be labeled.

[0223] However, this method yields a large number of image clusters to be labeled, and labeling errors accumulate. To reduce the number of image clusters to be labeled and decrease labeling errors, in some embodiments, image clusters to be labeled are selected from the image clusters corresponding to face images, including:

[0224] Based on the facial features in the image clusters corresponding to the facial images, determine the correlation between the image clusters corresponding to the facial images;

[0225] Based on the correlation, select image clusters to be labeled from the image clusters corresponding to the facial images.

[0226] The server can calculate the cluster similarity between two image clusters corresponding to two face images based on the central facial features of the image clusters corresponding to the face images (or, randomly select a facial feature from the image clusters corresponding to the face images to calculate the cluster similarity between the two image clusters corresponding to the face images). It can also calculate the label (ID) popularity of the image clusters corresponding to the face images based on the number of face images in the image clusters corresponding to the face images (i.e., the number of times the label corresponding to the image clusters corresponding to the face images appears) and the total number of face images. Finally, based on the label popularity and cluster similarity of the image clusters corresponding to the face images, the server determines the correlation between the image clusters corresponding to the face images.

[0227] Optionally, the label popularity, cluster similarity, and correlation of the image clusters corresponding to the face images satisfy the following relationship:

[0228]

[0229] cost(up,uq) This represents the image cluster corresponding to the face image. u p Image clusters corresponding to facial images u q The degree of correlation between them info_u p This represents the image cluster corresponding to the face image. u p The popularity of the tag info_u q This represents the image cluster corresponding to the face image. u q The popularity of the tag.

[0230] In this embodiment, after obtaining the image clusters corresponding to the face images, the correlation between the image clusters corresponding to the face images is determined based on the facial features of the face images in the image clusters corresponding to the face images. Then, based on the correlation, image clusters to be labeled are selected from the image clusters corresponding to the face images. The greater the correlation, the greater the probability that the image clusters corresponding to the two face images are similar. Then, the image clusters corresponding to the two face images are used as image clusters to be labeled, thereby reducing the number of image clusters to be labeled.

[0231] Optionally, based on the correlation, image clusters to be labeled are selected from the image clusters corresponding to the face images. This can be either the minimum sum of the correlation between the selected image clusters, or the second minimum sum of the correlation between the selected image clusters. This embodiment does not limit this.

[0232] In other embodiments, based on the degree of correlation, image clusters to be labeled are selected from the image clusters corresponding to the face images, including:

[0233] A data structure is constructed based on the image clusters corresponding to the face images and their correlation. The data structure includes multiple nodes and multiple edges. Each node represents the image cluster corresponding to the face image, and the weight of the edge represents the correlation between the image clusters corresponding to the face images.

[0234] Based on the weights of the edges in the data structure, the data structure is adjusted in a circular manner to obtain a data structure that does not contain circular connections.

[0235] Node groups are selected from data structures that do not contain circular connections, and the node groups correspond to the image clusters to be labeled.

[0236] For example, the data to be constructed can be like Figure 14 As shown, the data structure at this point contains a ring structure. Figure 14 If edge 1, edge 2, and edge 3 form a loop, then edge 1 can be removed to obtain a data structure that does not contain loop connections. A data structure that does not contain loop connections can also be understood as a tree-like data structure.

[0237] The method of adjusting the data structure in a circular manner based on the weights of the edges in the data structure to obtain a data structure without circular connections can be selected according to the actual situation. For example, Prim's algorithm, Kruskal's algorithm, or shortest path algorithm can be used. This embodiment does not limit it.

[0238] It should be understood that when the sum of the correlations between the final selected image clusters to be labeled is minimized, the data structure that does not contain circular connections can be understood as a minimum spanning tree.

[0239] In this embodiment, a data structure is constructed based on the image clusters corresponding to facial images and their correlation. The data structure includes multiple nodes and multiple edges. Each node represents an image cluster corresponding to a facial image, and the weight of the edge represents the correlation between the image clusters corresponding to facial images. Then, based on the edge weights in the data structure, a circular adjustment is performed on the data structure to obtain a data structure without circular connections. Finally, node groups are selected from the data structure without circular connections. The node groups correspond to the image cluster groups to be labeled. By using the data structure and based on the correlation, the image cluster groups to be labeled are obtained, which not only reduces the number of image cluster groups to be labeled but also obtains the image cluster groups to be labeled more accurately.

[0240] In other embodiments, to improve the recall rate of the face recognition model, candidate image clusters can be selected from the image clusters corresponding to the face images using active learning methods. Then, based on these candidate image clusters, the image clusters to be labeled are determined. Active learning refers to obtaining relatively "difficult" classifiable sample data through machine learning methods, having it manually reviewed, and then retraining the supervised or semi-supervised learning model with the manually reviewed data. This improves the model's performance and integrates human experience into the machine learning model.

[0241] Active learning methods can be such as Figure 15 As shown, the face recognition model periodically clusters historically recognized face images (clustering is performed again over extended time spans, referring to periodic clustering). Then, the clustered image clusters undergo a purification operation (which may include quality filtering) to obtain purified image clusters. These purified image clusters are then preprocessed to obtain preprocessed image clusters. Next, the preprocessed image clusters are manually labeled to obtain the labeling results. Finally, the labeled results are post-processed (post-processing refers to merging preprocessed image clusters belonging to the same object into a single image cluster) to obtain the training set.

[0242] In active learning methods, the difficulty of distinguishing data is typically determined through uncertainty sampling (US), diversity sampling (DS), expected model change (EMC), query-by-committee (QBC), and density-weighted methods.

[0243] When determining the distinguishing difficulty of data using density weighting, the correlation between image clusters corresponding to face images is determined based on facial features within those clusters. Based on this correlation, image clusters to be labeled are selected from these clusters, including:

[0244] Based on the facial features of the face image, the density weight of the image cluster corresponding to the face image is determined. The density weight is used to represent the difficulty of distinguishing the face image.

[0245] From the image clusters corresponding to the face images, filter out the image clusters corresponding to the density weights that meet the preset weights to obtain candidate image clusters;

[0246] The correlation between candidate image clusters is determined based on facial features in the candidate image clusters.

[0247] Based on the degree of relevance, select the image clusters to be labeled from the candidate image clusters.

[0248] At this point, a data structure is constructed based on the image clusters corresponding to the face images and their correlation. The data structure includes multiple nodes and multiple edges. Each node represents an image cluster corresponding to a face image, and the weight of the edge represents the correlation between the image clusters corresponding to the face images. This can include:

[0249] A data structure is constructed based on candidate image clusters and their correlation. The data structure includes multiple nodes and multiple edges. Each node represents a candidate image cluster, and the weight of the edge represents the correlation between candidate image clusters.

[0250] Density weight refers to the fact that if a piece of data is outlier or deviates significantly from most other data, it is not suitable as a sample. Instead, denser, harder-to-distinguish data is more valuable. In other words, density weight can be used to represent the difficulty of distinguishing facial images. The larger the density weight, the harder it is to distinguish facial images. The smaller the density weight, the easier it is to distinguish facial images.

[0251] For example, such as Figure 16 As shown, image cluster c3 and image cluster c1 are not face image clusters of the same object, image cluster c3 and image cluster c2 are not face image clusters of the same object, while image cluster c1 and image cluster c2 are face image clusters of the same object. Therefore, the density weight of image cluster c3 is less than the density weight of image cluster c1 and less than the density weight of image cluster c2.

[0252] In this embodiment, the difficulty of differentiation is represented by the density weight of the image cluster corresponding to the face image based on the facial features of the face image. Then, based on the density weight, the image clusters corresponding to the density weights that meet the preset weights are selected from the image clusters corresponding to the face image to obtain candidate image clusters. This makes the recall rate of the target face recognition model higher when the face recognition model is trained based on the candidate image clusters.

[0253] The process of determining the density weights of the image clusters corresponding to the face images based on their facial features can be as follows:

[0254] Calculate the similarity between the central features of the image clusters corresponding to each face image and the facial features of the face images;

[0255] Based on the facial features of the face image, determine the first prediction probability and the second prediction probability corresponding to the face image;

[0256] The density weights of the face image are determined based on the first and second predicted probabilities and the similarity.

[0257] The first prediction probability and the second prediction probability can be two of the prediction probabilities of the face recognition model for the face image. For example, the first prediction probability can be the highest probability that the face recognition model predicts the face image, and the second prediction probability can be the second highest probability that the face recognition model predicts the face image. Alternatively, the first prediction probability can be the second highest probability that the face recognition model predicts the face image, and the second prediction probability can be the third highest probability that the face recognition model predicts the face image. This embodiment does not impose any limitations on this.

[0258] Optionally, the density weights of the face image can be obtained by substituting the first predicted probability, the second predicted probability, and the similarity into the following formula:

[0259]

[0260]

[0261] x ID Let x(u) represent the density weight of facial feature x, x(u) represent the central feature of each image cluster, u represent the u-th image cluster, and U represent the number of image clusters. argmax x This represents the operation of finding the maximum value of the independent variable. sim This represents the similarity operation. Indicates the exponential parameter. argmin x This represents the operation of finding the minimum value of the independent variable. Indicates the first predicted probability. This indicates the category corresponding to the first predicted probability. Indicates the second predicted probability. The category corresponding to the second predicted probability The facial feature that minimizes the difference between the first and second predicted probabilities is, in other words, It can be obtained through the principle of edge sampling.

[0262] Active learning methods have achieved good results on closed-set tasks, where the training set and the test set contain samples of the same class. For example, if the training set includes samples of class A, class B, and class C, then the test set also includes samples of class A, class B, and class C.

[0263] In other words, if the training and test sets of the face recognition model are closed sets, the recall rate of the face recognition model will be high after training it on difficult-to-distinguish face images selected through active learning.

[0264] However, in practical applications, due to the high difficulty in obtaining samples, such as in the medical or aerospace fields, the test set may contain samples of categories not found in the training set. This type of problem is known as the Open Set Recognition problem.

[0265] Face recognition is an open-set recognition problem. Face images contain a lot of noise. If we directly select the image clusters to be labeled from the image clusters corresponding to the face images using active learning methods, and then label the face images according to the image clusters to be labeled, the resulting labeled face images will affect the recall rate of the face recognition model trained on the labeled face images.

[0266] Therefore, in order to reduce the impact on the recall rate of the face recognition model, in some embodiments, the density weight of the image cluster corresponding to the face image is determined based on the facial features of the face image. From the image cluster corresponding to the face image, image clusters corresponding to the density weights that meet the preset weights are selected to obtain candidate image clusters, including:

[0267] Get the tag popularity of image clusters of facial images;

[0268] The initial image clusters are obtained by filtering out image clusters of facial images that meet the preset popularity of tags.

[0269] The density weights of the initial image clusters are determined based on the facial features of the face images in the initial image clusters.

[0270] From the initial image clusters, select the image clusters that meet the preset weight density weights to obtain candidate image clusters.

[0271] Tag (ID) popularity can be calculated as the ratio between the number of times a tag appears and the total number of face images, i.e.: ID popularity = number of times an ID appears / total number of face images.

[0272] In this embodiment, instead of directly selecting candidate image clusters from the image clusters corresponding to facial images, the active learning method is improved. The improvement can be... Figure 17 The improved part 1 shown first filters out image clusters with preset tag popularity from the image clusters of face images based on tag popularity, and then determines the density weight of the initial image clusters based on the facial features of the face images in the initial image clusters, and obtains candidate image clusters based on the density weight.

[0273] Since high label popularity indicates that the samples corresponding to that label are relatively easy to obtain, the difference between the categories in the training set and the test set is smaller. Therefore, the difficulty of obtaining face images in the initial screening image cluster is lower, which means that the difficulty of obtaining face images in the candidate image cluster is lower, thereby reducing the impact on the recall rate of the face recognition model.

[0274] After obtaining the image clusters to be labeled, the server can label the face images according to the image clusters to obtain labeled face images. Optionally, labeling the face images according to the image clusters to obtain labeled face images includes:

[0275] Extract multiple modal information corresponding to facial images from the image clusters to be labeled;

[0276] Display the image clusters and modal information to be labeled;

[0277] The system receives annotation information from users, who annotate facial images in a cluster of images to be labeled based on modal information, and obtains labeled facial images.

[0278] Modal information can refer to the attribute information of a facial image. This attribute information can be inherent to the facial image itself, or it can be attribute information from the video in which the facial image is located. For example, attribute information can be at least one of the following: image details, caption information, video IP (the video containing the facial image), character information, watermark information, and person information. Displaying multiple modal information allows users to annotate based on these multiple modalities, achieving multimodal annotation and thus improving annotation accuracy.

[0279] For example, such as Figure 18As shown, set A and set B are image clusters to be labeled. Set A comes from video A, and set B comes from video B. The modal information of facial images in set A, set B, set A, and set B is displayed.

[0280] It should be understood that the multimodal annotation process in this embodiment can be understood as an improvement on the annotation method in active learning, such as... Figure 15 As shown, in related technologies, the annotation methods in active learning are simply done manually and do not display multiple modal information. However, in this embodiment, as... Figure 17 As shown, multiple modal information is displayed so that manual annotation can be performed based on the multiple modal information.

[0281] S302. Extract features from the training samples to obtain the facial features corresponding to the training samples, and fuse the facial features and the central features of the target clusters corresponding to the labels to obtain the fused features of the training samples.

[0282] One approach is to extract features from the training samples using the feature extraction layer in the face recognition model to obtain the corresponding facial features. Alternatively, the initial facial features corresponding to the training samples can be extracted using the feature extraction layer in the face recognition model, and then the initial facial features can be mapped using the feature mapping layer in the face recognition model to obtain the corresponding facial features for the training samples.

[0283] In some embodiments, the feature mapping layer in the face recognition model includes a first fully connected layer and a first activation layer. The feature mapping layer in the face recognition model maps initial facial features to obtain facial features corresponding to the training samples, including:

[0284] The initial facial features are mapped by the first fully connected layer in the face recognition model to obtain the dimensionality-reduced facial features corresponding to the training samples.

[0285] By using the first activation layer in the face recognition model, nonlinear feature mapping is performed on the dimensionality-reduced facial features to obtain the facial features corresponding to the training samples.

[0286] In this embodiment, the initial facial features are dimensionality-reduced and mapped by the first fully connected layer, thereby reducing the computational load of extracting the nonlinear features of the initial facial features. Then, the dimensionality-reduced facial features are mapped nonlinearly to obtain the facial features corresponding to the training samples. That is, the facial features include the nonlinear features of the initial facial features.

[0287] In other embodiments, to address the overfitting problem, the feature mapping layer in the face recognition model can be a residual structure. In this case, the face recognition model can be as follows: Figure 19As shown, the feature mapping layer in the face recognition model also includes a second fully connected layer. Through the first activation layer in the face recognition model, non-linear feature mapping is performed on the dimensionality-reduced facial features to obtain the facial features corresponding to the training samples, including:

[0288] By using the first activation layer in the face recognition model, nonlinear feature mapping is performed on the dimensionality-reduced face features to obtain the nonlinear face features corresponding to the training samples.

[0289] The second fully connected layer in the face recognition model is used to perform feature upsizing mapping on nonlinear face features, and the upsizing face features corresponding to the training samples are used.

[0290] Based on the upgraded facial features and the initial facial features, the facial features corresponding to the training samples are determined.

[0291] In other embodiments, the face recognition model further includes a central mapping layer, and the method further includes:

[0292] Obtain the initial center features of the target cluster corresponding to the label;

[0293] By using the center mapping layer of the face recognition model, the initial center features are mapped to obtain the center features of the target cluster corresponding to the label.

[0294] The structure of the central mapping layer can be the same as that of the feature mapping layer, as detailed in the description of the feature mapping layer above. Figure 19 This embodiment is not limited here.

[0295] In some embodiments, the parameters of the central mapping layer can be the same as the parameters of the feature mapping layer. That is, when training the face recognition model, the parameters of the feature mapping layer can be obtained, and then the parameters of the feature mapping layer can be shared with the central mapping layer. In this case, the face recognition model can be an adaptive face recognition model.

[0296] An adaptive face recognition model can refer to the process of automatically adjusting the processing methods, processing order, processing parameters, boundary conditions, or constraints based on the features of the face image during processing and analysis, so as to adapt them to the statistical distribution and structural features of the face image being processed, in order to achieve the best processing results. The adaptive process is a process of continuously approximating the target, which can be represented by a mathematical model.

[0297] S303. Based on the fusion features and facial features, adjust the central features of the target cluster to obtain the adjusted central features of the target cluster.

[0298] One approach is to determine the predicted labels of the training samples based on the fusion features, and then adjust the central features of the target cluster based on the facial features and the predicted labels to obtain the adjusted central features of the target cluster. Alternatively, one approach is to determine the predicted labels of the training samples based on the fusion features, and then adjust the central features of the target cluster based on the facial features, the predicted labels, and the initial central features of the target cluster to obtain the adjusted central features of the target cluster.

[0299] When the predicted labels of training samples are determined based on the fusion features, and then the central features of the target cluster are adjusted based on facial features and the predicted labels to obtain the adjusted central features of the target cluster, the adjusted central features of the target cluster include:

[0300] Obtain the first direction information corresponding to the facial features, and obtain the second direction information of the initial facial features;

[0301] Based on the first direction information, the second direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0302] In this embodiment, the central features of the target cluster are adjusted based on the first direction information, the second direction information, and the predicted label to obtain the adjusted central features of the target cluster. This allows for the adjustment of the central features of the target cluster based on the direction information, i.e., the boundary information, thereby constraining the central features of the target cluster and obtaining the adjusted central features of the target cluster.

[0303] Alternatively, based on facial features and predicted labels, the central features of the target cluster can be adjusted to obtain the adjusted central features of the target cluster, including:

[0304] Obtain the first orientation information corresponding to the facial features, and obtain the third orientation information of the central features of the target cluster;

[0305] Based on the first direction information, the third direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0306] In this embodiment, the central features of the target cluster are adjusted based on the first direction information, the third direction information, and the predicted label to obtain the adjusted central features of the target cluster. This allows for the adjustment of the central features of the target cluster based on the direction information, i.e., the boundary information, thereby constraining the central features of the target cluster and obtaining the adjusted central features of the target cluster.

[0307] In other embodiments, the process of adjusting the center features of the target cluster based on the first direction information, the second direction information, and the predicted label to obtain the adjusted center features of the target cluster can be as follows:

[0308] Obtain the direction weights corresponding to the first direction information and the second direction information;

[0309] Based on the direction weights corresponding to the first direction information, the first direction information is adjusted to obtain the adjusted first direction information; and based on the direction weights corresponding to the second direction information, the second direction information is adjusted to obtain the adjusted second direction information.

[0310] Based on the adjusted first direction information and the adjusted second direction information, determine the initial facial features and the first angle information between the facial features;

[0311] Based on the first included angle information and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0312] Alternatively, the process of adjusting the central features of the target cluster based on the first directional information, the second directional information, and the predicted label to obtain the adjusted central features of the target cluster can also be as follows:

[0313] Obtain third-party directional information of the central features of the target cluster;

[0314] Based on the first direction information, the second direction information, the third direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0315] The process of adjusting the central features of the target cluster based on the first direction information, the second direction information, the third direction information, and the predicted label to obtain the adjusted central features of the target cluster can be described as follows:

[0316] Based on the first direction information and the second direction information, determine the initial facial features and the first angle information between the facial features;

[0317] Based on the second direction information and the third direction information, determine the second included angle information between the initial facial features and the central features of the target cluster;

[0318] Based on the first included angle information, the second included angle information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0319] In other embodiments, the process of adjusting the center features of the target cluster based on the first included angle information, the second included angle information, and the predicted label to obtain the adjusted center features of the target cluster can be as follows:

[0320] Determine the angle difference between the first included angle information and the second included angle information;

[0321] Based on the predicted label, the angle difference, and the label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0322] It should be understood that the central features of the target cluster can be the initial central features of the target cluster, or they can be features obtained by mapping the initial central features of the target cluster.

[0323] When the central feature of the target cluster is the same as the initial central feature of the target cluster, the predicted label, the angle difference, the label, and the adjusted central feature of the target cluster satisfy the following relationship:

[0324]

[0325] kada' This represents the adjusted central features of the target cluster. N Indicates the category of the training samples. i This indicates the number of training samples. Indicates training samples i The tag, Indicates training samples i Predicted labels, b_func This represents a function for calculating the included angle. Indicates training samples i The initial facial features, Indicates training samples i Facial features, Kold Indicates training samples i The initial central characteristics of the target cluster.

[0326] In this embodiment, the center features of the target cluster are constrained by the difference between the first included angle information and the second included angle information, so that when the face recognition model is trained based on the compatibility loss value obtained from the adjusted center of the target cluster, the face recognition model can achieve a better convergence state.

[0327] S304. The adjusted center features of the facial features and the target cluster are fused to obtain the adjusted fused features of the training samples.

[0328] The process of fusing facial features and the adjusted central features of target clusters to obtain the adjusted fused features of training samples can be referenced to the process of fusing facial features and the central features of target clusters corresponding to labels to obtain the fused features of training samples. The specific implementation is not limited here.

[0329] It should be understood that after transforming the central features of the target cluster into adjusted central features of the target cluster, the face recognition model can be called a new face recognition model. At this point, the face features used in the fusion process of face features and the adjusted central features of the target cluster can refer to the face features obtained by feature extraction from the training samples using the new face recognition model. That is to say, at this time, Can refer to The relationship between the face recognition model and the new face recognition model can be as follows: Figure 18 As shown.

[0330] S305. Based on the adjusted fusion features, determine the compatibility loss value between the fusion features and the adjusted fusion features, and determine the classification loss value of the face recognition model.

[0331] The compatibility loss between the fused features and the adjusted fused features can also be understood as the compatibility loss between the facial features and the adjusted fused features. Specifically, the adjusted fused features and facial features undergo a first fusion process to obtain the target fused features of the training samples. Then, based on the target fused features, the compatibility loss between the fused features and the adjusted fused features is calculated.

[0332] Alternatively, the first fusion process can be performed on the adjusted fusion features and the initial facial features to obtain the target fusion features of the training samples, and then the compatibility loss value between the fusion features and the adjusted fusion features can be calculated based on the target fusion features.

[0333] The first fusion process is the process of determining the difference between the adjusted fused features and the initial facial features. The first fusion process can be a multiplicative fusion process, or it can be a subtractive fusion process.

[0334] When the first fusion process is a subtractive fusion process, the adjusted fusion features can be logarithmically processed first to obtain simplified adjusted fusion features. Then, the simplified adjusted fusion features are subtracted from the initial facial features to obtain the target fusion features of the training samples. Finally, the target fusion features are divided by the number of categories in the training set to obtain the compatibility loss value. At this point, the adjusted fusion features, the initial facial features, and the target fusion features satisfy the following relationship:

[0335]

[0336] This represents the compatibility loss value between the fused features and the adjusted fused features. y i Indicates training samples i The target cluster it belongs to.

[0337] In related technologies, during the process of recognizing facial images using a facial recognition model, the model is updated based on the facial images. However, the facial features extracted by the updated facial recognition model are incompatible with the features in the facial database (incompatibility means that the distance between the facial features extracted by the updated model and the features in the facial database cannot be directly calculated). This necessitates re-extracting features from the facial image in the database using the updated facial recognition model, which is time-consuming and results in slow recognition speed.

[0338] For example, such as Figure 20 As shown, the facial features in the query collection extracted by the updated facial recognition model are incompatible with the facial features in the gallery collection extracted by the facial recognition model. This results in the need to update the gallery database by updating the facial recognition model. When the gallery volume is large, the equipment and time costs of database updating are high.

[0339] In this embodiment, the central features of the target clusters corresponding to facial features and labels are fused to obtain the fused features of the training samples. This allows the central features of the target clusters to be adjusted based on the fused features and facial features, resulting in adjusted central features of the target clusters. This achieves boundary constraints on the center of the target clusters. After fusing the facial features and the adjusted central features of the target clusters to obtain the adjusted fused features of the training samples, the compatibility loss value between the fused features and the adjusted fused features can be determined. Based on the compatibility loss value, the face recognition model is trained to correct the central features of the target clusters, thus preserving the information of the central features of the target clusters. This ensures that the facial features obtained by the target face recognition model are compatible with those obtained by the face recognition model, eliminating the need to re-extract features from the face images in the face database using the target face recognition model. This reduces the time for feature re-extraction, improves recognition speed, and lowers time and equipment costs.

[0340] Furthermore, compared to directly calculating the compatibility loss value based on the central features of the target cluster, the embodiments of this application adjust the central features of the target cluster and then calculate the compatibility loss value based on the adjusted central features of the target cluster, which can enable the face recognition model to achieve a better convergence state.

[0341] S306. Based on the classification loss value and the compatibility loss value, train the face recognition model to obtain the updated target face recognition model.

[0342] The server can directly add the classification loss value and the compatibility loss value to obtain the target loss value of the face recognition model. Alternatively, the server can obtain the classification weight corresponding to the classification loss value and the compatibility weight corresponding to the compatibility loss value, then multiply the classification weight by the classification loss value to obtain the adjusted classification loss value, multiply the compatibility weight by the compatibility loss value to obtain the adjusted compatibility loss value, and finally add the adjusted classification loss value and the adjusted compatibility loss value to obtain the target loss value.

[0343] After obtaining the target loss value, if the target loss value is equal to or greater than the preset loss value, the parameters of the face recognition model are updated according to the target loss value (the parameters of the parameter mapping layer may not be updated), and the process returns to perform the step of extracting features from the training samples to obtain the face features corresponding to the training samples. At this time, it is not necessary to perform the step of fusing the face features and the central features of the target clusters corresponding to the labels to obtain the fused features of the training samples, and then adjusting the central features of the target clusters according to the fused features and the face features to obtain the adjusted central features of the target clusters. This embodiment will not elaborate on this step.

[0344] Optionally, the classification weights and compatibility weights can be fixed during training, or the classification weights and compatibility loss weights can be dynamically changed during training. That is, during the training of the face recognition model, the classification weights and compatibility loss weights are learned simultaneously. In this case, while updating the parameters of the face recognition model according to the target loss value, the classification weights and compatibility loss weights are also updated.

[0345] It should be noted that when applying the target face recognition model for face recognition, the feature mapping layer and center mapping layer can be removed from the model. This prevents the introduction of additional parameters during face recognition. In this case, the process of recognizing the face image using the target face recognition model can be as follows:

[0346] The face image to be identified is obtained, and the features of the face image to be identified are extracted through the feature extraction layer in the target face recognition model to obtain the face features corresponding to the face image to be identified.

[0347] The classification layer in the target face recognition model classifies the facial features to be recognized, thus obtaining the recognition result of the face image to be recognized.

[0348] After obtaining the target face recognition model, the update gain method is used to evaluate the compatibility index of the target face recognition model. The evaluation formula can be as follows:

[0349]

[0350] Indicates the update gain. This represents the target face recognition model. Let Q represent the face recognition model, Q represent the query set, and D represent the registration set (database set). This represents the true positive rate (TPR) when facial features are extracted from facial images in the query set using a target face model and facial features are extracted from facial images in the registration set using a face recognition model. The true positive rate when facial features are extracted from facial images in the registry and facial features are extracted from facial images in the query set using a facial recognition model. This represents the true positive rate when facial features are extracted from facial images in the registration set and facial features are extracted from facial images in the query set using the target face recognition model.

[0351] Tables 1 and 2 show the performance of the face recognition models obtained through various update methods. The training sets in Table 1 and Table 2 are different. `modelv0` represents the face recognition model before the update, `ours` represents the target face recognition model, `model*` represents the face recognition model without compatibility updates, `adapter wo boundary` represents the adaptive model without boundary constraints, `adapter wo residual` represents the adaptive model without residual structure, `adapter 0.5beta` represents the target face recognition model with classification and compatibility weights of 0.5, and `adapter learnable beta` represents the target face recognition model with dynamically changing classification and compatibility weights.

[0352]

[0353] Table 1

[0354]

[0355] Table 2

[0356] As shown in Table 1, the accuracy and update gain of the target face recognition model are both higher than those of the face recognition model obtained through related techniques. Table 2 shows that the true positive rate and update gain of the target face recognition model with fixed classification and compatibility weights are higher than those with dynamically changing classification and compatibility weights; the true positive rate and update gain of the target face recognition model with residual structure are higher than those without residual structure; and the true positive rate of the adaptive model without boundary constraints is higher than that of the adaptive model without residual structure (the importance of residual structure and boundary constraints is obtained through ablation study).

[0357] As described above, in this embodiment, a training set for updating the face recognition model is obtained. The training set includes at least one labeled training sample, which is a face image historically recognized by the face recognition model. Feature extraction is performed on the training sample to obtain the face features corresponding to the training sample. The face features and the center features of the target cluster corresponding to the label are fused to obtain the fused features of the training sample. This allows the center features of the target cluster to be adjusted based on the fused features and the face features, resulting in the adjusted center features of the target cluster. The fusion of the face features and the adjusted center features of the target cluster then yields the training model. After adjusting and fusing the features of the samples, the compatibility loss value between the fused features and the adjusted fused features can be determined, as well as the classification loss value of the face recognition model. Based on the classification loss value and the compatibility loss value, the face recognition model is trained to obtain the updated target face recognition model. This ensures that the face features obtained through the target face recognition model are compatible with those obtained through the face recognition model, thereby eliminating the need to re-extract features from the face images in the face database using the target face recognition model, reducing the time spent on feature re-extraction, and improving the recognition speed.

[0358] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.

[0359] Please see Figure 21 , Figure 21 The method for constructing a training set for an updated face recognition model provided in this application embodiment may include the following steps:

[0360] S2101. The server obtains historical face images recognized by the face recognition model and extracts key points from the historical face images.

[0361] S2102. The server uses the first comprehensive quality model to determine the first initial quality score corresponding to the historical face image based on key points, and uses the second comprehensive quality model to determine the second initial quality score corresponding to the historical face image based on key points, and obtains the preset screening threshold.

[0362] S2103. If both the first initial quality score and the second initial quality score are less than or equal to the preset screening threshold, the server will use the first initial quality score as the quality score corresponding to the historical face image.

[0363] S2104. If both the first initial quality score and the second initial quality score are greater than the preset screening threshold, the server will use the second initial quality score as the quality score corresponding to the historical face image.

[0364] S2105. The server will use the historical face images corresponding to the quality scores that meet the preset score threshold as the face images that the face recognition model has historically recognized.

[0365] In active learning methods, the purification operation, or quality filtering, is performed manually. However, in this embodiment, the quality scores corresponding to historical face images are determined using a first comprehensive quality model and a second comprehensive quality model, automating the data mining process and improving the quality filtering process in active learning methods. For example... Figure 17 Improvement 2 in the model improves the quality of training samples in the training set for updating the face recognition model, thereby increasing the recall rate of the face recognition model.

[0366] S2106. The server clusters face images using a face recognition model to obtain the image clusters in which the face images are located.

[0367] S2107. The server obtains the tag popularity of the image clusters of face images, and filters out the image clusters of face images that meet the preset popularity of the tag popularity to obtain the initial screened image clusters.

[0368] S2108. The server determines the density weight of the initial screening image cluster based on the facial features of the face images in the initial screening image cluster, and selects the image clusters corresponding to the density weights that meet the preset weights from the initial screening image clusters to obtain candidate image clusters.

[0369] S2109. The server determines the correlation between candidate image clusters based on the facial features corresponding to the face images in the candidate image clusters, and constructs a tree structure based on the candidate image clusters and the correlation. The tree structure includes multiple nodes and multiple edges, where each node represents a candidate image cluster and the weight of the edge represents the correlation between the candidate image clusters.

[0370] S21010. The server adjusts the tree structure according to the weight of the edges in the tree structure to obtain the minimum spanning tree, and selects node groups from the minimum spanning tree. The node groups correspond to the image clusters to be labeled.

[0371] S21011. The server extracts multiple modal information corresponding to the face images in the image clusters to be labeled, and displays the image clusters to be labeled and the modal information.

[0372] S21012. The server receives the annotation information from the user, which is based on the modal information, to annotate the face images in the image cluster to be annotated, and obtains the annotated face images.

[0373] In this embodiment, multiple modal information is displayed, which effectively suppresses interference caused by multiple makeup looks, multiple angles, and age ranges, making it easier for users to determine whether the face images in the image cluster to be labeled are images of the same person. This effectively improves the labeling efficiency and accuracy, reduces the labeling cost, and increases the accuracy from 85% to 99%, while improving the labeling efficiency by 5 times.

[0374] This application's embodiments display multiple modal information for user annotation, representing an improvement on the annotation operation in active learning methods. For example, ... Figure 17 Improvement part 3.

[0375] S21013. The server constructs a training set for updating the face recognition model based on the labeled face images. The training set includes at least one labeled training sample.

[0376] For details on the specific implementation and corresponding beneficial effects of this embodiment, please refer to the above-described model update method embodiment. This example will not be repeated here.

[0377] Please see Figure 22 , Figure 22 The face recognition model update method provided in this application embodiment may include:

[0378] S2201. The server extracts features from the training samples through the feature extraction layer in the face recognition model to obtain the initial face features corresponding to the training samples.

[0379] S2202. The server performs feature dimensionality reduction mapping on the initial face features through the first fully connected layer in the face recognition model to obtain the dimensionality-reduced face features corresponding to the training samples. Then, it performs nonlinear feature mapping on the dimensionality-reduced face features through the first activation layer in the face recognition model to obtain the nonlinear face features corresponding to the training samples.

[0380] S2203. The server performs feature upscaling mapping on nonlinear face features through the second fully connected layer in the face recognition model, trains the upscaled face features corresponding to the training samples, and determines the face features corresponding to the training samples based on the upscaled face features and the initial face features.

[0381] S2204. The server obtains the initial central features of the target cluster corresponding to the label, and performs dimensionality reduction mapping on the initial central features of the target cluster through the third fully connected layer in the face recognition model to obtain the dimensionality-reduced initial central features.

[0382] S2205. The server performs nonlinear feature mapping on the initial central features after dimensionality reduction through the second activation layer in the face recognition model to obtain nonlinear initial central features, and then performs feature up-dimensional mapping on the nonlinear initial central features through the fourth fully connected layer in the face recognition model to obtain the initial central features after dimensionality up-dimensionality.

[0383] S2206. The server determines the central characteristics of the target cluster based on the initial central characteristics after dimensionality upgrade and the initial central characteristics of the target cluster.

[0384] S2207. The server uses the feature fusion layer in the face recognition model to fuse the face features and the central features of the target cluster to obtain the fused features of the training samples. Then, through the classification layer in the face recognition model, the server determines the predicted label of the training samples based on the fused features.

[0385] S2208, The server obtains the first direction information corresponding to the face features, the second direction information of the initial face features, and the third direction information of the initial center features of the target cluster.

[0386] S2209. The server determines the first angle information between the initial face feature and the face feature based on the first direction information and the second direction information, and determines the second angle information between the initial face feature and the central feature of the target cluster based on the second direction information and the third direction information.

[0387] S22010, The server determines the angle difference between the first angle information and the second angle information, and adjusts the central features of the target cluster based on the predicted label, the label, and the angle difference to obtain the adjusted central features of the target cluster.

[0388] S22011. The server performs fusion processing on the face features and the adjusted center features of the target cluster to obtain the adjusted fused features of the training samples.

[0389] S22012. The server performs a first fusion process on the adjusted fusion features and the initial face features to obtain the target fusion features of the training samples.

[0390] S22013. Based on the target fusion features, the server determines the compatibility loss value between the fusion features and the adjusted fusion features, as well as the classification loss value of the face recognition model.

[0391] S22014. The server trains the face recognition model based on the classification loss value and the compatibility loss value to obtain the updated target face recognition model.

[0392] For details on the specific implementation and corresponding beneficial effects of this embodiment, please refer to the above-described model update method embodiment. This example will not be repeated here.

[0393] To facilitate better implementation of the model update method provided in the embodiments of this application, the embodiments of this application also provide an apparatus based on the above-described model update method. The meanings of the terms used are the same as in the above-described model update method, and specific implementation details can be found in the descriptions in the method embodiments.

[0394] For example, such as Figure 23 As shown, the model update device may include:

[0395] The acquisition module 2301 is used to acquire the training set of the updated face recognition model. The training set includes at least one labeled training sample, which is a face image recognized by the face recognition model in the past.

[0396] The extraction module 2302 is used to extract features from the training samples to obtain the facial features corresponding to the training samples, and to fuse the facial features and the central features of the target clusters corresponding to the labels to obtain the fused features of the training samples.

[0397] The adjustment module 2303 is used to adjust the center features of the target cluster based on the fusion features and facial features to obtain the adjusted center features of the target cluster.

[0398] The fusion module 2304 is used to fuse facial features and the adjusted center features of the target cluster to obtain the adjusted fused features of the training samples.

[0399] Module 2305 determines the compatibility loss value between the fused features and the adjusted fused features, as well as the classification loss value of the face recognition model, based on the adjusted fused features.

[0400] Model 2306 is trained using classification loss and compatibility loss values ​​to obtain the updated target face recognition model.

[0401] Optionally, adjustment module 2303 is specifically used to perform:

[0402] Based on the fusion features, determine the predicted labels for the training samples;

[0403] Based on facial features and predicted labels, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0404] Optionally, the extraction module 2302 is specifically used to perform:

[0405] The initial facial features corresponding to the training samples are obtained by extracting features from the feature extraction layer in the face recognition model.

[0406] The initial facial features are mapped using the feature mapping layer in the face recognition model to obtain the facial features corresponding to the training samples.

[0407] Optionally, adjustment module 2303 is specifically used to perform:

[0408] Obtain the first direction information corresponding to the facial features, and obtain the second direction information of the initial facial features;

[0409] Based on the first direction information, the second direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0410] Optionally, adjustment module 2303 is specifically used to perform:

[0411] Obtain third-party directional information of the central features of the target cluster;

[0412] Based on the first direction information, the second direction information, the third direction information, and the predicted label, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0413] Optionally, the fusion module 2304 is specifically used to perform:

[0414] The adjusted fused features and the initial facial features are subjected to a first fusion process to obtain the target fused features of the training samples;

[0415] Based on the target fusion features, determine the compatibility loss value between the fusion features and the adjusted fusion features.

[0416] Optionally, the model update device also includes:

[0417] The mapping module is used for execution:

[0418] Obtain the initial center features of the target cluster corresponding to the label;

[0419] By using the center mapping layer of the face recognition model, the initial center features are mapped to obtain the center features of the target cluster corresponding to the label.

[0420] Optionally, module 2301 is specifically used to perform:

[0421] Obtain facial images previously recognized by the facial recognition model;

[0422] Select image clusters to be labeled from the image clusters corresponding to the face images, and label the face images according to the image clusters to be labeled to obtain labeled face images;

[0423] Based on labeled facial images, determine the training set for updating the facial recognition model.

[0424] Optionally, module 2301 is specifically used to perform:

[0425] Obtain historical facial images recognized by the facial recognition model;

[0426] Extract key points from historical facial images and determine the corresponding quality scores based on the key points;

[0427] Historical face images that meet the preset score threshold and have the corresponding quality scores are used as the face images recognized by the face recognition model in the past.

[0428] Optionally, module 2301 is specifically used to perform:

[0429] Obtain multiple quality models of historical facial images;

[0430] Based on key points, multiple initial quality scores are determined for historical facial images using a quality model.

[0431] Based on the initial quality score, determine the quality score corresponding to the historical face image.

[0432] Optionally, historical facial images have multiple quality evaluation dimensions, each corresponding to a dimensional quality model. Accordingly, the acquisition module 2301 is specifically used to perform:

[0433] Using a dimensional quality model, initial quality scores for each quality evaluation dimension of historical facial images are determined based on key points.

[0434] The initial quality scores are weighted to obtain the quality scores corresponding to the historical face images.

[0435] Optionally, the model update device also includes:

[0436] The training module is used to perform:

[0437] Obtain the first training set of the dimensional quality model to be trained. The first training set includes multiple first training samples.

[0438] Extract the key points of the first training sample, and determine the pairwise loss value and / or anchor loss value of the quality model of the dimension to be trained based on the key points of the sample.

[0439] The dimensionality quality model is trained based on the pairwise loss value and / or anchor loss value to obtain the dimensionality quality model.

[0440] Optionally, the training module is specifically used to perform:

[0441] Obtain the anchor label corresponding to the first training sample. The anchor label represents the level of the quality evaluation dimension. There are at least three anchor labels in the first training set.

[0442] Based on the key points, determine the predicted score of the first training sample;

[0443] Based on the anchor label, occlusion prediction score, and the score range corresponding to the anchor label, determine the anchor loss value of the dimensional quality model to be trained.

[0444] Optionally, the multiple quality models of the historical facial images include a first comprehensive quality model and a second comprehensive quality model, and accordingly, the acquisition module 2301 is specifically used to perform:

[0445] Based on key points, the first initial quality score corresponding to the historical face image is determined using the first comprehensive quality model.

[0446] The second comprehensive quality model is used to determine the second initial quality score corresponding to the historical face image based on the key points.

[0447] The quality scores corresponding to historical face images are selected from the first initial quality score and the second initial quality score.

[0448] Optionally, module 2301 is specifically used to perform:

[0449] Get the preset filtering threshold;

[0450] If both the first initial quality score and the second initial quality score are less than or equal to the preset screening threshold, then the first initial quality score will be used as the quality score corresponding to the historical face image.

[0451] If both the first initial quality score and the second initial quality score are greater than the preset screening threshold, then the second initial quality score will be used as the quality score corresponding to the historical face image.

[0452] Optionally, module 2301 is specifically used to perform:

[0453] Based on the facial features in the image clusters corresponding to the facial images, determine the correlation between the image clusters corresponding to the facial images;

[0454] Based on the correlation, select image clusters to be labeled from the image clusters corresponding to the facial images.

[0455] Optionally, module 2301 is specifically used to perform:

[0456] A data structure is constructed based on the image clusters corresponding to the face images and their correlation. The data structure includes multiple nodes and multiple edges. Each node represents the image cluster corresponding to the face image, and the weight of the edge represents the correlation between the image clusters corresponding to the face images.

[0457] Based on the weights of the edges in the data structure, the data structure is adjusted in a circular manner to obtain a data structure that does not contain circular connections.

[0458] Node groups are selected from data structures that do not contain circular connections, and the node groups correspond to the image clusters to be labeled.

[0459] Optionally, module 2301 is specifically used to perform:

[0460] Extract multiple modal information corresponding to facial images from the image clusters to be labeled;

[0461] Display the image clusters and modal information to be labeled;

[0462] The system receives annotation information from users, based on modal information, to annotate facial images in a cluster of images to be annotated, and obtains the annotated facial images.

[0463] In practice, each of the above modules can be implemented as an independent entity or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation methods and corresponding beneficial effects of each of the above modules, please refer to the previous method embodiments, which will not be repeated here.

[0464] This application also provides an electronic device, which may be a server or a terminal, etc. Figure 24 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:

[0465] The electronic device may include components such as a processor 2401 with one or more processing cores, a memory 2402 with one or more computer-readable storage media, a power supply 2403, and an input unit 2404. Those skilled in the art will understand that... Figure 24 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0466] The processor 2401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes computer programs and / or modules stored in the memory 2402, and calls data stored in the memory 2402 to perform various functions and process data. Optionally, the processor 2401 may include one or more processing cores; preferably, the processor 2401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 2401.

[0467] The memory 2402 can be used to store computer programs and modules. The processor 2401 executes various functional applications and data processing by running the computer programs and modules stored in the memory 2402. The memory 2402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 2402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 2402 may also include a memory controller to provide the processor 2401 with access to the memory 2402.

[0468] The electronic device also includes a power supply 2403 that supplies power to the various components. Preferably, the power supply 2403 can be logically connected to the processor 2401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 2403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0469] The electronic device may also include an input unit 2404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0470] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 2401 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 2402 according to the following instructions, and the processor 2401 runs the computer programs stored in the memory 2402 to realize various functions, such as:

[0471] Obtain the training set for updating the face recognition model. The training set includes at least one labeled training sample, which is a face image that the face recognition model has previously recognized.

[0472] Feature extraction is performed on the training samples to obtain the facial features corresponding to the training samples. The facial features and the central features of the target clusters corresponding to the labels are then fused to obtain the fused features of the training samples.

[0473] Based on the fusion features and facial features, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0474] The adjusted center features of facial features and target clusters are fused to obtain the adjusted fused features of the training samples;

[0475] Based on the adjusted fusion features, determine the compatibility loss value between the fusion features and the adjusted fusion features, and determine the classification loss value of the face recognition model;

[0476] The face recognition model is trained based on the classification loss value and the compatibility loss value to obtain the updated target face recognition model.

[0477] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the detailed description of the model update method above, which will not be repeated here.

[0478] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0479] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the model update methods provided in embodiments of this application. For example, the computer program can execute the following steps:

[0480] Obtain the training set for updating the face recognition model. The training set includes at least one labeled training sample, which is a face image that the face recognition model has previously recognized.

[0481] Feature extraction is performed on the training samples to obtain the facial features corresponding to the training samples. The facial features and the central features of the target clusters corresponding to the labels are then fused to obtain the fused features of the training samples.

[0482] Based on the fusion features and facial features, the central features of the target cluster are adjusted to obtain the adjusted central features of the target cluster.

[0483] The adjusted center features of facial features and target clusters are fused to obtain the adjusted fused features of the training samples;

[0484] Based on the adjusted fusion features, determine the compatibility loss value between the fusion features and the adjusted fusion features, and determine the classification loss value of the face recognition model;

[0485] The face recognition model is trained based on the classification loss value and the compatibility loss value to obtain the updated target face recognition model.

[0486] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the previous embodiments, which will not be repeated here.

[0487] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0488] Since the computer program stored in the computer-readable storage medium can execute the steps in any of the model update methods provided in the embodiments of this application, the beneficial effects that any of the model update methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0489] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned model update method.

[0490] The above provides a detailed description of a model update method, apparatus, device, storage medium, and program product provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A model updating method characterized by, The method comprises the following steps: obtaining a training set for updating a face recognition model, the training set comprising at least one labeled training sample, the training sample being a face image historically recognized by the face recognition model; performing feature extraction on the training sample to obtain face features corresponding to the training sample, and performing fusion processing on the face features and center features of a target cluster corresponding to the label to obtain fused features of the training sample; determining a predicted label of the training sample according to the fused features, obtaining first direction information corresponding to the face features, obtaining second direction information of an initial face feature corresponding to the training sample, and obtaining third direction information of the center features of the target cluster; adjusting the center features of the target cluster according to the first direction information, the second direction information, the third direction information, and the predicted label to obtain adjusted center features of the target cluster; performing fusion processing on the face features and the adjusted center features of the target cluster to obtain adjusted fused features of the training sample; determining a compatibility loss value between the fused features and the adjusted fused features according to the adjusted fused features, and determining a classification loss value of the face recognition model; training the face recognition model according to the classification loss value and the compatibility loss value to obtain an updated target face recognition model of the face recognition model; wherein the obtaining of the training set for updating the face recognition model comprises: obtaining face images historically recognized by the face recognition model; determining cluster similarity between picture clusters corresponding to the face images according to face features in the picture clusters; calculating label heat of the picture clusters corresponding to the face images according to the number of face images in the picture clusters and the total number of the face images; determining correlation degrees between the picture clusters corresponding to the face images according to the cluster similarity and the label heat, and selecting a to-be-labeled picture cluster group from the picture clusters corresponding to the face images according to the correlation degrees; labeling the face images according to the to-be-labeled picture cluster group to obtain labeled face images, and determining a training set for updating the face recognition model according to the labeled face images; the feature extraction on the training sample to obtain face features corresponding to the training sample comprises: performing feature extraction on the training sample through a feature extraction layer in the face recognition model to obtain initial face features corresponding to the training sample; performing feature mapping on the initial face features through a feature mapping layer in the face recognition model to obtain face features corresponding to the training sample.

2. The model update method according to claim 1, characterized by, the adjustment of the center features of the target cluster according to the first direction information, the second direction information, the third direction information, and the predicted label to obtain adjusted center features of the target cluster comprises: determining first angle information between the initial face features and the face features according to the first direction information and the second direction information; determine second included angle information between the initial face feature and the center feature of the target cluster according to the second direction information and the third direction information; adjust the center feature of the target cluster according to the first included angle information, the second included angle information and the predicted label, to obtain an adjusted center feature of the target cluster.

3. The model update method according to claim 2, characterized by, The adjusting the center feature of the target cluster according to the first included angle information, the second included angle information and the predicted label, to obtain an adjusted center feature of the target cluster, includes: determining an included angle difference between the first included angle information and the second included angle information; adjusting the center feature of the target cluster according to the predicted label, the included angle difference and the label, to obtain an adjusted center feature of the target cluster.

4. The model update method according to claim 1, characterized by, The face recognition model further includes a center mapping layer, and the method further includes: obtaining an initial center feature of a target cluster corresponding to the label; mapping the initial center feature through the center mapping layer of the face recognition model, to obtain the center feature of the target cluster corresponding to the label.

5. A model updating method characterized by, including: obtaining face images historically recognized by the face recognition model; determining cluster similarities between picture clusters corresponding to the face images according to face features in the picture clusters; calculating label hotness of the picture clusters corresponding to the face images according to a number of face images in the picture clusters corresponding to the face images and a total number of the face images; determining correlation degrees between the picture clusters corresponding to the face images according to the cluster similarities and the label hotness, and filtering out a to-be-labeled picture cluster group from the picture clusters corresponding to the face images according to the correlation degrees; labeling the face images according to the to-be-labeled picture cluster group, to obtain labeled face images, and determining a training set for updating the face recognition model according to the labeled face images; training the face recognition model according to the training set, to obtain an updated target face recognition model of the face recognition model.

6. The model update method according to claim 5, characterized by, The obtaining of the face images historically recognized by the face recognition model includes: obtaining historical face images historically recognized by the face recognition model; extracting key points of the historical face images, and determining quality scores corresponding to the historical face images according to the key points; taking the historical face images corresponding to quality scores satisfying a preset score threshold as the face images historically recognized by the face recognition model.

7. The model updating method according to claim 6, characterized by, The determining of the quality scores corresponding to the historical face images according to the key points includes: obtaining multiple quality models of the historical face images; determining multiple initial quality scores of the historical face images according to the key points through the quality models; determining the quality scores corresponding to the historical face images according to the initial quality scores.

8. The model updating method according to claim 7, characterized by, The historical face images have multiple quality evaluation dimensions, and each quality evaluation dimension corresponds to a dimension quality model; The determining of the multiple initial quality scores of the historical face images according to the key points through the quality models, and the determining of the quality scores corresponding to the historical face images according to the initial quality scores, include: The dimension quality model is used to determine initial quality scores of the historical face image for each quality evaluation dimension according to the key points. The initial quality scores are weighted to obtain a quality score corresponding to the historical face image.

9. The model updating method according to claim 8, characterized by, Before the step of determining initial quality scores of the historical face image for each quality evaluation dimension according to the key points through the dimension quality model, the method further includes: obtaining a first training set for training a dimension quality model, the first training set including a plurality of first training samples; extracting sample key points of the first training samples, and determining pair loss values and / or anchor loss values of the dimension quality model to be trained according to the sample key points; training the dimension quality model to be trained according to the pair loss values and / or the anchor loss values to obtain a dimension quality model.

10. The model updating method according to claim 7, characterized by, The plurality of quality models of the historical face image include a first comprehensive quality model and a second comprehensive quality model, and the step of determining a plurality of initial quality scores of the historical face image according to the key points through the quality model, and determining a quality score corresponding to the historical face image according to the initial quality scores includes: determining a first initial quality score corresponding to the historical face image according to the key points through the first comprehensive quality model; determining a second initial quality score corresponding to the historical face image according to the key points through the second comprehensive quality model; screening the quality score corresponding to the historical face image from the first initial quality score and the second initial quality score.

11. The model updating method according to claim 10, characterized by, The step of screening the quality score corresponding to the historical face image from the first initial quality score and the second initial quality score includes: obtaining a preset screening threshold; if the first initial quality score and the second initial quality score are both less than or equal to the preset screening threshold, taking the first initial quality score as the quality score corresponding to the historical face image; if the first initial quality score and the second initial quality score are both greater than the preset screening threshold, taking the second initial quality score as the quality score corresponding to the historical face image.

12. The model update method according to claim 5, characterized by, The step of screening a to-be-labeled picture cluster group from the picture clusters corresponding to the face image according to the correlation degrees includes: constructing a data structure according to the picture clusters corresponding to the face image and the correlation degrees, the data structure including a plurality of nodes and a plurality of edges, each node representing a picture cluster corresponding to a face image, and a weight of each edge representing a correlation degree between picture clusters corresponding to face images; performing ring adjustment on the data structure according to the weights of the edges in the data structure to obtain a data structure not containing ring connections; screening a node group from the data structure not containing ring connections, the node group corresponding to the to-be-labeled picture cluster group.

13. The model updating method according to claim 5, characterized by, The step of labeling the face image according to the to-be-labeled picture cluster group to obtain a labeled face image includes: extracting a plurality of modal information corresponding to face images in the to-be-labeled picture cluster group; displaying the to-be-labeled picture cluster group and the modal information. Receiving user according to the modal information, the face image in the face image cluster group to be marked is marked the marking information, obtains the face image with label.

14. A model updating apparatus characterized by comprising: Comprise: The acquisition module is used for acquiring a training set for updating a face recognition model, the training set comprising at least one training sample with label, the training sample being a face image historically recognized by the face recognition model; The extraction module is used for extracting features of the training sample to obtain face features corresponding to the training sample, and fusing the face features and center features of a target cluster corresponding to the label to obtain fused features of the training sample; The adjustment module is used for determining a predicted label of the training sample according to the fused features, acquiring first direction information corresponding to the face features, acquiring second direction information of initial face features corresponding to the training sample, and acquiring third direction information of the center features of the target cluster, adjusting the center features of the target cluster according to the first direction information, the second direction information, the third direction information and the predicted label to obtain adjusted center features of the target cluster; The fusion module is used for fusing the face features and the adjusted center features of the target cluster to obtain adjusted fused features of the training sample; The determination module is used for determining a compatible loss value between the fused features and the adjusted fused features according to the adjusted fused features, and determining a classification loss value of the face recognition model; The training model is used for training the face recognition model according to the classification loss value and the compatible loss value to obtain an updated target face recognition model of the face recognition model; The acquisition module is specifically used for: Acquiring face images historically recognized by the face recognition model; Determining cluster similarity between picture clusters corresponding to the face images according to face features in the picture clusters; Calculating label heat of the picture clusters corresponding to the face images according to the number of face images in the picture clusters and the total number of the face images; Determining correlation degrees between the picture clusters corresponding to the face images according to the cluster similarity and the label heat, and screening a face image cluster group to be marked from the picture clusters corresponding to the face images according to the correlation degrees; Marking the face images according to the face image cluster group to be marked to obtain face images with label, and determining a training set for updating the face recognition model according to the face images with label; The extraction module is specifically used for: Extracting features of the training sample through a feature extraction layer in the face recognition model to obtain initial face features corresponding to the training sample; 15. An electronic device, comprising: Mapping the initial face features through a feature mapping layer in the face recognition model to obtain face features corresponding to the training sample. The device comprises a processor and a memory, the memory stores a computer program, and the processor is used for running the computer program in the memory to execute the model updating method in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is suitable for being loaded by the processor to execute the model updating method in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Unsupervised pedestrian re-identification method based on pseudo-label self-correction

    CN112507901A

  • Online face clustering method and system

    WO2020232697A1