Face model generation method, face image data processing method and device

By combining the feature extraction and classification results of the face model to be trained and the trained face model to generate loss data, the distance and compatibility between the new and old feature vectors are optimized, which solves the problems of resource consumption and low efficiency in the iterative update of face feature vectors, and achieves more efficient compatibility and lower cost.

CN121010993APending Publication Date: 2025-11-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410640901.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-22
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies consume a lot of resources and are inefficient when iterating and updating facial feature vectors. Furthermore, forcibly constraining the old and new feature vectors can easily lead to model non-convergence, and the additional cost of using a conversion module is high.

Method used

By acquiring sample face image data, feature extraction and classification are performed using the first face model to be trained and the second face model that has already been trained. Loss data is generated to update the parameters of the first face model so that it can converge. The prior information of the old model is combined to optimize the distance and compatibility of the new features.

Benefits of technology

It improves the compatibility between new and old face feature vectors, reduces model training costs and resource consumption, increases iteration speed and efficiency, and avoids the cost of repeatedly updating existing content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010993A_ABST
    Figure CN121010993A_ABST
Patent Text Reader

Abstract

The invention provides a face model generation method, a face image data processing method and a face model generation device, which are applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like, and the face model generation method comprises the following steps: performing face feature extraction on sample face image data based on a first face model; obtaining a first face feature extraction result, and classifying the first face feature extraction result to obtain a first face feature classification result; performing feature extraction on the sample face image data based on a second face model to obtain a second face feature extraction result, and classifying the first face feature extraction result and the second face feature extraction result to obtain a second face feature classification result; generating first loss data, second loss data and third loss data; and updating model parameters of the first face model according to the first loss data, the second loss data and the third loss data. According to the invention, compatible iteration between vectors of different versions can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, specifically relating to a method for generating a face model, a method for processing face image data, and an apparatus. Background Technology

[0002] Faces are a crucial part of visual perception. Information streams contain a large amount of content, including people, which often involves face understanding and vectorization. Face understanding and vectorization are fundamental to various downstream tasks and applications. Because content in information streams is continuously produced, there is a large amount of existing content. During face vectorization algorithm iterations, the facial features in this large amount of existing content need to be recalculated and updated. However, this recalculation and update process is resource-intensive and inefficient. Research shows that compatibility between feature versions can avoid the aforementioned drawbacks of high resource consumption.

[0003] Related technologies typically impose strict constraints on pairs of new and old facial feature vectors to make them close enough for compatible training. However, the feature vector space is complex, and such strict constraints can easily lead to model non-convergence. Other technologies use a transformation module to convert new facial feature vectors into those of older ones for compatible training. However, this approach requires introducing a new model, redeploying the service when updating the model, and the transformation module itself is complex and costly. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a method for generating a face model, a method for processing face image data, and an apparatus.

[0005] On the one hand, this application proposes a method for generating a face model, the method comprising:

[0006] Acquire sample face image data, a first face model, and a second face model; the first face model is the model to be trained, and the second face model is the model that has already been trained;

[0007] Based on the first face model, facial features are extracted from the sample facial image data to obtain a first facial feature extraction result, and the first facial feature extraction result is classified to obtain a first facial feature classification result; based on the second face model, features are extracted from the sample facial image data to obtain a second facial feature extraction result, and the first facial feature extraction result and the second facial feature extraction result are classified to obtain a second facial feature classification result;

[0008] First loss data is generated based on the first face feature classification result, second loss data is generated based on the second face feature classification result, and third loss data is generated based on the first face feature classification result and the second face feature classification result.

[0009] The model parameters of the first face model are updated based on the first loss data, the second loss data, and the third loss data, so that the first face model can converge and the trained first face model is obtained.

[0010] Wherein, the first loss data is the loss caused by increasing the distance between positive sample face image data and negative sample face image data in the sample face image data, the second loss data is the loss caused by compatibility optimization of the second face model, and the third loss data is the loss caused by shortening the distance between the first face feature classification result and the second face feature classification result belonging to the same feature cluster category.

[0011] On the other hand, this application proposes a method for processing facial image data, the method comprising:

[0012] Acquire the face image data to be processed;

[0013] The target face feature classification result is obtained by classifying the face image data to be processed based on the first face model that has been trained;

[0014] The first face model that has been trained is generated based on the face model generation method described in any of the above embodiments.

[0015] On the other hand, this application proposes a face model generation apparatus, the apparatus comprising:

[0016] The data acquisition module is used to acquire sample face image data, a first face model, and a second face model; the first face model is a model to be trained, and the second face model is a model that has already been trained.

[0017] The extraction and classification module is used to extract facial features from the sample facial image data based on the first facial model to obtain a first facial feature extraction result, and to classify the first facial feature extraction result to obtain a first facial feature classification result; and to extract features from the sample facial image data based on the second facial model to obtain a second facial feature extraction result, and to classify the first facial feature extraction result and the second facial feature extraction result to obtain a second facial feature classification result.

[0018] The loss generation module is used to generate first loss data based on the first face feature classification result, generate second loss data based on the second face feature classification result, and generate third loss data based on the first face feature classification result and the second face feature classification result.

[0019] An update module is used to update the model parameters of the first face model according to the first loss data, the second loss data, and the third loss data, so that the first face model converges and a trained first face model is obtained; wherein, the first loss data is the loss caused by increasing the distance between positive sample face image data and negative sample face image data in the sample face image data, the second loss data is the loss caused by compatibility optimization of the second face model, and the third loss data is the loss caused by shortening the distance between the first face feature classification result and the second face feature classification result belonging to the same feature cluster category.

[0020] On the other hand, this application proposes a face image data processing apparatus, the apparatus comprising:

[0021] The module for acquiring face image data to be processed is used to acquire face image data to be processed.

[0022] The target face feature classification result generation module is used to classify the face image data to be processed based on the trained first face model to obtain the target face feature classification result;

[0023] The first face model that has been trained is generated based on the face model generation method described in any of the above embodiments.

[0024] On the other hand, this application proposes an electronic device for generating a face model or processing face image data. The electronic device includes a processor and a memory. The memory stores at least one instruction or at least one program. The processor loads and executes the at least one instruction or at least one program to implement the face model generation method or face image data processing method as described above.

[0025] On the other hand, this application proposes a computer-readable storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the face model generation method or face image data processing method as described above.

[0026] On the other hand, this application proposes a computer program product, including a computer program that, when executed by a processor, implements the face model generation method or face image data processing method as described above.

[0027] This application proposes a face model generation method, a face image data processing method, and an apparatus. The method uses a first face model to be trained to extract face features from sample face image data, obtaining a first face feature extraction result. It then classifies the first face feature extraction result to obtain a first face feature classification result. A second face model, already trained, is used to extract features from sample face image data, obtaining a second face feature extraction result. The first and second face feature extraction results are then classified to obtain a second face feature classification result. First loss data is generated based on the first face feature classification result. Second loss data is generated based on the second face feature classification result. Third loss data is generated based on both the first and second face feature classification results. Finally, the model parameters of the first face model are updated based on the first, second, and third loss data to achieve convergence, resulting in a trained first face model. Therefore, on the one hand, the first loss data can be generated based on the first face feature extraction results generated by the first face model (new model) to be trained, thus fully considering the optimization of the distance between the generated new features, improving the training accuracy of the first face model and its compatibility with the second face model, thereby improving the compatibility between the new face feature vector and the old face feature vector; on the other hand, it makes full use of the prior information of the second face feature classification results output by the already trained second face model (old model), without introducing additional parameters, reducing the cost of model training and the consumption of system resources, enabling the model to achieve a better convergence state and better compatibility. At the same time, in order to further optimize compatibility with the old model, the second face feature classification results output by the already trained second face model can be used as... The system generates a second loss data set, thereby simultaneously considering the distance optimization between new features generated by the first face model and the compatibility optimization of the already trained second face model during the model optimization process. This further improves the training accuracy of the first face model and its compatibility with the second face model, and further enhances the compatibility between the new face feature vector and the old face feature vector. On the other hand, based on the compatibility constraints, a third loss data set can be generated based on the classification results of the first and second face features. This allows the features generated by the model to be trained to directly calculate the distance with the features generated by the already trained model, thereby better achieving the compatibility between the first and second face models, and thus better achieving compatible iteration between different versions of content vectors.

[0028] The face image data processing method provided in this application classifies the face image data to be processed based on a trained first face model to obtain a target face feature classification result. This makes the target face feature vector compatible with the old face feature vectors output by the old model, improving the compatibility between the new and old face feature vectors. After the face feature extraction vectorization model is iterated and upgraded, it is no longer necessary to use the trained new first face model to update the face-related vector representations in the existing content library. This effectively avoids the huge equipment and time costs incurred when updating the library when the number of existing features is large, accelerates the development of the entire business and reduces the consumption of system resources, effectively improves the R&D efficiency and iteration speed of face feature vectors, and reduces the cost of updating the face feature vector library. Attached Figure Description

[0029] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram illustrating the implementation environment of a face model generation method according to an exemplary embodiment.

[0031] Figure 2 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 1 .

[0032] Figure 3 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 2 .

[0033] Figure 4 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 3 .

[0034] Figure 5 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 4 .

[0035] Figure 6 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 5 .

[0036] Figure 7 This is a flowchart illustrating a face image data processing method according to an exemplary embodiment.

[0037] Figure 8 This is a block diagram illustrating a face model generation apparatus according to an exemplary embodiment.

[0038] Figure 9 This is a block diagram illustrating a face image data processing apparatus according to an exemplary embodiment.

[0039] Figure 10 This is a hardware structure block diagram of a server provided according to an exemplary embodiment. Detailed Implementation

[0040] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI technology is a multidisciplinary field, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Pre-trained models, also known as large models or foundational models, can be widely applied to downstream tasks across various AI directions after fine-tuning. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0041] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models represent the latest development in deep learning, integrating all of these techniques.

[0042] The following describes the technical terms used in the embodiments of this application:

[0043] Embedding, literally meaning "embedding," is essentially a mapping from semantic space to vector space, while preserving the semantic relationships of the original samples in the vector space as much as possible. For example, two semantically similar words are also likely to be located close to each other in the vector space. Embedding uses a low-dimensional vector to represent an object, which could be a word, an image, an article, or a face, etc. The property of embedding vectors is that vectors with close distances correspond to objects with similar meanings. For example, the distance between embeddings (e.g., "Zhang San's face") and (e.g., "Cartoon Zhang San") will be very close, but the distance between embeddings (e.g., "Zhang San's face") and (e.g., "Li Si's face") will be farther. Embedding represents objects from another space and even reveals the potential relationships between objects. The ability of embedding to encode items with low-dimensional vectors while retaining their meaning is very suitable for deep learning. For example, embedding is a powerful tool for handling sparse features. In recommendation scenarios, there are numerous category and ID-type features. Extensive use of one-hot encoding leads to extremely sparse sample feature vectors. Furthermore, the structural characteristics of deep learning are not conducive to processing sparse feature vectors. Therefore, all deep learning recommendation models typically use embedding layers to transform sparse high-dimensional feature vectors into dense low-dimensional feature vectors. Additionally, embedding can integrate a large amount of valuable information and is itself an extremely important feature vector, such as facial feature vectors.

[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0045] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the present application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0046] Figure 1 This is a schematic diagram illustrating the implementation environment of a face model generation method according to an exemplary embodiment. For example... Figure 1 As shown, the implementation environment may include at least terminal 01 and server 02. The terminal 01 and server 02 may be directly or indirectly connected through wired or wireless communication. This embodiment of the application does not impose any limitations on this.

[0047] Specifically, server 02 can be used to acquire sample face image data and train a first model based on the sample face image data to obtain a trained first face model. Optionally, server 02 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0048] Specifically, the terminal 01 can be used to collect sample facial image data. The terminal 01 may include, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc.

[0049] The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0050] It should be noted that, Figure 1 This is just one example. Other implementation environments may also be included in other scenarios.

[0051] It should be noted that in the specific implementation of this application, user information, such as facial image data and other related data, is involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0052] Figure 2 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 1 This method can be used for Figure 1In the implementation environment described in this specification, the method operation steps are as illustrated in the embodiments or flowcharts. However, based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server product execution, the method can be executed sequentially according to the embodiments or drawings, or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown in the embodiments or drawings... Figure 2 As shown, the method may include:

[0053] S101. Obtain sample face image data, a first face model, and a second face model; the first face model is the model to be trained, and the second face model is the model that has been trained.

[0054] Optionally, the sample face image data can be various types of face image data, including real face image data and non-real face image data. The real face image data can be a real face image of an object, and the non-real face image data can be generated based on the real face image data. The non-real face image data and the real face image data have different image styles.

[0055] Optionally, in terms of image source, the sample face image can be from various publicly available face images, images from open face image datasets, images extracted from video frames of short videos by image matting, images generated by AIGC, or images of different styles obtained by style processing of the original face image.

[0056] For example, real facial image data can be used to generate non-real facial image data through AIGC (Generative Artificial Intelligence), which is a technique that uses artificial intelligence methods based on generative adversarial networks, large pre-trained models, etc., to learn from and recognize existing data and generate relevant content with appropriate generalization capabilities. Alternatively, real facial image data can be stylized to obtain facial images of different styles, such as cartoons, thus obtaining non-real facial image data.

[0057] Optionally, the first face model is a new model to be trained, and the second face model is an existing model that has already been trained. The structures of the first and second face models can be the same or different. For example, both the first and second face models can use a pre-trained ArcFace+resnet50 structure. ArcFace is a face recognition loss function based on cosine distance. The working principle of the ArcFace loss function is: normalizing the feature vector and weights, and adding an angular interval m to the angular space θ. The angular interval has a more direct impact on the angle than the cosine interval. resnet50 is a deep convolutional neural network composed of residual units.

[0058] S103. Based on the first face model, perform face feature extraction on the sample face image data to obtain the first face feature extraction result, and classify the first face feature extraction result to obtain the first face feature classification result; based on the second face model, perform feature extraction on the sample face image data to obtain the second face feature extraction result, and classify the first face feature extraction result and the second face feature extraction result to obtain the second face feature classification result.

[0059] Optionally, sample face image data can be input into a first face model. This first face model extracts face features from the sample face image data, yielding a first face feature extraction result. This first face feature extraction result can be considered a new feature extracted by the new model. Then, the first face feature extraction result is classified using a face classifier in the first face model to obtain a first face feature classification result. This first face feature classification result can be considered the final embedded feature extracted by the new model, i.e., the new face feature vector embedding.

[0060] Optionally, sample face image data can be input into a second face model. This second face model extracts face features from the sample face image data to obtain a second face feature extraction result. This second face feature extraction result can be considered as the old features extracted by the old model. Then, the face classifier in the second face model classifies the first and second face feature extraction results to obtain a second face feature classification result. This second face feature classification result can be considered as the final embedded features extracted by the old model, i.e., the old face feature vector embedding.

[0061] It is evident that in the process of classifying face images using the old second face model, it is necessary to comprehensively consider both the old features extracted by the old model and the new features extracted by the new model. This allows for the generation of second loss data based on the features extracted by the old model, and third loss data based on the features extracted by the old and new models. This enables the simultaneous optimization of the distance between the new features generated by the first face model and the compatibility optimization of the trained second face model during model optimization. This further improves the training accuracy of the first face model and its compatibility with the second face model, as well as the compatibility between the new face feature vector and the old face feature vector. Furthermore, it allows for direct distance calculation between the new features generated by the new model and the old features generated by the old model, thus better achieving compatibility between the first and second face models, and ultimately better achieving compatible iteration between content vectors of different versions.

[0062] S105. Generate first loss data based on the first face feature classification result, generate second loss data based on the second face feature classification result, and generate third loss data based on the first face feature classification result and the second face feature classification result; wherein, the first loss data is the loss caused by increasing the distance between positive sample face image data and negative sample face image data in the sample face image data, the second loss data is the loss caused by compatibility optimization of the second face model, and the third loss data is the loss caused by shortening the distance between the first face feature classification result and the second face feature classification result belonging to the same feature cluster category.

[0063] Optionally, after obtaining the first and second face feature classification results, the first loss data can be generated based on the first face feature extraction results generated by the new model. This fully considers the optimization of the distance between the newly generated features, increasing the distance between positive and negative face image data in the sample face image data, thereby improving the training accuracy of the first face model and its compatibility with the second face model. Simultaneously, to fully utilize the prior information of the second face feature classification results output by the old model without introducing additional parameters, the model training cost is reduced, resulting in better convergence and compatibility. Furthermore, to maintain compatible training between the first and second face feature extraction results, the second loss data can be generated based on the second face feature classification results output by the old model. This allows for simultaneous consideration of distance optimization between the new features generated by the first face model and compatibility optimization of the already trained second face model during model optimization, further improving the training accuracy of the first face model and its compatibility with the second face model. Meanwhile, based on the compatibility constraints, a third loss data is generated based on the classification results of the first and second face features. This makes it easier for the features generated by the model to be trained to directly calculate the distance with the features generated by the already trained model, thereby better achieving the compatibility of the first and second face models, and further achieving better compatibility iteration between different versions of content vectors.

[0064] S107. Update the model parameters of the first face model based on the first loss data, the second loss data, and the third loss data, so that the first face model can converge and obtain the trained first face model.

[0065] Optionally, after obtaining the first loss data, the second loss data, and the third loss data, the model parameters of the second face model can be frozen, and the first face model can be trained based on the first loss data, the second loss data, and the third loss data. During the training process, the model parameters of the first face model are continuously updated until the first face model converges, resulting in a trained first face model. This trained first face model is used to process face image data to obtain the target face feature classification result. This target face feature classification result can be considered as an embedded feature, i.e., the target face feature vector Embedding.

[0066] Therefore, on the one hand, it can generate first loss data based on the first face feature extraction results generated by the first face model to be trained, thereby fully considering the optimization of the distance between the generated new features, improving the training accuracy of the first face model and its compatibility with the second face model; on the other hand, it fully utilizes the prior information of the second face feature classification results output by the already trained second face model (old model), without introducing additional parameters, reducing the cost of model training and the consumption of system resources, enabling the model to achieve a better convergence state and better compatibility. Simultaneously, to further optimize compatibility with the old model, second loss data can be generated based on the second face feature classification results output by the already trained second face model, thereby achieving... During model optimization, the distance optimization between new features generated by the first face model and the compatibility optimization of the trained second face model are considered simultaneously. This further improves the training accuracy of the first face model and its compatibility with the second face model, thereby further improving the compatibility between the new face feature vector and the old face feature vector. On the other hand, based on the compatibility constraints, a third loss data can be generated based on the classification results of the first and second face features. This allows the features generated by the model to be trained to directly calculate the distance with the features generated by the trained model, thereby better achieving the compatibility between the first and second face models, and thus better achieving compatible iteration between different versions of content vectors.

[0067] As facial features continue to iterate, especially with the development of new styles and AIGC (AI-generated content), such as the generation of stylized content like anime and Hanfu (traditional Han clothing) photos, it is necessary to perform similarity retrieval and matching between real facial image data and anime facial image data. Therefore, for actual information flow scenarios, facial image data of different types and styles (including AIGC generation, anime style, Hanfu photo style, etc.) can be introduced as compatible iterative samples. Triples can be constructed by the correspondence between real facial image data and facial image data of different styles of corresponding cartoons. Figure 3 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 2 ,like Figure 3 As shown, in an optional embodiment, in step S101 above, acquiring sample face image data includes:

[0068] S1011. Obtain real face image data of the target object, non-real face image data of the target object, and face image data of other objects; the non-real face image data is generated based on the real face image data, and the non-real face image data and the real face image data have different image styles.

[0069] S1013. Construct anchor sample face image data based on real face image data, construct positive sample face image data based on non-real face image data of the target object, and construct negative sample face image data based on face image data of other objects.

[0070] S1015. Construct triples based on positive sample face image data, negative sample face image data, and anchor point sample face image data to obtain sample face image data.

[0071] Optionally, the target object can be a person, and the real face image of the target object can be publicly available face image data of the target object, or an image obtained from an open face image dataset, or an image extracted from a short video frame through image matting. The non-real face image of the target object can be an image generated by AIGC from the real face image data of the target object, or a cartoon-type image obtained by stylizing the real face image data of the target object. The other object can be a person other than the target object, and the face image data of the other object can be real face image data of the other object, or non-real face image data of the other object. For example, the real face image data of the target object can be the real face image of Zhang San, the non-real face image data of the target object can be the non-real face image data of Zhang San, and the face image data of the other object can be the real face image data of Li Si or the non-real face image data of Li Si.

[0072] Optionally, in steps S1013-S1015 above, triplets can be constructed based on the correspondence between real face image data, non-real face image data, and face image data of other objects. Specifically, anchor sample face image data can be constructed based on real face image data, positive sample face image data can be constructed based on non-real face image data of the target object, and negative sample face image data can be constructed based on face image data of other objects. Anchor sample face image data matches positive sample face image data, and anchor sample face image data does not match negative sample face image data. Then, triplets are constructed based on positive sample face image data, negative sample face image data, and anchor sample face image data to obtain sample face image data. For example, anchor sample face image data is constructed based on Zhang San's real face image data, positive sample face image data is constructed based on Zhang San's non-real face image data, and negative sample face image data is constructed based on Li Si's face image data.

[0073] Therefore, different types and styles of face data, including AIGC-generated face data, can be introduced as compatible and iterative samples for real-world information flow scenarios. By constructing triples through the correspondence between real face image data and face image data of different styles of corresponding cartoons, the trained first face model can perform similarity retrieval and matching of real and non-real face image data in information flow scenarios, thereby improving the efficiency and accuracy of similarity retrieval.

[0074] In other embodiments, the sample image data may not be in the form of a triple. For example, the sample image data may include positive sample image data and negative sample image data. The positive sample image data may be real face image data of the target object, and the negative sample image data may be face image data of other objects.

[0075] It should be noted that step S103 above can be implemented in a variety of ways, and no specific limitation is made.

[0076] Figure 4 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 3 ,like Figure 4 As shown, in an optional embodiment, in step S103 above, the extraction of facial features from the sample facial image data based on the first facial model to obtain a first facial feature extraction result, and the classification of the first facial feature extraction result to obtain a first facial feature classification result, may include:

[0077] S103-1-1. Input the sample face image into the first face model.

[0078] S103-1-3. Based on the first face feature extraction layer in the first face model, face features are extracted from the sample face image to obtain the first face feature extraction result.

[0079] S103-1-5. Based on the first face feature classifier in the first face model, classify the first face feature extraction results to obtain the first positive sample face vector corresponding to the positive sample face image data, the first negative sample face vector corresponding to the negative sample face data, and the first anchor point sample face vector corresponding to the anchor point sample face image data.

[0080] S103-1-7. Determine the first positive sample face vector, the first negative sample face vector, and the first anchor point sample face vector as the first face feature classification result.

[0081] Figure 5 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 4 ,like Figure 5As shown, the first face model may include a first face feature extraction layer, a first hidden feature layer, and a first face feature classifier. Sample face images (including positive sample face image data, negative sample face image data, and anchor point sample face image data) are input into the first face model. Firstly, face features are extracted from the sample face images based on the first face feature extraction layer in the first face model, resulting in the first face feature extraction result, i.e., the new features. The first face feature extraction result includes the face feature extraction results corresponding to the positive sample face image data, the negative sample face image data, and the anchor point sample face image data.

[0082] Next, the hidden vector of the first face feature extraction result is extracted based on the first feature hidden layer in the first face model.

[0083] Next, the hidden vectors extracted based on the first face feature classifier in the first face model are classified to obtain the first positive sample face vector corresponding to the positive sample face image data, the first negative sample face vector corresponding to the negative sample face data, and the first anchor point sample face vector corresponding to the anchor point sample face image data. These first positive sample face vectors, first negative sample face vectors, and first anchor point sample face vectors are the new first face model's preset-dimensional embedding features (e.g., 512-dimensional embedding features) extracted from the sample face image data, i.e., the extracted new face feature vectors.

[0084] Therefore, features can be extracted from triples based on the new model, so that the first loss data can be generated based on the first face feature extraction results generated by the first face model. This fully considers the optimization of the distance between the generated new features, making the distance between the anchor point and the positive sample features as small as possible, while increasing the distance between the anchor point and the negative sample features. This improves the training accuracy of the first face model and its compatibility with the second face model, thereby improving the compatibility between the face feature vector and the old face feature vector.

[0085] In other embodiments, if the sample image data is not in the form of triples, for example, if the sample image data includes positive sample image data and negative sample image data, the positive sample image data and negative sample image data can be input into the first face model, and the first face model can process the positive sample image data and negative sample image data to obtain the vectors corresponding to the positive sample image data and negative sample image data respectively, thereby obtaining the first face feature classification result.

[0086] Figure 6 This is a flowchart illustrating a face model generation method according to an exemplary embodiment. Figure 5 ,like Figure 6 As shown, in an optional embodiment, in step S103 above, the above-mentioned feature extraction of sample face image data based on the second face model to obtain a second face feature extraction result, and the classification of the first face feature extraction result and the second face feature extraction result to obtain a second face feature classification result, may include:

[0087] S103-3. Based on the second face model, feature extraction is performed on the sample face image data to obtain the second face feature extraction result, and feature mapping processing is performed on the second face feature extraction result to obtain the mapped feature, and the first face feature extraction result and the mapped feature are classified to obtain the second face feature classification result.

[0088] Optionally, forward-compatible training can be understood as mapping learning at the feature level. Sample face image data can be input into a second face model for old feature extraction, resulting in the second face feature extraction result. The old features generated by the old model (i.e., the second face feature extraction result) are then used for feature mapping to obtain mapped features. These mapped features are then forward-compatiblely trained with the new features from the new model (i.e., the first face feature extraction result) to obtain the second face feature classification result, which is the old face feature vector embedding output by the old model. The advantage of forward compatibility is that compatible training does not affect the training of the new model, thus ensuring that the new model is not limited by the feature distribution of the old model during training.

[0089] In an optional embodiment, continue as follows Figure 6 As shown, in the above steps S103-3, the above-mentioned feature extraction of sample face image data based on the second face model to obtain the second face feature extraction result, and feature mapping processing of the second face feature extraction result to obtain the mapped feature, and classification of the first face feature extraction result and the mapped feature to obtain the second face feature classification result, may include:

[0090] S103-3-1. Input the sample face image into the second face model.

[0091] S103-3-3. Based on the second face feature extraction layer in the second face model, face features are extracted from the sample face image to obtain the second face feature extraction result.

[0092] S103-3-5. Based on the feature mapping layer in the second face model, the feature extraction results of the second face are processed by feature mapping to obtain the mapped features.

[0093] S103-3-7. Based on the second face feature classifier in the second face model, classify the first face feature extraction results and mapping features to obtain the second positive sample face vector corresponding to the positive sample face image data, the second negative sample face vector corresponding to the negative sample face data, and the second anchor point sample face vector corresponding to the anchor point sample face image data.

[0094] S103-3-9. Determine the second positive sample face vector, the second negative sample face vector, and the second anchor sample face vector as the second face feature classification result.

[0095] Optionally, continue as follows Figure 5 As shown, the first face model may include a second face feature extraction layer, a second face feature classifier, and a feature mapping layer. Sample face images (including positive sample face image data, negative sample face image data, and anchor point sample face image data) can be input into the second face model. First, face features are extracted from the sample face images based on the second face feature extraction layer in the second face model, resulting in the second face feature extraction result, i.e., the old features. The second face feature extraction result includes the face feature extraction results corresponding to the positive sample face image data, the negative sample face image data, and the anchor point sample face image data.

[0096] Next, based on the feature mapping layer in the second face model, the feature extraction results of the second face are processed by feature mapping to obtain mapped features. These mapped features include the mapped features corresponding to positive sample face image data, the mapped features corresponding to negative sample face image data, and the mapped features corresponding to anchor point sample face image data.

[0097] Next, based on the second face feature classifier in the second face model, the extraction results and mapping features of the first face feature are classified to obtain the second positive sample face vector corresponding to the positive sample face image data, the second negative sample face vector corresponding to the negative sample face data, and the second anchor point sample face vector corresponding to the anchor point sample face image data. These second positive sample face vectors, second negative sample face vectors, and second anchor point sample face vectors are the preset-dimensional embedding features extracted from the sample face image data by the old second face model, i.e., the extracted old face feature vector embedding.

[0098] Therefore, the old features generated by the old model (i.e., the second face feature extraction result) can be mapped to mapped features through a feature mapping layer. These mapped features and the new features from the new model (i.e., the first face feature extraction result) can then be forward-compatiblely trained, thus not affecting the training of the new model and preventing it from being limited by the feature distribution of the old model during training. Furthermore, during the face image classification process using the old second face model, both the old features extracted by the old model and the new features extracted by the new model need to be considered comprehensively. This allows for the generation of second loss data based on the features extracted by the old model, as well as the generation of second loss data based on the features extracted by the old model. The features extracted by the new model and the features of the first face model are used to generate a third loss data, thereby simultaneously considering the distance optimization between the new features generated by the first face model and the compatibility optimization of the trained second face model during the model optimization process. This further improves the training accuracy of the first face model and its compatibility with the second face model, thereby improving the compatibility between the new face feature vector and the old face feature vector. It also enables the new features generated by the new model to directly calculate the distance with the old features generated by the old model, thereby better realizing the compatibility between the first face model and the second face model, and thus better realizing the compatibility iteration between different versions of content vectors.

[0099] In an optional embodiment, the feature mapping layer includes an initial feature mapping layer and a transformed feature mapping layer. The transformed feature mapping layer is obtained by initializing the initial feature mapping layer, and the initial feature mapping layer and the transformed feature mapping layer share weights. In this embodiment, the transformed feature mapping layer is initialized based on the initial feature mapping layer. This initial feature mapping layer can be considered an adapter in the old second face model. By adjusting and fine-tuning this adapter, it is possible to avoid adjusting the old second face model for each task, reducing the number of parameters in the old second face model, thereby improving the efficiency of using the parameters of the old second face model and ultimately improving the training efficiency of the entire training process.

[0100] In an optional embodiment, in step S103-3-5 above, the feature mapping processing of the second face feature extraction result based on the feature mapping layer in the second face model to obtain mapped features may include:

[0101] The output vector is processed by the initial feature mapping layer to obtain the first initial mapping feature; the output vector is the vector output by the second face model during the face feature extraction process of the sample face image; and the second initial mapping feature is obtained by processing the second face feature extraction result by the transformation feature mapping layer.

[0102] The mapping features are obtained by fusing the first initial mapping features and the second initial mapping features.

[0103] Optionally, in the process of extracting facial features from sample facial images, the second face model can output a set of vectors in addition to the second face feature extraction result. These output vectors can be considered as a set of vectors that are not directly related to the second face feature extraction result and are output through the prior information of the old model.

[0104] In this embodiment, the output vector can first be input into the initial feature mapping layer for feature mapping processing to obtain the first initial mapping feature, which can be considered as the feature generated after adapter adjustment; at the same time, the second face feature extraction result extracted by the old model is input into the transform feature mapping layer for feature mapping processing to obtain the second initial mapping feature; finally, the first initial mapping feature and the second initial mapping feature are fused and connected to obtain the mapping feature.

[0105] Therefore, a new transform feature mapping layer can be obtained by initializing the adapter in the old model. The features are mapped by the adapter and the new transform feature mapping layer. The mapping results of the adapter and the new transform feature mapping layer are fused to obtain the mapped features. In this way, by adjusting and fine-tuning the adapter during the mapping process, it is possible to avoid adjusting the old second face model for each task, reduce the number of parameters of the old second face model, improve the efficiency of using the parameters of the old second face model, and thus improve the training efficiency of the entire training process.

[0106] In an optional embodiment, the transform feature mapping layer includes a first fully connected layer, a first non-linear connected layer, and a second fully connected layer. The above-described feature mapping processing of the second face feature extraction result based on the transform feature mapping layer to obtain the second initial mapped features includes:

[0107] The second face feature extraction result is input into the transform feature mapping layer.

[0108] The first fully connected feature is obtained by performing fully connected processing on the second face feature extraction result based on the first fully connected layer.

[0109] The first nonlinearly processed feature is obtained by performing nonlinear mapping on the first fully connected feature based on the first nonlinear connection layer.

[0110] The second fully connected feature is obtained by performing fully connected processing on the first nonlinear processing feature based on the second fully connected layer.

[0111] The second initial mapping feature is obtained by fusing the second face feature extraction result and the second fully connected feature.

[0112] Optionally, the transformed feature mapping layer can be considered as a detachable adaptive classification center corrector based on boundary constraints. The mapping process based on the transformed feature mapping layer can be considered as using a detachable adaptive classification center corrector based on boundary constraints to assist in achieving forward compatibility and correcting boundary classification. The transformed feature mapping layer includes a first fully connected layer, a first nonlinear connected layer, and a second fully connected layer. The first nonlinear connected layer is located between the first and second fully connected layers. Furthermore, the transformed feature mapping layer can be detached when the second face model is used, without introducing additional parameters, reducing computational costs, and ultimately enabling the model to achieve better convergence and better compatibility.

[0113] Continue as Figure 5 As shown, during the mapping process, the second face feature extraction result can be input into the transform feature mapping layer. The first fully connected layer then performs fully connected processing on the previously extracted second face feature result, essentially combining all previously extracted features to obtain the first fully connected feature. Next, a first non-linear connected layer performs non-linear mapping processing on the first fully connected feature to obtain the first non-linear processed feature. Then, a second fully connected layer performs fully connected processing on the previously extracted first non-linear processed feature, combining all previously extracted features to obtain the second fully connected feature. Finally, a residual structure connects and fuses the second face feature extraction result and the second fully connected feature to obtain the second initial mapped feature. The residual structure introduces skip connections, adding the input of the previous layer to the output of the current layer, making gradient propagation easier within the network.

[0114] Therefore, a detachable, boundary-constraint-based adaptive classification center corrector is used to assist in achieving forward compatibility and boundary classification correction. This fully utilizes the prior information of the old model's embedding, while the residual structure effectively suppresses overfitting, which is prone to occur during the embedding correction process. Specifically, a non-linear connection is added between the two fully connected layers, and the transformed feature mapping layer can be detached when the second face model is used, without introducing additional parameters, reducing computational costs, and ultimately enabling the model to achieve better convergence and compatibility. This embodiment uses the classification center of the face vector feature classifier for correction, achieving a better state within the original boundary constraints. The residual structure effectively solves overfitting while retaining the embedding information of the old model.

[0115] In an optional embodiment, the initial feature mapping layer includes a third fully connected layer, a second non-linear connected layer, and a fourth fully connected layer. The above-described feature mapping processing of the output vector based on the initial feature mapping layer to obtain the first initial mapped features includes:

[0116] The output vector is input into the initial feature mapping layer.

[0117] The third fully connected feature is obtained by performing fully connected processing on the second face feature extraction result based on the third fully connected layer.

[0118] The second nonlinear processing feature is obtained by performing nonlinear mapping on the third fully connected feature based on the second nonlinear connection layer.

[0119] The fourth fully connected feature is obtained by performing fully connected processing on the second nonlinear processing feature based on the fourth fully connected layer.

[0120] The first initial mapping feature is obtained by fusing the output vector and the fourth fully connected feature.

[0121] Optionally, continue as follows Figure 5 As shown, this initial feature mapping layer can also be considered as a detachable adaptive classification center corrector based on boundary constraints. The mapping process based on the initial feature mapping can be considered as using a detachable adaptive classification center corrector based on boundary constraints to assist in achieving forward compatibility and correcting boundary classification. This initial feature mapping includes a third fully connected layer, a second non-linear connected layer, and a fourth fully connected layer. The second non-linear connected layer is located between the third and fourth fully connected layers. Furthermore, this initial feature mapping layer can be detached when the second face model is referenced, without introducing additional parameters, reducing computational costs, and ultimately enabling the model to achieve a better convergence state and better compatibility.

[0122] During the mapping process, the output vector can be input into the initial feature mapping layer. This third fully connected layer performs a fully connected processing on the previously extracted output vector, essentially combining all previously extracted features to obtain the third fully connected feature. Next, a second non-linear connected layer performs a non-linear mapping on the third fully connected feature to obtain the second non-linear processed feature. Then, a fourth fully connected layer performs a fully connected processing on the previously extracted second non-linear processed feature, combining all previously extracted features to obtain the fourth fully connected feature. Finally, a residual structure is used to concatenate and fuse the output vector and the second fully connected feature to obtain the first initial mapped feature.

[0123] Therefore, a detachable, boundary-constraint-based adaptive classification center corrector is used to assist in achieving forward compatibility and boundary classification correction. This fully utilizes the prior information of the old model's embedding, while the residual structure effectively suppresses overfitting, which is prone to occur during the embedding correction process. Specifically, a non-linear connection is added between the two fully connected layers, and the transformed feature mapping layer can be detached when the second face model is used, without introducing additional parameters, reducing computational costs, and ultimately enabling the model to achieve better convergence and compatibility. This embodiment uses the classification center of the face vector feature classifier for correction, achieving a better state within the original boundary constraints. The residual structure effectively solves overfitting while retaining the embedding information of the old model.

[0124] In an optional embodiment, in step S105 above, generating the first loss data based on the first face feature classification result includes:

[0125] Determine the first difference between the first anchor point sample face vector and the first positive sample face vector, and determine the second difference between the first anchor point sample face vector and the first negative sample face vector.

[0126] First loss data is generated based on the first difference and the second difference.

[0127] In this embodiment, a first difference can be obtained by calculating the distance between the first anchor point sample face vector and the first positive sample face vector, and a second difference can be obtained by calculating the distance between the first anchor point sample face vector and the first negative sample face vector. The difference between the first difference and the second difference is calculated, and the sum of the difference and a preset interval function is calculated. Based on the magnitude of the sum and zero, the first loss data is calculated. For example, the distance can refer to Euclidean distance.

[0128] For example, the calculation process for the first loss data is as follows:

[0129] L(A,P,N)=max(||f A -f P || 2 -||f A -f N || 2 +τ,0);

[0130] Where L(A, P, N) refers to the first loss data, f A This refers to the face vector of the first anchor point sample, f P This refers to the first positive sample face vector, f N This refers to the first positive sample face vector, ||*||^2 is the distance calculation, and τ refers to the preset interval function.

[0131] The ultimate requirement for model compatibility training in this embodiment is that, after inputting a new triplet into the model, the distance between the first positive sample face vector generated by the new model and the second positive sample face vector generated by the old model is less than the distance between pairs of positive sample face feature vectors generated by the old model; and after inputting a new triplet into the model, the distance between the first negative sample face vector generated by the new model and the second negative sample face vector generated by the old model is greater than the distance between pairs of negative sample face feature vectors generated by the old model. Based on this, to meet the requirements of compatibility training, a first loss data is specifically introduced for the triplet, thereby fully considering the optimization of the distance between the new features generated by the new model, minimizing the distance between the anchor point and the positive sample features, while increasing the distance between the anchor point and the negative sample features, improving the training accuracy of the first face model and its compatibility with the second face model, and thus improving the compatibility between the face feature vectors and the old face feature vectors.

[0132] In an optional embodiment, in step S105 above, the second loss data determined as generated by compatibility training on the first and second face feature extraction results includes:

[0133] Get the preset interval function;

[0134] Determine the first similarity between the second positive sample face vector and the second anchor point sample face vector, and determine the second similarity between the second negative sample face vector and the second anchor point sample face vector;

[0135] The second loss data is generated based on the preset interval function, the first similarity, and the second similarity.

[0136] In this embodiment, to ensure model training compatibility, a second loss data can be introduced. This second loss data is calculated based on the form of Memory Bank + InfoNCE. Unlike the triplet-based training method, this Memory Bank + InfoNCE incorporates features from the old model (i.e., the second face feature classification result) as part of the loss calculation. This allows for simultaneous consideration of distance optimization between triplet samples and compatibility optimization with the old model during model optimization, thereby further improving the training accuracy of the first face model and its compatibility with the second face model. During training, the parameters of the old model are fixed and do not participate in gradient backpropagation; therefore, the second loss data can be considered a regularization constraint term.

[0137] Optionally, a preset interval function can be obtained, and the product between the second positive sample face vector and the second anchor sample face vector can be calculated. This product represents the first similarity between the second positive sample face vector and the second anchor sample face vector. Simultaneously, the product between the second negative sample face vector and the second anchor sample face vector can be calculated, and this product represents the second similarity between the second negative sample face vector and the second anchor sample face vector. Then, based on the preset interval function, the first similarity, and the second similarity, the loss generated by compatibility optimization of the second face model is determined, resulting in the second loss data.

[0138] The Memory Bank is a queue that acts as a negative sample pool, storing a fixed number of feature vectors (e.g., 1024 feature vectors) and adhering to a first-in, first-out (FIFO) logic. For each sample face image data, the features generated by the old model (i.e., the second negative sample face vector in the second face feature classification result) enter the Memory Bank, while one historical feature is discarded to maintain a consistent total number of features in the Memory Bank. The formula for calculating the second loss data can be as follows:

[0139]

[0140] Among them, L q This refers to the second loss data, where q is the face vector of the second anchor point sample, ki is the face vector of the second negative sample in the i-th memorybank, and τ refers to the preset interval function, k + q refers to the second positive sample face vector, the product of q and ki represents the similarity between the vectors, and K refers to the number of feature vectors stored in the memory bank queue.

[0141] In a specific embodiment, based on the above formula, generating the second loss data according to the preset interval function, the first similarity, and the second similarity may include:

[0142] Calculate the exponent of the quotient of the first similarity and the preset interval function to obtain the first exponent. Calculate the exponent of the quotient of each second negative sample face vector in the memory bank and the preset interval function to obtain the second exponent for each second negative sample face vector. Add the second exponents for each second negative sample face vector to obtain the third exponent. Calculate the quotient of the first and third exponents, and generate the second loss data based on the logarithm of this quotient.

[0143] Increasing the distance between negative samples reduces the overall loss, and decreasing the distance between two features generated by the new and old models at the anchor samples also reduces the overall loss. Ultimately, the compatibility of the new and old models is demonstrated during the loss optimization process. Furthermore, since the second loss data is calculated based on the features of the entire memory bank, and this second loss data is used to update the model parameters of the first model, it is equivalent to calculating the position of the new model's features within the old model's feature distribution based on the features of the entire memory bank. This improves training stability, the training accuracy of the first face model, and its compatibility with the second face model, thereby enhancing the compatibility between the face feature vectors and the old face feature vectors.

[0144] In an optional embodiment, in step S105 above, generating the third loss data based on the first face feature classification result and the second face feature classification result includes:

[0145] Clustering is performed on the second face feature classification results to obtain feature cluster categories.

[0146] A third loss data is generated based on the similarity between the first face feature classification result and the feature clustering category.

[0147] Optionally, to better achieve compatibility, adaptive boundary constraints can be added on top of the compatibility constraints. The main purpose is to effectively suppress the influence of bad cases from older models. The loss data involved in this adaptive boundary constraint is the compatible loss. It should be noted that the compatible loss constraint is not added to each embedding, but rather to the cluster centers of the feature clustering categories. Therefore, the second face feature classification results can be clustered first to obtain feature clustering categories and determine the cluster centers of these categories. Then, the similarity between the first face feature classification results and the feature clustering categories is calculated. Based on this similarity, a third loss data is generated, ensuring that during model convergence, the first face feature classification results belonging to the feature clustering category (i.e., the first face feature classification results whose distance to the cluster center of the feature clustering category is less than a preset distance) are concentrated in the cluster centers of the feature clustering category, and each embedding (including the first and second face feature classification results) has sufficient free space.

[0148] For example, the formula for calculating the compatible loss can be as follows:

[0149]

[0150]

[0151] in, This refers to the cluster center of the feature cluster category, x i This refers to sample face image data, y i This refers to the predicted vector results (including the first face feature classification result and the second face feature classification result). This refers to the logarithmic loss sum of the distances between the old face feature vector generated by the second face model (i.e., the second face feature classification result) and the new face feature vector generated by the first face model (the first face feature classification result), where N refers to the number of sample face image data. It refers to the set of clusters of old face feature vectors generated by the second face model. This refers to the log loss between the new facial feature vector extracted by the first face model and the vector of a certain category processed by the second face model.

[0152] In an optional embodiment, in step S107 above, updating the model parameters of the first face model based on the first loss data, the second loss data, and the third loss data includes:

[0153] Loss data analysis is performed on the first loss data, the second loss data, and the third loss data to obtain the target loss data.

[0154] Freeze the model parameters of the second face model and update the model parameters of the first face model based on the target loss data.

[0155] Optionally, loss data analysis can be performed on the first loss data, second loss data, and third loss data. For example, the target loss data can be obtained by summing the first loss data, second loss data, and third loss data. Alternatively, based on the importance of the first loss data, second loss data, and third loss data to the entire model training process, their respective first weights, second weights, and third weights can be assigned. The first product of the first loss data and the first weight is calculated, the second product of the second loss data and the second weight is calculated, and the third product of the third loss data and the third weight is calculated. The sum of the first product, the second product, and the third product is then calculated to obtain the target loss data.

[0156] Optionally, during training, the model parameters of the second face model can be frozen, meaning the parameters of the old second face model are fixed and do not participate in gradient backpropagation, while the model parameters of the first face model are updated based on the target loss data. Therefore, the first loss data can be generated based on the first face feature extraction result generated by the first face model to be trained, thus fully considering the optimization of the distance between the newly generated features and improving the training accuracy of the first face model and its compatibility with the second face model. At the same time, in order to maintain the compatibility training between the first face feature extraction result and the second face feature extraction result, the second loss data is generated based on the second face feature classification result output by the trained second face model. This realizes that the distance optimization between the newly generated features of the first face model and the compatibility optimization of the trained second face model are considered simultaneously during the model optimization process, thereby further improving the training accuracy of the first face model and its compatibility with the second face model. In addition, based on the compatibility constraints, the third loss data is generated based on the first face feature classification result and the second face feature classification result, so that the features generated by the model to be trained can directly calculate the distance with the features generated by the trained model, thereby better realizing the compatibility of the first face model and the second face model, and thus better realizing the compatibility iteration between different versions of content vectors.

[0157] This application also provides a method for processing facial image data. Figure 7 This is a flowchart illustrating a face image data processing method according to an exemplary embodiment, such as... Figure 7 As shown, the method includes:

[0158] S201. Obtain the face image data to be processed.

[0159] S203. Based on the trained first face model, classify the face image data to be processed to obtain the target face feature classification result; wherein, the trained first face model is generated by the face model generation method based on any of the above embodiments.

[0160] In this embodiment, after the first face model is trained, the face image data to be processed can be input into the trained first face model. The target face feature extraction layer in the trained first face model performs face feature extraction processing on the face image data to be processed, obtaining the target face feature extraction result. Based on the target face classifier in the trained first face model, the target face feature extraction result is classified to obtain the target face feature classification result. This target face feature classification result can be considered as an embedded feature, i.e., a target face feature vector. This target face feature vector is compatible with the old face feature vector output by the old model, improving the compatibility between the new face feature vector and the old face feature vector.

[0161] Optionally, similar to the sample face image data, the face image data to be processed can be in the form of triplets.

[0162] The following is a general description of the above training and application processes:

[0163] 1) Provides a first face model to be trained (i.e., the old model) and a second face model that has already been trained (i.e., the new model). The first face model may include a first face feature extraction layer and a first face feature classifier. The second face model includes a second face feature extraction layer, a feature mapping layer, and a second face feature classifier. The feature mapping layer includes an initial feature mapping layer and a transformed feature mapping layer. The transformed feature mapping layer is obtained by initializing the initial feature mapping layer, and the initial feature mapping layer and the transformed feature mapping layer share weights. The transformed feature mapping layer includes a first fully connected layer, a first non-linear connected layer, and a second fully connected layer. The initial feature mapping layer includes a third fully connected layer, a second non-linear connected layer, and a fourth fully connected layer.

[0164] 2) Obtain sample image data: Obtain real face image data of the target object, non-real face image data of the target object, and face image data of other objects; the non-real face image data is generated based on the real face image data, and the non-real face image data and the real face image data have different image styles; construct anchor sample face image data based on the real face image data, construct positive sample face image data based on the non-real face image data of the target object, and construct negative sample face image data based on the face image data of other objects; construct triples based on the positive sample face image data, negative sample face image data, and anchor sample face image data to obtain sample face image data.

[0165] 3) Input positive sample face image data, negative sample face image data, and anchor point sample face image data into the first face model. Extract face features from the sample face images based on the first face feature extraction layer in the first face model to obtain the first face feature extraction result. Classify the first face feature extraction result based on the first face feature classifier in the first face model to obtain the first positive sample face vector corresponding to the positive sample face image data, the first negative sample face vector corresponding to the negative sample face data, and the first anchor point sample face vector corresponding to the anchor point sample face image data. Determine the first positive sample face vector, the first negative sample face vector, and the first anchor point sample face vector as the first face feature classification result.

[0166] 4) Input positive sample face image data, negative sample face image data, and anchor point sample face image data into the second face model. Based on the second face feature extraction layer in the second face model, extract face features from the sample face images to obtain the second face feature extraction results. Based on the feature mapping layer in the second face model, perform feature mapping processing on the second face feature extraction results to obtain mapped features. Based on the second face feature classifier in the second face model, classify the first face feature extraction results and the mapped features to obtain the second positive sample face vector corresponding to the positive sample face image data, the second negative sample face vector corresponding to the negative sample face data, and the second anchor point sample face vector corresponding to the anchor point sample face image data. Determine the second positive sample face vector, the second negative sample face vector, and the second anchor point sample face vector as the second face feature classification results.

[0167] Specifically, the feature extraction results of the second face are processed by feature mapping based on the feature mapping layer in the second face model to obtain mapped features, including:

[0168] The output vector is processed by the initial feature mapping layer to obtain the first initial mapping feature; the output vector is the vector output by the second face model during the face feature extraction process of the sample face image; and the second initial mapping feature is obtained by processing the second face feature extraction result by the transformation feature mapping layer; the first initial mapping feature and the second initial mapping feature are fused to obtain the mapping feature.

[0169] The process involves performing feature mapping on the output vector based on an initial feature mapping layer to obtain a first initial mapping feature, including: inputting the second face feature extraction result into a transform feature mapping layer; performing fully connected processing on the second face feature extraction result based on a first fully connected layer to obtain a first fully connected feature; performing nonlinear mapping processing on the first fully connected feature based on a first nonlinear connected layer to obtain a first nonlinear processed feature; performing fully connected processing on the first nonlinear processed feature based on a second fully connected layer to obtain a second fully connected feature; and fusing the second face feature extraction result and the second fully connected feature to obtain a second initial mapping feature.

[0170] Specifically, the first initial mapping feature is obtained by performing feature mapping processing on the output vector based on the initial feature mapping layer, including: inputting the output vector into the initial feature mapping layer; performing fully connected processing on the output vector based on the third fully connected layer to obtain the third fully connected feature; performing nonlinear mapping processing on the third fully connected feature based on the second nonlinear connected layer to obtain the second nonlinear processed feature; performing fully connected processing on the second nonlinear processed feature based on the fourth fully connected layer to obtain the fourth fully connected feature; and fusing the output vector and the fourth fully connected feature to obtain the first initial mapping feature.

[0171] 5) Determine the first difference between the first anchor point sample face vector and the first positive sample face vector, and determine the second difference between the first anchor point sample face vector and the first negative sample face vector; generate first loss data based on the first and second differences. Obtain a preset interval function; determine the first similarity between the second positive sample face vector and the second anchor point sample face vector, and determine the second similarity between the second negative sample face vector and the second anchor sample face vector; generate second loss data based on the preset interval function, the first similarity, and the second similarity. Cluster the second face feature classification results to obtain feature cluster categories; generate third loss data based on the similarity between the first face feature classification results and the feature cluster categories.

[0172] 6) Perform loss data analysis on the first loss data, second loss data, and third loss data to obtain the target loss data; freeze the model parameters of the second face model, and update the model parameters of the first face model according to the target loss data so that the first face model can converge and obtain the trained first face model.

[0173] Implementing the embodiments of this application has the following beneficial effects:

[0174] 1) It can effectively support the compatibility iteration of face feature vectors. After the face feature vector model is upgraded, it is no longer necessary to use the newly trained first face model to refresh the face-related vector representations in the existing content library. This effectively avoids the huge equipment and time costs caused by refreshing the library when the amount of existing features is large, accelerates the development of the entire business and the consumption of system resources, effectively improves the R&D efficiency and iteration speed of face feature vector embedding, and reduces the cost of refreshing the face feature embedding library.

[0175] 2) It can effectively improve the matching and compatibility processing of face feature vector Embeeding between real face image data and non-real face image data (AIGC generated or image-processed cartoon faces). For example, the generation of many stylized contents such as anime and Hanfu photos requires similarity retrieval and matching between real face image data and non-real face image data, thus accelerating the application and implementation of related businesses.

[0176] 3) It enables the iteration and upgrading of the face feature vector Embedding model to be coordinated with the development process of the business itself. The new version of the face feature vector Embedding model can be dynamically upgraded by continuously adding new sample face image data, realizing the dynamic evolution of the model and effectively supporting various applications involving faces in video business, such as casting, face aggregation content dispersal, viewing only the content and clips of a certain character, and advertising, etc.

[0177] The face image data processing method provided in this application embodiment can be applied in at least the following scenarios:

[0178] 1) Application in role selection: During the role selection process for actors, video frames from programs the actor has previously participated in can be input into a trained first-face model. The first-face model can then be used to filter out segments from the programs the actor has previously participated in, reducing the screening cost and facilitating role selection.

[0179] 2) Advertising placement: For example, if an advertising spokesperson appears in a film or television drama, the advertising clips are usually placed in the clips where the advertising spokesperson appears, so that the two are naturally connected. To achieve this function, video frames from a film or television drama can be input into the trained first-person face model. The trained first-person face model can then filter out the clips where the advertising spokesperson appears, so that the clips where the advertising spokesperson appears can be naturally connected with the clips of the advertisement.

[0180] 3) "Watch Only TA" function: This function allows users to watch only the clips of a specific character in a specific episode of a TV series. In this case, video frames from a specific episode of a TV series can be input into the trained first face model, and the "clips of a specific character" can be extracted through the trained first face model.

[0181] 4) Retrieval of face images of different styles: When generating non-real face image data through AIGC, it is necessary to compare and retrieve real face image data and non-real face image data. At this time, real face image data and non-real face image data can be input into the trained first face model, and the comparison and retrieval functions can be realized through the trained first face model.

[0182] 5) Face Clustering and Dispersing: In recommendation scenarios, there may be situations where the same face appears in a series of videos. This phenomenon will affect the user's video viewing experience. Therefore, it is necessary to identify the faces in the video using a pre-trained first face model, add facial features, and disperse the videos with the same face to improve the diversity of recommendations.

[0183] Figure 8 This is a block diagram illustrating a face model generation apparatus according to an exemplary embodiment, such as... Figure 8 As shown, the face model generation device includes:

[0184] The data acquisition module 301 is used to acquire sample face image data, a first face model, and a second face model; the first face model is a model to be trained, and the second face model is a model that has been trained.

[0185] The extraction and classification module 303 is used to extract facial features from the sample facial image data based on the first facial model to obtain a first facial feature extraction result, and to classify the first facial feature extraction result to obtain a first facial feature classification result; to extract features from the sample facial image data based on the second facial model to obtain a second facial feature extraction result, and to classify the first facial feature extraction result and the second facial feature extraction result to obtain a second facial feature classification result;

[0186] The loss generation module 305 is used to generate first loss data based on the first face feature classification result, generate second loss data based on the second face feature classification result, and generate third loss data based on the first face feature classification result and the second face feature classification result.

[0187] The update module 307 is used to update the model parameters of the first face model according to the first loss data, the second loss data, and the third loss data, so that the first face model converges and a trained first face model is obtained; wherein, the first loss data is the loss caused by increasing the distance between positive sample face image data and negative sample face image data in the sample face image data, the second loss data is the loss caused by compatibility optimization of the second face model, and the third loss data is the loss caused by shortening the distance between the first face feature classification result and the second face feature classification result belonging to the same feature cluster category.

[0188] In an optional embodiment, the data acquisition module includes:

[0189] An image data acquisition unit is used to acquire real face image data of a target object, non-real face image data of the target object, and face image data of other objects; the non-real face image data is generated based on the real face image data, and the non-real face image data and the real face image data have different image styles;

[0190] The construction unit is used to construct anchor sample face image data based on the real face image data, construct positive sample face image data based on the non-real face image data of the target object, and construct negative sample face image data based on the face image data of other objects.

[0191] The triple generation unit is used to construct triples based on the positive sample face image data, the negative sample face image data, and the anchor sample face image data to obtain the sample face image data.

[0192] In an optional embodiment, the sample face image data includes positive sample face image data, negative sample face image data, and anchor point sample face image data; the extraction and classification module includes:

[0193] The input unit is used to input the sample face image into the first face model;

[0194] A face feature extraction unit is used to extract face features from the sample face image based on the first face feature extraction layer in the first face model, and obtain the first face feature extraction result;

[0195] The classification unit is used to classify the first face feature extraction result based on the first face feature classifier in the first face model to obtain the first positive sample face vector corresponding to the positive sample face image data, the first negative sample face vector corresponding to the negative sample face data, and the first anchor point sample face vector corresponding to the anchor point sample face image data.

[0196] The first face feature classification result generation unit is used to determine the first positive sample face vector, the first negative sample face vector, and the first anchor point sample face vector as the first face feature classification result.

[0197] In an optional embodiment, the loss generation module includes:

[0198] The difference determination unit is used to determine a first difference between the first anchor point sample face vector and the first positive sample face vector, and to determine a second difference between the first anchor point sample face vector and the first negative sample face vector.

[0199] The first loss generation unit is used to generate the first loss data based on the first difference and the second difference.

[0200] In an optional embodiment, the extraction and classification module includes:

[0201] The second face feature classification result generation unit is used to extract features from the sample face image data based on the second face model to obtain a second face feature extraction result, perform feature mapping processing on the second face feature extraction result to obtain a mapped feature, and classify the first face feature extraction result and the mapped feature to obtain a second face feature classification result.

[0202] In an optional embodiment, the sample face image data includes positive sample face image data, negative sample face image data, and anchor point sample face image data; the second face feature classification result generation unit includes:

[0203] An input subunit is used to input the sample face image into the second face model;

[0204] A face feature extraction subunit is used to extract face features from the sample face image based on the second face feature extraction layer in the second face model, and obtain the second face feature extraction result;

[0205] The mapping subunit is used to perform feature mapping processing on the feature extraction result of the second face based on the feature mapping layer in the second face model to obtain the mapped features;

[0206] The classification subunit is used to classify the first face feature extraction result and the mapping feature based on the second face feature classifier in the second face model, and obtain the second positive sample face vector corresponding to the positive sample face image data, the second negative sample face vector corresponding to the negative sample face data, and the second anchor point sample face vector corresponding to the anchor point sample face image data.

[0207] The vector determination subunit is used to determine the second positive sample face vector, the second negative sample face vector, and the second anchor point sample face vector as the second face feature classification result.

[0208] In an optional embodiment, the feature mapping layer includes an initial feature mapping layer and a transformed feature mapping layer, wherein the transformed feature mapping layer is obtained by initializing the initial feature mapping layer, and the initial feature mapping layer and the transformed feature mapping layer share weights; the mapping subunit includes:

[0209] The mapping feature generation subunit is used to perform feature mapping processing on the output vector based on the initial feature mapping layer to obtain a first initial mapping feature; the output vector is the vector output by the second face model during the face feature extraction process of the sample face image; and to perform feature mapping processing on the second face feature extraction result based on the transform feature mapping layer to obtain a second initial mapping feature;

[0210] A fusion subunit is used to fuse the first initial mapping feature and the second initial mapping feature to obtain the mapping feature.

[0211] In an optional embodiment, the transform feature mapping layer includes a first fully connected layer, a first nonlinear connected layer, and a second fully connected layer; the mapping feature generation subunit includes:

[0212] The mapping layer input subunit is used to input the second face feature extraction result into the transform feature mapping layer;

[0213] The first fully connected feature generation subunit is used to perform fully connected processing on the second face feature extraction result based on the first fully connected layer to obtain the first fully connected feature;

[0214] The first nonlinear processing feature generation subunit is used to perform nonlinear mapping processing on the first fully connected feature based on the first nonlinear connection layer to obtain the first nonlinear processing feature.

[0215] The second fully connected feature generation subunit is used to perform fully connected processing on the first nonlinear processing feature based on the second fully connected layer to obtain the second fully connected feature.

[0216] The second initial mapping feature generation subunit is used to fuse the second face feature extraction result and the second fully connected feature to obtain the second initial mapping feature.

[0217] In an optional embodiment, the initial feature mapping layer includes a third fully connected layer, a second non-linear connected layer, and a fourth fully connected layer; the mapping feature generation subunit includes:

[0218] An output vector input subunit is used to input the output vector into the initial feature mapping layer;

[0219] The third fully connected feature generation subunit is used to perform fully connected processing on the output vector based on the third fully connected layer to obtain the third fully connected feature.

[0220] The second nonlinear processing feature generation subunit is used to perform nonlinear mapping processing on the third fully connected feature based on the second nonlinear connection layer to obtain the second nonlinear processing feature.

[0221] The fourth fully connected feature generation subunit is used to perform fully connected processing on the second nonlinear processing feature based on the fourth fully connected layer to obtain the fourth fully connected feature.

[0222] The first initial mapping feature generation subunit is used to fuse the output vector and the fourth fully connected feature to obtain the first initial mapping feature.

[0223] In an optional embodiment, the loss generation module includes:

[0224] The preset interval function acquisition unit is used to acquire the preset interval function;

[0225] The similarity determination unit is used to determine the first similarity between the second positive sample face vector and the second anchor sample face vector, and to determine the second similarity between the second negative sample face vector and the second anchor sample face vector.

[0226] The second loss data generation unit is used to generate the second loss data based on the preset interval function, the first similarity, and the second similarity.

[0227] In an optional embodiment, the loss generation module includes:

[0228] Clustering unit, used to cluster the second face feature classification results to obtain feature cluster categories;

[0229] The third loss data generation unit is used to generate the third loss data based on the similarity between the first face feature classification result and the feature clustering category.

[0230] In an optional embodiment, the update module includes:

[0231] The analysis unit is used to perform loss data analysis on the first loss data, the second loss data, and the third loss data to obtain target loss data;

[0232] The freeze and update unit is used to freeze the model parameters of the second face model and update the model parameters of the first face model according to the target loss data.

[0233] Figure 9 This is a block diagram illustrating a face image data processing apparatus according to an exemplary embodiment, such as... Figure 9 As shown, the face image data processing device includes:

[0234] The face image data acquisition module 401 is used to acquire face image data to be processed;

[0235] The target face feature classification result generation module 403 is used to classify the face image data to be processed based on the trained first face model to obtain the target face feature classification result; wherein, the trained first face model is generated based on the face model generation method described in any of the above embodiments.

[0236] It should be noted that the device embodiments provided in this application are based on the same inventive concept as the method embodiments described above.

[0237] This application also provides an electronic device for generating a face model. The electronic device includes a processor and a memory. The memory stores at least one instruction or at least one program. The processor loads and executes the at least one instruction or at least one program to implement the face model generation method provided in any of the above embodiments.

[0238] This application also provides an electronic device for processing facial image data. The electronic device includes a processor and a memory. The memory stores at least one instruction or at least one program. The processor loads and executes the at least one instruction or at least one program to implement the facial image data processing method provided in any of the above embodiments.

[0239] Embodiments of this application also provide a computer-readable storage medium that can be disposed in a terminal to store at least one instruction or at least one program for implementing a face model generation method in the method embodiments. The at least one instruction or at least one program is loaded and executed by a processor to implement the face model generation method or face image data processing method provided in the above method embodiments.

[0240] Optionally, in the embodiments of this specification, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0241] The memory described in this specification can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for functions, etc.; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0242] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the face model generation method provided in the above-described method embodiments.

[0243] The face model generation method provided in this application can be executed on a terminal, computer terminal, server, or similar computing device. Taking running on a server as an example... Figure 10This is a hardware structure block diagram of a server according to an exemplary embodiment. For example... Figure 10 As shown, the server 500 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 510 (CPUs 510 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 530 for storing data, and one or more storage media 520 (e.g., one or more mass storage devices) for storing application programs 523 or data 522. The memory 530 and storage media 520 may be temporary or persistent storage. The program stored in the storage media 520 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 510 may be configured to communicate with the storage media 520 and execute the series of instruction operations stored in the storage media 520 on the server 500. Server 500 may also include one or more power supplies 560, one or more wired or wireless network interfaces 550, one or more input / output interfaces 540, and / or one or more operating systems 521, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0244] The input / output interface 540 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 500. In one example, the input / output interface 540 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 540 may be a radio frequency (RF) module for wireless communication with the Internet.

[0245] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 500 may also include... Figure 10 The more or fewer components shown, or having the same Figure 10 The different configurations shown.

[0246] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0247] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and server embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0248] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0249] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating a face model, characterized in that, The method includes: Acquire sample face image data, a first face model, and a second face model; the first face model is the model to be trained, and the second face model is the model that has already been trained; Based on the first face model, facial features are extracted from the sample facial image data to obtain a first facial feature extraction result, and the first facial feature extraction result is classified to obtain a first facial feature classification result; based on the second face model, features are extracted from the sample facial image data to obtain a second facial feature extraction result, and the first facial feature extraction result and the second facial feature extraction result are classified to obtain a second facial feature classification result; First loss data is generated based on the first face feature classification result, second loss data is generated based on the second face feature classification result, and third loss data is generated based on the first face feature classification result and the second face feature classification result. The model parameters of the first face model are updated based on the first loss data, the second loss data, and the third loss data, so that the first face model can converge and the trained first face model is obtained. Wherein, the first loss data is the loss caused by increasing the distance between positive sample face image data and negative sample face image data in the sample face image data, the second loss data is the loss caused by compatibility optimization of the second face model, and the third loss data is the loss caused by shortening the distance between the first face feature classification result and the second face feature classification result belonging to the same feature cluster category.

2. The face model generation method according to claim 1, characterized in that, The acquisition of sample face image data includes: Acquire real face image data of the target object, non-real face image data of the target object, and face image data of other objects; the non-real face image data is generated based on the real face image data, and the non-real face image data and the real face image data have different image styles; Anchor point sample face image data is constructed based on the real face image data, positive sample face image data is constructed based on the non-real face image data of the target object, and negative sample face image data is constructed based on the face image data of other objects. The sample face image data is obtained by constructing a triplet based on the positive sample face image data, the negative sample face image data, and the anchor point sample face image data.

3. The face model generation method according to claim 1, characterized in that, The sample face image data includes positive sample face image data, negative sample face image data, and anchor point sample face image data; the step of extracting face features from the sample face image data based on the first face model to obtain a first face feature extraction result, and classifying the first face feature extraction result to obtain a first face feature classification result, includes: The sample face image is input into the first face model; The sample face image is subjected to face feature extraction based on the first face feature extraction layer in the first face model to obtain the first face feature extraction result; The first face feature extraction result is classified based on the first face feature classifier in the first face model to obtain the first positive sample face vector corresponding to the positive sample face image data, the first negative sample face vector corresponding to the negative sample face data, and the first anchor point sample face vector corresponding to the anchor point sample face image data. The first positive sample face vector, the first negative sample face vector, and the first anchor point sample face vector are determined as the first face feature classification result.

4. The face model generation method according to claim 3, characterized in that, The step of generating first loss data based on the first facial feature classification result includes: Determine a first difference between the first anchor point sample face vector and the first positive sample face vector, and determine a second difference between the first anchor point sample face vector and the first negative sample face vector; The first loss data is generated based on the first difference and the second difference.

5. The face model generation method according to claim 1, characterized in that, The step of extracting features from the sample face image data based on the second face model to obtain a second face feature extraction result, and classifying the first face feature extraction result and the second face feature extraction result to obtain a second face feature classification result, includes: Based on the second face model, feature extraction is performed on the sample face image data to obtain the second face feature extraction result. The second face feature extraction result is then subjected to feature mapping processing to obtain the mapped feature. Finally, the first face feature extraction result and the mapped feature are classified to obtain the second face feature classification result.

6. The face model generation method according to claim 5, characterized in that, The sample face image data includes positive sample face image data, negative sample face image data, and anchor point sample face image data; the step of extracting features from the sample face image data based on the second face model to obtain a second face feature extraction result, performing feature mapping processing on the second face feature extraction result to obtain mapped features, and classifying the first face feature extraction result and the mapped features to obtain a second face feature classification result, includes: The sample face image is input into the second face model; The sample face image is subjected to face feature extraction based on the second face feature extraction layer in the second face model to obtain the second face feature extraction result; The feature extraction results of the second face are processed by feature mapping based on the feature mapping layer in the second face model to obtain the mapped features; Based on the second face feature classifier in the second face model, the first face feature extraction result and the mapping feature are classified to obtain the second positive sample face vector corresponding to the positive sample face image data, the second negative sample face vector corresponding to the negative sample face data, and the second anchor point sample face vector corresponding to the anchor point sample face image data. The second positive sample face vector, the second negative sample face vector, and the second anchor point sample face vector are determined as the second face feature classification result.

7. The face model generation method according to claim 6, characterized in that, The feature mapping layer includes an initial feature mapping layer and a transformed feature mapping layer. The transformed feature mapping layer is obtained by initializing the initial feature mapping layer, and the initial feature mapping layer and the transformed feature mapping layer share weights. The feature mapping processing of the second face feature extraction result based on the feature mapping layer in the second face model to obtain the mapped features includes: The output vector is processed by feature mapping based on the initial feature mapping layer to obtain the first initial mapping feature; the output vector is the vector output by the second face model during the face feature extraction process of the sample face image; And based on the transformation feature mapping layer, the second face feature extraction result is subjected to feature mapping processing to obtain the second initial mapping feature; The mapping feature is obtained by fusing the first initial mapping feature and the second initial mapping feature.

8. The face model generation method according to claim 7, characterized in that, The transform feature mapping layer includes a first fully connected layer, a first non-linear connected layer, and a second fully connected layer; the feature mapping processing of the second face feature extraction result based on the transform feature mapping layer to obtain the second initial mapped features includes: The result of the second facial feature extraction is input into the transform feature mapping layer; The second face feature extraction result is processed by a fully connected layer based on the first fully connected layer to obtain the first fully connected feature. Based on the first nonlinear connection layer, the first fully connected feature is subjected to nonlinear mapping processing to obtain the first nonlinear processed feature; Based on the second fully connected layer, the first nonlinear processing feature is fully connected to obtain the second fully connected feature; The second initial mapping feature is obtained by fusing the second face feature extraction result and the second fully connected feature.

9. The face model generation method according to claim 7, characterized in that, The initial feature mapping layer includes a third fully connected layer, a second non-linear connected layer, and a fourth fully connected layer; the feature mapping processing of the output vector based on the initial feature mapping layer to obtain the first initial mapped features includes: The output vector is input into the initial feature mapping layer; The output vector is processed by the third fully connected layer to obtain the third fully connected feature. Based on the second nonlinear connection layer, the third fully connected feature is subjected to nonlinear mapping processing to obtain the second nonlinear processed feature; Based on the fourth fully connected layer, the second nonlinear processing feature is fully connected to obtain the fourth fully connected feature. The first initial mapping feature is obtained by fusing the output vector and the fourth fully connected feature.

10. The face model generation method according to claim 6, characterized in that, The step of generating second loss data based on the second facial feature classification result includes: Get the preset interval function; Determine the first similarity between the second positive sample face vector and the second anchor sample face vector, and determine the second similarity between the second negative sample face vector and the second anchor sample face vector; The second loss data is generated based on the preset interval function, the first similarity, and the second similarity.

11. The face model generation method according to claim 1, characterized in that, The step of generating third loss data based on the first face feature classification result and the second face feature classification result includes: Cluster the second facial feature classification results to obtain feature cluster categories; The third loss data is generated based on the similarity between the first facial feature classification result and the feature clustering category.

12. The face model generation method according to any one of claims 1 to 12, characterized in that, The step of updating the model parameters of the first face model based on the first loss data, the second loss data, and the third loss data includes: Loss data analysis is performed on the first loss data, the second loss data, and the third loss data to obtain the target loss data; Freeze the model parameters of the second face model, and update the model parameters of the first face model according to the target loss data.

13. A method for processing facial image data, characterized in that, The method includes: Acquire the face image data to be processed; The target face feature classification result is obtained by classifying the face image data to be processed based on the first face model that has been trained; The first face model that has been trained is generated based on the face model generation method according to any one of claims 1 to 12.

14. A face model generation device, characterized in that, The device includes: The data acquisition module is used to acquire sample face image data, a first face model, and a second face model; the first face model is a model to be trained, and the second face model is a model that has already been trained. The extraction and classification module is used to extract facial features from the sample facial image data based on the first facial model to obtain a first facial feature extraction result, and to classify the first facial feature extraction result to obtain a first facial feature classification result; and to extract features from the sample facial image data based on the second facial model to obtain a second facial feature extraction result, and to classify the first facial feature extraction result and the second facial feature extraction result to obtain a second facial feature classification result. The loss generation module is used to generate first loss data based on the first face feature classification result, generate second loss data based on the second face feature classification result, and generate third loss data based on the first face feature classification result and the second face feature classification result. An update module is used to update the model parameters of the first face model according to the first loss data, the second loss data, and the third loss data, so that the first face model converges and a trained first face model is obtained; wherein, the first loss data is the loss caused by increasing the distance between positive sample face image data and negative sample face image data in the sample face image data, the second loss data is the loss caused by compatibility optimization of the second face model, and the third loss data is the loss caused by shortening the distance between the first face feature classification result and the second face feature classification result belonging to the same feature cluster category.

15. A facial image data processing device, characterized in that, The device includes: The module for acquiring face image data to be processed is used to acquire face image data to be processed. The target face feature classification result generation module is used to classify the face image data to be processed based on the trained first face model to obtain the target face feature classification result; The first face model that has been trained is generated based on the face model generation method according to any one of claims 1 to 12.