Cartoon Face Recognition Model Training Method, Related Devices and Storage Media

By using the feature distribution of normal face images to align cartoon face features in the cartoon face recognition model training, the problem of scarce training data of cartoon face recognition model is solved, and the accuracy of cartoon face recognition is improved.

CN115359528BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210968466.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2025-07-18
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

Due to the scarcity of training data, the existing cartoon face recognition model has low recognition accuracy and cannot meet actual business needs.

Method used

By obtaining the feature differences between the sample normal face image and the sample cartoon face image, the pre-trained face recognition model is used for feature extraction, and the target loss function value is determined based on the difference, and the model parameters are adjusted to realize the training of the cartoon face recognition model.

Benefits of technology

The recognition accuracy of the cartoon face recognition model is improved. By aligning the feature distribution between cartoon faces and normal faces in the feature space, the classification ability of normal faces is used to improve the recognition ability of cartoon faces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359528B_ABST
    Figure CN115359528B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training a cartoon face recognition model, related devices, and a storage medium. The method includes: obtaining a first sample image set, which includes sample normal face images and sample cartoon face images; respectively extracting features from the sample normal face images and the sample cartoon face images in the first sample image set based on a pre-trained face recognition model to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images; determining a target loss function value based on the difference between the normal face features and the cartoon face features; adjusting the model parameters of the face recognition model based on the target loss function value until a preset training end condition is satisfied to obtain a cartoon face recognition model; and the cartoon face recognition model is used to recognize cartoon face images. The present invention greatly improves the accuracy of cartoon face recognition of the cartoon face recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method for training a cartoon face recognition model, related devices, and a storage medium. Background Art

[0002] With the enhancement of brand protection awareness, the demand for auditing a large amount of cartoon data on the Internet is increasing day by day. Compared with manual auditing, a cartoon face recognition model based on machine learning can greatly improve the auditing efficiency of a large amount of cartoon data.

[0003] The cartoon face recognition model in the related art is obtained by performing face recognition training based on cartoon face data. However, due to the difficulty of collecting and annotating cartoon face data, the cartoon face data used for training the cartoon face recognition model is scarce, resulting in a low accuracy of cartoon face recognition of the trained cartoon face recognition model and being unable to meet the requirements in actual business. Summary of the Invention

[0004] To solve the problems of the prior art, embodiments of the present invention provide a method for training a cartoon face recognition model, related devices, and a storage medium. The technical solutions are as follows:

[0005] On the one hand, a method for training a cartoon face recognition model is provided. The method includes:

[0006] Obtain a first sample image set; the first sample image set includes sample normal face images and sample cartoon face images, and the ratio of the first quantity of the sample normal face images in the first sample image set to the second quantity of the sample cartoon face images in the first sample image set satisfies a preset condition;

[0007] Based on a pre-trained face recognition model, respectively perform feature extraction on the sample normal face images and the sample cartoon face images in the first sample image set to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images;

[0008] Based on the difference between the normal face features and the cartoon face features, determine a target loss function value;

[0009] Based on the target loss function value, adjust the model parameters of the face recognition model until a preset training end condition is satisfied to obtain a cartoon face recognition model; the cartoon face recognition model is used to recognize cartoon face images.

[0010] On the other hand, a device for training a cartoon face recognition model is provided. The device includes:

[0011] The first sample image acquisition module is used to acquire a first sample image set; the first sample image set includes sample normal face images and sample cartoon face images, and the ratio of the first quantity of the sample normal face images in the first sample image set to the second quantity of the sample cartoon face images in the first sample image set satisfies a preset condition;

[0012] The feature extraction module is used to respectively extract features from the sample normal face images and sample cartoon face images in the first sample image set based on a pre-trained face recognition model, and obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images;

[0013] The target loss determination module is used to determine a target loss function value based on the difference between the normal face features and the cartoon face features;

[0014] The parameter adjustment module is used to adjust the model parameters of the face recognition model based on the target loss function value until a preset training end condition is met to obtain a cartoon face recognition model; the cartoon face recognition model is used to recognize cartoon face images.

[0015] In an exemplary embodiment, the target loss determination module includes:

[0016] The first loss determination module is used to determine a first loss function value based on the difference between the normal face features and the cartoon face features;

[0017] The second loss determination module is used to determine a first face recognition result based on the normal face features, and determine a second loss function value based on the difference between the first face recognition result and first reference face information; the first reference face information is the face information annotated in the sample normal face image corresponding to the normal face features;

[0018] The third loss determination module is used to determine a second face recognition result based on the cartoon face features, and determine a third loss function value based on the difference between the second face recognition result and second reference face information; the second reference face information is the cartoon face information annotated in the sample cartoon face image corresponding to the cartoon face features;

[0019] The loss weighting module is used to perform weighted summation on the first loss function value, the second loss function value, and the third loss function value to obtain the target loss function value.

[0020] In an exemplary embodiment, the first loss determination module includes:

[0021] The first statistic matrix determination module is configured to determine a first statistic matrix based on the normal face features corresponding to the normal face images of each sample in the first sample image set;

[0022] The second statistic matrix determination module is configured to determine a second statistic matrix based on the cartoon face features corresponding to the cartoon face images of each sample in the first sample image set;

[0023] The statistic difference matrix determination module is configured to obtain a statistic difference matrix based on the difference between the first statistic matrix and the second statistic matrix;

[0024] The first loss determination sub-module is configured to determine the first loss function value based on the statistic differences in the statistic difference matrix.

[0025] In an exemplary embodiment, the feature extraction module includes:

[0026] The face recognition module is configured to input the normal face images and cartoon face images of the samples in the first sample image set into the face recognition model respectively for face recognition;

[0027] The intermediate feature extraction module is configured to obtain an intermediate feature map from the target access point of the face recognition model during the face recognition process, and obtain the normal face features corresponding to the normal face images of the samples and the cartoon face features corresponding to the cartoon face images of the samples;

[0028] Wherein, the target access point is any one of a plurality of preset access points, and the plurality of preset access points include the intermediate feature extraction layer of the face recognition model.

[0029] In an exemplary embodiment, the parameter adjustment module includes:

[0030] The parameter adjustment sub-module is configured to adjust the model parameters of the face recognition model based on the target loss function values corresponding to each of the preset access points until a preset training end condition is met, and obtain the target face recognition models corresponding to each of the preset access points;

[0031] The cartoon face recognition module is configured to perform face recognition on the sample cartoon face images respectively based on the target face recognition models corresponding to each of the preset access points, and obtain the cartoon face recognition results corresponding to each of the preset access points;

[0032] The model selection module is configured to select the cartoon face recognition model from the plurality of target face recognition models based on the cartoon face recognition results corresponding to each of the preset access points; wherein, the cartoon face recognition result of the cartoon face recognition model is better than the cartoon face recognition results of the remaining target face recognition models.

[0033] In an exemplary embodiment, the device further comprises:

[0034] A second sample image acquisition module, configured to acquire a second set of sample images; the second set of sample images includes sample normal face images;

[0035] A pre-training module, configured to perform face recognition training on a convolutional neural network based on the second set of sample images to obtain the face recognition model.

[0036] On the other hand, there is provided an electronic device, comprising a processor and a memory, where at least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the cartoon face recognition model training method in any of the above aspects.

[0037] On the other hand, there is provided a computer-readable storage medium, where at least one instruction or at least one segment of program is stored in the computer-readable storage medium, and the at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the cartoon face recognition model training method as described in any of the above aspects.

[0038] On the other hand, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the cartoon face recognition model training method in any of the above aspects.

[0039] In the embodiments of the present invention, by using sample normal face images (i.e., real face training data) to augment the cartoon face training data, and starting from the features extracted by the pre-trained face recognition model, based on the differences between the features corresponding to the real face training data and the features corresponding to the cartoon face training data, the target loss function value for fine-tuning the face recognition model is determined, so that the face recognition model can learn the common feature distribution between the cartoon face images and the normal face images, realizing the alignment of the feature space of the cartoon face images and the feature space of the normal face images, and using the classification ability of the normal face images in the feature space to improve the classification ability of the cartoon face images in the feature space, thereby greatly improving the cartoon face recognition accuracy of the fine-tuned face recognition model (i.e., the cartoon face recognition model). Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0041] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0042] Figure 2 is a schematic flowchart of a method for training a cartoon face recognition model provided by an embodiment of the present invention;

[0043] Figure 3a is a schematic flowchart of another method for training a cartoon face recognition model provided by an embodiment of the present invention;

[0044] Figure 3b is a schematic diagram of the feature distributions of cartoon face features and normal face features before and after fine-tuning provided by an embodiment of the present invention;

[0045] Figure 4 is a structural block diagram of a device for training a cartoon face recognition model provided by an embodiment of the present invention;

[0046] Figure 5 is a hardware structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0049] It is understandable that in the specific embodiments of the present application, data related to user information and the like are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0050] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an embodiment of the present invention. The implementation environment includes a terminal 110 and a server 120. Among them, the terminal 110 and the server 120 can be connected and communicate through a wired or wireless network.

[0051] The terminal 110 includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. A client software with image processing functions, such as an application program (abbreviated as App), is installed in the terminal 110. The application program can be an independent application program or a subroutine in an application program. In the embodiments of the present invention, the images processed by the image processing function may include cartoon face images.

[0052] The server 120 can provide background services for the application program in the terminal 110. Specifically, the background service can be a cartoon face recognition service for cartoon face images. Specifically, the server 120 can call a pre-trained cartoon face recognition model to perform cartoon face recognition on the cartoon face image, and then obtain a cartoon face recognition result, and return the cartoon face recognition result to the client 110.

[0053] Among them, the server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0054] In an exemplary implementation manner, both the terminal 110 and the server 120 can be node devices in a blockchain system, and can share the information obtained and generated with other node devices in the blockchain system to achieve information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain. The blockchain is composed of multiple blocks, and the adjacent blocks before and after have an association relationship, so that when the data in any block is tampered with, it can be detected by the next block, thereby avoiding the data in the blockchain from being tampered with and ensuring the security and reliability of the data in the blockchain.

[0055] Embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0056] Among them, Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0057] Among them, Computer Vision (CV) is a science that studies how to make machines "see". Further speaking, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition and measurement, and further perform graphics processing to make the computer process the images into a form more suitable for human eyes to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0058] The training of the cartoon face recognition model in the embodiments of the present invention is implemented based on machine learning. Machine Learning (ML) is a multi-disciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0059] Please refer to Figure 2 , which shows a schematic flowchart of a method for training a cartoon face recognition model provided by an embodiment of the present invention. This method can be applied to an electronic device, which can be a terminal or a server. It should be noted that this specification provides method operation steps as described in the embodiments or flowcharts, but based on routine or non-creative labor, there may be more or fewer operation steps. The order of steps listed in the embodiments is only one way among many execution orders of steps and does not represent the only execution order. In actual execution of a system or product, it can be executed in the order of the embodiments or as shown in the drawings, either sequentially or in parallel (for example, in an environment with parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method may include:

[0060] S201, obtain a first sample image set, which includes sample normal face images and sample cartoon face images.

[0061] Among them, the ratio of the first quantity of sample normal face images to the second quantity of sample cartoon face images in the first sample image set satisfies a preset condition.

[0062] Among them, a sample normal face image refers to a real face image that has been annotated. The annotation information of the sample normal face image is called the first reference face information in the embodiments of the present invention, which indicates the face information annotated in the sample normal face image.

[0063] A sample cartoon face image refers to a cartoon face image that has been annotated. The annotation information of the sample cartoon face image is called the second reference face information in the embodiments of the present invention, which indicates the cartoon face information annotated in the sample cartoon face image.

[0064] Among them, the preset condition can be that the ratio of the first quantity to the second quantity is a preset threshold value, which can be set according to actual experience. Generally, the closer the number of sample normal face images and sample cartoon face images in the first sample image set is, the more conducive it is to improving the accuracy of the subsequent fine-tuned cartoon face recognition model for cartoon face recognition. Based on this, in a possible design, the preset threshold value can be within the range of [0.8, 1.0]. For example, if the preset threshold value is 1.0, then the preset condition is that the ratio of the first quantity to the second quantity is 1.0, that is, the number of sample normal face images and sample cartoon face images in the first sample image set is equal.

[0065] S203. Respectively perform feature extraction on the sample normal face images and sample cartoon face images in the first sample image set based on the pre-trained face recognition model to obtain the normal face features corresponding to the sample normal face images and the cartoon face features corresponding to the sample cartoon face images.

[0066] Among them, the pre-trained face recognition model refers to a convolutional neural network model obtained by training for the face recognition task using a second sample image set, and the second sample image set includes sample normal face images. Based on this, in an exemplary embodiment, before the above step S203, the method further includes the process of training for the face recognition task to obtain the pre-trained face recognition model. Specifically, this process can include the following steps:

[0067] Obtain a second sample image set; the second sample image set includes sample normal face images;

[0068] Perform face recognition training on the convolutional neural network based on the second sample image set to obtain the face recognition model.

[0069] Among them, the second sample image set includes sample normal face images, that is, the second sample image set is mainly real face images, and the second sample image set can be from an open-source data set used to implement the real face recognition task.

[0070] The processing of the input sample normal face image by a Convolutional Neural Network (CNN) can include operations such as convolution calculation, non-linear activation function (Relu) calculation, pooling calculation, etc. to extract the face features in the input image. Then, based on the extracted face features, face recognition is performed to obtain the face recognition result, and the pre-training loss value is determined by using the pre-training loss function in combination with the difference between the face recognition result and the annotation information in the input image. Based on this pre-training loss value, the network parameters of the convolutional neural network are adjusted, and the iterative training is continued based on the adjusted network parameters until the pre-training end condition is met. The convolutional neural network corresponding to the network parameters at the end of pre-training is used as the pre-trained face recognition model in the embodiments of the present invention.

[0071] Among them, the pre-training end condition can be that the number of pre-training iterations reaches the set pre-training iteration number threshold, or the pre-training loss value is less than the set pre-training loss threshold. The pre-training loss function can be a classification function, such as Softmax, various Softmax types with margin, etc.

[0072] The face features can include but are not limited to at least one of the following: face shape feature, nose feature, eye feature, eyebrow feature, mouth feature, ear feature, and hair feature. In a specific implementation, the face features can be characterized by the key points of the face.

[0073] Among them, when adjusting the network parameters of the convolutional neural network based on the pre-training loss value, it can be carried out in the way of gradient descent. The specific gradient descent method can be stochastic gradient method, stochastic gradient method with momentum term, Adagard (Adaptive gradient), or Adam (Adaptive moment estimation), etc. The embodiments of the present invention do not make specific limitations on the gradient descent method.

[0074] In the above implementation manner, by pre-training the face recognition model based on the second sample image set including the sample normal face image, since the sample normal face image is easily obtained and has a large data volume in practical applications, the sufficiency of the training data during pre-training can be ensured. Furthermore, it can be ensured that the feature distribution of the normal face obtained based on the pre-trained face recognition model has the characteristics of a large distance between categories and a small distance within categories, with a large feature discrimination degree, and thus strong recognizability.

[0075] In the embodiment of the present invention, after the pre-training is completed to obtain the pre-trained face recognition model, the face recognition model is used to extract features from the sample normal face images and sample cartoon face images in the first sample image set respectively, so as to obtain the normal face features corresponding to the sample normal face images and the cartoon face features corresponding to the sample cartoon face images.

[0076] S205. Determine the target loss function value based on the difference between the normal face features and the cartoon face features.

[0077] S207. Adjust the model parameters of the face recognition model based on the target loss function value until the preset training end condition is met to obtain the cartoon face recognition model.

[0078] Among them, the cartoon face recognition model is used to perform cartoon face recognition on cartoon face images.

[0079] Among them, the preset training end condition can be that the number of iterations reaches the preset iteration number threshold, or the target loss function value is less than the preset loss function value.

[0080] Among them, the adjustment of the model parameters of the face recognition model can adopt the gradient descent method. The specific gradient descent method can be the stochastic gradient method, the stochastic gradient method with momentum term, Adagard (Adaptive gradient), or Adam (Adaptive moment estimation), etc.

[0081] In the embodiment of the present invention, the target loss function value is determined based on the difference between the normal face features and the cartoon face features, that is, it can represent the relevant alignment loss between the normal face features and the cartoon face features. By adjusting the model parameters of the face recognition model, the relevant alignment loss of the two features is constrained and optimized. When the optimization meets the preset training end condition, the feature distribution of the normal face features extracted by the face recognition model is consistent with the feature distribution of the cartoon face features, that is, the alignment of the normal face features and the cartoon face features is realized, thereby prompting the face recognition model to learn the common feature distribution between the cartoon face images and the normal face images.

[0082] In specific implementation, the relevant alignment loss between the normal face features and the cartoon face features can be measured by the statistical distance of the features. The goal of constraint optimization is that the features extracted by the face recognition model have almost the same statistics, so that the fine-tuned face recognition model can learn the common feature distribution between the normal face images and the cartoon face images.

[0083] Based on this, in an exemplary embodiment, such as Figure 3aAs shown, when implementing the above step S205, it may include:

[0084] S301, determine a first loss function value based on the difference between the normal face feature and the cartoon face feature.

[0085] S303, determine a first face recognition result based on the normal face feature, and determine a second loss function value based on the difference between the first face recognition result and the first reference face information.

[0086] Wherein, the first reference face information is the face information annotated in the sample normal face image corresponding to the normal face feature.

[0087] S305, determine a second face recognition result based on the cartoon face feature, and determine a third loss function value based on the difference between the second face recognition result and the second reference face information.

[0088] S307, perform weighted summation on the first loss function value, the second loss function value, and the third loss function value to obtain a target loss function value.

[0089] Wherein, the second reference face information is the cartoon face information annotated in the sample cartoon face image corresponding to the cartoon face feature.

[0090] Specifically, the normal face feature and the cartoon face feature may come from the feature map output by the intermediate feature extraction layer of the face recognition model. The first face recognition result is the result obtained by mapping the normal face feature output by the intermediate feature extraction layer through the fully connected layer of the face recognition model. Similarly, the second face recognition result is also the result obtained by mapping the cartoon face feature output by the intermediate feature extraction layer through the fully connected layer of the face recognition model.

[0091] Wherein, the first loss function value is used to measure the distance between the feature distribution of the normal face feature and the feature distribution of the cartoon face feature. The distance can be calculated using the statistical quantity of the feature (such as the second-order statistical quantity). When the statistical quantity corresponding to the normal face feature extracted by the face recognition model is consistent with the statistical quantity corresponding to the cartoon face feature extracted, the generalization of the normal face feature to the space of the cartoon face feature can be achieved.

[0092] Both the second loss function value and the third loss function value can be calculated using a classification function, such as Softmax, Softmax of various plus margin types, etc.

[0093] In a specific implementation manner, when implementing the above step S301, it may include the following steps:

[0094] Determine the first statistic matrix based on the normal face features corresponding to each sample normal face image in the first sample image set;

[0095] Determine the second statistic matrix based on the cartoon face features corresponding to each sample cartoon face image in the first sample image set;

[0096] Obtain the statistic difference matrix based on the difference between the first statistic matrix and the second statistic matrix;

[0097] Determine the value of the first loss function based on the statistic differences in the statistic difference matrix.

[0098] Among them, the statistics involved in the first statistic matrix and the second statistic matrix may include variance, covariance, etc. The value of the first loss function can be calculated based on the multi-order norm of the statistic difference, such as the second-order norm.

[0099] Taking the statistic involved in the statistic matrix as covariance and the value of the first loss function as the second-order norm of the statistic difference as an example, the first statistic matrix (hereinafter referred to as the first covariance matrix) can be expressed by the following formula (1):

[0100]

[0101] where C n represents the first covariance matrix; n n represents the number of sample normal face images in the first sample image set; D n represents the normal face feature; 1 represents the unit matrix.

[0102] Similarly, the second statistic matrix (hereinafter referred to as the second covariance matrix) can be expressed by the following formula (2):

[0103]

[0104] where C c represents the second covariance matrix; n c represents the number of sample cartoon face images in the first sample image set; D c represents the cartoon face feature; 1 represents the unit matrix.

[0105] Then the covariance difference matrix can be expressed as: C n -C c , and the elements (i.e., covariance differences) in this covariance difference matrix are the differences between the elements in the first covariance matrix and the second covariance matrix.

[0106] Then the value of the first loss function can be expressed by the following formula (3):

[0107]

[0108] Among them, F3 represents the first loss function value; d represents the dimension of the feature; ||·|| F represents the F-norm, which is defined as the square root of the sum of the squares of each element in the matrix. Then represents the square of the F-norm.

[0109] In the above embodiment, the corresponding statistic matrices are determined based on the normal face features and the cartoon face features respectively, and then the first loss function value is determined based on the difference between the statistic matrices, which is conducive to the subsequent alignment of the normal face features and the cartoon face features.

[0110] Exemplarily, in the above step S307, the target loss function value can be expressed by the following formula (4):

[0111] L total = αF1 + βF2 + γF3 (4)

[0112] Among them, L total represents the target loss function value; F1 represents the second loss function value; F2 represents the third loss function value; F3 represents the first loss function value; α, β, and γ represent the weights of the corresponding loss function values. In specific implementations, the weights can be set according to actual experience. For example, in some designs, α, β, and γ can all take the value of 1.

[0113] In the above embodiment, the face recognition model is trained through the total loss function value, that is, the target loss function value, to fine-tune the pre-trained face recognition model, so that the face recognition model can learn the common distribution information of the normal face images and the cartoon face images in the joint training. Therefore, the characteristic distribution of the normal face with the characteristics of a large distance between categories, a small distance within categories, a large feature discrimination degree, and strong recognizability can be utilized to transfer the recognition ability of the normal face to the cartoon face recognition, improving the accuracy of the fine-tuned face recognition model (i.e., the cartoon face recognition model) for cartoon face recognition.

[0114] As can be seen from the above technical solutions of the embodiments of the present invention, the embodiments of the present invention do not need to modify the original face recognition network. The cartoon face features and the real face features of the face recognition model are aligned in the feature space. By adding relevant alignment loss constraints on the normal face features and the cartoon face features during the fine-tuning process, the lightweight of the training process is realized, which not only reduces the training cost but also improves the training efficiency.

[0115] In addition, after aligning the feature space of the cartoon face images with the feature space of the normal face images through relevant alignment loss constraints, the classification ability of the cartoon face images in the feature space is enhanced by using the classification ability of the normal face images in the feature space, so that the recognition accuracy of the cartoon face recognition model obtained by training is greatly improved.

[0116] As Figure 3b shown in the schematic diagram of the feature distributions of the cartoon face features and the normal face features before and after fine-tuning according to the embodiment of the present invention. The feature distribution of the cartoon face features without being adjusted by the embodiment of the present invention and the feature distribution of the normal face features (such as Figure 3b the left figure in), where the distance between categories of the feature distribution of the cartoon face features is small and the distinguishability is small due to insufficient training data, so the recognition accuracy is low; after being adjusted by the embodiment of the present invention, the cartoon face features are aligned with the normal face feature space. Since the amount of data of the normal face features is large, the characteristics of large inter-class distance and small intra-class distance can be ensured. The inter-class distance affects the false rejection rate, and the intra-class distance affects the accuracy. Thus, the recognition ability of the normal face after adjustment is transferred to the cartoon face, thereby improving the recognition accuracy of the cartoon face.

[0117] It can be seen that the cartoon face recognition model obtained by using the embodiment of the present invention learns the common feature distribution between the cartoon face images and the normal face images, and thus can improve the accuracy of cartoon face recognition by using the easily distinguishable characteristics of the normal face features.

[0118] In an exemplary embodiment, in order to further improve the recognition accuracy of the obtained cartoon face recognition model, the above step S203 may include the following steps when respectively extracting features from the sample normal face images and the sample cartoon face images in the first sample image set based on the face recognition model:

[0119] Respectively input the sample normal face images and the sample cartoon face images in the first sample image set into the face recognition model for face recognition;

[0120] During the face recognition process, obtain the intermediate feature maps from the target access points of the face recognition model, and obtain the normal face features corresponding to the sample normal face images and the cartoon face features corresponding to the sample cartoon face images;

[0121] Among them, the target access point is any one of multiple preset access points, and the multiple preset access points include the intermediate feature extraction layer of the face recognition model. That is to say, the target access point can be any intermediate feature extraction layer of the face recognition model, and this target access point is used to access the model to obtain the target feature maps output by the corresponding intermediate feature extraction layer during the model feature extraction process.

[0122] In the above embodiments, by setting a plurality of preset access points and obtaining the intermediate feature maps in the corresponding face recognition process for any access point, the intermediate feature maps of each preset access point can be obtained for the sample normal face images and the sample cartoon face images respectively. Furthermore, for each preset access point, the target loss function value corresponding to the preset access point can be obtained based on the difference between the normal face features and the cartoon face features corresponding to the preset access point. Thus, the trained face recognition model corresponding to the best preset access point can be selected from the plurality of preset access points based on the target loss function values of each preset access point as the cartoon face recognition model for online deployment, so as to improve the recognition accuracy of cartoon face images.

[0123] Based on this, in an exemplary embodiment, the above step S205 of adjusting the model parameters of the face recognition model based on the target loss function value until the preset training end condition is satisfied to obtain the cartoon face recognition model may include:

[0124] Adjusting the model parameters of the face recognition model respectively based on the target loss function values corresponding to each of the preset access points until the preset training end condition is satisfied, to obtain the target face recognition models corresponding to each of the preset access points;

[0125] Performing face recognition on the sample cartoon face images respectively based on the target face recognition models corresponding to each of the preset access points, to obtain the cartoon face recognition results corresponding to each of the preset access points;

[0126] Selecting the cartoon face recognition model from the plurality of target face recognition models based on the cartoon face recognition results corresponding to each of the preset access points.

[0127] Wherein, the cartoon face recognition result of the cartoon face recognition model is better than the cartoon face recognition results of the remaining target face recognition models.

[0128] Specifically, for each preset access point, the pre-trained face recognition model can be fine-tuned based on the target loss function value corresponding to the preset access point, and then the target face recognition models corresponding one by one to the preset access points can be obtained. Further, performing the cartoon face recognition task respectively using each target face recognition model to obtain the cartoon face recognition results corresponding to each target face recognition model, so that the target face recognition model with the best cartoon face recognition result can be used as the final cartoon face recognition model for online deployment, so as to improve the cartoon face recognition accuracy of cartoon face images online.

[0129] In an exemplary embodiment, after fine-tuning to obtain a cartoon face recognition model, the cartoon face recognition model can be directly deployed online to perform cartoon face recognition tasks. Based on this, the method may further include: obtaining a to-be-processed cartoon face image, inputting the to-be-processed cartoon face image into the cartoon face recognition model for cartoon face recognition processing, and obtaining a cartoon face recognition result. Among them, the cartoon face recognition model is trained by using the cartoon face recognition model training method of the embodiments of the present invention.

[0130] In the above embodiment, the cartoon face recognition model has learned the common feature distribution of normal face features and cartoon face features, so that the recognition ability of normal faces is transferred to cartoon face recognition, improving the recognition accuracy of cartoon face images.

[0131] Corresponding to the cartoon face recognition model training methods provided in the above several embodiments, the embodiments of the present invention also provide a cartoon face recognition model training device. Since the cartoon face recognition model training device provided in the embodiments of the present invention corresponds to the cartoon face recognition model training methods provided in the above several embodiments, the implementation manners of the foregoing cartoon face recognition model training methods are also applicable to the cartoon face recognition model training device provided in this embodiment and will not be described in detail in this embodiment.

[0132] Please refer to Figure 4 , which shows a schematic structural diagram of a cartoon face recognition model training device provided by the embodiments of the present invention. The device has the function of implementing the cartoon face recognition model training method in the above method embodiments. The function can be implemented by hardware or by hardware executing corresponding software. As Figure 4 shown, the cartoon face recognition model training device 400 may include:

[0133] A first sample image acquisition module 410, configured to acquire a first sample image set; the first sample image set includes sample normal face images and sample cartoon face images, and the ratio of the first quantity of the sample normal face images in the first sample image set to the second quantity of the sample cartoon face images in the first sample image set satisfies a preset condition;

[0134] A feature extraction module 420, configured to respectively perform feature extraction on the sample normal face images and sample cartoon face images in the first sample image set based on a pre-trained face recognition model, to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images;

[0135] A target loss determination module 430, configured to determine a target loss function value based on the difference between the normal face features and the cartoon face features;

[0136] A parameter adjustment module 440 is configured to adjust the model parameters of the face recognition model based on the target loss function value until a preset training end condition is satisfied, obtaining a cartoon face recognition model; the cartoon face recognition model is used to recognize cartoon face images.

[0137] In an exemplary embodiment, the target loss determination module 430 includes:

[0138] A first loss determination module is configured to determine a first loss function value based on the difference between the normal face features and the cartoon face features;

[0139] A second loss determination module is configured to determine a first face recognition result based on the normal face features, and determine a second loss function value based on the difference between the first face recognition result and the first reference face information; the first reference face information is the face information annotated in the sample normal face image corresponding to the normal face features.

[0140] A third loss determination module is configured to determine a second face recognition result based on the cartoon face features, and determine a third loss function value based on the difference between the second face recognition result and the second reference face information; the second reference face information is the cartoon face information annotated in the sample cartoon face image corresponding to the cartoon face features.

[0141] A loss weighting module is configured to perform weighted summation on the first loss function value, the second loss function value, and the third loss function value to obtain the target loss function value.

[0142] In an exemplary embodiment, the first loss determination module includes:

[0143] A first statistic matrix determination module is configured to determine a first statistic matrix based on the normal face features corresponding to the sample normal face images in the first sample image set;

[0144] A second statistic matrix determination module is configured to determine a second statistic matrix based on the cartoon face features corresponding to the sample cartoon face images in the first sample image set;

[0145] A statistic difference matrix determination module is configured to obtain a statistic difference matrix based on the difference between the first statistic matrix and the second statistic matrix;

[0146] A first loss determination sub-module is configured to determine the first loss function value based on the statistic differences in the statistic difference matrix.

[0147] In an exemplary embodiment, the feature extraction module 420 includes:

[0148] A face recognition module for respectively inputting the sample normal face images and sample cartoon face images in the first sample image set into the face recognition model for face recognition;

[0149] An intermediate feature extraction module for obtaining an intermediate feature map from a target access point of the face recognition model during the face recognition process, to obtain a normal face feature corresponding to the sample normal face image and a cartoon face feature corresponding to the sample cartoon face image;

[0150] Wherein, the target access point is any one of a plurality of preset access points, and the plurality of preset access points include an intermediate feature extraction layer of the face recognition model.

[0151] In an exemplary embodiment, the parameter adjustment module 440 includes:

[0152] A parameter adjustment sub-module for respectively adjusting the model parameters of the face recognition model based on the target loss function values corresponding to each of the preset access points until a preset training end condition is satisfied, to obtain a target face recognition model corresponding to each of the preset access points;

[0153] A cartoon face recognition module for respectively performing face recognition on the sample cartoon face images based on the target face recognition models corresponding to each of the preset access points, to obtain cartoon face recognition results corresponding to each of the preset access points;

[0154] A model screening module for selecting the cartoon face recognition model from the plurality of target face recognition models based on the cartoon face recognition results corresponding to each of the preset access points; wherein, the cartoon face recognition result of the cartoon face recognition model is better than the cartoon face recognition results of the remaining target face recognition models.

[0155] In an exemplary embodiment, the device further includes:

[0156] A second sample image acquisition module for acquiring a second sample image set; the second sample image set includes sample normal face images;

[0157] A pre-training module for performing face recognition training on a convolutional neural network based on the second sample image set to obtain the face recognition model

[0158] It should be noted that, when the device provided in the above embodiment realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0159] An embodiment of the present invention provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement any one of the cartoon face recognition model training methods provided in the above method embodiments.

[0160] The memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0161] The method embodiments provided in the embodiments of the present invention can be executed on a computer terminal, a server, or a similar computing device, that is, the above electronic device can include a computer terminal, a server, or a similar computing device. Figure 5 is a hardware structure block diagram of an electronic device for running a cartoon face recognition model training method provided by an embodiment of the present invention, as Figure 5As shown, the server 500 can vary significantly depending on configuration or performance, and may include one or more central processing units (CPUs) 510 (the processor 510 may include, but is not limited to, processing devices such as a microprocessor MCU or a field-programmable gate array FPGA), a memory 530 for storing data, and one or more storage media 520 for storing application programs 523 or data 522 (such as one or more mass storage devices). Among them, the memory 530 and the storage media 520 can be transient storage or persistent storage. The programs stored in the storage media 520 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processor 510 can be set to communicate with the storage media 520 and execute a series of instruction operations in the storage media 520 on the server 500. The server 500 may also include one or more power supplies 560, one or more wired or wireless network interfaces 550, one or more input / output interfaces 540, and / or one or more operating systems 521, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0162] The input / output interface 540 can be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the server 500. In one example, the input / output interface 540 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 540 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0163] Those of ordinary skill in the art can understand that Figure 5 the structure shown is only illustrative and does not limit the structure of the above electronic device. For example, the server 500 may also include more or fewer components than Figure 5 shown, or have a different configuration from Figure 5 shown.

[0164] Embodiments of the present invention also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one segment of program related to implementing a method for training a cartoon face recognition model. The at least one instruction or the at least one segment of program is loaded and executed by the processor to implement any one of the methods for training a cartoon face recognition model provided by the above method embodiments.

[0165] An embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the cartoon face recognition model training method in any of the above aspects.

[0166] Optionally, in this embodiment, the above storage medium may include but is not limited to: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0167] It should be noted that: the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0168] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0169] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0170] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for training a cartoon face recognition model, characterized in that The method includes: Obtaining a first sample image set; the first sample image set includes sample normal face images and sample cartoon face images, and the ratio of the first quantity of the sample normal face images in the first sample image set to the second quantity of the sample cartoon face images in the first sample image set satisfies a preset condition; Based on a pre-trained face recognition model, respectively extracting features from the sample normal face images and sample cartoon face images in the first sample image set to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images; Based on the difference between the normal face features and the cartoon face features, determining a first loss function value; based on the normal face features, determining a first face recognition result, and based on the difference between the first face recognition result and first reference face information, determining a second loss function value; the first reference face information is the face information annotated in the sample normal face image corresponding to the normal face features; Based on the cartoon face features, determining a second face recognition result, and based on the difference between the second face recognition result and second reference face information, determining a third loss function value; the second reference face information is the cartoon face information annotated in the sample cartoon face image corresponding to the cartoon face features; Performing weighted summation on the first loss function value, the second loss function value, and the third loss function value to obtain a target loss function value; Based on the target loss function value, adjusting the model parameters of the face recognition model until a preset training end condition is satisfied to obtain a cartoon face recognition model; the cartoon face recognition model is used to recognize cartoon face images.

2. The method according to claim 1, wherein The determining the first loss function value based on the difference between the normal face features and the cartoon face features includes: Based on the normal face features corresponding to the sample normal face images in the first sample image set, determining a first statistic matrix; Based on the cartoon face features corresponding to the sample cartoon face images in the first sample image set, determining a second statistic matrix; Based on the difference between the first statistic matrix and the second statistic matrix, obtaining a statistic difference matrix; Based on the statistic differences in the statistic difference matrix, determining the first loss function value.

3. The method according to claim 1, characterized in that The respectively extracting features from the sample normal face images and sample cartoon face images in the first sample image set based on a pre-trained face recognition model to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images includes: Respectively inputting the sample normal face images and sample cartoon face images in the first sample image set into the face recognition model for face recognition; During the face recognition process, obtaining intermediate feature maps from a target access point of the face recognition model to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images; Among them, the target access point is any one of a plurality of preset access points, and the plurality of preset access points include the intermediate feature extraction layer of the face recognition model.

4. The method according to claim 3, wherein Adjusting the model parameters of the face recognition model based on the target loss function value until a preset training end condition is satisfied to obtain a cartoon face recognition model includes: Respectively adjusting the model parameters of the face recognition model based on the target loss function values corresponding to each of the preset access points until a preset training end condition is satisfied, to obtain the target face recognition models corresponding to each of the preset access points; Performing face recognition on the sample cartoon face images respectively based on the target face recognition models corresponding to each of the preset access points, to obtain the cartoon face recognition results corresponding to each of the preset access points; Selecting the cartoon face recognition model from the plurality of target face recognition models based on the cartoon face recognition results corresponding to each of the preset access points; wherein, the cartoon face recognition result of the cartoon face recognition model is better than the cartoon face recognition results of the remaining target face recognition models.

5. The method according to any one of claims 1 to 4, characterized in that, Before respectively performing feature extraction on the sample normal face images and the sample cartoon face images in the first sample image set based on a pre-trained face recognition model, the method further includes: Obtaining a second sample image set; the second sample image set includes sample normal face images; Performing face recognition training on the convolutional neural network based on the second sample image set to obtain the face recognition model.

6. A cartoon face recognition model training device, characterized in that The device includes: A first sample image acquisition module, configured to acquire a first sample image set; the first sample image set includes sample normal face images and sample cartoon face images, and the ratio of the first quantity of the sample normal face images in the first sample image set to the second quantity of the sample cartoon face images in the first sample image set satisfies a preset condition; A feature extraction module, configured to respectively perform feature extraction on the sample normal face images and the sample cartoon face images in the first sample image set based on a pre-trained face recognition model, to obtain normal face features corresponding to the sample normal face images and cartoon face features corresponding to the sample cartoon face images; A target loss determination module, configured to determine a first loss function value based on the difference between the normal face features and the cartoon face features; determine a first face recognition result based on the normal face features, and determine a second loss function value based on the difference between the first face recognition result and the first reference face information; the first reference face information is the face information annotated in the sample normal face image corresponding to the normal face features; determine a second face recognition result based on the cartoon face features, and determine a third loss function value based on the difference between the second face recognition result and the second reference face information; the second reference face information is the cartoon face information annotated in the sample cartoon face image corresponding to the cartoon face features; perform weighted summation on the first loss function value, the second loss function value, and the third loss function value to obtain a target loss function value; A parameter adjustment module, configured to adjust model parameters of the face recognition model based on the target loss function value until a preset training end condition is met to obtain a cartoon face recognition model; the cartoon face recognition model is used to recognize cartoon face images.

7. An electronic device, characterized in that, It includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the cartoon face recognition model training method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the cartoon face recognition model training method according to any one of claims 1 to 5.