Method for supervised learning based on face recognition

By adopting a method based on facial recognition supervision learning in the face recognition system, the recognition failure caused by feature changes is solved, the generalization ability of the model is improved, and user privacy is protected through homomorphic encryption, ensuring the accuracy and security of learning records.

CN120220205APending Publication Date: 2025-06-27LIANKE YUNCHUANG (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510245999.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems with recognition failure in facial recognition, especially feature changes caused by age changes, lighting changes, beard changes, etc., and data enhancement methods are insufficient, model generalization capabilities are poor, and there is a risk of user privacy leakage.

Method used

The method based on facial recognition supervision learning is adopted, and the system architecture is implemented, including real-name authentication in the initialization stage, first face recognition verification, face recognition analysis and update, face recognition during the learning process and final result processing. This method generates diverse enhanced samples, improves the diversity of training data, improves the generalization ability of the model, and protects user privacy through homomorphic encryption.

Benefits of technology

It effectively solves the problem of identification failure caused by feature changes, improves the generalization ability of the model, ensures the accuracy and security of learning records, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220205A_ABST
    Figure CN120220205A_ABST
Patent Text Reader

Abstract

The invention discloses a method for supervised learning based on face recognition, and the method comprises the steps: binding personnel information, carrying out the real-name authentication, and storing an identity card photo into a database; when video learning is started, first-time face recognition occurs to verify whether the person is a real-name authentication person of a current account; comparing the information of the first-time face recognition with a database identity card photo, and detecting whether the current learner is a real-name person of the account; a face randomly appearing in the learning process is compared with a second target face in the database, and whether the current student is learning is detected; learning can be continued after the recognition is passed within the specified time, and the learning record is reserved; when it is detected that a human face is shielded by a shielding object, the system carries out recognition again; if no face is detected, the learning record is emptied, and the video progress starts from the beginning. According to the method, diversified enhanced samples can be generated, the diversity of training data is improved, and the generalization ability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of learning supervision, and more particularly to a method for supervised learning based on face recognition. Background Art

[0002] With the continuous progress and development of society and the urgent demand for fast and effective automatic identity verification, it is difficult for traditional video learning systems to distinguish whether the current learning person is the actual person, and fraud is likely to occur. Therefore, as an emerging biometric identification technology, face recognition technology has gradually been introduced into the education field due to its characteristics of fast speed, accuracy, convenience and security.

[0003] In the prior art, face recognition of the current face image is usually performed according to the face data in a preset face database. However, due to the slow changes in illumination, pose, beard or age, the current face image does not match the face data in the preset face database, resulting in the failure of user face recognition.

[0004] Chinese Patent CN117373082A discloses a face recognition method, device, equipment and storage medium, which obtains the detection result of face recognition of the current face image through a preset face database; when it is determined that the detection result is that the current face image matches the first target face data in the preset face database, determines the previous face image that matches the first target face data; obtains the current matching score and the previous matching score when the previous face image and the current face image respectively match the first target face data; and when the current matching score and the previous matching score meet the preset difference condition, updates the first target face data in the preset face database according to the current face image. However, this method does not fully consider the problem of face recognition failure caused by factors such as slow age changes and changes in face features at different learning stages in the learning scenario. Its update mechanism is relatively single, and it is only updated based on the matching score condition, lacking adaptability to various complex situations in the learning process.

[0005] Moreover, traditional data augmentation methods are often relatively simple, only performing simple image copying or basic transformations, unable to fully explore the potential features of the data, and difficult to generate diverse and high-quality augmented data. This leads to insufficient diversity of the training data, limited generalization ability of the model when facing complex and variable actual face data, and prone to overfitting phenomena, affecting the accuracy and stability of face recognition.

[0006] In the process of face recognition, existing technologies usually do not effectively encrypt face image data during the transmission and inference stages, or the encryption methods are not perfect enough, posing a risk of user privacy leakage. Especially in learning supervision and management systems involving a large amount of user data, once face data is leaked, it will pose a serious threat to the security of users' personal information.

[0007] Moreover, in terms of model optimization, existing technologies often lack effective coping strategies for complex face changes in actual application scenarios. The model cannot automatically adjust its structure and parameters according to different scenarios (such as special situations like low light and wearing masks), resulting in a decrease in recognition accuracy in these specific scenarios and being unable to meet diverse actual needs. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a method for face recognition supervised learning, which can generate diverse enhanced samples, improve the diversity of training data, enhance the generalization ability of the model, and facilitate the supervision of the learning process.

[0009] To solve the above technical problems, the technical solutions adopted by the present invention are as follows.

[0010] A method for face recognition supervised learning, implemented according to the system architecture, includes the following steps: S1. Initialization stage: Bind personal information, conduct real-name authentication, and store the ID card photo in the database; S2. First face recognition: When entering video learning, the first face recognition will occur to verify whether it is the person who has passed real-name authentication for the current account. If the face recognition fails, the user cannot enter video learning; S3. Face recognition analysis and update: Compare the information of the first face recognition with the ID card photo in the database to detect whether the current learning person is the person with real name for this account. If the face recognition fails, check whether there is an object blocking the current photo. If there is an object blocking, give a prompt; if there is no object blocking, conduct re-real-name authentication; if the face recognition passes, store the current face in the database as the second target face; S4. Face recognition during the learning process: Compare the randomly appearing face during the learning process with the second target face in the database to detect whether the current student is learning; S5. Final result processing: If the recognition is passed within the specified time, learning can continue and the learning record is retained; when it is detected that there is a face and an object is blocking it, the system conducts re-recognition; when no face is detected, the learning record is cleared and the video progress starts from the beginning.

[0011] Further optimizing the technical solution, the system architecture includes a physical layer, a data layer, a model layer, and an application layer; Physical layer: Defines the actual physical environment in which the system operates and is responsible for acquiring face image data; Data layer: Responsible for storing, managing, and processing raw data and processed data; Model layer: Responsible for performing the core functions in supervised learning, including model training, inference, and optimization; Application layer: Oriented towards users and responsible for integrating face recognition technology into applications.

[0012] Further optimizing the technical solution, the data layer includes two sub-modules: preprocessing and feature extraction. The preprocessing module is responsible for performing operations such as denoising, alignment, and enhancement on the collected face images to improve the accuracy of subsequent feature extraction and classification recognition; the feature extraction module extracts discriminative feature vectors from the preprocessed images.

[0013] Further optimizing the technical solution, the method by which the feature extraction module extracts key face feature points is as follows: a1. Perform facial organ localization by using a pre-trained face detection and feature point localization model; a2. Calculate distances and angles. Use the Euclidean formula to calculate distances, that is, for two points P1(a1, b1) and P2(a2, b2), the distance d between them is: ; Calculate angles using the angle formula between vectors, that is, for two vectors and , the angle θ between them is: where represents the dot product of the vectors, and and represent the magnitudes of the vectors respectively.

[0014] a3. To improve the robustness of recognition, preprocess the calculated distances; a4. Use the preprocessed distance vector as the feature vector for subsequent classification or matching operations.

[0015] Further optimizing the technical solution, the model layer includes a neural network structure, a training process, and an inference process; The neural network structure is responsible for receiving external data and performing feature extraction and transformation to generate the final prediction result; The training process refers to the neural network updating the weights in the network through the backpropagation algorithm and optimization algorithm to minimize the loss function and improve the prediction accuracy of the model; The inference process refers to, after training is completed, inputting new input data into the trained model and obtaining the prediction result.

[0016] Further optimize the technical solution. The model inference adopts homomorphic encryption inference. The face image data is homomorphically encrypted on the client side and then sent to the server side. The face recognition model on the server side performs inference operations on the encrypted data. During the entire inference process, the model and the data always remain in an encrypted state. The specific steps are as follows: b1. Preliminary preparation stage: including the selection of homomorphic encryption scheme, determination and parameter extraction of the face recognition model, and the construction of the client and the server. b2. Data encryption stage: including face image preprocessing, polynomial representation and encryption of feature vectors, and sending of encrypted data. b3. Model inference stage: including receiving encrypted data and loading the model, inference calculation based on homomorphic encryption operation rules, convolution layer operation, fully connected layer calculation, non-linear activation function processing, and return of encrypted results. b4. Result decryption stage: including receiving and decrypting the encrypted result, displaying and applying the recognition result, optimizing the model using a generative adversarial network, and adjusting the model structure based on feedback.

[0017] Further optimize the technical solution. The model optimization uses a pre-trained generative adversarial network (GAN) or convolutional neural network (CNN) to automatically generate new enhanced data according to the characteristics of the input data. GAN consists of a generator and a discriminator. The generator generates enhanced data according to the characteristics of the input data; the discriminator judges the similarity between the generated data and the real data; during the training process, the generator is continuously optimized to generate more realistic enhanced data; CNN enhances the data by performing image processing on the original data.

[0018] Further optimize the technical solution. The specific method of the model optimization is as follows: c1. GAN data preparation: Collect a face image dataset containing various lighting conditions, expressions, and postures, and divide the dataset into a training set and a validation set to ensure that the training set is large enough for the model to learn the distribution characteristics of face data. c2. Construct the GAN model architecture: The generator adopts a deep convolutional generative adversarial network, and the discriminator adopts a conventional convolutional neural network architecture. c3. Train the GAN model: Set the training parameters; in each round of training, randomly select a batch of real face images from the training set as positive samples; according to the output of the discriminator, calculate the loss functions of the generator and the discriminator; during the training process, regularly save the images generated by the generator and observe the quality change of the generated images. c4. CNN Preprocessing Data: Collect a face image dataset, normalize it to the same size, and label the images; use bilinear interpolation to rotate the images with the center of the image as the rotation center; use bilinear interpolation to scale the images; randomly determine the upper left coordinates and the cropping size of the cropping area, and add the cropped images to the augmented dataset; merge the augmented dataset with the original dataset, and divide it into a training set, a validation set, and a test set.

[0019] Due to the above technical solutions, the technical progress achieved by the present invention is as follows.

[0020] A method based on face recognition supervised learning provided by the present invention ensures that the account is being learned by the current real-name person by collecting face information, and saves the second target face when first entering the learning, ensuring that the randomly occurring face recognition during the learning process will not fail due to the slow change of age, resulting in the mismatch between the current face image and the real-name ID card, thereby affecting the learning record. The present invention can generate diverse augmented samples, improve the diversity of training data, and improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is the system architecture diagram of the present invention; Figure 2 is the flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0023] A method based on face recognition supervised learning is implemented according to the system architecture. The system architecture diagram of the present invention is as Figure 1 shown, including a physical layer, a data layer, a model layer, and an application layer.

[0024] Physical layer: mainly involves the level of hardware devices, including devices, servers, and computing resources. This layer defines the actual physical environment in which the system runs. It is responsible for obtaining face image data. Capturing the current face image through the mobile phone camera to complete the collection of face images or videos. The collected data is stored in the database for subsequent processing and analysis.

[0025] Data layer: responsible for storing, managing, and processing raw data and processed data. Data is the core of supervised learning, especially in face recognition, the management of raw images and label data is crucial. It includes two sub-modules: preprocessing and feature extraction. The preprocessing module is responsible for performing operations such as denoising, alignment, and enhancement on the collected face images to improve the accuracy of subsequent feature extraction and classification recognition; the feature extraction module extracts discriminative feature vectors from the preprocessed images.

[0026] The method for the feature extraction module to extract key facial feature points is as follows: a1. Perform facial organ localization by using a pre-trained face detection and feature point localization model. The facial organs include eyes, the tip of the nose, the corners of the mouth, etc.

[0027] a2. Calculate distances and angles. Use the Euclidean formula to calculate distances. That is, for two points P1(a1, b1) and P2(a2, b2), the distance d between them is: ; Use the angle formula between vectors to calculate angles. That is, for two vectors and , the angle θ between them is:

[0028] where represents the dot product of vectors, and and represent the magnitudes of the vectors respectively; a3. To improve the robustness of recognition, preprocess the calculated distances. For example, each element can be divided by the same feature distance (such as the distance between the key point closest to the center of the eyebrows and the key point closest to the vertex of the chin) to convert the distance relationship into a proportional relationship to the length of the face.

[0029] a4. Use the preprocessed distance vectors as feature vectors for subsequent classification or matching operations. These feature vectors can be input into a classification model for classification or used to calculate the distance from the templates stored in the database.

[0030] Model layer: Responsible for performing the core functions in supervised learning, including model training, inference, and optimization. This layer contains the deep learning model and its training process. Usually includes a neural network structure, a training process, and an inference process. The neural network structure is responsible for receiving external data and performing feature extraction and transformation to generate the final prediction result; the training process refers to the neural network updating the weights in the network through the backpropagation algorithm and optimization algorithm to minimize the loss function and improve the prediction accuracy of the model. The inference process refers to after training is completed, inputting new input data into the trained model and obtaining the prediction result.

[0031] Model inference adopts homomorphic encryption inference. The face image data is homomorphically encrypted on the client side and then sent to the server side. The face recognition model on the server side performs inference operations on the encrypted data. During the entire inference process, the model and data always remain in an encrypted state. For example, adopting a polynomial-based homomorphic encryption scheme, representing the feature vector of the face image in polynomial form for encryption, and also performing homomorphic encryption processing on the weights of the model. During the inference process, through the operation rules of homomorphic encryption (such as addition and multiplication homomorphism), the model can perform linear and non-linear operations on the encrypted data, and finally obtain the encrypted recognition result. The client then decrypts the encrypted result, thus realizing face recognition inference under privacy protection. It specifically includes the following steps: b1. Preparatory stage Selection of homomorphic encryption scheme: Select a suitable polynomial-based homomorphic encryption scheme. The BFV (Brakerski-Fan-Vercauteren) scheme supports additive and multiplicative homomorphic operations over integers, is suitable for handling discrete data cases, and is often used in scenarios with high privacy requirements.

[0032] Determination of face recognition model and parameter extraction: Select a mature face recognition model architecture, a traditional face recognition model based on feature extraction and matching. For the selected model, extract its weight parameters, and these weight parameters will participate in subsequent homomorphic encryption processing and inference operations on encrypted data. For a convolutional neural network model, obtain parameters such as the convolutional kernel weights and biases of each layer.

[0033] Establishment of client and server: Build a communication architecture between the client and the server to ensure that the client can send encrypted data to the server, the server can receive and perform corresponding inference calculations, and then return the encrypted result to the client.

[0034] b2. Data encryption stage Preprocessing of face images: After obtaining the face image on the client side, first perform conventional preprocessing operations, crop the face area to ensure that only valid facial information is included, normalize the image to a unified size (such as 128×128 pixels), and perform grayscale processing (if the model supports grayscale image input) or normalize the color channels (if it is a color image). Then convert the face image into a feature vector representation through a suitable feature extraction method, using principal component analysis (PCA) to extract the feature vector or obtaining the feature vector through the feature extraction layer of a deep learning model.

[0035] Polynomial representation and encryption of feature vectors: Represent the extracted face image feature vector in polynomial form. Assume the feature vector is , a polynomial can be constructed (where \(x\) is a formal variable). Then, this polynomial is encrypted using the selected homomorphic encryption scheme to generate the ciphertext \(Enc(P(x))\). At the same time, the weight parameters of the face recognition model are also encrypted according to the same homomorphic encryption scheme. For the weight matrix of a certain layer in the model, its elements are formed into polynomial form one by one or according to certain rules and then encrypted to obtain the encrypted weight representation \(Enc(W)\).

[0036] Encrypted data transmission: The client sends the encrypted face image feature vector data (existing in the form of encrypted polynomials) to the server through a secure network communication protocol.

[0037] b3. Model inference stage Encrypted data reception and model loading: The server receives the encrypted face image feature vector data sent by the client. At the same time, the server loads the pre-encrypted face recognition model (its weight and other parameters have been processed by homomorphic encryption) and prepares to perform inference operations on the encrypted data.

[0038] Inference calculation based on homomorphic encryption operation rules: According to the calculation process of the face recognition model, linear and non-linear operations are performed on the encrypted data using the operation rules of homomorphic encryption (additive homomorphicity and multiplicative homomorphicity).

[0039] Convolution layer operation: In ordinary convolution operations, a certain element of the output feature map is the result of weighted summation of the input feature map and the convolution kernel (involving multiplication and addition operations). In the homomorphic encryption scenario, for the ciphertext \(Enc(X)\) of the encrypted input feature map and the ciphertext \(Enc(W)\) of the encrypted convolution kernel, the product ciphertext \(Enc(X)\times Enc(W)\) of their corresponding elements is calculated using multiplicative homomorphicity, and then these product ciphertexts are accumulated using additive homomorphicity to obtain the encrypted result \(Enc(Y)\) of the convolution layer output. The entire process is carried out in the ciphertext state without decryption.

[0040] Fully connected layer calculation: The calculation of the fully connected layer is similar to matrix multiplication and addition operations. Similarly, using the operation rules of homomorphic encryption, multiplication and addition operations are performed on the encrypted input vector (obtained from the encrypted output of the previous layer) and the encrypted weight matrix of the fully connected layer to obtain the encrypted output result of the fully connected layer.

[0041] Non-linear activation function processing: For some non-linear activation functions that can be processed by homomorphic encryption schemes (such as some schemes can process the ReLU function through methods such as polynomial approximation), the encrypted intermediate results are non-linearly transformed according to the corresponding homomorphic encryption calculation method to obtain the encrypted result after activation, and continue to be passed to the next layer for calculation in the encrypted state.

[0042] Through layers of homomorphic encryption operations, the encrypted face recognition results are finally obtained. The model and data of the entire process always remain encrypted. The server side will not access any plaintext data, effectively protecting the user's privacy.

[0043] Encrypted result return: The server sends the encrypted face recognition result back to the client through a secure network communication protocol.

[0044] b4. Result decryption stage Receiving and decrypting encrypted results: The client receives the encrypted face recognition results returned by the server, and then uses the decryption algorithm of the homomorphic encryption scheme corresponding to the encryption process to decrypt the encrypted results, restore the plaintext form of the recognition results, and use the decryption function matching the BFV scheme to decrypt them.

[0045] Display and application of recognition results: The client displays the decrypted recognition results, such as displaying the recognized person's identity information (if it is an identity recognition task), or determining whether it is an authorized user (if it is an application scenario such as access control), etc., thereby completing the face recognition reasoning task while achieving privacy protection.

[0046] Generate adversarial networks for model optimization: Model optimization uses pre-trained generative adversarial networks (GANs) or convolutional neural networks (CNNs) to automatically generate new enhanced data based on the characteristics of the input data. GANs consist of a generator and a discriminator. The generator generates enhanced data based on the characteristics of the input data; the discriminator determines the similarity between the generated data and the real data; during the training process, the generator is continuously optimized to generate more realistic enhanced data; CNN enhances data by performing image processing on the original data. The input data generates diverse enhanced samples through these models to increase the diversity of the training data, thereby improving the generalization ability of the model and avoiding the simple replication of traditional methods.

[0047] The specific method of model optimization is: c1. GAN data preparation: Collect a dataset of facial images under various lighting conditions, expressions, and postures. For example, obtain images from a recognized face database and preprocess them into a uniform size, such as 128×128 pixel grayscale or color images. Divide the dataset into a training set and a validation set, and ensure that the training set is large enough for the model to learn the distribution characteristics of facial data; c2. Build the GAN model architecture: The generator uses a deep convolutional generative adversarial network, taking a 100-dimensional random noise vector as input and gradually mapping it to a 128×128×3 tensor through multiple transposed convolutional layers. The first transposed convolutional layer may expand the input noise vector into a feature map of 4×4×1024, and subsequent transposed convolutional layers gradually increase the size of the feature map and reduce the number of channels, such as 8×8×512, 16×16×256, etc., until the target size is reached. Batch normalization is used after each transposed convolutional layer to stabilize the training process. The activation function can use ReLU, except for the last layer, where the Tanh activation function is used to map the pixel value range to [-1,1].

[0048] The discriminator adopts a conventional convolutional neural network architecture: the input is a 128×128×3 face image, and features are extracted through multiple convolutional layers. The initial convolutional layer may have 64 3×3 convolutional kernels with a stride of 2 for downsampling the image and extracting preliminary features. Subsequent convolutional layers increase the number of convolutional kernels (such as 128, 256, etc.) and appropriately adjust the stride. The LeakyReLU activation function is used after each convolutional layer, and the last layer is a fully connected layer that outputs a scalar representing the probability that the input image is a real face image.

[0049] c3. Training the GAN model Set the training parameters: learning rate (the learning rates of both the generator and the discriminator are set to 0.0002), number of training epochs (100 epochs), and batch size (64).

[0050] In each round of training, a batch of real face images is randomly selected from the training set as positive samples, and the generator generates a batch of fake face images by inputting a random noise vector through the generator network. The real face images and the generated fake face images are sent to the discriminator for judgment, and the discriminator outputs the probability that each image is a real face.

[0051] According to the output of the discriminator, calculate the loss functions of the generator and the discriminator. For example, for the generator, the adversarial loss is adopted, and the goal is to minimize the probability that the generated image is judged as fake by the discriminator; for the discriminator, the binary cross-entropy loss is adopted, and the goal is to correctly distinguish real and fake face images. Use the backpropagation algorithm to update the weights of the generator and the discriminator. Use the Adam optimizer to update the weights and adjust the network parameters according to the calculated gradients, so that the images generated by the generator become more realistic and the discrimination ability of the discriminator becomes stronger.

[0052] During the training process, the generated images of the generator are saved regularly to observe the quality changes of the generated images. For example, the generated face images are saved every 10 epochs, and the images are checked for similarity to the features such as the shape and texture of the face. The trained generator can be used to generate new face images, and these generated images are added to the original face dataset for subsequent training of the face recognition model, thereby increasing the diversity of the data.

[0053] c4.CNN preprocess data Collect a face image dataset and normalize it to the same size, such as a grayscale image of 96×96 pixels. Label the images, for example, label the identity tags corresponding to each face. Rotation augmentation: For each face image, randomly generate a rotation angle, and the angle range can be set to [-30, 30] degrees. Using the center of the image as the rotation center, perform rotation using the bilinear interpolation method.

[0054] Scaling augmentation: Randomly generate a scaling factor, and the range can be [0.8, 1.2]. Use the bilinear interpolation method to scale the image. Cropping augmentation: Randomly determine the upper-left coordinates and the cropping size of the cropping area. The upper-left coordinates (x, y) of the cropping area can be randomly selected within the range of [0, 0.2] of the image width and height, and the cropping size can be [0.6, 1] times the size of the original image. Add the cropped image to the augmented dataset.

[0055] Merge the augmented dataset with the original dataset, divide the training set, validation set, and test set, and select a suitable face recognition model architecture, such as a deep convolutional neural network (such as ResNet, VGG, etc.). Use the merged dataset to train the face recognition model, and update the weights of the model through the backpropagation algorithm so that the model can better learn the features of the face and improve the accuracy and generalization ability of face recognition.

[0056] Model structure adjustment based on feedback: A technology for adaptively adjusting the model structure according to the feedback information in the actual application of the model. For example, in the actual deployment of a face recognition system, collect feedback data such as the recognition accuracy, false recognition rate, and rejection rate of the model. According to this data, if it is found that the performance of the model deteriorates in certain specific scenarios (such as low light, wearing a mask), automatically adjust the structure of the model, such as adding layers for low-light feature extraction or adopting alternative feature extraction strategies in the mask occlusion area. This adaptive structure adjustment can be achieved by dynamically adding or modifying network layers and adjusting the connection methods between layers, enabling the model to continuously optimize its own structure according to the actual application scenario.

[0057] Application layer: It finally faces users and is responsible for integrating face recognition technology into applications. The management end configures the requirements of each region and institution; the learning end accepts the configuration of the management end and applies it flexibly.

[0058] The flowchart of the present invention is as Figure 2 shown and includes the following steps: S1. Initialization stage: Bind personal information. When entering the app, it will prompt for real-name authentication, and store the ID card photo in the database.

[0059] S2. First face recognition: When entering video learning, the first face recognition will occur to verify whether it is the person who has passed real-name authentication for the current account. If the face recognition fails, the user cannot enter video learning.

[0060] Face recognition method: When starting face recognition, it will capture the current face image and compare it with the first target face in the database. By extracting the face features of both sides, a score is calculated. When the score meets the preset difference condition, the recognition passes.

[0061] S3. Face recognition analysis and update: Compare the information of the first face recognition with the ID card photo in the database to detect whether the current learning person is the person who has passed real-name authentication for this account. If the face recognition fails, first detect whether there is an object blocking the current photo. If there is an object blocking, a prompt will be given. If there is no object blocking, re-perform real-name authentication to exclude the face recognition failure caused by unclear photo taking during the previous real-name authentication. If the face recognition passes, store the current face in the database as the second target face.

[0062] S4. Face recognition during the learning process: Compare the randomly appearing face during the learning process with the second target face in the database to detect whether the current student is studying and prevent students from swiping courses.

[0063] When comparing with the second target face instead of the real-name authentication photo, the most recent photo is taken for comparison, eliminating a large part of uncertain factors, minimizing the error, and ensuring the normal preservation of learning records. This face recognition process has a time limit, which can be set through the Pc management end.

[0064] S5. Final result processing: If the recognition passes within the specified time, learning can continue and the learning record is retained.

[0065] When the recognition fails within the specified time, when a face is detected, it is judged whether there is an object blocking. If there is, a prompt will be given, and at the same time, the page countdown restarts and the system performs re-recognition; if no face is detected, it indicates that the student is not watching the video, then the countdown is cleared, the learning record is cleared, and the video progress starts from the beginning.

[0066] When the present invention first enters video learning, by comparing the current face with the ID card photo, if the matching score meets the difference condition, the second target face is updated. This mechanism avoids the failure of face recognition during the video learning process due to factors such as age changes, ensuring the accuracy and security of learning records. The second target face, as the face recognition comparison benchmark during the learning process, can effectively handle random face recognition detections during the learning process, preventing the learning progress record from being affected by changes in face features.

Claims

1. A method based on supervised learning of face recognition, characterized in that: According to the system architecture, it includes the following steps: S1. Initialization phase: bind personnel information, conduct real-name authentication, and store ID card photos in the database; S2. First face recognition: When you enter the video learning, the first face recognition will appear to verify whether you are the real-name authenticated person of the current account. If the face recognition fails, you cannot enter the video learning; S3. Face recognition analysis and update: compare the information of the first face recognition with the ID card photo in the database to detect whether the current learner is the real-name person of this account. If the face recognition fails, check whether the current photo is blocked by an obstruction. If so, prompt it. If not, re-authenticate. If the face recognition passes, store the current face in the database as the second target face. S4. Face recognition during learning: randomly appearing faces during learning are compared with the second target face in the database to detect whether the current student is learning; S5. Final result processing: If the recognition is passed within the specified time, learning can continue and the learning record will be retained; if a face is detected and blocked by an obstruction, the system will re-recognize; if no face is detected, the learning record will be cleared and the video progress will start from the beginning.

2. The method according to claim 1, characterized in that: The system architecture includes a physical layer, a data layer, a model layer and an application layer; Physical layer: defines the actual physical environment in which the system operates and is responsible for acquiring facial image data; Data layer: responsible for storing, managing and processing raw data and processed data; Model layer: responsible for performing core functions in supervised learning, including model training, inference, and optimization; Application layer: user-oriented, responsible for integrating face recognition technology into applications.

3. The method based on face recognition supervised learning according to claim 2, characterized in that: The data layer includes two sub-modules: preprocessing and feature extraction. The preprocessing module is responsible for denoising, aligning, enhancing and other operations on the collected face images to improve the accuracy of subsequent feature extraction and classification recognition; the feature extraction module extracts a distinguishing feature vector from the preprocessed image.

4. The method based on supervised learning of face recognition according to claim 3, characterized in that: The method for extracting key feature points of a face by the feature extraction module is: a1. Facial organ localization can be performed by using pre-trained face detection and feature point localization models; a2. Calculation of distance and angle. Use Euclidean distance calculation. For two points P1 (a1, b1) and P2 (a2, b2), the distance d between them is: ; The angle between two vectors is calculated using the angle formula and , the angle θ between them is: in represents the dot product of vectors, and and They represent the modulus of the vector respectively; a3. In order to improve the robustness of recognition, the calculated distance is preprocessed; a4. Use the preprocessed distance vector as a feature vector for subsequent classification or matching operations.

5. The method based on face recognition supervised learning according to claim 2, characterized in that: The model layer includes a neural network structure, a training process and an inference process; The neural network structure is responsible for receiving external data and performing feature extraction and transformation to generate the final prediction results; The training process refers to the neural network updating the weights in the network through the back-propagation algorithm and the optimization algorithm to minimize the loss function and improve the prediction accuracy of the model; The inference process refers to the process of passing new input data into the trained model after training is completed and obtaining the prediction results.

6. The method based on supervised learning of face recognition according to claim 5, characterized in that: The model reasoning adopts homomorphic encryption reasoning. The face image data is homomorphically encrypted on the client and then sent to the server. The face recognition model on the server performs reasoning operations on the encrypted data. The model and data remain encrypted throughout the reasoning process. The specific steps include: b1. Preliminary preparation stage: including the selection of homomorphic encryption scheme, determination of face recognition model and parameter extraction, and establishment of client and server; b2. Data encryption stage: including face image preprocessing, feature vector polynomial representation and encryption, and encrypted data transmission; b3. Model inference stage: including encrypted data reception and model loading, inference calculation based on homomorphic encryption operation rules, convolution layer calculation, fully connected layer calculation, nonlinear activation function processing, and encrypted result return; b4. Result decryption stage: including receiving and decrypting encrypted results, displaying and applying recognition results, optimizing the model by generating adversarial networks, and adjusting the model structure based on feedback.

7. The method based on face recognition supervised learning according to claim 6, characterized in that: The model optimization uses a pre-trained generative adversarial network (GAN) or convolutional neural network (CNN) to automatically generate new enhanced data according to the characteristics of the input data. The GAN consists of a generator and a discriminator. The generator generates enhanced data according to the characteristics of the input data. The discriminator determines the similarity between the generated data and the real data; During the training process, the generator is continuously optimized to generate more realistic augmented data; CNN enhances the data by performing image processing on the original data.

8. The method based on supervised learning of face recognition according to claim 7, characterized in that: The specific method of model optimization is: c1. GAN data preparation: Collect a dataset of facial images under various lighting conditions, expressions, and postures, divide the dataset into a training set and a validation set, and ensure that the training set is large enough to allow the model to learn the distribution characteristics of facial data; c2. Build the GAN model architecture: the generator uses a deep convolutional generative adversarial network, and the judge uses a conventional convolutional neural network architecture; c3. Training the GAN model: set training parameters; in each round of training, randomly select a batch of real face images from the training set as positive samples; calculate the loss function of the generator and the discriminator based on the output of the discriminator; during the training process, regularly save the images generated by the generator and observe the changes in the quality of the generated images; c4.CNN preprocessing data: collect face image datasets, normalize them to the same size, and annotate the images; rotate them using the bilinear interpolation method with the center of the image as the rotation center; scale the images using the bilinear interpolation method; randomly determine the upper left corner coordinates and cropping size of the cropped area, and add the cropped image to the enhanced dataset; The augmented dataset is merged with the original dataset and divided into training set, validation set and test set.

Citation Information

Patent Citations

  • Face recognition method and device, equipment and storage medium

    CN117373082A