Virtual anchor identification method based on emotion perception and semi-supervised comparative learning

Through the method based on emotion perception and semi-supervised contrast learning, deep 3D convolutional neural network and KL divergence optimization, accurate identification between virtual anchors and real anchors is achieved, and identification difficulties and algorithm vulnerability problems in the existing technology are solved, which improves the robustness of recognition and anti-fraud capabilities.

CN120412059AActive Publication Date: 2025-08-01COMMUNICATION UNIVERSITY OF CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510913878.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

It is difficult to effectively identify virtual anchors and real anchors in the existing technology, and it shows algorithm fragility under dynamic confrontational attacks, and there are risks of social trust mechanisms and information manipulation risks.

Method used

Using a method based on emotion perception and semi-supervised contrast learning, facial areas are extracted through the facial recognition model, and pseudo-labels are generated using the emotional sense knowledge distinction model, and feature representation and recognition are performed through the virtual anchor representation learning network. Combined with deep 3D convolutional neural network and KL divergence optimization, the distinction between virtual anchors and real anchors is achieved.

Benefits of technology

It improves the accuracy and robustness of virtual anchor identification, and can effectively train under some tag conditions to prevent the illegal use and information manipulation of virtual anchors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412059A_ABST
    Figure CN120412059A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual anchor identification method based on emotion perception and semi-supervised contrast learning, and relates to the technical field of virtual video identification, and the method comprises the following steps: carrying out the random sampling processing of continuous video frames, and obtaining sequence frames; performing face recognition on the preprocessed sequence frame, and cutting the recognized face; the obtained face image sequence is input into an emotion perception recognition module for emotion recognition, and a non-label face image sequence is marked with a pseudo label; constructing a contrast learning data set; enabling the comparative learning data set to pass through a virtual anchor representation learning network, enabling the face representation to pass through a designed virtual anchor representation comparative learning network, and enabling the learned feature representation to pass through a virtual anchor recognition network to obtain a prediction result; according to the virtual anchor identification method and system based on emotion perception and semi-supervised comparative learning provided by the invention, the identification of the virtual anchor can be carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virtual video recognition, and specifically relates to a virtual anchor recognition method based on emotion perception and semi-supervised contrast learning. Background Art

[0002] As an innovative carrier of the digital culture industry, the technological evolution of virtual anchors is strongly coupled with the breakthroughs in the field of artificial intelligence. Driven by deep generation technologies such as Generative Adversarial Network (GAN) and Neural Radiance Field (NeRF), virtual anchors have broken through the limitations of traditional 3D modeling and achieved a realistic effect with a micro-expression modeling accuracy of 0.1 mm and a naturalness MOS score of over 4.5 for speech synthesis. At present, the scale of China's virtual anchor industry has exceeded hundreds of billions, and has deeply penetrated into scenarios such as e-commerce live broadcasts, news broadcasts, and education and popular science. Its intelligent production and broadcasting system can significantly improve the content production efficiency.

[0003] At present, the rapid development of virtual anchors has become the focus of public attention, but there are also some problems. First, the iterative upgrade of deep synthesis technology has made the perceptual boundary between virtual anchors and real human anchors increasingly blurred, and traditional recognition methods based on single-modal feature analysis can no longer effectively handle multi-modal coupled deep virtual content. Second, the cognitive confusion effect caused by high-fidelity virtual anchors has exceeded the critical threshold of human visual perception, resulting in significant risks to the social trust mechanism. Therefore, virtual anchors may spread inaccurate information and may also be used to manipulate public opinion or for other illegal activities such as fraud and defamation. Finally, the existing supervised learning framework shows algorithm vulnerability in dealing with dynamic adversarial attacks, especially in the context of the rapid evolution of generative models, there are serious defects in generalization ability. Therefore, constructing a semi-supervised contrast learning system integrating emotion perception mechanism has become the key technical breakthrough direction for realizing the reliable recognition of virtual anchors. There is no related technology for virtual anchor recognition in the prior art. Summary of the Invention

[0004] The purpose of the present invention is to provide a virtual anchor recognition method based on emotion perception and semi-supervised contrast learning, which can determine whether the anchor in the video is a virtual anchor.

[0005] The technical solution of the present invention is as follows:

[0006] A virtual anchor recognition method based on emotion perception and semi-supervised contrast learning includes the following steps:

[0007] Sampling continuous video frames to obtain a sequence of frames;

[0008] Extracting the facial region from the sequence of frames through a facial recognition model to obtain a sequence of face images;

[0009] Input a sequence of unlabeled face images into an emotion perception and recognition model for emotion recognition to obtain pseudo-labels;

[0010] Input the sequence of face images into a virtual anchor representation learning network to obtain feature representations; the virtual anchor representation learning network is trained through a contrastive learning dataset, and the contrastive learning dataset includes a sequence of labeled face images and a sequence of face images with obtained pseudo-labels;

[0011] Input the feature representations into a virtual anchor recognition network to obtain recognition results.

[0012] Furthermore, before extracting the face regions from the sequence of frames through a face recognition model to obtain a sequence of face images, an image enhancement processing step for the sequence of frames is also included.

[0013] Specifically, the sampling process of the continuous video frames to obtain a sequence of frames includes:

[0014] In the live video stream, sample at a random starting point to generate a sequence of continuous original frames with a fixed duration;

[0015] For the sequence of continuous original frames, perform random frame sampling according to a preset discard rate.

[0016] Specifically, the extraction of face regions from the sequence of frames through a face recognition model to obtain a sequence of face images includes:

[0017] Perform scaling and / or normalization processing on the sequence of frames to obtain input sequence frames;

[0018] Input the input sequence frames into a face recognition model to obtain face bounding boxes;

[0019] Crop the face regions through the face bounding boxes, and adjust all face regions to the same size to obtain a sequence of face images.

[0020] Specifically, the input of the sequence of unlabeled face images into an emotion perception and recognition model for emotion recognition to obtain pseudo-labels includes:

[0021] Obtain a global feature vector for the overall face image through a global feature extraction model , is the dimension of the global feature vector;

[0022] Divide the overall face image into n*n local region blocks, and perform local feature extraction on each local region block respectively to obtain local feature vectors ;

[0023] Calculate the similarity score between the local feature vectors and the global feature vector to obtain the position weight score of each local region block;

[0024] After guiding the local vectors with the position weight scores, the local vectors are then fused with the global feature vectors to generate a fused feature map;

[0025] The clustering algorithm is used to perform unsupervised clustering on the fused feature map, and the obtained classification result is used as the pseudo-label of the unlabeled face image sequence.

[0026] Preferably, the calculation of the similarity score between the local feature vector and the global feature vector specifically uses dot product similarity and normalization operations to perform the similarity score calculation, and the formula is:

[0027] ,

[0028] where, is the similarity score between the i-th local feature vector and the global feature vector, is the global feature vector, is the i-th local feature vector, is the globally learnable weight parameter, is the locally learnable weight parameter, and T is the transpose operation.

[0029] Specifically, after guiding the local vectors with the position weight scores, the local vectors are then fused with the global feature vectors to generate a fused feature map, including:

[0030] Each local feature vector is weighted to obtain a weighted local feature vector ,

[0031]

[0032] where, is the weighted i-th local feature vector;

[0033] According to the originally divided n×n grid positions, the n×n weighted local feature vectors are arranged in corresponding positions to form a low-resolution feature map, and the channel dimension is adjusted through 1×1 convolution to generate a reorganized feature map ;

[0034] The global feature vector is converted into an updated global feature map with the same dimension as the reorganized feature map through a fully connected layer ;

[0035] The reorganized feature map is concatenated with the updated global feature map along the channel dimension to generate a fused feature map ;

[0036] Gradually upsample the fused feature map through multi-layer deconvolution to the resolution of the original face image.

[0037] Preferably, the virtual anchor representation learning network adopts a deep 3D convolutional neural network, including a first convolutional layer, a second convolutional layer, a third convolutional layer, an Inception module, and a Softmax layer connected in sequence; the first convolutional layer uses a 7×7×7 convolutional kernel; the second convolutional layer uses a 1×1×1 convolutional kernel; the third convolutional layer uses a 3×3×3 convolutional kernel; MaxPooling max pooling layers are provided after the first convolutional layer and after the third convolutional layer; the Inception module has 3 convolutional branches, the first convolutional branch is a 1×1×1 convolution, the second convolutional branch is a 1×1×1 convolution followed by a 3×3×3 convolution, and the third convolutional branch is a 1×1×1 convolution followed by a 5×5×5 convolution. Finally, the outputs of the three branches are concatenated and merged to form a single feature representation; the feature representation output by the Inception module is normalized by the Softmax function to obtain an empirical distribution.

[0038] Preferably, the training of the virtual anchor representation learning network through contrastive learning data set includes:

[0039] Construct a contrastive learning data set, where the labeled face image sequence is denoted as , and the face image sequence with pseudo-labels is denoted as ; in and , for each pair of sequence data , if the sequences x i and the sequence x j are both face image sequences of real anchors or both face image sequences of virtual anchors, they are marked as 1, which is a positive sample pair, denoted as , if the sequences x i and the sequence x j are a face image sequence of a real anchor and a face image sequence of a virtual anchor, they are marked as 0, which is a negative sample pair, denoted as ;

[0040] Input the contrastive learning data set into the virtual anchor representation learning network to obtain the empirical distribution of each pair of sequence data. Each pair of sequence data in the contrastive learning data set is the data pair input to the virtual anchor representation learning network and shares the same parameters; after the training of the contrastive learning network, the parameters of the deep 3D convolutional neural network are fixed;

[0041] Specifically, the empirical distribution of the face image sequence is , Represents the virtual anchor representation learning network;

[0042] Calculate the KL divergence for the paired sequence data of the empirical distribution:

[0043]

[0044] Wherein, represents the KL divergence value, represents the marginal parameter, and its random value range is (0, 1); is a positive sample pair, represents a negative sample pair.

[0045] Specifically, the loss function of the virtual anchor representation learning network is:

[0046]

[0047] Wherein, represents the logarithm of the sequence data P in represents the logarithm of the sequence data P in

[0048] After adopting the above solution, the beneficial effects of the present invention are as follows:

[0049] The technology of the present invention can identify virtual anchors and real anchors in video frames, filling the gap in this field for use in anti-fraud, anti-defamation and other fields. Compared with the existing virtual video recognition technology, the present invention focuses on the different changes in facial muscles of virtual anchors and real anchors in terms of emotion when assigning pseudo-labels, and learns the facial features of the anchors through a location-based autoencoder without using manual features. In this way, the local feature differences can be amplified, making the pseudo-label classification more accurate. The KL divergence and loss function designed by the present invention enable the model to continuously optimize the discriminant features of the data with pseudo-labels, and improve the discriminant performance with the contrast loss as the guide. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the flowchart of the core steps in the specific embodiment of the present invention;

[0051] Figure 2 is the schematic diagram of the core steps in the specific embodiment of the present invention;

[0052] Figure 3 is the schematic diagram of the random sampling process in S1 of the specific embodiment of the present invention;

[0053] Figure 4 is the structural diagram of the Inception module in the present invention; To avoid infringement, Figure 2The special shape used for the portrait blocks the face area. Detailed implementation manners

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0055] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.

[0056] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or having more advantages than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for the purpose of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid unnecessary details from obscuring the description of the present invention. Therefore, the present invention is not intended to be limited to the illustrated embodiments, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0057] In order to automatically identify virtual anchors, this specific implementation provides a virtual anchor identification method based on emotion perception and semi-supervised contrast learning. After randomly sampling a live video or a historical video to form a sequence of frames and preprocessing them, the pre-trained SSD face detection network extracts the face representation of each frame. By constructing a contrast learning data set, the face representation passes through a designed virtual anchor representation learning network, and then the learned feature representation passes through a virtual anchor identification network to obtain a prediction result, so as to determine whether the anchor is a virtual anchor. Through a semi-supervised processing method, the present invention can be well trained according to some labeled data, thereby obtaining an identification result.

[0058] The method of the present invention includes a training part and a testing part. For the training part, as Figure 1 and Figure 2 shown, it includes the following steps:

[0059] S1. Randomly sample continuous video frames to obtain a sequence of frames; the continuous video frames include virtual anchor video frames and real anchor video frames.

[0060] S1.1. In the live video stream, sample at a random starting point to generate a continuous sequence of original frames of a fixed duration.

[0061] S1.2. For the continuous sequence of original frames, perform random frame sampling according to a preset discard rate to eliminate redundant video frames and reduce computational complexity. For the live anchor video (frame rate of 30fps, resolution of 720*1280), as Figure 3 shown, randomly sample a continuous sequence with a fixed length of 10 seconds as the original sequence, and then randomly discard some frames with a length of 2 seconds in the original sequence to form the final sequence of frames, with a size of 240*720*1280, that is, the size of one sample is 240*720*1280.

[0062] The above S1.1 and S1.2 represent the sub-steps of S1, and the same applies to the following text.

[0063] S2. Perform image enhancement processing on the sequence of frames; image enhancement processing can improve the robustness of features. Common image enhancement processing includes noise injection, sharpening, cropping, translation, flipping, etc. In this specific implementation, common image enhancement processing can be used.

[0064] S3. Extract the facial regions from the sequence of frames through a pre-trained facial recognition model, and resize the extracted facial regions to obtain a sequence of face images; in this specific implementation, the facial recognition model is a two-dimensional model. Therefore, in actual processing, the size of each picture in the sequence of frames is adjusted, and then the adjusted pictures are stitched together to obtain a sequence of face images. S3 specifically includes the following steps:

[0065] S31. Perform scaling and normalization processing on the preprocessed sequence of frames to obtain the input sequence of frames, so as to construct input features that conform to the input specifications of the deep neural network.

[0066] Specifically, uniformly scale the input sequence of frames to the standard size of 720×720 pixels, with an overall size of 240*720*720, and then perform normalization processing to normalize the pixel values to the [0, 1] interval, and then subtract the average value of each pixel value of the current frame image from each pixel value of each frame image.

[0067] S32. Input the input sequence frames into the face recognition model to obtain face bounding boxes. The face recognition model uses an existing SSD model (with a VGG16 backbone network). The input of the SSD model should be a two-dimensional image, and the specific input size is 720*720. It is necessary to split the sequence frames with an overall size of 240*720*720 obtained in S31 in the first dimension. That is, split the 240-frame sequence into multiple sub-batches, that is, set the batch size batch_size = 16, and then input them into the model in turn to reduce the video memory occupancy. The SSD model in this specific implementation has been pre-trained on a large-scale face dataset. The large-scale face dataset can use both public datasets and non-public datasets. For example, select the public datasets CelebA and WIDER FACE. In the SSD model, multi-level feature maps are generated through the backbone network (VGG16). The sizes of these feature maps are 38×38, 19×19, 10×10, 5×5, 3×3, and 1×1 respectively. The multi-level feature maps can capture face features at different scales. Based on the preset aspect ratio (set to 1:1 in this specific implementation) and scale parameters, prior boxes are generated for each layer of the feature map, and their coordinate offsets, confidence levels, and class labels are predicted. Then, detection boxes with a confidence level lower than the dynamic threshold (set to 0.7) are removed, and the detection boxes within a single frame are de-duplicated according to the intersection over union (set IoU≥0.5), and the result with the highest confidence level is retained to obtain a set of refined and accurate face bounding boxes. Specifically, the SSD model is an existing technology. The above description only describes the technical key points. The model details and model code are all public technologies, and those skilled in the art can know them.

[0068] S33. Crop the face regions through the face bounding boxes, and adjust all face regions to the same size to obtain a sequence of face images with a size of 240*300*300.

[0069] The sequence objects processed in the above S1-S3 include labeled sequence data and unlabeled sequence data. The labeled sequence data comes from virtual anchor video frames and real anchor video frames. Similarly, the unlabeled sequence data comes from virtual anchor video frames and real anchor video frames. For the unlabeled sequence of face images, pseudo-labels need to be obtained, that is, enter S4, and for the labeled sequence of face images, directly enter S5.

[0070] S4. Input the obtained unlabeled sequence of face images into the emotion perception and recognition model for emotion recognition to obtain pseudo-labels.

[0071] The facial muscles of virtual anchors and real anchors show different changes in emotion expressions. Therefore, a position-based autoencoder is used to learn the facial features of the anchors. S4 includes the following operations:

[0072] S41. Obtain the global feature vector for the overall face image through the global feature extraction model ; The global feature extraction model here includes multiple layers of 3D convolution. Through 3D convolution, the spatio-temporal dimensions are gradually compressed, and finally a global feature vector with a dimension of 1*1*1*128 is generated. The settings of each convolutional layer are as follows: the first 3D convolutional layer, the first 3D max pooling layer, the second 3D convolutional layer, the second 3D max pooling layer, the third 3D convolutional layer, and the first adaptive 3D average pooling layer. The convolutional kernels of the first 3D convolutional layer and the second 3D convolutional layer are 3*3*3, the stride is 2, and the padding is 1. The convolutional kernels of the first 3D max pooling layer and the second 3D max pooling layer are 2*2*2, the stride is 2. The convolutional kernel of the third 3D convolutional layer is 3*3*3, the stride is 1, and the padding is 1.

[0073] S42. Divide the overall face image into n*n local region blocks. In this specific embodiment, n is taken as 3, and 3*3 local region blocks are used for illustration; local feature vectors are obtained after performing the local feature extraction model on each local region block , where 3*3 is for each image in the face image sequence. Therefore, the divided local region blocks as a whole are also a sequence. The settings of each convolutional layer are as follows: the fourth 3D convolutional layer, the third 3D max pooling layer, the fifth 3D convolutional layer, the sixth 3D convolutional layer, and the second adaptive 3D average pooling. The convolutional kernels of the fourth 3D convolutional layer, the fifth 3D convolutional layer, and the sixth 3D convolutional layer are 3*3*3, the stride is 2, and the padding is 1. The convolutional kernel of the third 3D max pooling layer is 2*2*2, the stride is 2.

[0074] S43. Calculate the similarity score between the local feature vector and the global feature vector to obtain the position weight score of each local region block. Specifically, the dot product similarity and normalization operations are used to perform the similarity score calculation. The formula is:

[0075] ,

[0076] where, is the similarity score between the i-th local feature vector and the global feature vector, is the global feature vector, is the i-th local feature vector, is the globally learnable weight parameter, is the locally learnable weight parameter, and T is the transpose operation.

[0077] S44. After guiding the local vector through the position weight score, the local vector is then fused with the global feature vector to generate a fused feature map ; The specific content of S44 includes:

[0078] S441. Weight each local feature vector to obtain the weighted local feature vector , specifically:

[0079]

[0080] Among them, is the i-th weighted local feature vector; this operation can amplify the local region features with high correlation with the global feature and suppress the irrelevant regions.

[0081] S442. According to the positions of the originally divided 3×3 grids, arrange the 3×3 weighted local feature vectors in the corresponding positions to form a low-resolution feature map, and adjust the channel dimension through 1×1×1 convolution to generate a recombined feature map .

[0082] S443. Convert the global feature vector into an updated global feature map with the same dimension as the recombined feature map .

[0083] S444. Concatenate the recombined feature map and the updated global feature map along the channel dimension to generate a fused feature map ; the fused feature map combines global information and local information and can improve the reconstruction quality.

[0084] Gradually upsample the fused feature map to the resolution of the original face image through multiple layers of transposed convolution; in this specific implementation, 5 layers of 3D transposed convolution layers are used, including the first 3D transposed convolution layer, the second 3D transposed convolution layer, the third 3D transposed convolution layer, the fourth 3D transposed convolution layer, and the fifth 3D transposed convolution layer. The convolution kernels of the first 3D transposed convolution layer and the second 3D transposed convolution layer are 1*4*4, the stride is 1*2*2, and the padding is 0*1*1. The convolution kernel of the third 3D transposed convolution layer is 1*5*5, the stride is 1*3*z, and the padding is 0*1*1. The convolution kernels of the fourth 3D transposed convolution layer and the fifth 3D transposed convolution layer are 1*6*6, the stride is 1*5*5, and the padding is 0*1*1; multiple layers of transposed convolution gradually increase the resolution, and finally adjust the number of channels to 3. The last layer uses the activation function to constrain the output value to the range of [0,1], that is, gradually upsample to restore the high-resolution face image.

[0085] The loss function of the S41-S44 process is: , which constrains the pixel and semantic consistency between the fused feature map sequence and the original face image sequence g.

[0086] S45. Use the clustering algorithm for the fused feature map Perform unsupervised clustering, and use the obtained classification results as the pseudo-labels of the unlabeled face image sequences. Here, any common unsupervised clustering algorithm can be used for the clustering algorithm. In this specific implementation, the KMeans clustering algorithm is adopted, and the number of clusters is 2, which are "0" or "1" respectively. In this way, each clustered class is assigned a pseudo-label. Since virtual anchors and real anchors are very similar visually and cannot be directly distinguished using binary clustering for the entire image, and the changes in facial muscles of virtual anchors and real anchors are different in terms of emotions, thus operations S41 - S44 are performed. By learning the facial features of the anchors through a location-based autoencoder, in this way, the local feature differences can be amplified.

[0087] S5. Construct a contrastive learning dataset ; The entire contrastive learning dataset contains some labeled real anchor sequences and some unlabeled virtual anchor sequences. Specifically, it includes labeled real anchor face sequences, unlabeled real anchor face sequences, labeled virtual anchor face sequences, and unlabeled virtual anchor face sequences; The labeled face image sequences are denoted as (including labeled real anchor face sequences and labeled virtual anchor face sequences), and the originally unlabeled but pseudo-labeled face image sequences obtained through S4 are denoted as (including the originally unlabeled real anchor face sequences and the originally unlabeled virtual anchor face sequences); In and , for each pair of sequence data , if both sequences x i and sequence x j are both real anchor face image sequences or both are virtual anchor face image sequences, they are marked as 1, which is a positive sample pair, denoted as , if sequence x i and sequence x j are a real anchor face image sequence and a virtual anchor face image sequence, they are marked as 0, which is a negative sample pair, denoted as , it should be noted that sequence x i and sequence x j either both come from , or both come from .

[0088] S6. Input the contrastive learning dataset into the virtual anchor representation learning network to obtain the empirical distribution of each pair of sequence data. Each pair of sequence data in the contrastive learning dataset is the data pair input to the virtual anchor representation learning network and shares the same parameters.

[0089] The virtual anchor representation learning network uses a deep 3D convolutional neural network, including a first convolutional layer, a second convolutional layer, a third convolutional layer, an Inception module, and a Softmax layer connected in sequence; the first convolutional layer uses a 7×7×7 convolutional kernel for extracting low-level information; the second convolutional layer uses a 1×1×1 convolutional kernel for reducing the feature size and total number of parameters; the third convolutional layer uses a 3×3×3 convolutional kernel for extracting high-level information; MaxPooling max pooling layers are provided after the first convolutional layer and after the third convolutional layer; the Inception module is an existing module, such as Figure 4 shown, for extracting temporal features of different scales and enhancing the representation ability of the model. The Inception module includes 3 convolutional branches. The first convolutional branch is a 1×1×1 convolution, the second convolutional branch first passes through a 1×1×1 convolution and then through a 3×3×3 convolution, and the third convolutional branch first passes through a 1×1×1 convolution and then through a 5×5×5 convolution. Then, the outputs of the three branches are concatenated and merged to form a single feature representation. The Inception module is used to extract temporal features of different scales and enhance the representation ability of the model.

[0090] The feature representation output by the Inception module is normalized by the Softmax function to obtain the empirical distribution of the face image sequence of , denotes the virtual anchor representation learning network.

[0091] Calculate the KL divergence for the empirical distribution of the paired sequence data :

[0092]

[0093] where denotes the KL divergence value, denotes the random value range of the marginal parameter (0, 1); is a positive sample pair, denotes a negative sample pair; the positive sample pair represents a dataset of similar sequence frames, and the negative sample pair represents a dataset of dissimilar sequence frames, that is, making the representation values of similar sequence frames larger and the representation values of dissimilar frames smaller.

[0094] The loss function of the virtual anchor representation learning network is:

[0095]

[0096] where denotes the logarithm of the sequence data P in denotes the logarithm of the sequence data P in;

[0097] After training the contrastive learning network for the virtual anchor representation, the parameters of the 3D deep convolutional neural network are fixed. The virtual anchor representation learning network obtains better feature representations through contrastive learning, and then directly classifies the obtained better feature representations.

[0098] S7. Identify the output of the virtual anchor representation learning network through the virtual anchor recognition network; after the training of the virtual anchor representation contrastive learning network in S6, the parameters of the 3D deep convolutional neural network are fixed; specifically, the virtual anchor recognition network adopts a feedforward neural network, followed by a function. The virtual anchor recognition network is used to classify the output of the representation learning network. A conventional classification network can be used. In this specific implementation, a two-layer feedforward neural network is used for classification, and the cross-entropy loss function is used to obtain the distributions of different categories, so as to automatically determine whether the anchor is a virtual anchor.

[0099] In the present invention, the SSD face recognition model, the virtual anchor representation learning network, and the virtual anchor recognition network are trained separately.

[0100] The test part (i.e., the real-time recognition part) of the present invention includes the following steps:

[0101] S100. Sample the continuous video frames to obtain a sequence of frames; the continuous video frames here are unknown frames, that is, it is not known whether they are virtual anchor video frames or real anchor video frames.

[0102] S200. Extract the face regions from the sequence of frames through the trained face recognition model to obtain a sequence of face images.

[0103] S300. Input the sequence of face images into the trained virtual anchor representation learning network to obtain feature representations.

[0104] S400. Input the feature representations into the virtual anchor recognition network to obtain the recognition results.

[0105] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0106] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning, characterized in that, It includes the following steps: Sampling continuous video frames to obtain a sequence of frames; Extracting facial regions from the sequence of frames through a facial recognition model to obtain a sequence of face images; Inputting the sequence of unlabeled face images into an emotion perception recognition model for emotion recognition to obtain pseudo-labels; Inputting the sequence of face images into a virtual anchor representation learning network to obtain feature representations; the virtual anchor representation learning network is trained through a contrastive learning dataset, and the contrastive learning dataset includes a sequence of labeled face images and a sequence of face images with obtained pseudo-labels; Inputting the feature representations into a virtual anchor recognition network to obtain recognition results.

2. The virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning according to claim 1, wherein Before the step of extracting facial regions from the sequence of frames through the facial recognition model to obtain a sequence of face images, it further includes a step of performing image enhancement processing on the sequence of frames.

3. The virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning according to claim 1, characterized in that, The step of sampling continuous video frames to obtain a sequence of frames includes: In a live video stream, sampling at a random starting point to generate a continuous sequence of original frames with a fixed duration; Performing random frame sampling on the continuous sequence of original frames according to a preset discard rate.

4. The virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning according to claim 1, wherein The step of extracting facial regions from the sequence of frames through the facial recognition model to obtain a sequence of face images includes: Performing scaling and / or normalization processing on the sequence of frames to obtain an input sequence of frames; Inputting the input sequence of frames into a facial recognition model to obtain face bounding boxes; Cropping facial regions through the face bounding boxes, and adjusting all facial regions to the same size to obtain a sequence of face images.

5. The virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning according to claim 1, characterized in that, The step of inputting the sequence of unlabeled face images into an emotion perception recognition model for emotion recognition to obtain pseudo-labels includes: The global feature extraction model is used to obtain a global feature vector from the overall face image , where is the dimension of the global feature vector; Divide the overall face image into local region blocks of n*n, and after separately extracting local features for each local region block, obtain local feature vectors ; Calculating similarity scores between local feature vectors and global feature vectors to obtain position weight scores for each local region block; After guiding the local vectors through the position weight scores, the local vectors are fused with the global feature vectors to generate a fused feature map; Using a clustering algorithm to perform unsupervised clustering on the fused feature map, and taking the obtained classification results as pseudo-labels for the sequence of unlabeled face images.

6. The virtual anchor recognition method based on emotion perception and semi-supervised contrast learning according to claim 5, wherein The specific calculation of similarity scores between local feature vectors and global feature vectors is to perform the similarity score calculation by using dot product similarity and normalization operations, and the formula is: ; wherein, is the similarity score between the i-th local feature vector and the global feature vector, is the global feature vector, is the i-th local feature vector, is the globally learnable weight parameter, is the locally learnable weight parameter, and T is the transpose operation.

7. The virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning according to claim 5, characterized in that, After guiding the local vectors through the position weight scores, the local vectors are fused with the global feature vectors to generate a fused feature map includes: Weight each local feature vector to obtain the weighted local feature vector , ; Among them, is the weighted i-th local feature vector; According to the n×n grid positions divided originally, arrange the n×n weighted local feature vectors in the corresponding positions to form a low-resolution feature map, and adjust the channel dimension through 1×1 convolution to generate a recombined feature map ; Convert the global feature vector into an updated global feature map of the same dimension as the reconstructed feature map through a fully connected layer ; Recombinant feature map And the updated global feature map Are concatenated along the channel dimension to generate a fused feature map ; Gradually upsample the fused feature map through multi-layer deconvolution to the resolution of the original face image.

8. The virtual anchor recognition method based on emotion perception and semi-supervised contrastive learning according to claim 1, wherein The virtual anchor representation learning network adopts a deep 3D convolutional neural network, including a first convolutional layer, a second convolutional layer, a third convolutional layer, an Inception module, and a Softmax layer connected in sequence; the first convolutional layer uses a 7×7×7 convolutional kernel; the second convolutional layer uses a 1×1×1 convolutional kernel; the third convolutional layer uses a 3×3×3 convolutional kernel; MaxPooling max pooling layers are provided after the first convolutional layer and after the third convolutional layer; the Inception module has 3 convolutional branches, the first convolutional branch is a 1×1×1 convolution, the second convolutional branch is a 1×1×1 convolution followed by a 3×3×3 convolution, the third convolutional branch is a 1×1×1 convolution followed by a 5×5×5 convolution, and finally the outputs of the three branches are concatenated and merged to form a single feature representation; The feature representation output by the Inception module is normalized by the Softmax function to obtain an empirical distribution.

9. The virtual anchor recognition method based on emotion perception and semi-supervised contrast learning according to claim 1, wherein The training of the virtual anchor representation learning network through the contrastive learning dataset includes: Construct a contrastive learning dataset, where the labeled face image sequence is denoted as , and the face image sequence with pseudo-labels is denoted as ; In and , for each pair of sequence data , if the sequences x i and the sequence x j are both face image sequences of real live streamers or both face image sequences of virtual live streamers, they are marked as 1, which is a positive sample pair, denoted as , if the sequences x i and the sequence x j are a face image sequence of a real live streamer and a face image sequence of a virtual live streamer, they are marked as 0, which is a negative sample pair, denoted as ; Input the contrastive learning dataset into the virtual anchor representation learning network to obtain the empirical distribution of each pair of sequence data. Each pair of sequence data in the contrastive learning dataset is the data pair input into the virtual anchor representation learning network and shares the same parameters. After training by the contrastive learning network, the parameters of the deep 3D convolutional neural network are fixed; Specifically, the face image sequence has an empirical distribution of , denotes the virtual anchor representation learning network; Calculate the KL divergence for the empirical distribution of the paired sequence data : ; Among them, represents the KL divergence value, represents the marginal parameter, and its random value range is (0, 1); is a positive sample pair, represents a negative sample pair.

10. The virtual anchor recognition method based on emotion perception and semi-supervised contrast learning according to claim 9, wherein , the loss function of the virtual anchor representation learning network is: ; Among them, represents the logarithm of the sequence data P in represents the logarithm of the sequence data P in.

Citation Information

Patent Citations

  • Facial parameter identification method and device, electronic equipment and storage medium

    CN112818772A

  • Heartbeat anomaly detection method based on semi-supervised graph contrast learning

    CN115099351A

  • Virtual anchor management system and method

    CN115344891A

  • Virtual anchor distinguishing method and system based on optical action data

    CN117523678A

  • Face emotion recognition method based on dual-stream convolutional neural network

    US20190311188A1