User feature detection method and device

By collecting user feature data in the target application and using pre-trained detection models for detection, the security threats brought by AI face change and sound change are solved, and automated detection and convenient security protection are achieved.

CN120047803APending Publication Date: 2025-05-27LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125254.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology cannot effectively protect the security threats caused by AI face change and sound change, making it difficult for users to realize that the other party uses false information, and the existing solutions have a high threshold and poor practicality.

Method used

Provide a user feature detection method, by collecting user feature data (including face and voice data) in the target application, using a pre-trained detection model to detect, determine whether the user uses fake user features, and outputs alarm information.

Benefits of technology

It realizes automatic detection of AI forgery features in video calls, virtual meetings and other scenarios, improves computer security, lowers user operation thresholds, and provides convenient and easy-to-use AI face-changing and sound-changing security protection solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047803A_ABST
    Figure CN120047803A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a user feature detection method and device, and the method comprises the steps: responding to the operation of a target application, collecting a plurality of pieces of user feature data in the target application, the target application is related to communication, and the user feature data comprises at least one of the face data and voice data of a user; each piece of collected user feature data is detected based on a pre-trained detection model, a detection result is obtained, and the detection model is used for determining whether the input user feature data includes forged user features; and determining whether the user uses the fake user features based on the plurality of detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer recognition, and particularly to a method and device for detecting user characteristics. Background Art

[0002] With the development of artificial intelligence technology, a large number of AI face-swapping and voice-changing tools have emerged. These tools can obtain the face information and voice information of a target user based on a small amount of face images and voice data of the target user, replace the original face in a video with the target face, or replace the voice data during a video or real-time call with the target voice data. Such technologies pose a serious security threat to scenarios such as video calls, virtual meetings, and virtual communications. An attacker can use AI face-swapping and voice-changing tools to change their own face or voice during a video or communication process to that of a specific person (such as a user's relative or friend) to commit illegal and criminal acts such as fraud against the user, seriously threatening the user's personal and property safety.

[0003] To solve the above problems, there are currently some existing works related to AI forged face and AI forged voice detection algorithms. These works have achieved the classification of AI face-swapped videos and normal face videos, or forged voices and real voices through machine learning or neural network models. However, in actual use scenarios, it is difficult to accurately distinguish AI face-swapping and voice-changing from normal faces and voices by human eyes and ears. Therefore, it is difficult for users to realize that the other party is using false information and that they are facing a security threat, so they will not actively use the AI face-swapping detection algorithm for security protection. And if the user does not actively start the detection algorithm, the detection algorithm cannot start itself. At the same time, the existing solutions are limited to the design of the algorithm and have not formed an AI face-swapping and voice-changing security protection solution suitable for computers and easy to use. The user has a high usage threshold and poor practicability. Summary of the Invention

[0004] The embodiments of the present application provide a method for detecting user characteristics, including:

[0005] In response to the running of a target application, collect multiple user characteristic data in the target application, where the target application is related to communication, and the user characteristic data includes at least one of the user's face data and voice data;

[0006] Based on a pre-trained detection model, detect each piece of collected user characteristic data to obtain a detection result, where the detection model is used to determine whether the input user characteristic data includes forged user characteristics;

[0007] Based on multiple detection results, determine whether the user uses forged user characteristics.

[0008] In some embodiments, the method further includes:

[0009] Detect the current application window and determine whether the currently running application window is the target application, where the target application includes at least one application;

[0010] If the current application window is the target application, capture multiple user images displayed in the current application window as the captured display images, or capture the user voice data in the current application as the captured voice data.

[0011] In some embodiments, determining whether the user uses forged user features based on multiple detection results includes:

[0012] Determine the number of the first detection results among the multiple detection results, where the first detection results characterize that the user feature data includes forged user features;

[0013] Calculate and determine the ratio of the first detection results to the multiple detection results;

[0014] Based on the ratio and a preset reference value, calculate and determine whether the user uses forged user features. If the user includes forged user features, output an alarm message.

[0015] In some embodiments, the method further includes:

[0016] If the multiple user feature data includes the feature data of multiple users, respectively detect whether the user features of different users include forged user features to obtain the detection results corresponding to different users;

[0017] When it is determined based on the detection results that the target user includes the forged user features, output an alarm message including the user feature data of the target user.

[0018] In some embodiments, the pre-training of the detection model includes:

[0019] Input the user feature data into the detection model, and the detection model respectively outputs the inherent features and forged features of the user. The user feature data includes the original user feature data and the user feature data including forged features, and the user feature data including forged features is generated based on at least one forgery method;

[0020] Input the forged features into the style processing module to form forged features with the same style as the inherent features;

[0021] After fusing the inherent features and the forged features after style processing, obtain the synthesized user features;

[0022] Determine the similarity between the input user features and the synthesized user features;

[0023] If the similarity does not meet the preset conditions, a pre-training process with iterative combination of the loss function is performed until the preset conditions are met, and the trained detection model is output.

[0024] In some embodiments, the detection model includes a first encoder and a second encoder, and the detection model outputs the inherent features and forged features of the user respectively, including:

[0025] Based on the first encoder, extract the inherent features corresponding to the face data in the input user feature data;

[0026] Based on the second encoder, extract the forged features corresponding to the face data in the input user feature data.

[0027] In some embodiments, the loss function includes:

[0028] The first loss function, which is used to enable the detection model to extract the inherent features and forged features that can restore the original user feature data;

[0029] The second loss function, which is used to make the inherent features and forged features extracted by the detection model stable and not change with the change of the forgery method.

[0030] In some embodiments, the loss function further includes:

[0031] The third loss function, which is used to make the forged features corresponding to different user feature data extracted by the detection model close in distance and far from the forged features of real faces.

[0032] In some embodiments, the detection model further includes a decoder, and after fusing the inherent features with the forged features processed by the style, the synthesized user features are obtained, including:

[0033] Based on the decoder, the inherent features are fused with the forged features processed by the style to obtain the synthesized user features.

[0034] Another embodiment of the present application provides a user feature detection device, including:

[0035] An acquisition module, configured to acquire multiple user feature data in a target application in response to the operation of the target application, where the target application is related to communication, and the user feature data includes at least one of user's face data and voice data;

[0036] A detection module, configured to detect each acquired user feature data based on a pre-trained detection model to obtain a detection result, where the detection model is used to determine whether the input user feature data includes forged user features;

[0037] A determination module, configured to determine whether the user uses forged user features based on multiple detection results.

[0038] Other features and advantages of the present application will be described in the following specification, and part of them will be obvious from the specification, or understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings.

[0039] The technical solutions of the present application will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0040] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings required for the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flowchart of the user feature detection method in the embodiment of the present application.

[0042] Figure 2 It is a schematic application flowchart of the user feature detection method in the embodiment of the present application.

[0043] Figure 3 It is a schematic diagram of the training and detection process of the detection model in the embodiment of the present application.

[0044] Figure 4 It is a structural block diagram of the user feature detection device in the embodiment of the present application. Detailed Embodiments

[0045] Next, the specific embodiments of the present application will be described in detail with reference to the drawings, but it is not a limitation of the present application.

[0046] It should be understood that various modifications can be made to the embodiments disclosed herein. Therefore, the following specification should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope of the present disclosure.

[0047] The drawings included in the specification and constituting a part of the specification illustrate the embodiments of the present disclosure, and are used together with the general description of the present disclosure given above and the detailed description of the embodiments given below to explain the principles of the present disclosure.

[0048] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given as non - limiting examples with reference to the accompanying drawings.

[0049] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application, which have the features as described in the claims and thus are all within the protection scope defined thereby.

[0050] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present disclosure will become more apparent in view of the following detailed description.

[0051] Hereinafter, specific embodiments of the present disclosure will be described with reference to the accompanying drawings; however, it should be understood that the disclosed embodiments are merely examples of the present disclosure, which can be implemented in various ways. Well - known and / or repetitive functions and structures are not described in detail to avoid obscuring the present disclosure with unnecessary or redundant details. Therefore, the specific structural and functional details disclosed herein are not intended to be limiting, but are merely used as a basis and representative basis for the claims to teach those skilled in the art to use the present disclosure in substantially any suitable detailed structure in various ways.

[0052] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present disclosure.

[0053] Next, embodiments of the present application will be described in detail with reference to the accompanying drawings.

[0054] As Figure 1 shown, an embodiment of the present application provides a user feature detection method, including:

[0055] S1: In response to the running of a target application, collect a plurality of user feature data in the target application, the target application being related to communication, and the user feature data including at least one of the user's face data and voice data;

[0056] S2: Detect each piece of the collected user feature data based on a pre - trained detection model to obtain a detection result, the detection model being used to determine whether the input user feature data includes forged user features;

[0057] S3: Determine whether the user uses forged user features based on the plurality of detection results.

[0058] In this embodiment, the target application is related to communication. For example, it can be a virtual meeting application, a video chat application, a network call (including virtual call, mobile network call) application (non-video), etc. The system will monitor the startup of the target application or the startup of a specified function of the target application, such as the video function or the call function. When it monitors the startup of the target application or its specified function, or when it monitors the connection of a video or a call, and the avatar of the other party is output on the screen and the voice of the other party is output from the speaker, at this time, the system will automatically trigger the detection function, and the collection and detection and identification of user feature data can be automatically performed without manual operation. In this way, not only can the user operation be simplified, but also the security of the user can be guaranteed more effectively and in a more timely manner.

[0059] Further, after the system in this embodiment starts the detection function, that is, after starting the detection program, the program will determine the type of data to be collected according to the started function or the type of application, such as video, call, or the information displayed on the current display screen, the information output by the speaker, and collect data such as display images and audio data. Or it can also be default to collect display images and audio data simultaneously. After the program determines the data to be collected, it can collect images of the portrait / avatar located in the application session box on the screen in real time or periodically (the program will identify and only collect images containing face data, or filter after collection and only retain images with face data), or collect the audio data output by the speaker to obtain the corresponding face images and voice data. The face images and voice data are both the user feature data described above. When two dialog boxes are simultaneously displayed on the screen and different portraits / avatars are respectively displayed in the two session boxes, the program will collect the data in both session boxes at this time and identify the data corresponding to different session boxes collected, so as to facilitate the identification of different portraits during subsequent detection. Or the program can also output a prompt for the user to input the identity identifiers of different portraits / avatars. In addition, since some applications simultaneously display the video image of our user and the video image of the other user on the screen, in order to avoid the program from detecting the display image of our user, the program can compare the collected display image with the user feature data of our user pre-stored. When it is determined that the collected display image shows the portrait / avatar of our user, the session box corresponding to the collected display image is determined and marked, so that the program no longer collects the display data in this session box or does not detect the display data corresponding to this session box. For audio data, the program will collect the voice data played by the speaker. Inevitably, the microphone will also simultaneously collect the voice data output by our user. At this time, the program will distinguish and identify our user and the other user according to the audio transmission line and mark based on the identification result, or perform voice recognition of our user on the collected audio data according to the user feature data of our user pre-stored, so as to filter out the voice data of the other user based on the identification result.

[0060] When performing detection, the program inputs the collected user feature data into a pre-trained detection model, which performs unified detection and outputs a detection result on whether the input user feature data includes forged user feature data, such as whether it includes forged face data or forged voice data. To avoid misjudgment, the program in this embodiment separately detects the multiple collected data inputs based on the detection model to obtain multiple detection results, and then combines the multiple detection results for comprehensive analysis and calculation to finally determine whether most of the collected user feature data is forged user feature data. If so, it indicates that the other user has used, or continues to use, forged user features for this communication. At that time, the program can output a prompt message to prompt the user that the other user has used forged user features, or directly output an alarm message to the other party, or output an alarm message to our user, or even directly pause the communication. The specific method is not unique and supports user-defined settings. For example, once it is determined that the other party has used forged user features, the communication is paused and an alarm is issued, etc., so that the program directly executes the indicated content based on the user's preset instructions during operation.

[0061] The above solution in this embodiment can significantly improve computer security, automatically detect AI forged features, trigger the security threat detection function for users automatically in common computer communication scenarios such as video calls and virtual meetings, and then provide security threat detection for users to ensure user safety. Secondly, since the above steps do not require any active operation by the user during the implementation process, and only require the user to passively receive information when the alarm message is output, the solution of this embodiment belongs to an integrated deployment solution. There is no need for the user to collect pictures by themselves and call the detection algorithm, and the use threshold is low. Moreover, it will not affect any functions of the user's normal use of the computer throughout the process, and truly realizes non-intrusive detection. In addition, the solution of this embodiment can also meet the privacy and practical needs. Because the detection in this embodiment is actually triggered by the scenario, and the security detection is automatically triggered only when a specified scenario is monitored, such as sensitive scenarios such as video calls, virtual meetings, and virtual calls, so it will not cause waste of power consumption. Moreover, the program in this embodiment only needs to detect the collected picture data or audio data, and the occupied storage space is small. That is, the overall computing resources and storage resources required by the solution of this embodiment are low, and it can be directly deployed locally for convenient use.

[0062] In one embodiment, the method further includes:

[0063] S4: Detect the current application window to determine whether the currently running application window is the target application, where the target application includes at least one application;

[0064] S5: If the current application window is the target application, capture multiple user images displayed in the current application window as the collected face data, or capture the user voice data in the current application as the collected voice data.

[0065] Exemplarily, when the system performs monitoring, it detects whether the application window displayed on the current screen is the application window of the target application. If so, it activates the detection function. Therefore, if the target application is running in the background and its application window is not displayed on the current screen, even if the target application is in the startup state or running mode, the detection program will not be triggered. Only when the target application is running in the foreground and its application window is displayed on the screen, the system will respond to the running of the target application and activate the detection function. After activating the detection function, the detection program will determine the program type, then determine the type of data to be collected, or simultaneously start the data collection for the displayed image and output audio. For example, the program periodically or continuously captures the user images displayed in the current application window within a period of time to obtain multiple user images for subsequent detection. And the program captures or records the voice data of the other user output from the speaker, earphone, etc. The collection of this voice data can also be periodic or continuous within a period of time, and the specific collection method is not unique. After obtaining the user image / display image containing the face data of the other party and the voice data of the other user, it can be input into the detection model for detection either alternatively or simultaneously according to the detection requirements.

[0066] In another embodiment, determining whether the user includes forged user features based on multiple detection results includes:

[0067] S6: Determine the quantity of the first detection results among the multiple detection results, where the first detection results indicate that the user feature data includes forged user features;

[0068] S7: Calculate and determine the proportion of the first detection results relative to the multiple detection results;

[0069] S8: Based on the proportion value and a preset reference value, calculate and determine whether the user uses forged user features. If the user uses forged user features, output an alarm message.

[0070] In this embodiment, the program collects user feature data of the other party user (the same user) at different times (the number of acquisitions or the number of images acquired can set a threshold, and it is necessary to collect user feature data with a quantity meeting the corresponding threshold), such as the above-mentioned user images and voice data, and then inputs them into the detection model to obtain the detection results corresponding to each input user feature data output by the detection model. After the program obtains the detection results of different user feature data at different times, it will count the number of first detection results whose detection results are forged user features. Assume that the detection results output by the detection model are all binary numbers, that is, all are 0 or 1, where 1 indicates that the corresponding user feature data is forged user feature data, and 0 indicates that the corresponding user feature data is non-forged user feature data. The program calculates the value of the number of 1s, and at the same time counts all the detection results, that is, the total number of 0s and 1s, and then calculates the ratio of the value of the number of 1s to the total number of 0s and 1s, and compares the calculated ratio with a preset reference value. If the comparison result meets the requirements, such as the ratio is greater than the reference value, it indicates that the other party user continuously uses forged user features to interact with our user. At this time, an alarm message is output. On the contrary, if the ratio is less than the reference value, it indicates that the other party user does not use or continuously uses forged user features, or rather, the other party user does not use forged user features most of the time. At this time, continue to monitor, or stop monitoring, or extend the detection period, etc. (for example, if the previous data acquisition period is 1 min, and it is determined that the other party user does not use forged user features after monitoring for a period of time, then extend the data acquisition period to 2 - 5 min, etc.), which can be specifically determined according to the default configuration or the user's custom configuration. Among them, the reference value can be but is not limited to being calculated and determined according to the ROC curve during the pre-training process of the detection model. It can also be determined according to historical experience values, etc., and the method is not unique.

[0071] If there are multiple other party users, the method further includes:

[0072] S9: If multiple pieces of the user feature data include the feature data of multiple users, respectively detect whether the user features of different users include forged user features to obtain the detection results corresponding to different users;

[0073] S10: When it is determined based on the detection results that the target user includes the forged user features, output an alarm message including the user feature data of the target user.

[0074] For example, if the target application supports multi-person video communication or multi-person non-video communication, and it is determined through detection that the current application scenario is multi-person simultaneous online video communication or non-video communication, such as when the target application generates multiple conversation boxes corresponding to different other users, etc., at this time, the program can determine that it involves multi-person communication based on these multiple different conversation boxes of the same type, or when it is determined through the analysis of the detected user feature data that it involves multi-person communication, such as detecting multiple different face data, or detecting multiple voice data with different voiceprint characteristics, etc., then the user feature data of different other users are collected respectively and their identities are marked. For example, the user feature data is marked by using the number of the conversation box, or the user is asked to input the corresponding identity mark, etc. Then, the detection model is used to detect the user feature data of different other users respectively, and multiple detection results corresponding to each other user are obtained. Next, the program analyzes and calculates the multiple detection results of each other user, and then determines the target user (other user) who uses or continuously uses forged user features and their corresponding marks. Finally, an alarm message is generated and output based on the mark of the target user, so that our user can know which other users among the multiple other users use or continuously use forged user features.

[0075] The forged user feature data described in the above embodiments, such as forged face, forged voice, etc., does not include noise reduction and beautification factors within the normal range, but refers to that the whole face and all voice features are virtually forged and do not belong to the normal optimization or beautification range.

[0076] In actual application, taking the detection of forged face as an example, as shown in the figure, the overall process can but is not limited to include:

[0077] Such as Figure 2As shown, the foreground window of the monitoring device determines whether the foreground application is a target communication application such as a video call or a virtual meeting by reading the window name and the status of the voice session at certain time intervals (such as 10 seconds). If the above sensitive applications are detected, functions such as primaryScreen (main display) and grabWindow (grab screen) in the Qt framework (application development framework) are used to intercept the image of the specified application window according to the handle input and convert it into a QImage object, and then save it as an image locally. This step is executed every certain time interval (such as 5s). Then, functions in the Dlib library (open-source C++ library) are used to perform face detection on the intercepted image of the application window. If a face is detected, the time interval of data collection in the previous step is reduced (such as from 5s to 500ms). When the number of intercepted face pictures reaches a certain number (such as 32), the data collection step is paused, and all the pictures are marked as a group, and the number of pictures in the group is recorded as N. Face localization and interception are performed on the pictures in this group to unify the size of each face picture for subsequent detection by the detection model. That is, the input image data is preprocessed according to the input requirements of the detection model. Then, a group of preprocessed face pictures obtained in the above steps are input into the detection model to obtain the detection results of each picture output by the detection model. In this embodiment, each detection result output by the detection model is a binary number, where 1 represents a forged face and 0 represents a non-forged face. The program can judge whether the other user has undergone AI face swapping, that is, whether a forged face is used to communicate with our user, according to the ratio between the number of 1s output in a group of pictures and the total number of pictures N. Specifically, the program will compare the ratio with a preset reference value C. When the ratio is greater than the preset reference value, it is considered that the face of this user is an AI-swapped face. The preset reference value is determined according to the ROC curve during the training of the detection model. If the final detection result is that the other user uses a forged face, a pop-up warning message will be sent to the user. If the detection result is not, the above monitoring and detection steps will continue to be executed at a certain time interval (such as 60 seconds, and the detection interval is extended compared with the above 5s and 500ms).

[0078] In one embodiment, the pre-training of the detection model includes:

[0079] S11: Input user feature data into the detection model, and the detection model respectively outputs the inherent features and forged features of the user. The user feature data includes original user feature data and user feature data containing forged features, and the user feature data containing forged features is generated based on at least one forgery method;

[0080] S12: Input the forged features into the style processing module to form forged features with the same style as the inherent features;

[0081] S13: After fusing the inherent features and the forged features processed by the style, the synthesized user features are obtained.

[0082] S14: Determine the similarity between the input user features and the synthesized user features.

[0083] S15: If the similarity does not meet the preset conditions, a pre-training process with iterative combination of the loss function is performed until the preset conditions are met, and the trained detection model is output.

[0084] For example, taking the detection of forged faces as an example, such as the detection of AI face swapping, the user feature data input into the detection model in this embodiment is actually training data, which includes user feature data corresponding to different users. The user feature data of each user includes real user feature data and user feature data corresponding to this user forged by a forgery method. To enhance the training effect, the forged user feature data used in this embodiment needs to be generated by various attack methods respectively. Additionally, methods such as image compression, horizontal flipping, rotation, Gaussian blur, brightness adjustment, and contrast adjustment can be used to perform image transformation on the user feature data, that is, the face image, and the transformed image is also used for training to improve the detection accuracy of the detection method for forged faces after different compression encodings and image beautifications.

[0085] In this embodiment, the detection model is composed of an encoder and a decoder. For example, the detection model includes a first encoder, a second encoder, and a decoder. During training, the encoder is connected to the encoder through a style transfer module (style processing module) as a whole for training. When the detection model outputs the inherent features and forged features of the user respectively, it includes:

[0086] S11: Based on the first encoder, extract the inherent features corresponding to the face data in the input user feature data.

[0087] S12: Based on the second encoder, extract the forged features corresponding to the face data in the input user feature data.

[0088] Then, the two encoders encode the extracted features into hidden layer vectors, and the decoder restores the hidden layer vectors into face images. The purpose of using the decoder is to provide a loss function for optimizing the encoder to make the extracted features more accurate. After the training is completed and the detection model is put into use, only the encoder is used during detection. A binary classifier is formed by adding a fully connected layer and a softmax layer after the original encoder to detect AI face-swapped pictures based on this classifier, that is, to detect forged faces.

[0089] In this embodiment, the encoders all use the Xception model (it can also be other model architectures, which is only used as an exemplary illustration here and is not a feature limitation) as the basic architecture. The Xception model mainly consists of parts such as the Conv2D layer, the BatchNorm layer, the SeparableConv2d layer, the MaxPooling2D layer, the ResNet layer, and the ReLU loss function. From the above two steps of feature extraction, it can be seen that in this embodiment, in order to decouple the inherent features and forged features of the face (face data), two Xception models with exactly the same structure but different parameters are used as encoders to respectively extract and output the inherent features and forged features of the face data. During the detection process in the non-training stage, the encoder will connect the output forged features to the fully connected layer and the softmax layer to output the final binary classification result, that is, the detection result.

[0090] The decoder in this embodiment is used to fuse the inherent features and the forged features after style processing to obtain the synthesized user features. For example, the decoder in this embodiment also uses the Xception model as the basic architecture. To ensure that the output image has the same size as the input image, the decoder changes the downsampling process in the Xception model to an upsampling process. Further, the input of the decoder is composed of two parts: the inherent features and the forged features of the real face data and the forged face data corresponding to the same user. It uses AdaIN (Adaptive Instance Normalize, style transfer module / style processing) to input the forged features as the style of the inherent features, that is, to perform style transfer of the forged features to the inherent features, and then fuses the inherent features and the forged features after style transfer into one feature, that is, to obtain the fused user features.

[0091] Specifically, the inherent features and the forged features in this embodiment are generated by the first encoder and the second encoder by processing the input face images. These two encoders cannot identify forged images and real images. The two encoders are only used to respectively extract the inherent features and the forged features of the obtained face images. That is, when a real face is input into the two encoders, the inherent features and the forged features corresponding to the real face image will be obtained. The forged features are the features used to distinguish forged features, that is, the features used to distinguish whether they are forged features, and they will be affected by the forgery method. Different forgery methods will have an impact on the extracted forged features. The inherent features are the features used to describe the inherent features of the face. Due to different facial features of different people, their corresponding inherent features are also different.

[0092] Before training the detection model, the parameters of the encoder are the same, and it cannot accurately extract the inherent features and the forged features. Therefore, it needs to be trained. The training process and the testing process can refer toFigure 2 and Figure 3 As shown in Figure 3 , first, training data is prepared, including original user feature data and user feature data containing forged features, that is, real user feature data and forged user feature data corresponding to the same user. The forged feature data can be forged by various forging methods. Then, it is input into the first encoder and the second encoder to obtain multiple inherent features and multiple forged features corresponding to the same user respectively. After that, the style of the inherent features is determined, and then the style transfer module is used to transfer the style of the inherent features to the forged features, so that the forged features are consistent with the style of the inherent features. The style is similar to the filter style. When the style transfer is completed, the inherent features and the forged features after style transfer are feature fused to form fused features, and then input into the decoder. The decoder decodes to obtain a synthesized face image generated corresponding to the fused features. Then, the similarity between the synthesized face image and the image in the original input encoder is calculated and fed back to the loss function. The loss function adjusts the parameters of the encoder and continues iterative training until the number of iterations reaches a specified value or the similarity always meets the requirements. When the training of the encoder is completed and the parameters of the encoder are determined, it can be connected to a fully connected layer, etc. to form a binary classifier, which is used for forged face detection of the input image.

[0093] Based on the above content, it can be seen that the training process of the encoder is actually determined by the loss function. Specifically, it tells the encoder in the form of a mathematical expression in which direction to train the parameters, and accordingly adjusts the parameters of the encoder. In this embodiment, three loss functions are designed in total, including:

[0094] The first loss function is used to enable the detection model to extract inherent features and forged features that can restore the original user feature data;

[0095] The second loss function is used to make the inherent features and forged features extracted by the detection model stable and not change with the change of the forging method.

[0096] The third loss function is used to make the forged features corresponding to different user feature data extracted by the detection model close in distance and far from the forged features of real faces.

[0097] Specifically, assume that the first encoder E inh is used to extract the inherent features f inh of the face, and the second encoder E att is used to extract the forged features f att, to ensure that the features extracted by the encoder contain the important parts of the face, and the inherent features provide the basic face information and the forged features provide the relevant features for whether the face is forged or not, this embodiment designs a first loss function to measure whether the inherent features and the forged features can restore the original image. If the image can be restored, it means that these two features contain most of the information of this picture and meet the requirements, and the encoder needs to be trained in this direction. The first loss function L rec is as follows:

[0098] L rec = ||x, D(AdaIN(f inh (x), f att (x)))|| 2

[0099] where AdaIN is the style transfer module, D is the encoder, and x is the original image.

[0100] The second loss function is used to measure the similarity between the inherent features and the similarity between the forged features of the forged face pictures when forging the same face with different forging methods (each forging method will obtain a forged face picture). If the former distance is small, it means that the inherent features of the same face are basically unchanged, so it means that the inherent features extracted by the first encoder are stable and related to the user's face. When the user's face remains unchanged, the inherent features remain unchanged, and only when the user changes, the inherent features will change, and this state meets the requirements; if the latter distance is small, it means that for the face pictures generated by different forging methods, the forged features extracted by the second encoder are basically unchanged, proving that the extracted forged features are stable and will not change due to the change of the forging method, and this state also meets the requirements. Through the second loss function, the training direction of the encoder can be constrained by the crossbeam result to make the extracted features more accurate. The second loss function L exc includes:

[0101]

[0102] fig k_mn = D(AdaIN(f inh (x k_m ), f att (x k_n )))

[0103] fig k_nm = D(AdaIN(f inh (x k_n ), f att (x k_m )))

[0104] where m, n are different forging methods, xk_m and x k_n are forged face data obtained by using different forgery methods for the same face data k, where K is a face image set, and fig k_mn , fig k_nm are synthetic face images corresponding to different forgery methods respectively.

[0105] Both the first loss function and the second loss function need to combine the fusion features to optimize the encoder.

[0106] To ensure that the forged features can be used to distinguish between forged faces and real face images, it is necessary to aggregate the forged features of different forged faces in the feature space and keep them away from the forged features of normal faces. This characteristic distribution conforms to the characteristics of diverse styles and difficult aggregation of normal faces in real situations. Therefore, a third loss function L a is designed. That is, the third loss function is used to measure the distance between the forged features of multiple forged faces and the distance between the forged features of forged faces and the forged features of real faces. If the former distance is small and the latter distance is large, it indicates that the forged features are different between real faces and forged faces. Then, the forged features at this time can be used to distinguish real faces and forged faces, so it meets the training requirements. The third loss function L a includes:

[0107]

[0108] where X a is the forged face data set, X b is the real face data set, x i and x j are different forged face samples, and x z is the real face sample.

[0109] The training direction of the encoder is determined by the above three loss functions. Finally, the features extracted by the two trained encoders meet the requirements and can accurately extract the inherent features and forged features required to identify forged faces.

[0110] During training, although mainly the encoder is trained, it is still trained as a whole with the encoder and the decoder. Therefore, its loss function is L = L rec + βL exc + αL a , where α and β are weight hyperparameters. After training is completed, the output layer of the encoder is connected to the fully connected layer and the softmax layer to output the final binary classification result, that is, the detection result. The parameters of the fully connected layer are trained separately while freezing the parameters of the encoder.

[0111] Such as Figure 4As shown in the figure, another embodiment of the present application also provides a user feature detection device 100, including:

[0112] An acquisition module, configured to collect multiple user feature data in a target application in response to the operation of the target application, where the target application is related to communication (video and / or audio), and the user feature data includes at least one of face data and voice data;

[0113] A detection module, configured to detect each piece of collected user feature data based on a pre-trained detection model to obtain a detection result, where the detection model is used to determine whether the user feature data includes a forged user feature detection result;

[0114] A determination module, configured to determine whether the user includes a forged user feature based on multiple detection results.

[0115] In some embodiments, the device further includes:

[0116] Detect the current application window, and determine whether the currently running application window is the target application, where the target application includes at least one application;

[0117] If the current application window is the target application, intercept multiple user images displayed in the current application window as the collected display images, or intercept the user voice data in the current application as the collected voice data.

[0118] In some embodiments, determining whether the user includes a forged user feature based on multiple detection results includes:

[0119] Determine the number of first detection results among the multiple detection results, where the first detection result indicates that the user feature data includes a forged user feature;

[0120] Calculate and determine the ratio of the first detection result to the multiple detection results;

[0121] Based on the ratio and a preset reference value, calculate and determine whether the user includes a forged user feature. If the user includes a forged user feature, output an alarm message.

[0122] In some embodiments, the detection module is further configured to, when the multiple pieces of user feature data include the feature data of multiple users, respectively detect whether the user features of different users include forged user features to obtain the detection results corresponding to different users;

[0123] The device further includes an output module, configured to output an alarm message including the user feature data of the target user when it is determined based on the detection result that the target user includes the forged user feature.

[0124] In some embodiments, the pre-training of the detection model includes:

[0125] Input user feature data into the detection model, and the detection model respectively outputs the inherent features and forged features of the user. The user feature data includes original user feature data and user feature data containing forged features, and the user feature data containing forged features is generated based on at least one forgery method;

[0126] Input the forged features into the style processing module to form forged features with the same style as the inherent features;

[0127] After fusing the inherent features and the forged features processed by the style, obtain the synthesized user features;

[0128] Determine the similarity between the input user features and the synthesized user features;

[0129] If the similarity does not meet the preset conditions, then perform an iterative pre-training process in combination with the loss function until the preset conditions are met, and output the trained detection model.

[0130] In some embodiments, the detection model includes a first encoder and a second encoder. The detection model respectively outputs the inherent features and forged features of the user, including:

[0131] Extract the inherent features corresponding to the face data in the input user feature data based on the first encoder;

[0132] Use the second encoder to extract the forged features corresponding to the face data in the input user feature data.

[0133] In some embodiments, the loss function includes:

[0134] A first loss function, which is used to enable the detection model to extract inherent features and forged features that can restore the original user feature data;

[0135] A second loss function, which is used to make the inherent features and forged features extracted by the detection model stable and not change with the change of the forgery method.

[0136] In some embodiments, the loss function further includes:

[0137] A third loss function, which is used to make the forged features corresponding to different user feature data extracted by the detection model close in distance and far from the forged features of real faces.

[0138] In some embodiments, the detection model further includes a decoder. After fusing the inherent features with the forged features processed by the style, the synthesized user features are obtained, including:

[0139] Based on the decoder, the inherent features are fused with the forged features processed by the style to obtain the synthesized user features.

[0140] Another embodiment of the present application further provides an electronic device, including:

[0141] One or more processors;

[0142] A memory configured to store one or more programs;

[0143] When the one or more programs are executed by the one or more processors, the one or more processors implement the user feature detection method described in any one of the above.

[0144] Furthermore, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the program is executed by a processor, the user feature detection method described above is implemented. It should be understood that each solution in this embodiment has the corresponding technical effects in the above method embodiment, and will not be elaborated here.

[0145] Furthermore, an embodiment of the present application further provides a computer program product. The computer program product is tangibly stored on a computer-readable medium and includes computer-readable instructions. When the computer-executable instructions are executed, at least one processor executes the user feature detection method in the above embodiments.

[0146] It should be noted that the computer storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable medium can, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access storage medium (RAM), a read-only storage medium (ROM), an erasable programmable read-only storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only storage medium (CD-ROM), an optical storage medium, a magnetic storage medium, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program configured to be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, antenna, optical cable, RF, etc., or any suitable combination of the above.

[0147] In addition, those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0148] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing the process Figure 1One process or multiple processes and / or boxes Figure 1 A system of functions specified in one box or multiple boxes.

[0149] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction system that implements the functions specified in Figure 1 One process or multiple processes and / or boxes Figure 1 One box or multiple boxes.

[0150] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present application is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.

Claims

1. A user feature detection method, comprising: In response to the running of a target application, collecting a plurality of user characteristic data in the target application, wherein the target application is related to communication, and the user characteristic data includes at least one of face data and voice data of the user; Detecting each collected user feature data based on a pre-trained detection model to obtain a detection result, wherein the detection model is used to determine whether the input user feature data includes a forged user feature; It is determined whether the user uses a fake user feature based on the plurality of detection results.

2. The user feature detection method according to claim 1, further comprising: Detecting a current application window to determine whether the currently running application window is a target application, wherein the target application includes at least one application; If the current application window is the target application, multiple user images displayed in the current application window are captured as collected face data, or user voice data in the current application is captured as collected voice data.

3. The user feature detection method according to claim 1, wherein determining whether the user uses a forged user feature based on a plurality of detection results comprises: determining the number of first detection results among the plurality of detection results, the first detection results representing that the user feature data includes a forged user feature; Calculate and determine the proportion of the first detection result relative to the plurality of detection results; Based on the proportion value and the preset reference value, it is calculated to determine whether the user uses a fake user feature, and if the user uses a fake user feature, an alarm message is output.

4. The user feature detection method according to claim 1, further comprising: If the plurality of user feature data include feature data of a plurality of users, detecting whether the user features of different users include forged user features respectively, and obtaining the detection results corresponding to the different users; When it is determined based on the detection result that the target user uses the forged user feature, an alarm message including user feature data of the target user is output.

5. According to the user feature detection method of claim 1, the pre-training of the detection model comprises: Inputting user feature data into a detection model, the detection model outputting the inherent features and forged features of the user respectively, the user feature data comprising original user feature data and user feature data containing forged features, the user feature data containing forged features being generated based on at least one forging method; Inputting the forged feature into a style processing module to form a forged feature consistent with the inherent feature style; After fusing the inherent features with the forged features after style processing, the synthesized user features are obtained; Determining the similarity between the input user features and the synthesized user features; If the similarity does not meet the preset conditions, an iterative pre-training process is performed in combination with a loss function until the preset conditions are met, and the trained detection model is output.

6. The user feature detection method according to claim 5, wherein the detection model comprises a first encoder and a second encoder, and the detection model outputs the user's inherent features and forged features respectively, including: Extracting inherent features corresponding to face data from the input user feature data based on the first encoder; The second encoder is used to extract forged features corresponding to face data from input user feature data.

7. The user feature detection method according to claim 5, wherein the loss function comprises: a first loss function, wherein the first loss function is used to enable the detection model to extract inherent features and forged features that can restore original user feature data; The second loss function is used to stabilize the inherent features and forged features extracted by the detection model and does not change with the change of the forgery method.

8. The user feature detection method according to claim 5, wherein the loss function further comprises: A third loss function, wherein the third loss function is used to make the forged features corresponding to different user feature data extracted by the detection model close in distance and away from the forged features of the real face.

9. The user feature detection method according to claim 5, wherein the detection model further comprises a decoder, and the synthesized user feature is obtained by fusing the inherent feature with the forged feature after style processing, comprising: The inherent features are fused with the forged features after style processing based on the decoder to obtain the synthesized user features.

10. A user feature detection device, comprising: A collection module, configured to collect a plurality of user characteristic data in the target application in response to the running of the target application, wherein the target application is related to communication, and the user characteristic data includes at least one of face data and voice data of the user; A detection module, used to detect each collected user feature data based on a pre-trained detection model to obtain a detection result, wherein the detection model is used to determine whether the input user feature data includes a forged user feature; The determination module is used to determine whether the user uses a forged user feature based on multiple detection results.