A fraud warning system, method, device and medium for video calls

By designing a fraud warning system for video calls, the facial and voiceprint features are extracted and reorganized in real time, and fraud warnings are carried out in combination with security detection models, the problem of difficulty in preventing video call fraud in existing technologies is solved, and real-time detection and early warning of AI face swaps and AI sound meter is achieved.

CN119181381BActive Publication Date: 2025-06-06湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411256462.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-06-06
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively prevent video call fraud using artificial intelligence face-changing technology, especially when verifying in the event, criminals can use technical means to verify cracking or false verification.

Method used

A fraud warning system for video calls is designed, including a data storage module, a feature extraction module, a call exception monitoring module, a feature restructuring module, a model training module and a fraud warning module. The system extracts and recombines the facial and voiceprint features in real time, and combines the security detection model to conduct fraud warnings.

Benefits of technology

Real-time detection and early warning of AI face swaps and AI sound swaps during video calls has been realized, the ability to prevent fraudulent behavior has been improved, and property losses caused by AI face swaps and AI sound swaps are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181381B_ABST
    Figure CN119181381B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of call security technology, and specifically to a fraud warning system, method, device and medium for video calls, wherein the system includes a data storage module, a feature extraction module, a call anomaly monitoring module, a feature reorganization module, a model training module and a fraud warning module. The advantage of the present invention is that the present invention can detect whether the face and voice of a person in a video call have undergone AI face-changing and AI voice-onomatopoeia. Different from the prior art of post-detection of video calls, the present invention discloses a fraud warning system that can monitor video calls in real time, can detect AI face-changing and AI voice-onomatopoeia in a timely manner and give a warning, and the present invention can be applied to any scenario where video calls can be made, and is used to prevent property losses caused by AI face-changing and AI voice-onomatopoeia fraud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of call security technology, and in particular to a fraud early warning system, method, device and medium for video calls. Background Art

[0002] With the advent of the big data era, people's daily life and learning are identified and tracked by big data, and even more so, their personal preferences can be accurately understood, and customers can be accurately classified and recommended. With the increasing transparency of personal information, cases of personal information leakage and fraud using personal information are common, and the methods of criminals are endless, making it difficult for people to guard against them. Among them, the use of technical means to commit telecommunications fraud, resulting in property losses, is everywhere. Criminals use artificial intelligence technology to replace their voices and appearances with the victims' family and friends, and use this as a condition for property fraud. This is very unfriendly to those groups who are not familiar with technical means. They lack the knowledge of emerging technical means and cannot take correct countermeasures in time, resulting in huge property losses. Therefore, it is urgent to study a fraud warning method for artificial intelligence face-changing technology.

[0003] For this kind of property fraud pretending to be relatives and friends, more preventive measures start from information authentication. To this end, the academic community has conducted a lot of research on the uniqueness of biometrics, and reduced the possibility of identity fraud by verifying biometrics such as fingerprints, faces, palm prints, irises, and voiceprints. However, in most cases, the above methods need to be verified when the client is known, and it is a pre-verification, while most frauds require in-process verification, and the criminals of fraud will also use technical means to verify and crack or pseudo-verify. This requires the study of a solution that can apply the above biometric recognition technology to in-process verification to further strengthen the prevention of fraud.

[0004] In summary, there is an urgent need for a fraud warning system, method, device and medium for video calls to solve the problems in the prior art. Summary of the invention

[0005] The present invention aims to provide a fraud warning system, method, device and medium for video calls. The specific technical scheme is as follows:

[0006] A fraud warning system for video calls, including a data storage module, a feature extraction module, a call anomaly monitoring module, a feature reorganization module, a model training module and a fraud warning module:

[0007] Data storage module: including a first memory and a second memory, wherein the first memory stores the face and voiceprint features extracted during each system operation; the second memory stores the fusion features extracted during each system operation;

[0008] Feature extraction module: performs real-time feature extraction on the video call, obtains the facial features and voiceprint features of the video call, stores the facial features and voiceprint features in the first memory when the video call does not involve abnormalities, and sends the facial features and voiceprint features to the feature reconstruction module when the video call involves abnormalities;

[0009] Abnormal call monitoring module: query the bearer device information of both parties in the video call, monitor the video call process in real time, and determine whether the video call process involves abnormalities;

[0010] Feature Recombination Module: Recombines facial features and voiceprint features to form recombined facial features and voiceprint features, and then calls the feature training module to standardize the recombined facial features and voiceprint features and transmit them to the fraud warning module;

[0011] Model training module: calling the facial features and voiceprint features in the first memory to establish a data set, dividing the data set into a training set, a test set and a verification set, building a pre-trained converter model, training the pre-trained converter model based on the training set, the test set and the verification set, and obtaining a security detection model;

[0012] Fraud warning module: calls the security detection model to perform security detection on the facial features and voiceprint features after feature reorganization, outputs fraud warning results, and performs fraud warning based on the fraud warning results.

[0013] Preferably, in the feature extraction module, the image feature extraction algorithm is implemented by the Hog feature algorithm, and the process is as follows:

[0014] Gamma correction of positioning images for video calls;

[0015] Calculate the gradient value of each pixel in the positioning image to obtain the gradient map. The calculation formula is:

[0016]

[0017] Among them, g represents the total gradient strength value, g x represents the horizontal gradient, g y represents the vertical gradient; θ represents the gradient direction, which takes an absolute value and ranges from 0 to 180°;

[0018] Calculate the gradient histogram and perform sample normalization to obtain facial features.

[0019] Preferably, in the feature extraction module, the process of the voiceprint recognition algorithm is as follows:

[0020] Extracting frame-level features of speech frames through a time-delay neural network layer;

[0021] The frame-level features of all frames in the input sequence are averaged and their standard deviations are taken through the statistical pooling layer;

[0022] concatenating the mean and standard deviation as segment-level features;

[0023] The segment-level features are classified through a standard feed-forward network to obtain the voiceprint features.

[0024] Preferably, in the call anomaly monitoring module, real-time monitoring of the video call process is specifically: locating the IP addresses and login device information of both parties in the video call, and when the current login location and login device information are different from the commonly used login location and login device information, the video call is judged to be abnormal.

[0025] Preferably, in the feature recombination module, the feature recombination process is as follows:

[0026] Adding random noise to the features extracted in the aforementioned feature extraction module;

[0027] Randomly select a part of the noise-contaminated features, divide the features before and after adding noise into blocks evenly according to the feature size and locate the coordinates. Rely on the coordinate position to achieve one-to-one correspondence between the feature areas before and after the noise pollution. The similarity of the features before and after the noise pollution in the adjacent area is used for positioning. The error term ε is added to the calculation of the similarity. The specific formula for similarity calculation is as follows:

[0028] If the facial feature f' is contaminated by noise i And the noise-contaminated voiceprint feature v' j All of them cannot be recognized, at this time f' i 、v' j If it is 0, the similarity range of the pollution feature is calculated according to the characteristics of the adjacent areas of the pollution feature, and the interval value is taken. The similarity calculation formula is as follows:

[0029] f' in =f in +ε;f' i ∈f' in ;

[0030] v' jn =v in +ε;v' j ∈v' jn ;

[0031] Among them, f in represents the features of the neighboring area where the facial features are initially extracted; f'in Represents the neighboring area features of the face features polluted by noise; v in Indicates the features of the neighboring area where the voiceprint features are initially extracted; v' jn Represents the neighboring area features of the voiceprint features after being contaminated by noise;

[0032] If the facial feature f' is contaminated by noise i And the noise-contaminated voiceprint feature v' j can be identified, then f' i 、v' j Not 0, the similarity calculation formula is:

[0033] f' i =f i +ε;

[0034] v' j =v j +ε;

[0035] Among them, f i represents the extracted facial features; v j represents the extracted voiceprint features;

[0036] The Euclidean distance is used to calculate the similarity of different features in the same area before and after adding noise. The specific formula is as follows:

[0037]

[0038] Among them, f(f i ,f' i ) indicates f i and f' i The similarity between j ,v' j ) represents vj and v' j The similarity between them; m represents the number of polluted facial features in the extracted facial features; n represents the number of polluted voiceprint features in the extracted voiceprint features;

[0039] The smaller the distance, the higher the similarity. The position with low similarity is located as a noise contaminated position.

[0040] After locating the noise point, the position feature is removed, and the feature is reorganized by calling the feature of the corresponding position area in the first memory. The reorganization process is as follows:

[0041] Similarly, the original feature coordinates of the corresponding position in the first memory are located by calculating the feature matching degree before and after the noise pollution, and then the original features are added to the corresponding area to obtain new features. After that, the security detection model is called to perform similarity detection on the feature fusion degree.

[0042] Preferably, in the model training module, the model training process is as follows:

[0043] Build the network structure of the Transformer model;

[0044] The binary cross entropy loss function is used to estimate the similarity between the actual output probability and the expected output probability. The probabilities predicted for the positive and negative classes are p and 1-p. The specific formula is as follows:

[0045]

[0046] Among them, L represents the binary cross entropy loss, y i represents the label of sample i, the positive class is 1 and the negative class is 0; p i It represents the probability that sample i is predicted as the positive class. The smaller the cross entropy value, the closer the probability distribution of the two is.

[0047] Regularize the model;

[0048] To evaluate the model, the calculation formula is:

[0049]

[0050] Among them, Ac represents accuracy; Sn represents sensitivity; TP represents true positive; FP represents false positive; TN represents true negative; FN represents false negative;

[0051] Model optimization, using the adaptive matrix estimation optimization algorithm to optimize the model, the process is as follows:

[0052]

[0053] m t =β 1 m t-1 +(1-β 1 ) t ;

[0054]

[0055]

[0056]

[0057] Among them, g t represents the gradient; m t represents the first-order moment estimate of the gradient; v t represents the second-order moment estimate of the gradient; represents the corrected first-order moment estimate; represents the corrected second-order moment estimate; represents a dynamic constraint on the learning rate; θ represents the parameter to be solved; represents the gradient of parameter θ; f t (θ t-1 ) represents the objective function to be optimized; β 1 is the first-order moment attenuation coefficient; β 2 is the second-order moment attenuation coefficient; g t and is the gradient; ∈ represents a constant, which is used to prevent the denominator from being zero;

[0058] The optimized model is a safety detection model.

[0059] Preferably, in the fraud warning module, the process of security detection is as follows:

[0060] Data envelopment analysis is used to assign weights to the facial features and voiceprint features after feature reorganization, the facial features and voiceprint features are input into the security detection model, the feature fusion values ​​in the second memory are called for similarity comparison, and the fraud warning results are output.

[0061] In addition, the present invention also includes a fraud warning method for video calls, using the above-mentioned fraud warning system, the method comprising:

[0062] Abnormal call monitoring: Monitor the IP address, login device and other information of the video call and compare it with the commonly used IP address and login device information for abnormal monitoring. If an abnormality is detected, security detection is performed. If no abnormality is detected, feature extraction continues.

[0063] Security detection: By inputting facial features and voiceprint features into the security detection model, the fraud warning result is output. Based on the fraud warning result, it is determined whether the video call has AI face-swapping or AI voice-swapping. If there is no AI face-swapping or AI voice-swapping, the video call is safe and the system is automatically operated. If there is AI face-swapping or AI voice-swapping, it indicates that the video call may have a fraud risk, and a fraud warning message is issued to the video call user without any abnormality;

[0064] The system runs automatically: facial features and voiceprint features are collected through the feature extraction module, facial and voiceprint fusion features are collected through the model training module, the data storage module is continuously updated based on the facial features, voiceprint features and facial and voiceprint fusion features, and the security monitoring model is updated based on the data storage module.

[0065] In addition, the present invention also includes a computer device, including a memory and a processor;

[0066] The memory is used to store a computer program executable on the processor;

[0067] The processor is used to implement the steps of the above-mentioned fraud warning method when executing the computer program.

[0068] In addition, the present invention also includes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the fraud warning method as described above are implemented.

[0069] The application of the technical solution of the present invention has the following beneficial effects:

[0070] The present invention can detect whether the face and voice of a person in a video call have undergone AI face-swapping and AI voice-stuffing. Different from the prior art of post-detection of video calls, the present invention discloses a fraud early warning system that can monitor video calls in real time, and can detect AI face-swapping and AI voice-stuffing in time and give early warnings. The present invention can be applied to any scenario in which video calls can be made, and is used to prevent property losses caused by AI face-swapping and AI voice-stuffing fraud. In addition, after the call monitoring is abnormal, the system of the present invention will add random noise to the features extracted in real time during this operation. If one party to the call uses AI face-changing and voice-sounding, the accuracy of the face and voiceprint combined by the technology will be reduced by adding random noise, and then a part of the noise pollution points will be randomly located, and the feature parts polluted by the noise points will be deleted, and then the features at the same position in the storage database will be called to add them for feature recombination, thereby reducing the similarity between the AI ​​technology face-changing and voice-sounding and the real face and voiceprint. If one party in the video call uses AI face-changing and AI voice-sounding, the edge connection between the reorganized feature parts and the unreorganized parts will be less than the connection between the features at the same position in the inventory and other features, thereby affecting the similarity between the fusion feature values ​​of the face and voiceprint and the fusion feature values ​​in the inventory, thereby realizing the recognition of AI face-changing and voice-sounding.

[0071] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0073] Figure 1 It is a flow chart of the steps of the fraud early warning method in the preferred embodiment of the present invention.

[0074] Figure 2It is a schematic diagram of the process of feature recombination in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0075] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0076] Example:

[0077] like Figure 1 and Figure 2 As shown, this embodiment discloses a fraud warning system for video calls, including a data storage module, a feature extraction module, a call anomaly monitoring module, a feature reorganization module, a model training module and a fraud warning module:

[0078] Data storage module: includes a first memory (data storage module 1) and a second memory (data storage module 2). The first memory stores the face and voiceprint features extracted during each system operation, and calls the feature data in the first memory for feature reorganization after an abnormality is captured; the second memory stores the fused features extracted during each system operation, and performs a similarity comparison with the fused features after feature reorganization.

[0079] Feature extraction module: perform real-time feature extraction on the video call to obtain the facial features and voiceprint features of the video call. After feature extraction, if the call does not involve abnormalities, the extracted facial and voiceprint features are stored in the first memory after the call ends. If the call involves abnormalities, the extracted features enter the feature reorganization module for reorganization, and then call the feature training model to standardize the facial features and voiceprint features after feature reorganization and transmit them to the fraud warning module; establish a data set with all facial features and voiceprint features in the first memory, input the data set into the model training module, and output the feature fusion value as the data in the second memory;

[0080] Call anomaly monitoring module: The IP addresses, login devices and other bearer device information of both parties of the call are included in the call monitoring module, and the video call is monitored in real time for anomaly monitoring. If the video call is monitored to have information mismatch, the feature reorganization module is entered to process the extracted features, and then the fraud warning module is called for security detection. If there is no anomaly, the features extracted by the feature extraction module are stored in the first memory after the call ends;

[0081] Feature Recombination Module: After capturing an anomaly, random noise is added to the features extracted by the feature extraction module, and then some noise points are located and removed after being located. The historical feature data of the same position stored in the data storage module is called and added to the corresponding position to form new face and voiceprint features to be fused. The feature training module is then called to standardize the new face and voiceprint features and transmit them to the fraud warning module.

[0082] Model training module: divide the data set into a training set, a test set and a validation set, construct a pre-trained converter model, and train the pre-trained converter model based on the training set, the test set and the validation set to obtain a security detection model;

[0083] Fraud warning module: call the security detection model to perform security detection on the facial features and voiceprint features after feature reorganization, output fraud warning results, and make fraud warnings based on the fraud warning results. Specifically, the facial features and voiceprint features after reorganization are fused to obtain a first fused feature, the facial features and voiceprint features in the first memory are fused to obtain a second fused feature, the similarity between the first fused feature and the second fused feature is compared, and the fraud warning result is output.

[0084] It should be noted that the data storage module is divided into two modules. One part is used to store the face and voiceprint features of the feature extraction module (recorded as data storage module 1), and the extracted face and voiceprint features are only extracted in real time but not stored in real time. Instead, they are stored after the call is confirmed to be safe; the other part is used to store the features after the fusion of face and voiceprint features (recorded as data storage module 2), and this part of the data is fused based on the face and voiceprint features stored in the data storage module 1, and the fusion process is consistent with the model training module method.

[0085] It should be noted that the feature extraction is real-time extraction, and the face feature is extracted by an image feature extraction algorithm, and the face feature is denoted by f i (i=1,2,3……,m); the voiceprint feature is extracted by the voiceprint recognition algorithm, and the voiceprint feature is recorded as v j (j=1,2,3……,n).

[0086] Preferably, in the feature extraction module, the image feature extraction algorithm is implemented by the Hog feature algorithm, and the process is as follows:

[0087] Gamma correction of positioning images for video calls;

[0088] Calculate the gradient value of each pixel in the positioning image to obtain the gradient map. The calculation formula is:

[0089]

[0090] Among them, g represents the total gradient strength value, g x represents the horizontal gradient, g y represents the vertical gradient; θ represents the gradient direction, which takes an absolute value and ranges from 0 to 180°;

[0091] Calculate the gradient histogram and perform sample normalization to obtain facial features.

[0092] Preferably, in the feature extraction module, the voiceprint recognition algorithm is specifically an x-vector feature algorithm, and the process is as follows:

[0093] Extract the frame-level features of speech frames through the time delay neural network layer;

[0094] The statistical pooling layer is used to take the mean and standard deviation of the frame-level features of all frames in the input sequence;

[0095] concatenating the mean and standard deviation as segment-level features;

[0096] The segment-level features are classified through a standard feed-forward network to obtain the voiceprint features.

[0097] Furthermore, the abnormal monitoring specifically includes: locating the IP addresses, login devices and other information of both parties in the call, comparing the differences between the current login location, login device and other information and the commonly used login location, login device and other information, and realizing call monitoring.

[0098] Preferably, in the feature recombination module, the feature recombination process is as follows:

[0099] Add random noise to the features extracted in the feature extraction module. After adding noise, the face features contaminated by noise are recorded as f' i (i=1,2,3,……,m); the voiceprint feature polluted by noise is recorded as v' j (j=1,2,3,……,n);

[0100] A part of the noise-contaminated features is randomly selected, and the features before and after adding noise are evenly divided into blocks according to the feature size and the coordinates are located. The one-to-one correspondence of the feature areas before and after the noise pollution is achieved by relying on the coordinate position. In order to improve the similarity, the coordinates of the neighboring areas of each block area are taken into account. When the feature coordinate areas before and after the noise pollution cannot be accurately corresponded, the similarity of the neighboring areas is used for positioning. Considering the deviation of information loss, the error term ε is added in the calculation of the similarity. The specific formula for similarity calculation is as follows:

[0101] If the facial feature f' is contaminated by noise i And the noise-contaminated voiceprint feature v' j All of them cannot be recognized, at this time f' i 、v' j If it is 0, the similarity range of the pollution feature is calculated according to the characteristics of the adjacent areas of the pollution feature, and the interval value is taken. The similarity calculation formula is as follows:

[0102] f' in =f in +ε;f' i ∈f' in ;

[0103] v' jn =v in +ε;v' j ∈v' jn ;

[0104] Among them, f in represents the features of the neighboring area where the facial features are initially extracted; f' in Represents the neighboring area features of the face features polluted by noise; v in Indicates the features of the neighboring area where the voiceprint features are initially extracted; v' jn Represents the neighboring area features of the voiceprint features after being contaminated by noise;

[0105] If the facial feature f' is contaminated by noise i And the noise-contaminated voiceprint feature v' j can be identified, then f' i 、v' j Not 0, the similarity calculation formula is:

[0106] f' i =f i +ε;

[0107] v' j =v j +ε;

[0108] Among them, f i represents the extracted facial features; v j represents the extracted voiceprint features;

[0109] The Euclidean distance is used to calculate the similarity of different features in the same area before and after adding noise. The specific formula is as follows:

[0110]

[0111] Among them, f(f i ,f' i ) indicates f i and f' iThe similarity between j ,v' j ) indicates v j and v' j The similarity between them; m represents the number of polluted facial features in the extracted facial features; n represents the number of polluted voiceprint features in the extracted voiceprint features;

[0112] The smaller the distance, the higher the similarity. The position with low similarity is located as a noise contaminated position.

[0113] After locating the noise point, the position feature is removed, and the feature is reorganized by calling the feature of the corresponding position area in the data storage module 1. The reorganization process is as follows:

[0114] Similarly, the original feature coordinates of the corresponding position in the data storage module 1 are located by calculating the feature matching degree before and after the noise pollution, and then the original features are added to the corresponding area to obtain new features. After that, the security detection model is called to perform similarity detection on the feature fusion degree.

[0115] Preferably, in the model training module, the model training process is as follows:

[0116] Build the network structure of the Transformer model (preferred pre-trained transformer model in this embodiment);

[0117] Taking into account the imbalance problem of positive and negative samples in the acquired image and audio data (positive samples>negative samples), this embodiment adopts a cross entropy loss function to estimate the similarity between the actual output probability and the expected output probability. Specifically, the output results of the system are divided into two types of results: positive and negative. Therefore, the preferred cross entropy loss function of this embodiment is a binary cross entropy loss function. The probabilities predicted by the positive and negative classes are p and 1-p. The specific formula is as follows:

[0118]

[0119] Among them, L represents the binary cross entropy loss; y i represents the label of sample i, the positive class is 1 and the negative class is 0; p i It represents the probability that sample i is predicted as the positive class. The smaller the cross entropy value, the closer the probability distribution of the two is.

[0120] Regularize the model;

[0121] To evaluate the model, the calculation formula is:

[0122]

[0123] Among them, Ac represents accuracy; Sn represents sensitivity; TP represents true positive; FP represents false positive; TN represents true negative; FN represents false negative;

[0124] Model optimization, using the adaptive matrix estimation optimization algorithm to optimize the model, the process is as follows:

[0125]

[0126] m t =β 1 m t-1 +(1-β 1 ) t ;

[0127]

[0128]

[0129] Among them, g t represents the gradient; m t represents the first-order moment estimate of the gradient; v t represents the second-order moment estimate of the gradient; represents the corrected first-order moment estimate; represents the corrected second-order moment estimate; represents a dynamic constraint on the learning rate; θ represents the parameter to be solved; represents the gradient of parameter θ; f t (θ t-1 ) represents the objective function to be optimized; β 1 is the first-order moment attenuation coefficient; β 2 is the second-order moment attenuation coefficient; g t and is the gradient; ∈ represents a constant, which is used to prevent the denominator from being zero;

[0130] The optimized model is a safety detection model.

[0131] Preferably, in the fraud warning module, the process of security detection is as follows:

[0132] Data envelopment analysis (DEA) is used to assign weights to the facial features and voiceprint features in feature extraction, and the facial features and voiceprint features are input into the security detection model to output fraud warning results. The fraud warning results in this embodiment comprehensively consider the facial features and voiceprint features, and can better identify AI face-changing or AI imitation voice.

[0133] In addition, if Figure 1 As shown, this embodiment also discloses a fraud warning method for video calls, using the above-mentioned fraud warning system, the method includes:

[0134] Abnormal call monitoring: Monitor the IP address, login device and other information of the video call and compare it with the commonly used IP address and login device information for abnormal monitoring. If an abnormality is detected, security detection is performed. If no abnormality is detected, feature extraction continues.

[0135] Security detection: By inputting facial features and voiceprint features into the security detection model, the fraud warning result is output. Based on the fraud warning result, it is determined whether the video call has AI face-swapping or AI voice-swapping. If there is no AI face-swapping or AI voice-swapping, the video call is safe and the system is automatically operated. If there is AI face-swapping or AI voice-swapping, it indicates that the video call may have a fraud risk, and a fraud warning message is issued to the video call user without any abnormality;

[0136] The system runs automatically: facial features and voiceprint features are collected through the feature extraction module, facial and voiceprint fusion features are collected through the model training module, the data storage module is continuously updated based on the facial features, voiceprint features and facial and voiceprint fusion features, and the security monitoring model is updated based on the data storage module.

[0137] The embodiment of the present invention also discloses a computer device, including a memory and a processor;

[0138] The memory is used to store a computer program executable on the processor;

[0139] The processor is used to implement the steps of the above-mentioned fraud warning method when executing the computer program.

[0140] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the computer device.

[0141] The computer device may be a computing device such as a mobile phone, a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may also include an input / output device, a network access device, a bus, etc.

[0142] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0143] The memory can be used to store the computer program and / or module, and the processor implements the computer program by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0144] Wherein, if the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0145] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned fraud warning method are implemented.

[0146] Compared with the prior art, the system of the present invention can perform real-time detection and comparison when an abnormal event occurs, and promptly warn the victim. The system does not require the user to have professional anti-fraud knowledge and technical knowledge. The user only needs to embed the system and maintain system access rights. The system will start automatically and extract and update information in real time during operation, so that when an abnormality is captured, it can identify and warn more quickly and accurately.

[0147] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0148] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A fraud warning system for video calls, characterized in that: It includes data storage module, feature extraction module, call anomaly monitoring module, feature reorganization module, model training module and fraud warning module: Data storage module: including a first memory and a second memory, wherein the first memory stores the face and voiceprint features extracted during each system operation; The second memory stores the fusion features extracted during each system operation; Feature extraction module: performs real-time feature extraction on the video call, obtains the facial features and voiceprint features of the video call, stores the facial features and voiceprint features in the first memory when the video call does not involve abnormalities, and sends the facial features and voiceprint features to the feature reconstruction module when the video call involves abnormalities; Abnormal call monitoring module: query the bearer device information of both parties in the video call, monitor the video call process in real time, and determine whether the video call process involves abnormalities; Feature Recombination Module: Recombines facial features and voiceprint features to form recombined facial features and voiceprint features, and then calls the feature training module to standardize the recombined facial features and voiceprint features and transmit them to the fraud warning module; Model training module: calling the facial features and voiceprint features in the first memory to establish a data set, dividing the data set into a training set, a test set and a verification set, building a pre-trained converter model, training the pre-trained converter model based on the training set, the test set and the verification set, and obtaining a security detection model; Fraud warning module: calls the security detection model to perform security detection on the facial features and voiceprint features after feature reorganization, outputs fraud warning results, and performs fraud warning based on the fraud warning results.

2. The fraud early warning system according to claim 1, characterized in that: The feature extraction module includes an image feature extraction algorithm. The face features are extracted by the image feature extraction algorithm. The image feature extraction algorithm is implemented by the Hog feature algorithm. The process is as follows: Gamma correction of positioning images for video calls; Calculate the gradient value of each pixel in the positioning image to obtain the gradient map. The calculation formula is: Among them, g represents the total gradient strength value, g x represents the horizontal gradient, g y represents the vertical gradient; θ represents the gradient direction, which takes an absolute value and ranges from 0 to 180°; Calculate the gradient histogram and perform sample normalization to obtain facial features.

3. The fraud early warning system according to claim 2, characterized in that: The feature extraction module also includes a voiceprint recognition algorithm. The voiceprint feature is extracted by the voiceprint recognition algorithm. The process of the voiceprint recognition algorithm is as follows: Extracting frame-level features of speech frames through a time-delay neural network layer; The frame-level features of all frames in the input sequence are averaged and their standard deviations are taken through the statistical pooling layer; concatenating the mean and standard deviation as segment-level features; The segment-level features are classified through a standard feed-forward network to obtain the voiceprint features.

4. The fraud early warning system according to claim 3, characterized in that: In the call anomaly monitoring module, the real-time monitoring of the video call process is specifically: locating the IP addresses and login device information of both parties in the video call. When the current login location and login device information are different from the commonly used login location and login device information, the video call is judged to be abnormal.

5. The fraud early warning system according to claim 1, characterized in that: In the feature recombination module, the feature recombination process is as follows: Adding random noise to the features extracted in the aforementioned feature extraction module; Randomly select a part of the noise-contaminated features, divide the features before and after adding noise into blocks evenly according to the feature size and locate the coordinates. Rely on the coordinate position to achieve one-to-one correspondence between the feature areas before and after the noise pollution. The similarity of the features before and after the noise pollution in the adjacent area is used for positioning. The error term ε is added to the calculation of the similarity. The specific formula for similarity calculation is as follows: If the facial feature f′ is contaminated by noise i and the noise-contaminated voiceprint feature v′ j All of them cannot be identified, at this time f′ i , v′ j If it is 0, the similarity range of the pollution feature is calculated according to the characteristics of the adjacent areas of the pollution feature, and the interval value is taken. The similarity calculation formula is as follows: f′ in =f in +ε;f′ i ∈f′ in ; v′ jn =v in +ε;v′ j ∈v′ jn ; Among them, f in represents the features of the neighboring area where the facial features are initially extracted; f′ in Represents the neighboring area features of the face features polluted by noise; v in represents the neighboring area features of the initially extracted voiceprint features; v′ jn Represents the neighboring area features of the voiceprint features after being contaminated by noise; If the facial feature f′ is contaminated by noise i and the noise-contaminated voiceprint feature v′ j can be identified, then f′ i , v′ j Not 0, the similarity calculation formula is: f′ i =f i +e; v′ j =v j +ε; Among them, f i represents the extracted facial features; v j represents the extracted voiceprint features; The Euclidean distance is used to calculate the similarity of different features in the same area before and after adding noise. The specific formula is as follows: Among them, f(f i ,f′ i ) indicates f i and f′ i The similarity between j ,v′ j ) indicates v j and v′ j The similarity between them; m represents the number of polluted facial features in the extracted facial features; n represents the number of polluted voiceprint features in the extracted voiceprint features; The smaller the distance, the higher the similarity. The position with low similarity is located as a noise contaminated position. After locating the noise point, the position feature is removed, and the feature is reorganized by calling the feature of the corresponding position area in the first memory. The reorganization process is as follows: Similarly, the original feature coordinates of the corresponding position in the first memory are located by calculating the matching degree of the features before and after the noise pollution, and then the original features are added to the corresponding area to obtain new features, and then the security detection model is called to perform similarity detection on the fused features.

6. The fraud early warning system according to claim 1, characterized in that: In the model training module, the model training process is as follows: Build the network structure of the Transformer model; The binary cross entropy loss function is used to estimate the similarity between the actual output probability and the expected output probability. The probabilities predicted for the positive and negative classes are p and 1-p. The specific formula is as follows: Among them, L represents the binary cross entropy loss, y i represents the label of sample i, the positive class is 1 and the negative class is 0; p i It represents the probability that sample i is predicted as the positive class. The smaller the cross entropy value, the closer the probability distribution of the two is. Regularize the model; To evaluate the model, the calculation formula is: Among them, Ac represents accuracy; Sn represents sensitivity; TP represents true positive; FP represents false positive; TN represents true negative; FN represents false negative; Model optimization, using the adaptive matrix estimation optimization algorithm to optimize the model, the process is as follows: m t =β1m t-1 +(1-β1)g t ; v t =β2v t-1 +(1-β2)g t 2 ; Among them, g t represents the gradient; m t represents the first-order moment estimate of the gradient; v t represents the second-order moment estimate of the gradient; represents the corrected first-order moment estimate; represents the corrected second-order moment estimate; represents a dynamic constraint on the learning rate; θ represents the parameter to be solved; represents the gradient of parameter θ; f t (θ t-1 ) represents the objective function to be optimized; β1 is the first-order moment attenuation coefficient; β2 is the second-order moment attenuation coefficient; g t and is the gradient; ∈ represents a constant, which is used to prevent the denominator from being zero; The optimized model is a safety detection model.

7. The fraud early warning system according to claim 1, characterized in that: In the fraud warning module, the security detection process is as follows: Data envelopment analysis is used to assign weights to the facial features and voiceprint features after feature reorganization, the facial features and voiceprint features are input into the security detection model, the feature fusion values ​​in the second memory are called for similarity comparison, and the fraud warning results are output.

8. A fraud warning method for video calls, characterized in that: Using the fraud early warning system according to any one of claims 1 to 7, the method comprises: Call anomaly monitoring: Monitor the IP address and login device information of the video call and compare them with the common IP address and login device information for anomaly monitoring. If anomalies are detected, security detection is performed. If no anomalies are detected, feature extraction is continued. Security detection: By inputting facial features and voiceprint features into the security detection model, the fraud warning result is output. Based on the fraud warning result, it is determined whether the video call has AI face-swapping or AI voice-swapping. If there is no AI face-swapping or AI voice-swapping, the video call is safe and the system is automatically operated. If there is AI face-swapping or AI voice-swapping, it indicates that the video call may have a fraud risk, and a fraud warning message is issued to the video call user without any abnormality; The system runs automatically: facial features and voiceprint features are collected through the feature extraction module, facial and voiceprint fusion features are collected through the model training module, the data storage module is continuously updated based on the facial features, voiceprint features and facial and voiceprint fusion features, and the security detection model is updated based on the data storage module.

9. A computer device, characterized in that: including memory and processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the fraud warning method described in claim 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the fraud warning method according to claim 8 are implemented.

Citation Information

Patent Citations

  • Anti-fraud device and method integrating face recognition and voice recognition

    CN111860350A

  • Financial anti-network fraud system

    CN116523521A