Method, device, computer equipment and storage medium for user type identification

By extracting and analyzing the video data of the target user, determining whether it is a fraudulent user, the problem of low recognition accuracy of fraudulent user in the prior art is solved, and the effect of risk control is improved.

CN114155460BActive Publication Date: 2025-05-23PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111430024.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-05-23
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

The existing face-to-face review technology has low accuracy when identifying fraudulent users, especially for users who have been packaged or deliberately disguised.

Method used

By processing the video data to be detected by the target user, the audio data file and the image data file are extracted, and feature extraction is performed separately to obtain video features and voice features. Based on these characteristics, the reasonable value of the target user for each question and answer is determined. If the reasonable value is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user.

Benefits of technology

Improve the accuracy of identifying whether the target user is a fraudulent user, which is conducive to improving risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155460B_ABST
    Figure CN114155460B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence, and discloses a method, device, computer equipment, and storage medium for user type identification. The method includes: processing the video data to be detected of the target user to obtain the audio data file and image data file of the target user for each question and answer; extracting features from the image data file to obtain the video features of each moment in the first duration of the image data file, wherein the first duration includes the second moment and the third moment; extracting features from the audio data file to obtain the voice features of the second moment; determining the reasonable value of the target user for each question and answer based on the voice features and video features of the second moment, and the video features of the third moment; if the reasonable value is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user. The implementation of the embodiment of the present application can improve the accuracy of identifying whether the target user is a fraudulent user, which is conducive to improving risk control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a method, apparatus, computer device and storage medium for identifying user types. Background Art

[0002] Financial institutions usually need to conduct face-to-face interviews with applicants during the loan approval process. By collecting facial images or video data during the user's Q&A process, facial recognition technology is used to identify facial expressions to determine whether the applicant is the real person, whether there is any lying or fraudulent behavior, and then evaluate the applicant's willingness and ability to repay the loan to reduce the possibility of risk events. However, for users who are deliberately disguised through packaging, the existing face-to-face interview technology has a low accuracy rate in identifying fraudulent users. Summary of the invention

[0003] The embodiments of the present application provide a method, apparatus, computer device, and storage medium for identifying a user type, which can improve the accuracy of identifying whether a target user is a fraudulent user, thereby facilitating risk control.

[0004] In a first aspect, an embodiment of the present application provides a method for identifying a user type, wherein:

[0005] Processing the to-be-detected video data of the target user to obtain an audio data file and an image data file of the target user for each question and answer;

[0006] Performing feature extraction on the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration includes a second moment and a third moment;

[0007] Performing feature extraction on the audio data file to obtain speech features at the second moment;

[0008] Determining a reasonable value of each question and answer of the target user based on the voice feature and the video feature at the second moment and the video feature at the third moment;

[0009] If the reasonable value is greater than or equal to a preset threshold, the target user is determined to be a fraudulent user.

[0010] In a second aspect, an embodiment of the present application provides a device for identifying a user type, wherein:

[0011] A data processing unit, used to process the video data to be detected of the target user to obtain the audio data file and image data file of the target user for each question and answer;

[0012] a feature extraction unit, configured to extract features from the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration includes a second moment and a third moment;

[0013] Performing feature extraction on the audio data file to obtain speech features at the second moment;

[0014] a determination unit, configured to determine a reasonable value of the target user for each question and answer based on the voice feature and the video feature at the second moment and the video feature at the third moment;

[0015] If the reasonable value is greater than or equal to a preset threshold, the target user is determined to be a fraudulent user.

[0016] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a communication interface, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for some or all of the steps described in the first aspect of the embodiment of the present application.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.

[0018] Implementing the embodiments of the present application will have the following beneficial effects:

[0019] The above-mentioned method, device, computer equipment and storage medium for identifying user types are used to process the target user's video data to be detected, and after obtaining the target user's audio data files and image data files for each question and answer, feature extraction is performed on the image data file to obtain the video features of each moment in the first duration, and feature extraction is performed on the audio data file to obtain the voice features of the second moment. Among them, the first duration includes the second moment and the third moment. Then, based on the voice features and video features at the same moment (second moment) in the image data file and the audio data file, and the video features at the third moment in the image data file, the reasonable values ​​of the target user for each question and answer are determined. If one of the reasonable values ​​is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user. In this way, the user type is identified through the video features and voice features at the same moment, and the video features at the moment when there is no voice feature, the accuracy of identifying whether the target user is a fraudulent user can be improved, which is conducive to improving risk control. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0021] Figure 1 A schematic diagram of a system architecture provided for an embodiment of the present application;

[0022] Figure 2 A flowchart of a method for identifying a user type provided in an embodiment of the present application;

[0023] Figure 3 A schematic diagram of the structure of a device for identifying user types provided in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0026] The terms "first", "second", "third" and "fourth" etc. in the specification and claims of the present application and the drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0027] Reference to "embodiments" herein means that a particular feature, result, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0028] In order to better understand the technical solution of the embodiment of the present application, the system architecture that may be involved in the embodiment of the present application is first introduced. Figure 1 , a schematic diagram of a system architecture provided in an embodiment of the present application, the system architecture may include: an electronic device 101 and a server 102. The electronic device 101 and the server 102 may communicate via a network. The network communication may be based on any wired and wireless network, including but not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), and a wireless communication network, etc.

[0029] The embodiment of the present application does not limit the number of electronic devices and servers, and the server can provide services for multiple electronic devices at the same time. In the embodiment of the present application, the electronic device is mainly a device used by the salesperson of a financial institution to handle business, and can be used to transmit the video data to be detected by the target user collected by the face-to-face interview to the server through the network. The electronic device can be a personal computer (personal computer, PC), a laptop or a smart phone, and can also be an all-in-one machine, a handheld computer, a tablet computer (pad), a smart TV playback terminal, a car terminal or a portable device. The operating system of the electronic device on the PC side, such as an all-in-one machine, can include but is not limited to Linux system, Unix system, Windows series system (such as Windows xp, Windows 7, etc.), Mac OS X system (Apple computer operating system) and other operating systems. The electronic device on the mobile side, such as a smart phone, etc., can include but is not limited to Android system, IOS (Apple mobile phone operating system), Window system and other operating systems.

[0030] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The server can also be implemented as a server cluster consisting of multiple servers.

[0031] Financial institutions usually need to conduct face-to-face interviews with applicants during the loan approval process. By collecting facial images or video data during the user's Q&A process, facial recognition technology is used to identify facial expressions to determine whether the applicant is the real person, whether there is any lying or fraudulent behavior, and then evaluate the applicant's willingness and ability to repay the loan to reduce the possibility of risk events. However, for users who are deliberately disguised through packaging, the existing face-to-face interview technology has a low accuracy rate in identifying fraudulent users.

[0032] In order to solve the above problems, the embodiment of the present application provides a method for identifying user types, which can be applied to electronic devices or servers configured by financial institutions such as banks, securities, and insurance. By implementing this method, the accuracy of identifying whether the target user is a fraudulent user can be improved, which is conducive to improving risk control.

[0033] Please refer to Figure 2 , Figure 2 1 is a flow chart of a method for identifying a user type provided in an embodiment of the present application. Taking the method applied in a server as an example, the method may include the following steps S201-S205, wherein:

[0034] Step S201: Process the target user's video data to be detected to obtain the target user's audio data files and image data files for each question and answer.

[0035] In the embodiment of the present application, the target user refers to the object being asked questions during the face-to-face interview. It can be a loan applicant, or a credit card applicant, etc. The video data to be detected can be the dialogue information between the target user and the reviewer when applying for a credit card, loan, or other activities. The video data to be detected can include video clips in which the target user and the reviewer speak separately, or can include video clips in which the target user and the reviewer speak at the same time, or can include video clips in which neither the target user nor the reviewer speaks. The video data to be detected can be video data of the face-to-face interview scene recorded in real time by an electronic device with a camera function, or it can be pre-recorded and saved face-to-face interview video data.

[0036] The video data to be detected includes audio data and image data, wherein the audio data refers to the collected voice data input by the target user, and the image data refers to the collected picture data. In the embodiment of the present application, the audio data and image data in the video data to be detected can be separated by video editing software to obtain the audio data file and image data file of the target user for each question and answer.

[0037] In a possible implementation, step S201 may include the following steps: performing semantic recognition on the target user's video data to be detected to obtain a target video segment of the target user for each question and answer; and extracting an audio data file and an image data file of the target video segment.

[0038] During the face-to-face review process, the number of questions that the reviewer asks the target user can range from a dozen to dozens. The video data to be tested can be semantically recognized by the target user, and the start time of the reviewer's questions can be marked, so that the video data to be tested can be divided into target video segments of the target user for each question and answer. When asking questions to the target user, the reviewer can give a prompt word to start the question, such as "Question 1...", "Question 2...". For example, when semantic recognition is performed on the video data to be tested, if "Question 1" is detected, the corresponding timestamp at this time can be marked as the start time of the first question. If "Question 2" is detected, the corresponding timestamp at this time can be marked as the end time of the first question (i.e., the start time of the second question).

[0039] The embodiment of the present application does not limit the method of semantic recognition. In a possible implementation, a semantic recognition model can be used to perform semantic recognition on the video data to be detected of the target user. The semantic recognition model can be a bidirectional attention neural network model (bidirectional encoder representation from transformers, BERT), a recurrent neural network (RNN), a convolutional neural network (CNN), etc., and there is no limitation on this.

[0040] For different target users, the questions asked by the reviewer may be different, and the questions may be specifically based on the basic information of the target user. In an embodiment of the present application, the basic information of the target user may include: the identity information of the target user and the identity information of the associated user, and the associated user may be a spouse, relative, friend, etc.; the identity information includes name, ID number, contact number, gender, age, home address, occupation, and education, etc. The identity information of the target user and the associated user may be filled in by the target user when submitting the application for the service, or may be determined by the relevant information submitted in the application information, wherein the relevant information may include an ID card, marriage certificate, real estate certificate, household register information, etc. The face-to-face review questions may also be about career, assets, consumption, family, etc. In addition, the face-to-face review questions may also be generated by the intelligent question-and-answer system based on the basic information of the target user.

[0041] It can be seen that by cutting out the video clips of the target user for each question and answer as the target video clips, and then extracting the audio data file and image data file of the target user from the target video clips, the amount of calculation can be reduced and the processing efficiency can be improved.

[0042] Step S202: extracting features from the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration includes a second moment and a third moment.

[0043] In an embodiment of the present application, the image data file includes the image data of the target video clip, and the audio data file includes the audio data of the target video clip. The duration of the image data file and the duration of the audio data file may be equal to the duration of the target video clip. The duration of the target video clip containing the face of the target user may be referred to as the first duration. Since there may be a moment when the target user in the target video clip speaks without sound, the moment when the target user speaks in the target video clip is referred to as the second moment, and the moment when the target user does not speak in the target video clip is referred to as the third moment. That is, the duration of the target video clip (or image data file) containing the face of the target user (or video feature) may be referred to as the first duration, and the first duration may include the second moment when the target user makes a sound and the third moment when the target user does not make a sound. That is, the moment when the target video clip (or audio data file) contains the sound (or voice feature) of the target user may be referred to as the second moment.

[0044] For example, if the total length of the video data to be detected is 10 minutes, and the time period of the target video segment of the target user answering a certain question is from 1 minute 5 seconds to 1 minute 55 seconds, the first time period may be from 1 minute 10 seconds to 1 minute 50 seconds, the second time period may be 1 minute 10 seconds, and the third time period may be 1 minute 30 seconds. Alternatively, other division methods may be used, which are not limited thereto.

[0045] Video features can be facial action units (AUs), or they can be the angle of looking up or down, the angle of shaking the head left or right, the direction of the head, the direction of the eyeballs, etc. In the facial action coding system (FACS) summarized by Paul Ekman, humans have a total of 39 main AUs, where one AU represents a small group of muscle contraction codes on the face. Please refer to Table 1, which lists some main AUs. These AUs can be combined to identify the emotions of the target user. For example, AU4, AU5, AU7, and AU23 can be combined to represent anger. Please refer to Table 2 for details.

[0046] Table 1 Main action unit encoding

[0047]

[0048]

[0049] Table 2 Emotion calculation formula

[0050] mood Emotion calculation formula happy AU6+AU12 sad AU1+AU4+AU15 surprise AU1+AU2+AU5+AU26 Fear AU1+AU2+AU4+AU5+AU7+AU20+AU26 anger AU4+AU5+AU7+AU23 disgust AU9+AU15+AU16 Contempt AU12+AU14

[0051] The following takes AU as an example to introduce the process of extracting video features at each moment in the first duration of the image data file. In a possible implementation, step S202 may include the following steps: performing frame processing on the image data file to obtain a first video frame at each moment in the first duration of the image data file; performing key frame extraction on the first video frame to obtain a second video frame; performing facial feature extraction on the second video frame to obtain an action unit; and determining the video features at each moment in the first duration of the image data file based on the action unit.

[0052] In the embodiment of the present application, the frame processing of the image data file can refer to the frame processing process of the audio data file below, which will not be described in detail here. The key frame (I frame) is a frame in the compressed video that completely retains the image data file. When decoding the key frame, only the image data file of this frame is needed to complete the decoding. Since the similarity between the key frames in the second video frame is small, the second video frame can more comprehensively represent the first video frame.

[0053] In an embodiment of the present application, a face recognition algorithm can be used to extract face features from the second video frame to obtain an action unit. The face recognition algorithm can be one or more of a 3D convolutional neural network (3DCNN), a spatial temporal graph convolutional network (ST-GCN), a support vector machine (SVM), and the like, without limitation.

[0054] It can be seen that in the embodiment of the present application, the image data file is firstly processed by frame division, and then the key frame is extracted from the first video frame obtained by the frame division to obtain the second video frame. Finally, the face feature is extracted from the second video frame to obtain the action unit as the video feature. In this way, the processing efficiency and accuracy of the video feature extraction can be improved.

[0055] In a possible implementation, before extracting facial features from the second video frame to obtain the action unit, the resolution of the second video frame can be adjusted to make the size of the second video frame moderate. In this way, it is possible to avoid the second video frame data being too large and causing the data processing speed to be too slow, or to avoid the second video frame data being too small and causing the subsequent fraud detection accuracy to be too low.

[0056] Step S203: extracting features from the audio data file to obtain speech features at the second moment.

[0057] Voice features may include, but are not limited to, mel-frequency cepstral coefficients (MFCC), identity-vector (i-vector), text keywords, pitch, sound intensity, sound length, frequency band energy distribution, harmonic signal-to-noise ratio, short-term energy jitter, etc. Among them, text keywords can reflect the part-of-speech characteristics of the words used by the target user. Word features may include, but are not limited to, negative words, neutral words, and positive words. Furthermore, text keywords combined with other voice features or video features can help identify the emotions of the target user, thereby helping to improve the accuracy of fraud identification.

[0058] The following describes the extraction process of the speech features at the second moment by taking text keywords as an example. In a possible implementation, step S203 may include the following steps:

[0059] The audio data file is frame-processed to obtain speech frames at each moment in the first time length of the audio data file; the speech frames are preprocessed to obtain target speech frames at the second moment; speech recognition is performed on the target speech frames to obtain text data; the text data is segmented to obtain segmentation vocabulary and segmentation emotions; text keywords are selected from the segmentation vocabulary according to the segmentation emotions; and speech features at the second moment are obtained according to the text keywords.

[0060] Among them, frame is the smallest observation unit in the audio data file, and framing is the process of dividing according to the time sequence of the audio data file. In order to ensure that each frame signal remains stable in a sufficiently short time and that there are enough vibration cycles, the frame length is generally 20-50ms, and can be specifically 20ms, 30ms, 40ms, etc. In addition, in order to avoid excessive changes between two adjacent frames, the overlapping segmentation method can usually be used for framing processing, so that the transition between frames is smooth and their continuity is maintained. The overlapping part between frames is called frame shift, and the ratio of frame shift to frame length is generally taken as 0-1 / 2. In this way, by framing the audio data file, a relatively stable voice frame can be obtained, which is conducive to improving the accuracy of subsequent voice recognition or voiceprint recognition.

[0061] The target voice frame refers to the voice frame at the second moment, which can be understood as the voice frame corresponding to the moment when the target user makes a sound in the target video clip. The embodiment of the present application does not limit the preprocessing method of the voice frame. Before obtaining the voice frame at the second moment, the preprocessing may include one or more of dimensionality reduction and normalization. Specifically, singular value decomposition (SVD), principal component analysis (PCA), factor analysis (FA), independent component analysis (ICA) and other methods can be used for dimensionality reduction. Through dimensionality reduction, the most important features can be retained from the high-dimensional vector, and noise and unimportant features can be removed. The voice frame can also be normalized. For example, the min-max normalization method (min-max normalization) can be used to uniformly map the value of the voice frame to the [0, 1] interval. In this way, the normalized voice frame has a certain comparability in terms of numerical value, which can improve the accuracy of fraud identification.

[0062] Furthermore, the speech frames obtained after the framing process can be windowed and pre-emphasized to obtain target speech frames with better quality. Windowing can be used to eliminate signal discontinuities that may be caused at both ends of each frame, and can transform non-stationary speech signals into short-term stationary signals. Commonly used window functions include square windows, Hamming windows, and Hanning windows. Pre-emphasis can enhance the high-frequency part, which is used to filter out low frequencies and make high frequencies more prominent, so as to improve the signal-to-noise ratio.

[0063] In a possible implementation, preprocessing the speech frames to obtain the target speech frames at the second moment may include the following steps: calculating the frame energy of each frame in the speech frames; marking the speech frames whose frame energy is less than a preset threshold as silent frames; and removing the silent frames in the speech frames to obtain the target speech frames.

[0064] Among them, frame energy is the short-time energy of the speech signal, reflecting the data volume of the speech information of the speech frame. The frame energy can be used to determine whether the speech frame is a target speech frame or a silent frame. If the frame energy of the speech frame is less than the preset threshold, the frame energy is a silent frame. If the frame energy of the speech frame is greater than or equal to the preset threshold, the frame energy is a target speech frame. The preset threshold is a pre-set parameter, which can be set specifically based on historical experience, such as setting the preset threshold to 0.5, or it can be specifically analyzed and set based on the calculated frame energy of the speech frame.

[0065] It can be seen that by cutting out the silent frames in the speech frames to obtain the target speech frames, the speech frames at the third moment when the target user is not speaking can be filtered out. Then, feature extraction is performed on the target speech frames to obtain speech features. In this way, the efficiency and quality of speech feature extraction can be improved.

[0066] Specifically, the target speech frame can be converted into text data by using technologies such as automatic speech recognition (ASR). Then, the text data can be divided according to punctuation marks to obtain multiple sentence texts of different lengths, and then each sentence text is segmented to obtain segmentation vocabulary and segmentation emotions.

[0067] Among them, the word segmentation processing method can adopt a word segmentation method based on string matching, also known as a mechanical word segmentation method. For example, the forward maximum matching method is to segment the string in a segmented sentence from left to right; or, the reverse maximum matching method is to segment the string in a segmented sentence from right to left; or, the shortest path word segmentation method is to require the least number of words to be segmented in the string in a segmented sentence; or, the bidirectional maximum matching method is to perform word segmentation matching in both the forward and reverse directions. The word meaning segmentation method can also be used to segment each segmented sentence. The word meaning segmentation method is a word segmentation method for machine speech judgment, which uses syntactic information and semantic information to deal with ambiguous phenomena to segment words. In addition, the jieba word segmentation tool or the word2vec word vector model can be used to parse text data to obtain the part of speech (for example, the two major categories of nouns and verbs, as well as names of people, places, and institutions, or adverbs, noun verbs, etc.) and word meaning corresponding to each character or word in the text data. In addition, in a possible implementation, the segmentation sentiment of each segmentation vocabulary in the text data can also be obtained based on neural network models such as CNN and long short-term memory network (LSTM).

[0068] Segmented emotions can include positive emotions, neutral emotions, and negative emotions. Each segmented emotion has a preset emotion intensity. For example, when the segmented emotion expresses a negative emotion, the emotion intensity value of the segmented emotion can be set to "-3", "-2", "-1". Among them, the stronger the negative emotion, the greater the absolute value of the emotion intensity. When the segmented emotion expresses a neutral emotion, the emotion intensity value of the segmented emotion can be set to "0". When the segmented emotion expresses a positive emotion, the emotion intensity value of the segmented emotion can be set to "1", "2", "3". Among them, the stronger the negative emotion, the greater the absolute value of the emotion intensity.

[0069] There may be some negative words in the text data, and negative words may cause the segmentation sentiment to show opposite emotional characteristics. In specific applications, you can search whether the segmentation vocabulary contains negative words. If negative words are identified, the corresponding emotional intensity value of the segmentation sentiment is adjusted in the opposite direction. For example, the target user's answer is "The company's operating income this year has not increased compared to last year." In this sentence, the negative word "no" is retrieved. If the preset emotional intensity value corresponding to the segmentation sentiment is "+3", the emotional intensity value is adjusted in the opposite direction to "-3".

[0070] In addition, there may be some degree adverbs in the text data, such as "very", "very", etc. These degree adverbs can strengthen the emotional intensity value of the segmentation emotion. In specific applications, it can be determined whether the segmentation vocabulary contains degree adverbs. If it does, the emotional intensity value of the segmentation emotion can be adjusted according to the preset value of the degree adverb. For example, the preset value of the degree adverb can be ±0.5. When the segmentation emotion expresses a negative emotion, the preset value of the degree adverb can be -0.5; when the segmentation emotion expresses a positive emotion, the preset value of the degree adverb can be 0.5. According to the preset emotional intensity value of the segmentation emotion and whether it contains negative words or degree adverbs, the comprehensive value of the emotional intensity value of the segmentation emotion is calculated, the segmentation vocabulary is sorted according to the emotional intensity value of the segmentation emotion, and the text keywords are determined according to the sorting results. For example, segmentation vocabulary with emotional intensity values ​​of [-3, -1] and [1, 3] can be selected as text keywords.

[0071] In this way, when selecting text keywords, not only the preset emotion intensity value of the word segmentation emotion is taken into consideration, but also multiple features such as negative words and degree adverbs are taken into consideration, so that the selected text keywords can more accurately reflect the emotions of the target users, which is conducive to improving the accuracy of fraud identification in subsequent use.

[0072] It can be seen that after the audio data file is framed, the speech frames obtained by framing are then preprocessed to obtain the target speech frames. Then, speech recognition is performed on the target speech frames, and the text data obtained by speech recognition is segmented to obtain segmentation vocabulary and segmentation emotions. After that, text keywords are selected from the segmentation vocabulary according to the segmentation emotions. Finally, the speech features at the second moment are obtained according to the text keywords. In this way, by selecting text keywords from the segmentation vocabulary according to the segmentation emotions, the selected text keywords can more accurately reflect the emotions of the target user. Therefore, obtaining the speech features at the second moment according to the text keywords can also improve the accuracy of fraud identification in subsequent use.

[0073] MFCC has good robustness, conforms to the auditory characteristics of the human ear, and can still have good recognition performance when the signal-to-noise ratio is reduced. The following uses MFCC as an example to introduce the process of extracting speech features. Speech features include Mel-frequency cepstral coefficients. In a possible implementation, the speech features at the second moment are obtained according to the text keywords. Specifically, the following steps may be included: fast Fourier transform processing is performed on the target speech frame corresponding to the text keyword to obtain speech spectrum data; the speech spectrum data is input into the Mel filter to obtain Mel frequency data; the Mel frequency data is subjected to cepstral analysis to obtain Mel frequency cepstral coefficients; the text feature words and the Mel frequency cepstral coefficients are used as the speech features at the second moment.

[0074] In the embodiment of the present application, the target speech frame is a signal obtained after the audio data file is processed by framing, and it is still a time domain signal at this time. However, it is difficult to see the characteristics of the signal in the time domain signal, so it is necessary to convert the time domain signal into a frequency domain signal. Fast Fourier transform (fast Fourier transform, FFT) is a general term for the fast calculation of discrete Fourier transform (discreteFourier transform, DFT). After FFT processing, the time domain signal of the target speech frame can be converted into a frequency domain signal of speech spectrum data. Mel filter can smooth the speech spectrum data, and play a role in eliminating filtering, highlighting the resonance peak characteristics of speech. Finally, the Mel spectrum data is subjected to cepstrum analysis to obtain MFCC as a speech feature, and cepstrum analysis can be implemented using discrete cosine transform (discrete cosine transform, DCT). DCT is a transform related to Fourier transform, similar to DFT, but DCT only uses real numbers. DCT can be used to remove the correlation between signals of each dimension and map the signal to a low-dimensional space.

[0075] It can be seen that after Fourier transform processing is performed on the target audio frame, the obtained speech spectrum data is input into the Mel filter to obtain Mel spectrum data. Finally, the obtained Mel spectrum data is subjected to cepstrum analysis to obtain Mel frequency cepstrum coefficients. In this way, the obtained Mel frequency cepstrum coefficients can have good robustness, which is conducive to improving the accuracy of fraud identification.

[0076] In a possible implementation, after executing step S203, the following steps may also be included: inputting the basic information, voice features or video features of the target user into a preset blacklist database to determine whether the target user is a non-blacklist user; if so, executing step S204.

[0077] In the embodiment of the present application, the blacklist user refers to a user who is determined to have committed fraud. The preset blacklist database may be pre-stored in the electronic device, or stored in the server, and the electronic device obtains the preset database by accessing the server. The preset blacklist database may store the basic information of the blacklist user. The basic information of the blacklist user may include the identity information of the blacklist user and the identity information of the associated users of the blacklist user. Among them, the identity information may include name, ID number, contact number, gender, age, address, occupation and education, etc. The blacklist user and the identity information of the blacklist user may be filled in by the blacklist user when submitting the application for the service, or may be determined by the relevant information submitted in the application information, wherein the relevant information may include ID card, marriage certificate, property certificate, household register information, etc. The basic information of the blacklist user may also include overdue repayment information, transaction flow information, personal consumption installment contract, channel consumption installment contract or loan or borrowing declaration information, etc.

[0078] In a possible implementation, the basic information of the target user may be input into a preset blacklist database to determine whether the basic information of the target user matches the basic information of the blacklist user, and whether the target user is a blacklist user is determined based on the matching result.

[0079] In the embodiment of the present application, if the basic information of the target user matches the name, ID card and other identity information of the blacklist user, the target user is determined to be a blacklist user. If the basic information of the target user matches the name, ID card and other identity information of the associated user of the blacklist user, the target user is marked as a suspected blacklist user, and the voice features and video features of the suspected blacklist user can be identified later to determine whether the suspected blacklist user is a blacklist user. If the basic information of the target user is inconsistent with the basic information of the blacklist user, the target user is determined to be a non-blacklist user.

[0080] It can be seen that by inputting the basic information of the target user into the preset blacklist database and judging whether the target user is a blacklist user based on the recognition result, the accuracy of fraud recognition can be improved.

[0081] The preset blacklist database may also store a pre-trained blacklist voiceprint recognition model, voice features of blacklist users, and basic information of blacklist users. In a possible implementation, the following steps may be performed to determine whether the target user is a blacklist user: input the voice features of the target user into the blacklist voiceprint recognition model; determine whether the voice features of the target user match the voice features of the blacklist user; if the corresponding voice features are matched, determine that the target user is a blacklist user; otherwise, determine that the target user is a non-blacklist user.

[0082] Voiceprint recognition, also known as speaker recognition, is a technology that identifies the identity of a speaker through sound. In an embodiment of the present application, the blacklist voiceprint recognition model may include but is not limited to a Gaussian mixture model (GMM), SVM, deep neural network (DNN), etc. Among them, the blacklist voiceprint recognition model is a text-independent model and does not restrict the content of the input audio data file. In this way, the blacklist voiceprint recognition model can perform identity recognition based on any audio data file of the user, which is conducive to reducing the dependence on the audio data file.

[0083] It can be seen that by inputting the target user's voice features into the blacklist voiceprint recognition model and judging whether the target user is a blacklist user based on the recognition results, the accuracy of fraud identification can be improved.

[0084] The preset blacklist database may also store a pre-trained blacklist face recognition model, video features of blacklist users, and basic information of blacklist users. In a possible implementation, the following steps may be performed to determine whether the target user is a blacklist user: input the video features of the target user into the blacklist face recognition model; determine whether the video features of the target user match the video features of the blacklist user; if the matching degree of the video features of the target user and the video features of the blacklist user is greater than a preset threshold, determine that the target user is a blacklist user; otherwise, determine that the target user is a non-blacklist user.

[0085] In the embodiment of the present application, the blacklist face recognition model may include but is not limited to CNN, hidden Markov model (HMM), eigenface method (Eigenface), etc. The preset threshold can be determined according to the actual situation. For example, the preset threshold can be set to 80%. When the video features of the target user match the video features of the blacklist user by more than 80%, the target user is determined to be a blacklist user; otherwise, the target user is a non-blacklist user. In this way, by inputting the video features of the target user into the blacklist voiceprint face model, judging whether the target user is a blacklist user based on the recognition results, the accuracy of fraud identification can be improved.

[0086] It can be seen that by inputting the basic information, audio features and video features of the target user into the preset blacklist database, it is determined whether the target user is a blacklist user. If the target user is a blacklist user, it can be determined that the target user is a fraud user; if the target user is not a blacklist user, step S204 is executed. In this way, not only the efficiency of fraud identification can be improved, but also the accuracy of fraud identification can be improved.

[0087] Step S204: Determine reasonable values ​​of the target user for each question and answer based on the voice features and video features at the second moment and the video features at the third moment.

[0088] In the embodiment of the present application, the reasonable value is used to describe the consistency of the audio features and the video features when the target user answers each question. That is, the closer the audio features of the target user's answer are to the video features of the target user's answer, the greater the reasonable value. In a possible implementation, step S204 may include the following steps A1-A6, wherein:

[0089] A1: Clustering the speech features and video features at the second moment according to each of at least two clustering methods to obtain a feature set corresponding to the clustering method.

[0090] Among them, the clustering method can be at least two of the clustering methods such as k-means clustering algorithm (k-means), fuzzy C-means clustering algorithm (fuzzy c-means, FCM), density-based spatial clustering of applications with noise (density-based spatial clustering of applications with noise, DBSCAN), mean shift clustering algorithm, etc., and the embodiments of the present application do not limit this. The feature set can be audio features and video features that express the same emotion. For example, audio features and video features that express fear are used as a feature set; audio features and video features that express contempt are used as a feature set, and so on, which are not limited here.

[0091] The following description is made by taking the k-means clustering algorithm as an example. In a possible implementation manner, step A1 may specifically include steps A11-A15, wherein:

[0092] A11: Select at least two speech features or video features from the second moment as initial clustering centers.

[0093] A12: Calculate the similarity between each voice feature and video feature at the second moment and the initial cluster center. The similarity can be calculated based on the Euclidean distance.

[0094] A13: Clustering the voice features and video features at the second moment according to the calculated similarity and the preset minimum similarity.

[0095] Specifically, if the calculated similarity is greater than or equal to the preset minimum similarity, the voice features and video features are clustered with the initial cluster center. Otherwise, no clustering is performed. Further, all voice features and video features whose calculated similarities are less than the preset minimum similarity can be clustered into one category.

[0096] A14: Reselect the cluster center from the speech features and video features at the second moment after clustering, and re-cluster until the cluster center converges or reaches a specified number of iterations.

[0097] A15: The finally determined cluster center is determined as a feature set.

[0098] It can be seen that the clustering algorithm is used to form cluster centers of similar voice features and video features, and the final cluster centers obtained through iterative updates are determined as feature sets. This can reduce the errors caused by human subjective factors, making the obtained feature set more representative, which is conducive to improving the accuracy of fraud identification.

[0099] A2: Input the feature set corresponding to the clustering method into the clustering sub-model corresponding to the clustering method to obtain the similarity sub-value of the feature set corresponding to the clustering method.

[0100] In the implementation of this application, the clustering submodel is used to describe the consistency of the input feature set. Each clustering submodel can be trained based on a feature set of a clustering method, and the value obtained by each clustering submodel can be called the similarity subvalue of the feature set corresponding to its clustering method. The larger the similarity subvalue, the more similar the emotions of the clustered speech features and the emotions of the video features in the feature set are.

[0101] For example, the feature set of the first clustering method is input into the clustering submodel corresponding to the first clustering method, and the value of the output result is 80%, then the similarity subvalue of the feature set corresponding to the first clustering method is determined to be 80%. The feature set of the second clustering method is input into the clustering submodel corresponding to the second clustering method, and the value of the output result is 90%, then the similarity subvalue of the feature set corresponding to the second clustering method is determined to be 90%.

[0102] The present application embodiment does not limit the training method of the clustering sub-model. Taking the k-means clustering algorithm as an example, the following steps B1-B6 may be included, wherein:

[0103] B1: Obtain a video training set, and extract a first audio feature and a first video feature from the video training set;

[0104] B2: performing feature concatenation on the first audio feature and the first video feature to obtain a target feature;

[0105] B3: Randomly select at least two objects from the target feature as the first cluster center;

[0106] B4: Calculate the distance between each object in the target feature and the first cluster center, assign it to the first cluster center with the closest distance, and obtain the second cluster center;

[0107] B5: Calculate the average value of the second cluster center, and use the average value of the second cluster center to update the first cluster center;

[0108] B6: Repeat steps B4-B5 until the model converges or reaches the specified number of iterations, and use the trained model as the clustering sub-model corresponding to the k-means clustering algorithm.

[0109] Among them, the video training set can be the conversation information between the user and the approval personnel when applying for a credit card, loan or performing other activities. The video training set may include video clips in which the user and the approval personnel speak separately, or may include video clips in which the user and the approval personnel speak at the same time, or may include video clips in which neither the user nor the approval personnel speaks. The definitions of the first audio feature and the first video feature can refer to the audio features and video features in the previous text, and will not be repeated here. The target feature is obtained by splicing the first audio feature and the first video feature. For example, the dimension of the first audio feature is P dimension, and the dimension of the first video feature is Q dimension, then the dimension of the target feature obtained after splicing is P+Q dimension. The distance calculation formula can be the Euclidean distance in one-dimensional space, or other distance metrics, which are not limited here. The square error criterion can be used as the objective function, which is defined as follows:

[0110]

[0111] Where E is the sum of the squared errors of all objects in the audio features or video features in the video training set, p is a point in space, and m i It is cluster C i The average value of .

[0112] The clustering sub-model may be a model obtained by independent training as described above, or may be a part of a clustering model obtained by training a feature set based on different clustering methods, etc., which is not limited here.

[0113] A3: Perform weighted calculation on the similarity subvalues ​​of the feature set corresponding to the clustering method and the preset weights corresponding to the clustering method to obtain similarity values ​​of the speech features and the video features at the second moment.

[0114] The similarity values ​​of the speech features and video features at the second moment are used to describe the consistency of the audio features and video features at the second moment. The similarity values ​​of the speech features and video features at the second moment can be determined based on the similarity subvalues ​​of the feature set corresponding to the clustering method and the preset weights corresponding to the clustering method. Among them, the preset weights corresponding to the clustering method can be determined based on the accuracy of the clustering submodel corresponding to the clustering method, which is not limited here.

[0115] For example, at least two clustering methods include a first clustering method and a second clustering method. The preset weight corresponding to the first clustering method is 95%, and the similarity subvalue of the feature set corresponding to the first clustering method is 85%. The preset weight corresponding to the second clustering method is 85%, and the similarity subvalue of the feature set corresponding to the first clustering method is 80%. Then, the similarity value S of the speech feature and the video feature at the second moment can be:

[0116]

[0117] A4: Determine the target video feature of the first duration based on the video feature of each moment in the first duration of the image data file.

[0118] The target video feature of the first duration is used to describe the overall video feature of a single image data file. The target video feature can be obtained by clustering based on at least one of the above clustering methods, or by counting the video features at each moment in the first duration.

[0119] A5: Obtain the matching value between the video feature at the third moment and the target video feature.

[0120] In an embodiment of the present application, the matching value is used to describe the degree of matching between the video features of the target user at the third moment and the video features during the entire question-and-answer duration, or it can be understood as the consistency between the facial features of the target user at the third moment and the facial features corresponding to the first duration. The matching value can be obtained by clustering based on at least one of the above clustering methods. The video features at the third moment and the target video features can also be input into a pre-trained machine learning model to obtain the matching value between the video features at the third moment and the target video features. The machine learning model can be CNN, RNN, fully convolutional networks (FCN); it can also be one or more of LSTM, SVM and other models, without limitation.

[0121] A6: Determine the reasonable value of each question and answer for the target user based on the similarity value and the matching value.

[0122] In the embodiment of the present application, the reasonable value of the target user for each question and answer can be obtained by weighting based on the similarity value and the preset weight corresponding to the similarity value, and the matching value and the preset weight corresponding to the matching value. The preset weight between the similarity value and the matching value can be determined based on the time length between all the second moments and all the third moments, or based on the number of features between all the second moments and all the third moments, etc., and there is no limitation on this. For example, the number of features at the second moment is 95, and the number of features at the third moment is 80. The weights corresponding to the similarity value and the matching value can be expressed as follows:

[0123]

[0124]

[0125] Among them, P1 represents the number of features at the second moment, and P2 represents the number of features at the third moment; W1 represents the preset weight of the similarity value, and W2 represents the preset weight of the matching value. If the similarity value is 80% and the matching value is 70%, then a reasonable value P for a question and answer is obtained by weighted calculation based on W1 and W2:

[0126]

[0127] In a possible implementation, the feature set corresponding to each clustering method may be added to the historical training set, and the clustering sub-model corresponding to the clustering method may be continuously iteratively optimized to improve the accuracy of similarity sub-value calculation.

[0128] It can be seen that selecting the voice features and video features at the same moment (the second moment) as the feature set makes the feature selection more representative and can improve the accuracy of fraud identification. In addition, each feature set is a feature obtained based on a clustering algorithm, and then the similarity sub-value of the feature set corresponding to each clustering method is obtained based on the clustering sub-model corresponding to the clustering algorithm, which is conducive to improving the similarity value of the voice features and video features at the second moment. Based on the similarity value of the voice features and video features at the second moment, and the matching value between the video features at the third moment and the target video features, the reasonable value of the target user for each question and answer is determined, which helps to improve the accuracy of the fraud identification results of the target user.

[0129] Step S205: If the reasonable value is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user.

[0130] In an embodiment of the present application, the preset threshold may be determined based on the basic information and loan information of the target user, which is not limited here. Different preset thresholds may be set for different target users. For example, if the target user has a relatively good asset situation and a high level of education, the preset threshold of the target user is lower. Loan information may include loan amount, repayment period, and repayment method, etc. Loan information may be obtained through channels such as target user input, partner crawling, and credit investigation of the People's Bank of China. If the target user has no existing loans or has few existing loans, the preset threshold of the target user is lower. In addition, for target users with relatively good asset conditions and high education, if the loan amount applied for this time is too high, the preset threshold of the target user may also be a high-risk value. Similarly, for target users with poor asset conditions and low education, if the loan amount applied for this time is small, the preset threshold of the target user may also be a low-risk value.

[0131] The calculation of the reasonable value can refer to the above description, which will not be repeated here. For example, the preset threshold can be set to 60%, and if the calculated reasonable value is 75.4% (greater than the preset threshold), it can be determined that the target user is a fraudulent user.

[0132] Alternatively, in a possible implementation, after step S204, the following steps may also be included: if the reasonable value is less than a preset threshold, weighting is performed based on the preset weight and reasonable value of the target video segment to obtain a target reasonable value; if the target reasonable value is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user; or if the target reasonable value is less than the preset threshold, the target user is determined to be a non-fraudulent user.

[0133] In the embodiment of the present application, the preset weight of the target video clip can be determined according to the specific content of the face-to-face interview answer. For example, the preset weight of the target video clip whose face-to-face interview answer involves the basic information of the target user (such as name, ID number, contact number) can be set to 0.9; the preset weight of the target video clip whose face-to-face interview answer involves assets (such as at least one of real estate, car property, insurance, salary and bank flow) can be set to 0.85; the preset weight of the target video clip whose face-to-face interview answer involves consumption can be set to 0.8, etc.

[0134] The reasonable value of the target video clip can be determined by referring to the above description, which will not be repeated here. For example, the preset threshold can be set to 60%. If the calculated target reasonable value is 65% (greater than the preset threshold), it can be determined that the target user is a fraudulent user. If the calculated target reasonable value is 55% (less than the preset threshold), it can be determined that the target user is not a fraudulent user.

[0135] It can be seen that if the reasonable value is less than the preset threshold, the preset weight and reasonable value of the target video clip are weighted to obtain the target reasonable value. Then, according to the size of the target reasonable value and the preset threshold, it is judged whether the target user is a fraudulent user. In this way, the comprehensiveness and diversity of fraud identification can be improved, thereby improving the accuracy of fraud identification.

[0136] exist Figure 2In the method shown, the target user's video data to be detected is processed to obtain the target user's audio data files and image data files for each question and answer. Then, feature extraction is performed on the image data file to obtain the video features of each moment in the first duration of the image data file, wherein the first duration includes the second moment and the third moment, and feature extraction is performed on the audio data file to obtain the voice features of the second moment. Then, based on the voice features and video features at the same moment (second moment) in the image data file and the audio data file, and the video features at the third moment in the image data file, the reasonable values ​​of the target user for each question and answer are determined. If one of the reasonable values ​​is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user. In this way, the user type is identified through the video features and voice features at the same moment, and the video features at the moment when there is no voice feature, which can improve the accuracy of identifying whether the target user is a fraudulent user, which is conducive to improving risk control.

[0137] The method of the embodiment of the present application is described in detail above, and the device of the embodiment of the present application is provided below.

[0138] Please refer to Figure 3 , Figure 3 1 is a schematic diagram of a device for identifying user types provided in an embodiment of the present application. The device is applied to a server. Figure 3 As shown, the user type identification device 300 includes a data processing unit 301, a feature extraction unit 302 and a determination unit 303. The detailed description of each unit is as follows:

[0139] The data processing unit 301 is used to process the video data to be detected of the target user to obtain the audio data file and image data file of the target user for each question and answer;

[0140] The feature extraction unit 302 is used to perform feature extraction on the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration includes a second moment and a third moment; perform feature extraction on the audio data file to obtain voice features at the second moment;

[0141] The determination unit 303 is used to determine the reasonable value of the target user for each question and answer based on the voice features and video features at the second moment and the video features at the third moment; if the reasonable value is greater than or equal to a preset threshold, the target user is determined to be a fraudulent user.

[0142] In a possible implementation, the data processing unit 301 is specifically used to perform semantic recognition on the video data to be detected of the target user, obtain the target video segment of the target user for each question and answer; and extract the audio data file and image data file of the target video segment.

[0143] In a possible implementation, the determination unit 303 is further used to, if the reasonable value is less than a preset threshold, perform weighting based on the preset weight of the target video segment and the reasonable value to obtain a target reasonable value; if the target reasonable value is greater than or equal to the preset threshold, determine that the target user is a fraudulent user; or if the target reasonable value is less than the preset threshold, determine that the target user is a non-fraudulent user.

[0144] In a possible implementation, the feature extraction unit 302 is specifically used to perform frame processing on the image data file to obtain a first video frame at each moment in the first duration of the image data file; perform key frame extraction on the first video frame to obtain a second video frame; perform facial feature extraction on the second video frame to obtain an action unit; and determine the video features at each moment in the first duration of the image data file based on the action unit.

[0145] In a possible implementation, the feature extraction unit 302 is specifically used to perform frame processing on the audio data file to obtain speech frames at each moment in the first time length of the audio data file; preprocess the speech frames to obtain target speech frames at the second moment; perform speech recognition on the target speech frames to obtain text data; perform word segmentation processing on the text data to obtain word segmentation vocabulary and word segmentation emotions; select and determine text keywords from the word segmentation vocabulary according to the word segmentation emotions; and obtain speech features at the second moment according to the text keywords.

[0146] In a possible implementation, the feature extraction unit 302 is specifically used to perform fast Fourier transform processing on the target speech frame corresponding to the text keyword to obtain speech spectrum data; input the speech spectrum data into a Mel filter to obtain Mel frequency data; perform cepstral analysis on the Mel frequency data to obtain Mel frequency cepstral coefficients; and use the text keyword and the Mel frequency cepstral coefficients as speech features at the second moment.

[0147] In a possible implementation, the determination unit 303 is specifically used to cluster the voice features and video features at the second moment according to each of at least two clustering methods to obtain a feature set corresponding to the clustering method; input the feature set corresponding to the clustering method into the clustering sub-model corresponding to the clustering method to obtain a similarity sub-value of the feature set corresponding to the clustering method; perform weighted calculation on the similarity sub-value of the feature set corresponding to the clustering method and the preset weight corresponding to the clustering method to obtain a similarity value of the voice features and video features at the second moment; determine the target video features of the first time length based on the video features of each moment in the first time length of the image data file; obtain the matching value between the video features at the third moment and the target video features; and determine the reasonable value of the target user for each question and answer based on the similarity value and the matching value.

[0148] It should be noted that the implementation of each unit can also refer to Figure 2 The corresponding description of the method embodiment shown.

[0149] Please refer to Figure 4 , Figure 4 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 4 As shown, the computer device 400 includes a processor 401, a memory 402, and a communication interface 403, wherein the processor 401, the memory 402, and the communication interface 403 may be connected via a bus 405. The memory 402 stores a computer program 404, which is configured to be executed by the processor 401, and the computer program 401 includes instructions for executing the following steps:

[0150] Processing the to-be-detected video data of the target user to obtain an audio data file and an image data file of the target user for each question and answer;

[0151] Performing feature extraction on the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration includes a second moment and a third moment;

[0152] Performing feature extraction on the audio data file to obtain speech features at the second moment;

[0153] Determining a reasonable value of each question and answer of the target user based on the voice feature and the video feature at the second moment and the video feature at the third moment;

[0154] If the reasonable value is greater than or equal to a preset threshold, the target user is determined to be a fraudulent user.

[0155] In a possible implementation, in terms of processing the target user's to-be-detected video data to obtain the target user's audio data files and image data files for each question and answer, the computer program 404 specifically includes instructions for executing the following steps:

[0156] Perform semantic recognition on the video data to be detected of the target user to obtain the target video clips of the target user for each question and answer;

[0157] The audio data file and the image data file of the target video segment are extracted.

[0158] In a possible implementation, after determining the reasonable value of each question and answer based on the voice features and video features at the second moment and the video features at the third moment, the computer program 404 further includes instructions for executing the following steps:

[0159] If the reasonable value is less than a preset threshold, weighting is performed based on the preset weight of the target video segment and the reasonable value to obtain a target reasonable value;

[0160] If the target reasonable value is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user; or

[0161] If the target reasonable value is less than the preset threshold, the target user is determined to be a non-fraudulent user.

[0162] In a possible implementation manner, in the aspect of extracting features from the image data file to obtain video features at each moment in the first duration of the image data file, the computer program 404 specifically includes instructions for executing the following steps:

[0163] Performing frame processing on the image data file to obtain a first video frame at each moment in a first duration of the image data file;

[0164] Extracting key frames from the first video frame to obtain a second video frame;

[0165] Extracting facial features from the second video frame to obtain action units;

[0166] The video feature of each moment in a first duration of the image data file is determined based on the action unit.

[0167] In a possible implementation manner, in the aspect of extracting features from the audio data file to obtain the speech features at the second moment, the computer program 404 specifically includes instructions for executing the following steps:

[0168] Performing frame processing on the audio data file to obtain a speech frame at each moment in the first duration of the audio data file;

[0169] Preprocessing the speech frame to obtain a target speech frame at the second moment;

[0170] Performing speech recognition on the target speech frame to obtain text data;

[0171] Performing word segmentation processing on the text data to obtain word segmentation vocabulary and word segmentation sentiment;

[0172] Selecting text keywords from the segmentation vocabulary according to the segmentation sentiment;

[0173] The voice feature at the second moment is acquired according to the text keyword.

[0174] In a possible implementation, in the aspect of acquiring the speech feature at the second moment according to the text keyword, the computer program 404 specifically includes instructions for executing the following steps:

[0175] Performing fast Fourier transform processing on the target speech frame corresponding to the text keyword to obtain speech spectrum data;

[0176] Inputting the speech spectrum data into a Mel filter to obtain Mel frequency data;

[0177] Performing cepstrum analysis on the mel frequency data to obtain mel frequency cepstrum coefficients;

[0178] The text keywords and the Mel-frequency cepstral coefficients are used as speech features at the second moment.

[0179] In a possible implementation manner, in determining the reasonable value of the target user for each question and answer based on the voice feature and the video feature at the second moment and the video feature at the third moment, the computer program 404 specifically includes instructions for executing the following steps:

[0180] Clustering the speech features and the video features at the second moment according to each of at least two clustering methods to obtain a feature set corresponding to the clustering method;

[0181] Inputting the feature set corresponding to the clustering method into the clustering sub-model corresponding to the clustering method to obtain a similarity sub-value of the feature set corresponding to the clustering method;

[0182] Performing weighted calculation on the similarity subvalues ​​of the feature set corresponding to the clustering method and the preset weights corresponding to the clustering method to obtain similarity values ​​of the voice features and the video features at the second moment;

[0183] Determine a target video feature of the first duration based on the video feature of each moment in the first duration of the image data file;

[0184] Obtaining a matching value between the video feature at the third moment and the target video feature;

[0185] A reasonable value of the target user for each question and answer is determined based on the similarity value and the matching value.

[0186] Those skilled in the art will appreciate that for ease of description, Figure 4 Only one memory and processor are shown. In an actual terminal or server, there may be multiple processors and memories. The memory 402 may also be referred to as a storage medium or a storage device, etc., which is not limited in the present embodiment of the application.

[0187] It should be understood that in the embodiment of the present application, the processor 401 can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0188] It should also be understood that the memory 402 mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronize link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0189] It should be noted that when the processor 401 is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) is integrated in the processor.

[0190] It should be noted that the memory 402 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0191] The bus 405 may include, in addition to the data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as buses in the figure.

[0192] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software. The steps of the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in a processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.

[0193] In various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0194] Those skilled in the art will appreciate that the various illustrative logical blocks (ILBs) and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0195] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0196] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0197] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0198] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integration. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk), etc.

[0199] In the above embodiment, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of blockchain nodes, etc. For example, the blockchain may store basic information of blacklisted users in a preset blacklist database, voice features of blacklisted users, video features of blacklisted users, blacklist voiceprint recognition models, and blacklist face recognition models, etc. Alternatively, 3DCNN algorithms, ST-GCN algorithms, SVM algorithms, ASR algorithms, GMM algorithms, DNN algorithms, etc. may be stored, or k-means algorithms, FCM algorithms, DBSCAN algorithms, etc. in clustering algorithms may be stored, which are not limited here.

[0200] The blockchain referred to in the embodiments of the present application is a new application model of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods, each of which contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.

[0201] An embodiment of the present application also provides a computer storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement part or all of the steps of any one of the user type identification methods recorded in the above method embodiments.

[0202] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any one of the user type identification methods recorded in the above method embodiments.

[0203] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for identifying user types, It is characterized in that include: Processing the to-be-detected video data of the target user to obtain an audio data file and an image data file of the target user for each question and answer; Performing feature extraction on the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration is a duration during which the face of the target user is included in the video data to be detected, and the first duration includes a second moment during which the voice of the target user is included in the video data to be detected and a third moment during which the voice of the target user is not included in the video data to be detected; Performing feature extraction on the audio data file to obtain speech features at the second moment; Clustering the speech features and the video features at the second moment according to each of at least two clustering methods to obtain a feature set corresponding to the clustering method; Inputting the feature set corresponding to the clustering method into the clustering sub-model corresponding to the clustering method to obtain a similarity sub-value of the feature set corresponding to the clustering method; Performing weighted calculation on the similarity subvalues ​​of the feature set corresponding to the clustering method and the preset weights corresponding to the clustering method to obtain similarity values ​​of the voice features and the video features at the second moment; Determine a target video feature of the first duration based on the video feature of each moment in the first duration of the image data file; Obtaining a matching value between the video feature at the third moment and the target video feature; Determine a reasonable value of the target user for each question and answer based on the similarity value and the matching value; If the reasonable value is greater than or equal to a preset threshold, the target user is determined to be a fraudulent user.

2. The method according to claim 1, It is characterized in that The processing of the target user's video data to be detected to obtain the target user's audio data files and image data files for each question and answer includes: Perform semantic recognition on the video data to be detected of the target user to obtain the target video clip of the target user for each question and answer; The audio data file and the image data file of the target video segment are extracted.

3. The method according to claim 2, It is characterized in that After determining the reasonable value of the target user for each question and answer based on the voice feature and the video feature at the second moment and the video feature at the third moment, the method further includes: If the reasonable value is less than a preset threshold, weighting is performed based on the preset weight of the target video segment and the reasonable value to obtain a target reasonable value; If the target reasonable value is greater than or equal to the preset threshold, the target user is determined to be a fraudulent user; or If the target reasonable value is less than the preset threshold, the target user is determined to be a non-fraudulent user.

4. The method according to any one of claims 1 to 3, It is characterized in that The extracting features of the image data file to obtain the video features of each moment in the first duration of the image data file includes: Performing frame processing on the image data file to obtain a first video frame at each moment in a first duration of the image data file; Extracting key frames from the first video frame to obtain a second video frame; Extracting facial features from the second video frame to obtain action units; The video feature of each moment in a first duration of the image data file is determined based on the action unit.

5. The method according to any one of claims 1 to 3, It is characterized in that The extracting features from the audio data file to obtain the speech features at the second moment includes: Performing frame processing on the audio data file to obtain a speech frame at each moment in the first duration of the audio data file; Preprocessing the speech frame to obtain a target speech frame at the second moment; Performing speech recognition on the target speech frame to obtain text data; Performing word segmentation processing on the text data to obtain word segmentation vocabulary and word segmentation sentiment; Selecting text keywords from the segmentation vocabulary according to the segmentation sentiment; The voice feature at the second moment is acquired according to the text keyword.

6. The method according to claim 5, It is characterized in that The acquiring the voice feature at the second moment according to the text keyword includes: Performing fast Fourier transform processing on the target speech frame corresponding to the text keyword to obtain speech spectrum data; Inputting the speech spectrum data into a Mel filter to obtain Mel frequency data; Performing cepstrum analysis on the mel frequency data to obtain mel frequency cepstrum coefficients; The text keywords and the Mel-frequency cepstral coefficients are used as speech features at the second moment.

7. A device for identifying user type, It is characterized in that include: A data processing unit, used to process the video data to be detected of the target user to obtain the audio data file and image data file of the target user for each question and answer; a feature extraction unit, configured to perform feature extraction on the image data file to obtain video features at each moment in a first duration of the image data file, wherein the first duration is a duration in which the face of the target user is included in the video data to be detected, and the first duration includes a second moment in which the voice of the target user is included in the video data to be detected and a third moment in which the voice of the target user is not included in the video data to be detected; perform feature extraction on the audio data file to obtain voice features at the second moment; A determination unit is used to cluster the voice features and video features at the second moment according to each of at least two clustering methods to obtain a feature set corresponding to the clustering method; input the feature set corresponding to the clustering method into a clustering sub-model corresponding to the clustering method to obtain a similarity sub-value of the feature set corresponding to the clustering method; perform weighted calculation on the similarity sub-value of the feature set corresponding to the clustering method and a preset weight corresponding to the clustering method to obtain a similarity value of the voice features and video features at the second moment; determine the target video features of the first time length based on the video features of each moment in the first time length of the image data file; obtain a matching value between the video features at the third moment and the target video features; determine a reasonable value of the target user for each question and answer based on the similarity value and the matching value; if the reasonable value is greater than or equal to a preset threshold, determine that the target user is a fraudulent user.

8. A computer device, It is characterized in that The method comprises a processor, a memory and a communication interface, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the steps in any one of the methods of claims 1 to 6.

9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, which enables a computer to execute to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Improvements in or relating to doorlocks

    AU140051B

  • Facial examination risk control method and device, computer equipment and storage medium

    CN111429267A