Data processing method, device and program product

By collecting facial image or video feature data in financial transactions, and utilizing target recognition models and anomaly detection technology, combined with static and dynamic biometrics, the problem of low security in biometric technology has been solved, achieving higher transaction security.

CN120808459APending Publication Date: 2025-10-17INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510882395.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The security verification of financial services in existing technologies is low, and biometric technology is easily forged and cannot be recognized under duress.

Method used

The system collects feature data from facial images or videos, detects live objects using a target recognition model, combines static and dynamic biometrics for anomaly detection, and uses neural networks and autoregressive moving average models to predict abnormal behavior.

Benefits of technology

It improves the reliability of security verification of financial services, avoids security vulnerabilities in cases of forgery and coercion, and achieves higher transaction security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808459A_ABST
    Figure CN120808459A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device and a program product. Relates to the field of artificial intelligence or financial science and technology, and the method comprises the steps: collecting face feature data in a face image or a face video when a financial service is handled, inputting the face feature data into a target recognition model, and obtaining a recognition result, the recognition result is used for indicating whether a target object in the face image or the face video is a moving object; under the condition that the recognition result indicates that the target object in the face image or the face video is a moving object, collecting static biological characteristics of the target object to obtain static characteristic data, and collecting dynamic behavior characteristics of the target object to obtain dynamic characteristic data; and performing anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result. According to the method and the device, the problem of low security of security verification in a financial business handling process in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence or the field of financial technology, in particular, to a data processing method, device and program product. BACKGROUND

[0002] At present, in the process of handling business in a financial institution, the identity of a user needs to be verified to ensure the security of the transaction.

[0003] In the related art, the user's identity is mainly identified by using biometric technology (e.g., fingerprint, face, iris, etc.), but due to the fact that the biometric features in the related art have certain illegal counterfeit conditions (e.g., counterfeit fingerprints, masks, false irises, voice imitation, etc. to bypass the system verification), in addition, there are some users who may be in a threatened or deceived state, but the biometric method of the financial institution in the related art cannot identify this situation, resulting in low security of the transaction.

[0004] In view of the low security of the security verification in the process of handling financial business in the related art, an effective solution has not been proposed so far. SUMMARY

[0005] The main purpose of the present application is to provide a data processing method, device and program product to solve the problem of low security of security verification in the process of handling financial business in the related art.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a data processing method is provided. The method comprises: collecting facial feature data in a face image or face video when handling financial business, and inputting the facial feature data into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether a target object in the face image or face video is an active object, and the target recognition model comprises a neural network model for active object detection; in the case that the recognition result indicates that the target object in the face image or face video is an active object, collecting a static biometric feature of the target object to obtain static feature data, and collecting a dynamic behavior feature of the target object to obtain dynamic feature data; performing abnormality detection on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has abnormal behavior.

[0007] Further, the static feature data comprises at least one of a fingerprint feature of the target object, an iris feature of the target object, and a face feature of the target object, and the dynamic feature data comprises at least one of a voiceprint feature of the target object, a speech content of the target object, and a change of facial expression of the target object.

[0008] Further, the static feature data and the dynamic feature data are subjected to anomaly detection to obtain a target detection result, comprising: detecting whether the static feature data is fake, to obtain a first detection result; detecting whether the dynamic feature data is abnormal, to obtain a second detection result; inputting the static feature data and the dynamic feature data into a target prediction model to predict a behavior feature of the target object in a target time period, to obtain a predicted behavior feature, and determining a third detection result based on a difference between the predicted behavior feature and an actual behavior feature of the target object in the target time period, wherein the target prediction model comprises a model-trained autoregressive moving average model; and determining the target detection result based on the first detection result, the second detection result, and the third detection result.

[0009] Further, the face feature data is collected from a face image or a face video, comprising: collecting a face image or a face video of a target object to obtain first multimedia data; extracting a face image from the first multimedia data to obtain a target face image, and generating a texture feature image of the target face image; and extracting features from the target face image and the texture feature image to obtain the face feature data.

[0010] Further, the face feature data is collected from a face image or a face video, comprising: collecting a face image or a face video of a target object to obtain first multimedia data; extracting a face image from the first multimedia data to obtain a target face image, and generating a texture feature image of the target face image; and extracting features from the target face image and the texture feature image to obtain the face feature data.

[0011] Further, the target recognition model is obtained by: obtaining multimedia data of M real human faces and multimedia data of N false human faces, to obtain S second multimedia data, wherein M and N are positive integers, and S=N+M; extracting a human face feature in each of the S second multimedia data to obtain S human face feature images; generating an image label of each of the S human face feature images to obtain S target image labels, wherein the image label of each of the human face feature images comprises: image information in the human face feature image; performing binaryzation processing on each of the S target image labels to obtain S binary mask labels; and performing model training on an initial recognition model based on the S binary mask labels, and determining the initial recognition model after the training is completed as the target recognition model.

[0012] Further, the model training on the initial recognition model based on the S binary mask labels comprises: performing model training on the initial recognition model based on the S binary mask labels and a target loss function, and adjusting a hyperparameter in the target loss function in the model training process, wherein the target loss function comprises at least one of: an amplitude loss function, a frequency spectrum loss function, and a repeatability loss function, the amplitude loss function is used to obtain a similar part between data predicted by the initial recognition model and actually input data, the frequency spectrum loss is used to obtain a difference between the data predicted by the initial recognition model and the actually input data, and the repeatability loss function is used to amplify noise between the data predicted by the initial recognition model and the actually input data.

[0013] Further, after the abnormality detection on the static feature data and the dynamic feature data to obtain a target detection result, the method further comprises: in a case where the target detection result indicates that the target object has an abnormal behavior, generating a pre-warning prompt information, and performing risk assessment on a financial service handled by the target object to obtain a risk assessment result, wherein the risk assessment result is used to indicate whether the target object is in a threatened state.

[0014] In order to achieve the above object, according to another aspect of the present application, a data processing apparatus is provided. The apparatus comprises: a processing unit configured to collect facial feature data in a facial image or a facial video when a financial service is handled, and input the facial feature data into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether a target object in the facial image or the facial video is an active object, and the target recognition model comprises a neural network model configured to perform active object detection; a collection unit configured to collect a static biometric feature of the target object to obtain static feature data and collect a dynamic behavior feature of the target object to obtain dynamic feature data in a case where the recognition result indicates that the target object in the facial image or the facial video is an active object; and a detection unit configured to perform anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has an abnormal behavior.

[0015] Further, the static feature data comprises at least one of a fingerprint feature of the target object, an iris feature of the target object and a facial feature of the target object, and the dynamic feature data comprises at least one of a voiceprint feature of the target object, a voice content of the target object and a change condition of a facial expression of the target object.

[0016] Further, the detection unit comprises: a first detection subunit configured to detect whether the static feature data has a forgery condition to obtain a first detection result; a second detection subunit configured to detect whether the dynamic feature data has an abnormal change to obtain a second detection result; a processing subunit configured to input the static feature data and the dynamic feature data into a target prediction model to predict a behavior feature of the target object in a target time period to obtain a predicted behavior feature, and determine a third detection result based on a difference between the predicted behavior feature and an actual behavior feature of the target object in the target time period, wherein the target prediction model comprises a model-trained autoregressive moving average model; and a determination subunit configured to determine the target detection result based on the first detection result, the second detection result and the third detection result.

[0017] Further, the processing unit comprises: a collection subunit configured to collect a facial image or a facial video of a target object to obtain first multimedia data; a clipping subunit configured to clip a facial image in the first multimedia data to obtain a target facial image, and generate a texture feature image of the target facial image; and an extraction subunit configured to perform feature extraction on the target facial image and the texture feature image to obtain the facial feature data.

[0018] Further, the extraction subunit comprises: a fusion module configured to perform multi-modal feature fusion on the target face image and the texture feature image to obtain a fused feature image; and an extraction module configured to perform feature extraction on the fused feature image by using a feature extraction module and a down-sampling module to obtain the face feature data.

[0019] Further, the target recognition model is obtained by the following subunits: an acquisition subunit configured to acquire multimedia data of M real faces and multimedia data of N fake faces to obtain S second multimedia data, where M and N are positive integers and S = N + M; a first processing subunit configured to extract a face feature in each of the S second multimedia data to obtain S face feature images; a generation subunit configured to generate an image label of each of the S face feature images to obtain S target image labels, where the image label of each of the face feature images comprises image information in the face feature image; a second processing subunit configured to perform binaryzation processing on each of the S target image labels to obtain S binaryzation mask labels; and a training subunit configured to perform model training on an initial recognition model based on the S binaryzation mask labels, and determine the initial recognition model after the training is completed as the target recognition model.

[0020] Further, the training subunit comprises: a training module configured to perform model training on an initial recognition model based on the S binaryzation mask labels and a target loss function, and adjust a hyperparameter in the target loss function during the model training, where the target loss function comprises at least one of an amplitude loss function, a frequency spectrum loss function and a repeatability loss function, the amplitude loss function is used to acquire a similar part between data predicted by the initial recognition model and actually input data, the frequency spectrum loss is used to acquire a difference between the data predicted by the initial recognition model and the actually input data, and the repeatability loss function is used to amplify noise between the data predicted by the initial recognition model and the actually input data.

[0021] Further, the data processing apparatus further comprises: an evaluation unit configured to, after performing anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result, generate a pre-warning prompt information and perform risk evaluation on a financial service handled by the target object to obtain a risk evaluation result in a case where the target detection result indicates that the target object has an abnormal behavior, where the risk evaluation result is used to indicate whether the target object is in a threatened state.

[0022] According to another aspect of the present application, a computer readable storage medium is provided, comprising a stored executable program, wherein the computer readable storage medium is caused to perform the data processing method when the executable program is run.

[0023] According to another aspect of the present application, an electronic device is provided, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program is configured to perform the data processing method when run.

[0024] According to another aspect of the present application, a computer program product is provided, comprising computer instructions configured to implement the steps of the data processing method when executed by a processor.

[0025] In the embodiments of the present application, when a financial service is handled, the facial feature data in the face image or face video is collected, and the facial feature data is input into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether a target object in the face image or face video is an active object, and the target recognition model comprises a neural network model used for active object detection; in a case where the recognition result indicates that the target object in the face image or face video is an active object, static biometric features of the target object are collected to obtain static feature data, and dynamic behavior features of the target object are collected to obtain dynamic feature data; the static feature data and the dynamic feature data are subjected to abnormality detection to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has an abnormal behavior, thereby solving the technical problem of low security of security verification in the process of handling a financial service in the related art. In the present application, whether the target object is an active object is detected through the target recognition model, and the target object is subjected to abnormality detection through the static feature data and the dynamic feature data, thereby avoiding the case of low security caused by only verifying the identity of a user through biometric recognition technology in the related art, and thereby achieving the technical effect of improving the reliability of security verification of a financial service. BRIEF DESCRIPTION OF DRAWINGS

[0026] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The accompanying drawings in accordance with the present application comprise:

[0027] Figure 1 A hardware structure block diagram of a computer terminal for implementing the data processing method is shown;

[0028] Figure 2 A flowchart of the data processing method provided by the embodiments of the present application is shown;

[0029] Figure 3is a flowchart of security verification of a financial service provided according to an embodiment of the application;

[0030] Figure 4 is a flowchart of image processing provided according to an embodiment of the application;

[0031] Figure 5 is a schematic diagram of a data processing device provided according to an embodiment of the application;

[0032] Figure 6 is a structural block diagram of an electronic device according to an embodiment of the application. DETAILED DESCRIPTION

[0033] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0035] It should be noted that the data processing method and device thereof in the present application can be used in the field of artificial intelligence or financial technology, in the process of transaction or processing of financial business, in the case of security verification, and can also be used in any field other than the field of artificial intelligence or financial technology, in the process of transaction, in the case of security verification. The application field of the data processing method and device thereof in the present application is not limited.

[0036] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, face data, biometric data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal. For example, the system and related users or institutions are provided with an interface to provide the user with a corresponding operation portal for the user to choose to agree or refuse the automatic decision result; if the user chooses to refuse, the expert decision process is entered.

[0037] It should be noted that in the present application, the information and data authorized by the user or authorized by all parties can be viewed by the user in real time through the authorization interface, and the user has the right to withdraw authorization or delete data at any time. After withdrawing authorization, the system will terminate the related data processing within 24 hours.

[0038] The present application can be applied to various software products, control systems, client terminals (including but not limited to mobile client terminals, PC terminals, etc.) of various financial institutions. Taking a software product as an example, through the software product installed on the mobile client terminal, security verification can be performed when processing the business content of the financial institution (including but not limited to transfer, financial management, fund, payment, account inquiry, advertising, recommendation, etc.), and the security of the business can be improved.

[0039] Embodiment one

[0040] According to the embodiments of the present application, a method embodiment of a data processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0041] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the data processing method is shown. As shown in the figure, Figure 1As shown, the computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include, but not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or less components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .

[0042] It should be noted that the one or more processors 102 and / or other data processing circuits described above can be referred to herein as "data processing circuits" in general. The data processing circuit can be embodied in whole or in part as software, hardware, firmware or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or any one of the other elements incorporated into the computer terminal 10 (or mobile device) in whole or in part. As referred to in the embodiments of the present application, the data processing circuit serves as a processor to control (for example, selection of a variable resistance terminal path connected to an interface).

[0043] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the data processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned data processing method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory remotely disposed with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0044] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can connect to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module that is configured to communicate with the Internet wirelessly.

[0045] The display can be a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0046] In the above operating environment, the present application provides a data processing method as shown in Figure 2 Figure 2 is a flowchart of the data processing method according to the first embodiment of the present application.

[0047] In step S201, when a financial service is performed, facial feature data of a face image or a face video is collected, and the facial feature data is input into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether a target object in the face image or the face video is an active object, and the target recognition model includes a neural network model used for active object detection.

[0048] The financial service can include a service performed by a target object in a financial institution, and the target recognition model can be used to detect whether an object in a face image or a face video is an active object (i.e., whether it is a living body).

[0049] In this embodiment, when identity verification is required, the target object is informed that face information needs to be collected, and after the target object is verified for security when performing a service and the target object authorizes, a face video or a face image of an object performing a financial service can be collected by an image collection device, and facial features thereof are extracted to obtain facial feature data, and then the facial feature data can be input into a target recognition model to identify whether a target object in a face image or a face video is an active object.

[0050] In step S202, when the recognition result indicates that the target object in the face image or the face video is an active object, static biometric features of the target object are collected to obtain static feature data, and dynamic behavior features of the target object are collected to obtain dynamic feature data.

[0051] ​In a case where the identification result indicates that the target object in the face image or the face video is an active object, in a case where authorization of the target object is obtained, a static biometric feature (for example, a fingerprint, an iris, a facial feature, or the like) of the target object can be collected to obtain static feature data, and a dynamic behavior feature (for example, a dynamic interaction process with the target object, a feature of the target object, for example, a change in a voiceprint, an emotion, a speech content, and a facial expression in a process of answering a question, inputting a password through a touch screen, or signature verification, or the like) of the target object can be collected to obtain dynamic feature data.

[0052] In step S203, the static feature data and the dynamic feature data are subjected to anomaly detection to obtain a target detection result, where the target detection result is used to indicate whether the target object has an abnormal behavior.

[0053] In this embodiment, the static feature data can be subjected to anomaly detection. For example, the quality of a fingerprint image can be detected to avoid that a low-quality image affects identification, and blurring, scratches, and dirt can all indicate that the fingerprint information has been destroyed. A living body detection technology can also be used to determine whether the fingerprint is from a real and living person or is obtained from a fake fingerprint or a replica.

[0054] In this embodiment, the dynamic feature data can also be subjected to anomaly detection. For example, the emotional state of the target object (for example, a customer) can be determined by analyzing the voice features (such as a speech speed, a tone, a breathing pattern, or the like) of the target object. Whether the target object has a rapid tone or shows emotional fluctuations such as anxiety, panic, or the like during a conversation with a teller of a financial institution can also be determined by analyzing the context in the speech content to determine whether the customer is in a threatening situation.

[0055] In this embodiment, the static feature data and the dynamic feature data can also be input into a target prediction model to predict a behavior trajectory of the target object in a next time period, and whether the target object is in an abnormal state, that is, whether the target object has an abnormal behavior, can be determined according to whether an actual behavior trajectory of the target object in the next time period and the behavior trajectory predicted by the target prediction model are consistent.

[0056] Through the above steps, in this embodiment, whether the target object is an active object is detected by the target identification model, and the target object is subjected to anomaly detection by the static feature data and the dynamic feature data, thereby avoiding the case in the related art that only a biometric recognition technology is used to verify the identity of a user, and the security is low, thereby achieving the technical effect of improving the reliability of security verification of a financial service. Further, the technical problem of low security of security verification in a process of handling a financial service in the related art is solved.

[0057] Figure 3 is a flowchart of security verification of a financial service according to an embodiment of the present application, as shown inFigure 3 As shown, during the process of the customer conducting financial business at the counter, the face feature image / video can be acquired, and based on the face feature image / video, whether the feature value of the image in the face feature image / video is a live feature is identified. If it is not a live feature, the customer manager can be warned. If it is a live feature, the abnormal feature detection stage can be performed. Abnormal behavior analysis can be performed according to the feature value of the static feature and the feature value of the dynamic feature. Whether the feature value of the static feature and the feature value of the dynamic feature are abnormal feature values is identified. If there is no abnormal behavior, subsequent business can be completed. If there is abnormal behavior, the customer manager can be warned, and risk assessment can be performed. Before identifying whether the feature value of the image in the face feature image / video is a live feature, real face feature images / videos and fake face feature images / videos can be collected to train a neural network model to obtain a target recognition model, so that the target recognition model can perform live detection.

[0058] Optionally, in the data processing method provided in the embodiments of the present application, the static feature data includes at least one of the following: fingerprint features of the target object, iris features of the target object, and face features of the target object. The dynamic feature data includes at least one of the following: voiceprint features of the target object, voice content of the target object, and changes in facial expressions of the target object.

[0059] The static feature data described above can include: fingerprint features of the target object, iris features of the target object, and face features of the target object. The dynamic feature data of the target object can include: voiceprint features (such as pitch, speech speed, amplitude and volume of speech, and pause conditions) of the target object, voice content (such as keywords in the voice) of the target object, and changes in facial expressions (such as the positions of the eyebrows, eyes, and mouth of a person, and rapid changes in expressions) of the target object. In the present embodiment, by collecting the static feature data and the dynamic feature data of the target object, whether the target object is in a threatened state can be accurately identified, thereby achieving the technical effect of improving the reliability of security verification of transactions.

[0060] Optionally, in the data processing method provided in the embodiment of the present application, anomaly detection is performed on static feature data and dynamic feature data to obtain a target detection result, including: detecting whether the static feature data is forged to obtain a first detection result; detecting whether the dynamic feature data has abnormal changes to obtain a second detection result; inputting the static feature data and the dynamic feature data into a target prediction model to predict the behavioral characteristics of the target object in the target time period to obtain predicted behavioral characteristics, and determining a third detection result based on the difference between the predicted behavioral characteristics and the actual behavioral characteristics of the target object in the target time period, wherein the target prediction model includes: an autoregressive moving average model trained by the model; and determining the target detection result based on the first detection result, the second detection result, and the third detection result.

[0061] In this embodiment, anomaly detection can be performed on static biometric features in static feature data, that is, to detect whether the static biometric feature data is forged, or to detect whether the static feature data is generated when the target object is in a threatened state. For example, biometric recognition technology can be used to detect whether an individual is in an abnormal or potential threat state through static biometric features (such as fingerprints, irises, facial features, etc.). The detection can focus on the identification and verification of biometric features rather than behavioral patterns or emotional fluctuations. The core of static biometric anomaly detection is to determine whether there is a risk of impersonation or external threats by monitoring and comparing the user's identity information. The main function of static biometric features can be to verify identity, while anomaly detection can focus on whether there is a threat of tampering or illegal access during the identity authentication process. For example, fingerprints are a static biometric feature, and their uniqueness and stability are usually used for identity authentication. If it is detected that the user's fingerprint features are obviously mismatched with the preset fingerprint template, it can indicate that the fingerprint has been tampered with or forged. The anomaly may be caused by a threatened situation, for example, some illegal objects illegally cheat by copying or illegally forging fingerprints. Specific identification features may include: (1) Fingerprint image quality detection: The quality of the fingerprint image can be detected to avoid the influence of low-quality images on recognition. Blurred, scratched, or stained fingerprints may indicate that the fingerprint information has been destroyed. (2) Anti-forgery technology: For example, liveness detection technology is used to determine whether the fingerprint is from a real, living person, rather than a fake fingerprint or a replica.

[0062] The above feature detection is implemented based on relevant technical equipment. For example, a capacitive fingerprint identifier can be used at the counter of a financial institution. The identifier can obtain fingerprint features by sensing tiny current changes in the skin. It has strong anti-counterfeiting capabilities against fake fingerprints (such as silicone or photos). It has small size, high resolution, and is easier to obtain fine fingerprint information. After passing the biometric liveness detection stage, the fingerprint detector can be used to verify whether it is the person himself and whether there is any illegal identity forgery.

[0063] In this embodiment, the dynamic biometric features in the dynamic feature data can also be subjected to anomaly detection, for example, to identify whether the client is under threat through a dynamic interaction process. For example, the target object (e.g., the client) can be asked to perform some simple verification behaviors (such as answering questions, entering passwords through a touch screen, or performing signature verification) at the counter of the financial institution, while combining biometric recognition data to confirm whether the operation is in line with the normal mode. If the target object fails to answer questions normally or exhibits unnatural input behavior when conducting business at the counter, it can be further determined whether it is under threat, for example, the verification rigor can be automatically adjusted or other verification means can be required. The principle of dynamic biometric feature detection whether under threat is that, based on monitoring and analyzing the dynamic changes of the biometric features of the target object, abnormal patterns caused by psychological stress, panic, nervousness, or other threat situations are identified. Dynamic biometric features can fluctuate over time, and through monitoring of these changes, potential threatened or coerced states can be detected. Dynamic biometric features can refer to biometric information that changes dynamically over time and in different situations. Unlike static biometric features (such as fingerprints, irises, facial features, etc.), dynamic biometric features focus on behavioral patterns and physiological responses, which can change significantly over time, especially when emotional and psychological states change. Analyzing dynamic biometric features can include:

[0064] 1. Voiceprint recognition analysis: Voiceprint recognition technology can help identify the identity of the target object, and can also be combined with voice emotion analysis to determine whether the target object is under threat. For example, by analyzing the voice features of the target object (such as speech rate, tone, breathing pattern, etc.), the emotional state of the target object can be determined. If the target object speaks in a hurried manner or shows emotional fluctuations such as anxiety, panic, etc. during the conversation with the teller, it can be determined whether there is a risk, and additional verification procedures can be initiated, or the teller can be alerted that the target object may be in a threatened state, which can be identified in the following aspects: (1) Pitch: Pitch is an important feature of speech, and changes in human emotions (such as nervousness, anxiety, fear) often affect pitch. For example, when threatened, a person's voice may become higher and the tone more unstable. (2) Speech Rate: Individuals under threat may exhibit rapid speech. In a state of high anxiety, the speech rate may be too fast or too slow, reflecting the individual's psychological stress. (3) Volume of speech: When threatened, the target object may increase the volume or exhibit dramatic fluctuations in the volume of speech. For example, screaming, shouting, etc. can be manifested as a dramatic change in volume. (3) Pauses: In a threat situation, the target object may frequently pause or hesitate in speech due to fear or anxiety, affecting the fluency of speech.

[0065] 2. Emotion analysis and voice emotion recognition: Voiceprint recognition can not only determine the content of the voice, but also recognize the emotional changes of the target object. Through emotion analysis of the voice, it can be determined whether the target object shows anxiety, panic, fear, surprise, etc. These emotions are usually related to threats or coercion. Common emotional features can include: emotional feature extraction: emotion analysis extracts emotional components from the voice, such as high-pitched tone, rapid speech, emotional fluctuations, etc., to identify whether the target object is in a state of tension, panic, etc. Emotional category judgment: by comparing with the normal voice mode, it is judged whether the emotion in the voice is normal. For example, if the target object shows significant anxiety, urgency or unstable emotion in the voice, the system will judge that it may be in a threatening situation.

[0066] 3. Context analysis of voice content: In addition to the physiological and psychological characteristics of the voice, the content of the voice can also be analyzed to determine whether there are obvious threats or abnormal requests, and the context of the voice content can be analyzed to determine whether the target object is in a threatening situation. For example: keyword detection: key words in the voice content can be extracted through voice recognition technology, and if the target object mentions keywords such as "threat", "danger", "persecution" and other illegal nature, it can be judged as a potential threat signal. Context judgment: the language content of the target object can be combined to determine whether the scene described belongs to a threatening or coercive situation. For example, the target object may repeatedly repeat certain words or express a request for help when under threat.

[0067] 4. Facial expression changes: Facial expressions are a direct reflection of emotional state, especially in situations of tension, fear or anger, facial expressions will change significantly. For example, the position of a person's eyebrows, eyes, mouth, and rapid changes in expression can indicate a threatening situation. Specifically, (1) eye movement: the movement of the eyes can reflect a person's level of anxiety, frequent eye movement or tense eye contact may indicate unease. (2) Changes in the corners of the mouth: when anxious or fearful, the corners of the mouth will also change significantly. (3) Micro-expression of the face: micro-expression can reveal a person's true emotions in a specific situation, rapid facial muscle contractions are generally uncontrollable, and this feature is particularly pronounced when facing threats.

[0068] The target prediction model described above can be applied to the counter of a financial institution, and during the abnormal behavior recognition process of the object handling business, the target object's facial expressions, postures, hand movements, etc. can be monitored in real time through cameras and sensors. Through the abnormal behavior detection stage, the normal value of the object is obtained, a training set is formed, and a neural network model is trained to predict the next action of the object, and real-time analysis of those objects whose behavior deviates from the predicted value is defined as abnormal information.

[0069] The target prediction model described above can adopt a time series prediction algorithm. The algorithm model is trained based on the first n sample history on the "normal" value training set to predict the value of the next sample. In deployment, if the past history comes from a system working in "normal" condition, the prediction of the next sample value will be relatively accurate, close to the real sample value. If the past history sample comes from a system not running in "normal" condition, the predicted value will deviate from the actual value. In this case, the difference between the predicted sample value and the real sample value can be measured to identify abnormal event candidates. The counter mode is more suitable for ARMA algorithm (autoregressive moving average model), so the model type of the target prediction model in this embodiment can be an autoregressive model, and the specific formula is as follows:

[0070]

[0071] Wherein, the meanings of each index are as follows: X t represents the observation value of the time series at time t, c represents the constant term (intercept of the model), φ i represents the autoregressive coefficient of the AR part (i = 1, 2, 3,..., p), θ j represents the moving average coefficient of the MA part (j = 1, 2, 3,..., q), and ε t represents the random error (white noise), that is, the unpredictable error term, p represents the order of the autoregressive term (AR part), and q represents the order of the moving average term (MA part),

[0072] In this embodiment, ACF (autocorrelation function) and PACF (partial autocorrelation function) can be used to select the orders p and q of the model. The PACF cutoff corresponds to the order p of the AR part, and the ACF cutoff corresponds to the order q of the MA part.

[0073] After the target prediction model is trained according to the ARMA algorithm, the static feature data and the dynamic feature data can be input into the target prediction model, and the predicted value of the target object in the next time period can be obtained through the target prediction model. According to the actual predicted value obtained from the counter video picture, whether the target object shows unnatural behavior signs such as anxiety and tension can be analyzed. For the case of identifying abnormal behavior, an alarm can be triggered to remind the financial institution staff to conduct more detailed verification. The behavior analysis algorithm (i.e. ARMA algorithm) in this embodiment can capture subtle changes of the target object when conducting business at the counter, such as abnormal emotions, tension state or abnormal actions.

[0074] In this embodiment, in the abnormal behavior recognition process, biological recognition information can also be combined to determine whether the user is likely to be in danger (such as being coerced), and a deep learning model (corresponding to a target prediction model) is used to analyze the behavior pattern of the target object when conducting business at the counter, compare it with its historical behavior, and find potential abnormalities. For example, a target object usually conducts simple transactions at the counter, but today attempts to conduct large-value transfers or other unusual operations, which can be used for risk identification based on these changes. By comparing the target object's past behavior data (such as past transaction frequency, amount, type of business conducted, etc.), operations that do not conform to the regular pattern can be identified, and it can be determined whether the behavior is likely to be due to coercion or fraud in a state, achieving the technical effect of improving the safety of financial institution counter transactions.

[0075] Optionally, in the data processing method provided in the present application, the face feature data in the face image or face video is collected, including: collecting the face image or face video of the target object to obtain first multimedia data; intercepting the face image in the first multimedia data to obtain a target face image, and generating a texture feature image of the target face image; and performing feature extraction on the target face image and the texture feature image to obtain the face feature data.

[0076] In this embodiment, the target object can be notified that face information needs to be collected and the target object needs to be security verified when conducting business, and under the condition that the authorization of the target object is obtained, data containing face information can be obtained from real-time video streams captured by image collection devices or pre-prepared images. These data can be static face images or dynamic face videos, and first multimedia data can be obtained. The collected images or videos (i.e., first multimedia data) can also be preprocessed, such as adjusting brightness and contrast, to ensure the accuracy of subsequent processing steps. From the collected first multimedia data, the region containing the face can be identified and intercepted to generate a “target face image”. Subsequently, the target face image is further processed to generate its “texture feature image”, which helps more detailed feature extraction in subsequent steps.

[0077] For example, the detected face region is cropped from the original image and can be adjusted to a uniform size and format (e.g., 3x256x256) to meet the requirements of subsequent feature extraction. Through special algorithm processing, a "texture feature map" of the target face image is generated. The texture feature map can highlight the detailed features of the face, such as wrinkles, pores, and textures. It is very important for distinguishing real faces from fake objects (such as masks, photos), because the texture features of real faces are more natural and complex. Feature extraction can be performed on the "target face image" and "texture feature image" to capture unique facial features that can distinguish individual identities and biological features that can verify living bodies. Ensure that the extracted features can effectively distinguish the authenticity of the face and avoid fake attacks.

[0078] Optionally, in the data processing method provided in the embodiments of the present application, the target face image and the texture feature image are subjected to feature extraction to obtain face feature data, including: performing multi-modal feature fusion on the target face image and the texture feature image to obtain a fused feature image; and using a feature extraction module and a down-sampling module to perform feature extraction on the fused feature image to obtain the face feature data.

[0079] A texture feature map is generated from the target face image, which enhances the texture details of the face, such as the microstructure of the skin and pores, etc. These details have significant differences between real faces and fake objects, and are the key to living body detection.

[0080] The original target face image and the generated texture feature image are input into a multi-modal feature fusion module. The module can include multiple blocks (such as Block A and Block B), which cross-process and fuse the feature information of the two images through point-by-point convolution, channel-by-channel convolution, etc. Point-by-point convolution (Pointwise Convolution): can be called 1x1 convolution, which is used to adjust the depth of the feature map, i.e., change the number of feature channels without changing the spatial size. Depthwise separable convolution (Depthwise Separable Convolution): a more computationally efficient convolution method, which first performs depthwise convolution (each channel is independently convolved), and then performs point-by-point convolution to combine the information of each channel. Through the above operations, the information of the original image and the texture feature image can be integrated into a new "fused feature image", which contains richer and more discriminative facial features.

[0081] In the present embodiment, a feature extraction module and a down-sampling module can be used to perform feature extraction on the fused feature image.

[0082] The feature extraction module (e.g., Block A) can be used to further refine and enhance the facial features in the fused feature image. Feature extraction can be performed through multiple convolution layers, such as the residual modules in ResNet (Residual Network). The Block A module modifies the number of channels of the feature vector through point-by-point, channel-by-channel, and then point-by-point convolution, but keeps the image size unchanged, which helps to capture higher-level feature representations.

[0083] The down-sampling module (e.g., Block B) can reduce the size of the feature map while increasing the depth (number of channels) of the features. During the down-sampling process, channel-by-channel convolution with a down-sampling step can be used, which helps to focus on more local features while reducing the computational burden. The Block B module can also pass the feature map through average pooling and point-by-point convolution, and then add it to the feature vector of the main path to generate the down-sampled feature map.

[0084] Through the processing of the feature extraction module and the down-sampling module, the information of the "fused feature image" can be further refined, and the final "face feature data" contains enough details to support identity verification and liveness detection, while also optimizing the speed and efficiency of data processing.

[0085] Optionally, in the data processing method provided in the embodiments of the present application, the target recognition model is obtained by: obtaining multimedia data of M real faces and multimedia data of N false faces, obtaining S second multimedia data, wherein M and N are positive integers, and S = N + M; extracting the face features in each of the S second multimedia data to obtain S face feature images; generating an image label for each of the S face feature images to obtain S target image labels, wherein the image label for each face feature image includes image information in the face feature image; performing binaryzation processing on each of the S target image labels to obtain S binary mask labels; and performing model training on an initial recognition model based on the S binary mask labels, and determining the initial recognition model after the training is completed as the target recognition model.

[0086] In this embodiment, the multimedia data described above can include images and videos of human faces. In this embodiment, in the case of obtaining the authorized use of M users, multimedia data of M real human faces and multimedia data of N false human faces can be acquired to obtain S second multimedia data. Then, a face image can be cropped from the S second multimedia data, and a texture feature image of the face image can be generated. Each face feature image and the corresponding texture feature image are fused to obtain S face feature images. Then, image information of each face feature image in the S face feature images can be extracted. The image information in each face feature image constitutes a target image label corresponding to the face feature image. Each target image label in the S target image labels is binarized to obtain S binarization mask labels. Finally, the initial recognition model can be trained using the S binarization mask labels, and the initial recognition model after the training is completed is determined as the target recognition model.

[0087] For example, a texture feature image of each face feature image can be generated. Through multi-modal fusion of the original image (face feature image) and the texture feature image, the robustness of the living body detection algorithm can be improved, so that it can better cope with different attack methods and interference factors such as masks, wigs, changes in illumination, occlusions, etc. Among them, the original image or video can use a cross-platform computer vision and machine learning software library to detect a face, and the face is cropped and resized to 3x256x256. The feature texture graph module can better distinguish real faces from fraudulent items, thereby improving the accuracy of detection. The fusion of the original image and the feature texture graph can provide more image information effective for living body detection, so that the detection algorithm can better understand the morphology and structure of the face, further improving the detection accuracy. The intelligent anti-counterfeiting algorithm can judge whether the target object is a living body through a deep learning model combined with multi-modal information (such as face recognition, voiceprint recognition, fingerprint recognition, etc.), to avoid bypassing verification using photos, fake fingerprints, etc. The model combined with multi-modal information includes:

[0088] (1) Blink detection: detecting whether the user has a natural eye blink to prevent photos and videos from replacing the face.

[0089] (2) 3D depth perception: using a depth camera or a multi-camera system to obtain three-dimensional information of the face to determine whether the face is a real three-dimensional form.

[0090] (3) Facial expression analysis: analyzing the natural expression changes of the user, such as the dynamic changes of the lips, eyebrows, etc., to identify fake static photos or videos.

[0091] Figure 4 is a flowchart of image processing provided by an embodiment of the present application, such as Figure 4As shown, in the present embodiment, a lightweight convolution algorithm can be used for feature extraction, wherein the shortcut connection of the residual module in ResNet is used, namely Block A and Block B. The network constructed using such lightweight convolution module is easier to train and effectively prevents model degradation. Block A can be referred to as a feature extraction module. The input feature vector is sequentially subjected to point-by-point, channel-by-channel and point-by-point convolution and then added to the original input feature. The feature vector passing through this module can change the number of channels but not the size. Block B can be referred to as a down-sampling module. There are two branches. One input is sequentially subjected to point-by-point, channel-by-channel and point-by-point convolution, wherein the step length of the channel-by-channel convolution is 2. The other input is subjected to average pooling with a step length of 2 and point-by-point convolution. Finally, the outputs are added. This module will change the size of the features. Through the feature extraction of Block A and Block B, detailed data of image and texture features are obtained, and the image size is modified to a depth image and a binary image.

[0092] For real face images or videos and false face images or videos, respective depth image labels (corresponding to target image labels) are generated through the above feature extraction method, and then binary mask labels are generated by binarizing the depth image labels. All binary mask labels can be divided into a training set and a test set in a ratio of 8:2, which are respectively used for training and test evaluation of an FSN network (a kind of deep learning model). In this way, a target recognition model is obtained. During training, the FSN network can adopt an end-to-end manner.

[0093] Optionally, in the data processing method provided in the embodiments of the present application, the initial recognition model is trained based on the S binary mask labels, including: training the initial recognition model based on the S binary mask labels and a target loss function, and adjusting the hyperparameters in the target loss function during the model training process, wherein the target loss function includes at least one of the following: an amplitude loss function, a frequency spectrum loss function and a repeatability loss function. The amplitude loss function is used to obtain the similar part between the data predicted by the initial recognition model and the actually input data. The frequency spectrum loss is used to obtain the difference between the data predicted by the initial recognition model and the actually input data. The repeatability loss function is used to amplify the noise between the data predicted by the initial recognition model and the actually input data.

[0094] In the present embodiment, the binary mask label can be used to construct a loss function L f using a binary classification detection technology, and the specific formula is:

[0095] L f = λ1L m + λ2L s + λ3L r

[0096]

[0097] Wherein, λ1, λ2, λ3 are the loss weights of each loss function, which are adjusted as hyperparameters in model training.

[0098] Amplitude loss function L m : Hidden information F contained in the image in addition to image information, define F as the noise information of the image itself, through the loss function, the image noise part with the smallest amplitude can be obtained, and the most similar part between the model prediction data and the actual data is obtained.

[0099] Spectrum loss function L s : According to the low frequency content in the image field extracted, the low frequency information is minimized in the training so as to retain more high frequency features. The frequency domain conversion of the image can be realized by fast Fourier transform, the two-dimensional Fourier spectrum information f of the input image is obtained, the difference between the prediction data and the actual data of the model is obtained, and the error is minimized.

[0100] Repetitive loss function L r : The distribution of noise feature information on the image has strong repeatability, such features produce a larger amplitude in the high frequency band of the frequency domain image, based on this, in order to better amplify the image noise features between the prediction data and the actual data, the repetitive loss function can be used to maximize the high frequency information in the frequency domain image to encourage such repeatability, The function obtains the difference between the prediction distribution and the actual distribution, ξ(F) is the prediction information, and f is the actual information.

[0101] All binary mask labels are divided into training set and test set according to the preset proportion (for example, 8:2), which are used for training and testing evaluation of FSN network respectively. The FSN network adopts an end-to-end mode during training. The weight size in the loss function L f Can be set to λ1=0.05, λ2=0.001, λ3=0.1, so that the orders of magnitude of each component loss function are similar and comparable in size. Among them, the low-pass filter in the spectrum loss function L s And the high-pass filter in the repetitive loss function L r The frequency domain region size parameter f in the high-pass filter can be set to 50, according to the loss function, the difference between the real face image or video and the false face video is obtained, according to the threshold value, whether it is a living body is judged, and the technical effect of improving the reliability of the security verification of the target object in the business process of the financial institution is realized.

[0102] Optionally, in the data processing method provided in the embodiments of the present application, after the static feature data and the dynamic feature data are subjected to anomaly detection to obtain a target detection result, the method further includes: in a case where the target detection result indicates that the target object has an abnormal behavior, generating an early warning prompt information, and performing risk assessment on a financial service handled by the target object to obtain a risk assessment result, wherein the risk assessment result is used to indicate whether the target object is in a threatened state.

[0103] After the living body detection and the abnormal behavior recognition, once it is recognized that the target object has an abnormal behavior, an early warning information can be generated to notify the financial institution staff or the security system, so as to ensure that timely intervention can be ensured. The early warning mechanism can also be automatically triggered according to the abnormal detection result, and the early warning prompt containing the type of abnormal behavior, the detection time, the possible threat level and the like is generated. The early warning information can be sent in real time to the risk control department of the financial institution or the counter staff through an internal network or a special interface, so as to remind them to pay attention to the situation of a specific customer.

[0104] In the embodiments, the abnormal detection result and the service information can also be combined, and a preset risk assessment model is used to judge the risk level of the service operation. The risk assessment model can be based on historical data, pre-defined rules or machine learning algorithms, and can identify the correlation between abnormal behavior and high-risk transactions. The obtained risk assessment result will indicate whether the target object is likely to be in a threatened state and the severity of the threat. For example, if a customer is labeled as having an abnormal behavior and then performs a large amount of money transfer operation, it can be determined that he or she is likely to be in a state of coercion, and thus a high risk assessment result is given.

[0105] Based on the risk assessment result, appropriate security measures can be taken, such as stopping the transaction, requiring additional identity verification (such as secondary password input or telephone confirmation), notifying the security personnel of the financial institution to the scene or directly contacting the law enforcement agencies. The results of abnormal behavior and risk assessment will be recorded as the basis for improving the biological recognition algorithm and the threat detection strategy, so as to improve the overall security of the financial service.

[0106] For example, by analyzing the multi-dimensional data such as the biological recognition features, behavior patterns, voice emotions and the like of the customer in real time, the risk level of the target object can be assessed, and the background risk control system is linked, so that protection measures can be taken quickly when the customer is coerced. If it is detected that the biological features of an object are abnormal (such as abnormal tone when communicating with the bank staff), the risk assessment can be automatically started, and the predetermined response measures (such as direct alarm, notification of security personnel or automatic video recording, etc.) can be triggered according to the risk level.

[0107] In this embodiment, biometric technologies such as fingerprint, iris, and face recognition can be used to accurately identify the user's identity, greatly reducing the risk of being cracked or stolen compared to traditional password, card, and other authentication methods. Anti-forgery biometric algorithms can detect the authenticity of biometric features, identify fake masks, fake fingerprints, and other forgery methods, protect user account security, and reduce losses. Threat detection technology can monitor transactions and system activities in real time, quickly identify potential threats and issue warnings, and prevent malicious operations in a timely manner. Automated biometric identification and threat detection speed up business process handling, such as quickly verifying identity and risk during loan approval, improving business processing speed. The use of advanced technologies to ensure customer information security and transaction security helps financial institutions meet regulatory requirements and avoid regulatory risks.

[0108] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.

[0109] Embodiment Two

[0110] The embodiment of the present application also provides a data processing device. It should be noted that the data processing device of the embodiment of the present application can be used to execute the data processing method provided by the embodiment of the present application. The data processing device provided by the embodiment of the present application is introduced as follows.

[0111] According to the embodiment of the present application, a device for implementing the above-mentioned data processing method is also provided, as shown in Figure 5 The device comprises a processing unit 51, an acquisition unit 52, and a detection unit 53.

[0112] The processing unit 51 is configured to acquire facial feature data in a facial image or a facial video when a financial service is handled, and input the facial feature data into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether a target object in the facial image or the facial video is an active object, and the target recognition model comprises a neural network model used for active object detection.

[0113] The acquisition unit 52 is configured to acquire static biometric features of the target object to obtain static feature data and acquire dynamic behavior features of the target object to obtain dynamic feature data in a case where the recognition result indicates that the target object in the facial image or the facial video is an active object.

[0114] The detection unit 53 is configured to perform anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has an abnormal behavior.

[0115] In the data processing apparatus provided in the embodiments of the present application, the face feature data in the face image or the face video can be collected by the processing unit 51 when the financial service is handled, and the face feature data is input into the target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether the target object in the face image or the face video is an active object, and the target recognition model includes a neural network model used for active object detection. The collection unit 52 is used to collect the static biological features of the target object to obtain static feature data and collect the dynamic behavior features of the target object to obtain dynamic feature data in the case where the recognition result indicates that the target object in the face image or the face video is an active object. The detection unit 53 performs abnormality detection on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has abnormal behavior. Thus, the technical problem of low security of security verification in the process of handling the financial service in the related art is solved. In the embodiments, whether the target object is an active object is detected by the target recognition model, and the target object is abnormally detected by the static feature data and the dynamic feature data, which avoids the case that the security is low because only the biological recognition technology is used to verify the identity of the user in the related art, thereby achieving the technical effect of improving the reliability of security verification of the financial service.

[0116] Optionally, in the data processing apparatus provided in the embodiments of the present application, the static feature data includes at least one of the following: the fingerprint features of the target object, the iris features of the target object, and the face features of the target object, and the dynamic feature data includes at least one of the following: the voiceprint features of the target object, the voice content of the target object, and the change of the facial expression of the target object.

[0117] Optionally, in the data processing apparatus provided in the embodiments of the present application, the detection unit includes: a first detection subunit, configured to detect whether the static feature data is fake, to obtain a first detection result; a second detection subunit, configured to detect whether the dynamic feature data has abnormal fluctuation, to obtain a second detection result; a processing subunit, configured to input the static feature data and the dynamic feature data into a target prediction model to predict the behavior features of the target object in a target time period, to obtain predicted behavior features, and to determine a third detection result based on the difference between the predicted behavior features and the actual behavior features of the target object in the target time period, wherein the target prediction model includes a model-trained autoregressive moving average model; and a determination subunit, configured to determine the target detection result based on the first detection result, the second detection result, and the third detection result.

[0118] Optionally, in the data processing apparatus provided in the embodiments of the present application, the processing unit comprises: an acquisition subunit, configured to acquire a face image or a face video of a target object to obtain first multimedia data; a clipping subunit, configured to clip a face image in the first multimedia data to obtain a target face image, and generate a texture feature image of the target face image; and an extraction subunit, configured to perform feature extraction on the target face image and the texture feature image to obtain face feature data.

[0119] Optionally, in the data processing apparatus provided in the embodiments of the present application, the extraction subunit comprises: a fusion module, configured to perform multi-modal feature fusion on the target face image and the texture feature image to obtain a fusion feature image; and an extraction module, configured to perform feature extraction on the fusion feature image by using a feature extraction module and a down-sampling module to obtain the face feature data.

[0120] Optionally, in the data processing apparatus provided in the embodiments of the present application, the target recognition model is obtained by the following subunits: an acquisition subunit, configured to acquire multimedia data of M real faces and multimedia data of N false faces to obtain S second multimedia data, wherein M and N are positive integers, and S = N + M; a first processing subunit, configured to extract face features in each of the S second multimedia data to obtain S face feature images; a generation subunit, configured to generate an image label of each of the S face feature images to obtain S target image labels, wherein the image label of each face feature image comprises image information in the face feature image; a second processing subunit, configured to perform binaryzation processing on each of the S target image labels to obtain S binaryzation mask labels; and a training subunit, configured to perform model training on an initial recognition model based on the S binaryzation mask labels, and determine the initial recognition model after the training as the target recognition model.

[0121] Optionally, in the data processing apparatus provided in the embodiments of the present application, the training subunit comprises: a training module, configured to perform model training on the initial recognition model based on the S binaryzation mask labels and a target loss function, and adjust a hyperparameter in the target loss function in the model training process, wherein the target loss function comprises at least one of the following: an amplitude loss function, a frequency spectrum loss function, and a repetitiveness loss function, the amplitude loss function is used to obtain a similar part between data predicted by the initial recognition model and actually input data, the frequency spectrum loss is used to obtain a difference between the data predicted by the initial recognition model and the actually input data, and the repetitiveness loss function is used to amplify noise between the data predicted by the initial recognition model and the actually input data.

[0122] Optionally, in the data processing device provided in the embodiment of the present application, the data processing device also includes: an evaluation unit, which is used to perform anomaly detection on static feature data and dynamic feature data to obtain a target detection result, and further includes: when the target detection result indicates that the target object has abnormal behavior, generating early warning prompt information, and performing a risk assessment on the financial business handled by the target object to obtain a risk assessment result, wherein the risk assessment result is used to indicate whether the target object is in a threatened state.

[0123] It should be noted that the processing unit 51, acquisition unit 52, and detection unit 53 described above correspond to steps S201 to S203 in the first embodiment. The examples and application scenarios implemented by each unit and the corresponding steps are the same, but are not limited to the contents disclosed in the first embodiment. It should be noted that the modules or units described above can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be run as part of a device in the computer terminal 10 provided in the first embodiment.

[0124] Example 3

[0125] An embodiment of the present application may provide an electronic device, Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 Only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0126] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0127] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: collecting face feature data in a face image or a face video when a financial service is handled, and inputting the face feature data into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether a target object in the face image or the face video is an active object, and the target recognition model includes a neural network model used for active object detection; in a case where the recognition result indicates that the target object in the face image or the face video is an active object, collecting a static biological feature of the target object to obtain static feature data, and collecting a dynamic behavior feature of the target object to obtain dynamic feature data; performing abnormality detection on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has an abnormal behavior.

[0128] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: further, the static feature data includes at least one of the following: fingerprint features of the target object, iris features of the target object, and face features of the target object, and the dynamic feature data includes at least one of the following: voiceprint features of the target object, voice content of the target object, and change conditions of facial expressions of the target object.

[0129] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: performing abnormality detection on the static feature data and the dynamic feature data to obtain a target detection result, including: detecting whether the static feature data is fake to obtain a first detection result; detecting whether the dynamic feature data has abnormal changes to obtain a second detection result; inputting the static feature data and the dynamic feature data into a target prediction model to predict behavior features of the target object in a target time period to obtain predicted behavior features, and determining a third detection result based on a difference between the predicted behavior features and actual behavior features of the target object in the target time period, wherein the target prediction model includes a model-trained autoregressive moving average model; and determining the target detection result based on the first detection result, the second detection result, and the third detection result.

[0130] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: collecting face feature data in a face image or a face video, including: collecting a face image or a face video of a target object to obtain first multimedia data; intercepting a face image in the first multimedia data to obtain a target face image, and generating a texture feature image of the target face image; and performing feature extraction on the target face image and the texture feature image to obtain the face feature data.

[0131] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: feature extraction is performed on the target face image and the texture feature image to obtain face feature data, including: multi-modal feature fusion is performed on the target face image and the texture feature image to obtain a fused feature image; and the fused feature image is subjected to feature extraction by using a feature extraction module and a down-sampling module to obtain the face feature data.

[0132] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: the target recognition model is obtained by: obtaining multimedia data of M real faces and multimedia data of N false faces to obtain S second multimedia data, wherein M and N are positive integers, and S = N + M; extracting face features in each of the S second multimedia data to obtain S face feature images; generating an image label of each of the S face feature images to obtain S target image labels, wherein the image label of each face feature image includes: image information in the face feature image; performing binaryzation processing on each of the S target image labels to obtain S binary mask labels; and performing model training on an initial recognition model based on the S binary mask labels, and determining the initial recognition model after the training as the target recognition model.

[0133] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: the model training on the initial recognition model based on the S binary mask labels, including: performing model training on the initial recognition model based on the S binary mask labels and a target loss function, and adjusting hyperparameters in the target loss function during the model training, wherein the target loss function includes at least one of: an amplitude loss function, a frequency spectrum loss function, and a repeatability loss function, the amplitude loss function is used to obtain similar parts between data predicted by the initial recognition model and actually input data, the frequency spectrum loss is used to obtain differences between the data predicted by the initial recognition model and the actually input data, and the repeatability loss function is used to amplify noise between the data predicted by the initial recognition model and the actually input data.

[0134] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: after the abnormality detection on the static feature data and the dynamic feature data to obtain a target detection result, further including: in a case where the target detection result indicates that the target object has an abnormal behavior, generating a warning prompt information, and performing risk assessment on a financial service handled by the target object to obtain a risk assessment result, wherein the risk assessment result is used to indicate whether the target object is in a threatened state.

[0135] By adopting the embodiment of the application, whether the target object is an active object is detected through the target identification model, and the target object is abnormally detected through the static feature data and the dynamic feature data, thereby avoiding the situation that the security is low in the related art that only the identity of the user is verified through the biometric recognition technology, and thereby the technical effect of improving the reliability of the security verification of the financial service is achieved.

[0136] Those skilled in the art can understand that, Figure 6 The structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, and the like. Figure 6 It does not limit the structure of the electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, and the like) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 6 Figure 6 The structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, and the like.

[0137] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by a program instructing the related hardware of the terminal device, and the program can be stored in a computer readable storage medium, and the storage medium can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, and the like.

[0138] Embodiment Four

[0139] The embodiment of the application further provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the data processing method provided in the embodiment one.

[0140] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0141] The application further provides a computer program product, which, when executed on a data processing device, is adapted to execute the steps of the data processing method.

[0142] The serial numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0143] In the above embodiments of the application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0144] ​In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0145] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0146] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0147] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0148] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A data processing method, characterized in that: include: When handling financial business, facial feature data from a facial image or facial video is collected and input into a target recognition model to obtain a recognition result, wherein the recognition result is used to indicate whether the target object in the facial image or facial video is an active object, and the target recognition model includes: a neural network model for active object detection; If the recognition result indicates that the target object in the face image or face video is an active object, collecting static biometric features of the target object to obtain static feature data, and collecting dynamic behavioral features of the target object to obtain dynamic feature data; Anomaly detection is performed on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has abnormal behavior.

2. The data processing method according to claim 1, wherein: The static feature data includes at least one of the following: the fingerprint feature of the target object, the iris feature of the target object, and the facial feature of the target object; the dynamic feature data includes at least one of the following: the voiceprint feature of the target object, the speech content of the target object, and the changes in the facial expression of the target object.

3. The data processing method according to claim 1, wherein: Performing anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result includes: Detecting whether the static feature data is forged, and obtaining a first detection result; detecting whether the dynamic feature data has abnormal changes, and obtaining a second detection result; Inputting the static feature data and the dynamic feature data into a target prediction model, predicting the behavior characteristics of the target object in a target time period to obtain a predicted behavior characteristic, and determining a third detection result based on a difference between the predicted behavior characteristic and the actual behavior characteristics of the target object in the target time period, wherein the target prediction model includes: an autoregressive moving average model that has been trained by a model; The target detection result is determined based on the first detection result, the second detection result, and the third detection result.

4. The data processing method according to claim 1, wherein: Collect facial feature data from facial images or facial videos, including: Capturing a face image or face video of a target object to obtain first multimedia data; intercepting a facial image in the first multimedia data to obtain a target facial image, and generating a texture feature image of the target facial image; Feature extraction is performed on the target face image and the texture feature image to obtain the face feature data.

5. The data processing method according to claim 4, characterized in that: Performing feature extraction on the target face image and the texture feature image to obtain the face feature data includes: Performing multimodal feature fusion on the target face image and the texture feature image to obtain a fused feature image; A feature extraction module and a downsampling module are used to extract features from the fused feature image to obtain the facial feature data.

6. The data processing method according to claim 1, wherein: The target recognition model is obtained by: Acquire multimedia data of M real faces and multimedia data of N fake faces to obtain S second multimedia data, where M and N are positive integers and S = N + M; Extracting facial features from each of the S pieces of the second multimedia data to obtain S facial feature images; Generate an image label for each of the S facial feature images to obtain S target image labels, wherein the image label of each of the facial feature images includes: image information in the facial feature image; Binarizing each of the S target image labels to obtain S binary mask labels; The initial recognition model is trained based on the S binary mask labels, and the initial recognition model after training is determined as the target recognition model.

7. The data processing method according to claim 6, characterized in that: Performing model training on the initial recognition model based on the S binary mask labels includes: The initial recognition model is trained based on S of the binary mask labels and the target loss function, and the hyperparameters in the target loss function are adjusted during the model training process, wherein the target loss function includes at least one of the following: an amplitude loss function, a spectrum loss function, and a repeatability loss function. The amplitude loss function is used to obtain the similarity between the data predicted by the initial recognition model and the actual input data, the spectrum loss is used to obtain the difference between the data predicted by the initial recognition model and the actual input data, and the repeatability loss function is used to amplify the noise between the data predicted by the initial recognition model and the actual input data.

8. The data processing method according to claim 1, wherein: After performing anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result, the method further includes: When the target detection result indicates that the target object has abnormal behavior, an early warning prompt information is generated, and a risk assessment is performed on the financial business handled by the target object to obtain a risk assessment result, wherein the risk assessment result is used to indicate whether the target object is in a threatened state.

9. A data processing device, characterized in that: include: a processing unit configured to collect facial feature data from a facial image or facial video when handling financial transactions, and input the facial feature data into a target recognition model to obtain a recognition result, wherein the recognition result indicates whether a target object in the facial image or facial video is an active object, the target recognition model comprising: a neural network model for active object detection; a collection unit configured to, when the recognition result indicates that the target object in the face image or face video is an active object, collect static biometric features of the target object to obtain static feature data, and collect dynamic behavioral features of the target object to obtain dynamic feature data; The detection unit is used to perform anomaly detection on the static feature data and the dynamic feature data to obtain a target detection result, wherein the target detection result is used to indicate whether the target object has abnormal behavior.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Alarm operation control method and system based on face authentication

    CN122027341A