Internet enterprise multi-mode identity verification method and system

By using a dynamic risk-driven multimodal verification path and cross-modal biometric association decision-making, the problem of the inability to dynamically adjust verification strength in remote identity verification for Internet companies is solved, achieving a balance between security and user experience, and enhancing the defense against highly realistic attacks and the system's adaptability.

CN121125206APending Publication Date: 2025-12-12NAN JING OU YI TAI XIN XI KE JI YOU XIAN GONG SI

Patent Information

Application Number
CN202511224165.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing remote identity verification technologies used by internet companies cannot dynamically adjust verification intensity according to the real-time transaction environment. They suffer from operational redundancy in low-risk scenarios or insufficient defense in high-risk scenarios. Multimodal biometric verification lacks spatiotemporal correlation analysis, making it vulnerable to highly realistic synthetic attacks. Furthermore, their defense strategies are rigid and cannot evolve autonomously.

Method used

Employing a dynamic risk-driven multimodal verification path generation, cross-modal biometric association decision-making, and incremental learning mechanism, this system generates multidimensional risk feature vectors by acquiring user historical behavior, current transaction, and device environment data. It then analyzes these data in real time and generates personalized verification sequence instructions. Furthermore, it collects multimodal biometric packages for cross-modal association comparison and combines this with document information verification to achieve feature-level association analysis to generate verification decisions.

Benefits of technology

It achieves dynamic verification strength adjustment that adapts to risk, improves the security of identity verification and user experience, effectively resists deepfakes and liveness attacks, enhances the defense against new fraud patterns, and forms a closed-loop evolutionary fraud defense system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125206A_ABST
    Figure CN121125206A_ABST
Patent Text Reader

Abstract

The invention discloses an Internet enterprise multi-mode identity verification method and system, and belongs to the technical field of Internet enterprise security, and the method comprises the steps: obtaining user historical behavior data, current transaction request data, equipment environment parameters and initial biological signal data, carrying out the risk assessment, and generating a verification path; obtaining a personalized verification sequence instruction, collecting a user face dynamic video stream, a real-time voice stream and response action time sequence data, carrying out cross-modal association comparison with a user reference biological feature template, outputting a biological feature confidence matrix, carrying out association analysis in combination with the obtained structured identity feature vector and risk assessment, and obtaining a personalized verification result; and generating a verification decision feature vector to judge a verification result state, and obtaining a pass instruction, a rejection instruction or a manual auditing request instruction. According to the method, dynamic risk-driven multi-modal verification path generation, cross-modal biological feature association decision and incremental learning mechanisms are adopted, so that the optimal balance between security and user experience can be realized in a complex network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of internet enterprise security technology, and in particular to a method and system for multimodal identity verification for internet enterprises. Background Technology

[0002] The rapid development of internet companies' businesses places higher demands on remote identity verification technology, requiring the verification of user identity authenticity in scenarios without physical contact. Existing technologies typically rely on a combination of biometric recognition and document information verification.

[0003] Current mainstream solutions employ static threshold control for the verification process, such as pre-setting fixed rules to trigger face or voiceprint verification, and combining optical character recognition technology to extract document information for comparison. Some solutions introduce multimodal biometric data collection, but each modality's data is processed independently, and the verification result is determined through weighted fusion.

[0004] However, existing technologies are slow to respond to risks and cannot dynamically adjust the verification intensity according to the real-time trading environment, resulting in redundant operations in low-risk scenarios or insufficient defense in high-risk scenarios; multimodal biometric verification is carried out in isolation and lacks spatiotemporal correlation analysis, making it vulnerable to highly realistic synthetic attacks; defense strategies are rigid and cannot evolve autonomously based on new fraud patterns, relying on manual rule updates. Summary of the Invention

[0005] To address the aforementioned issues, this invention provides a multimodal identity verification method and system for internet enterprises. It employs dynamic risk-driven multimodal verification path generation, cross-modal biometric association decision-making, and incremental learning mechanisms, enabling the optimal balance between security and user experience in complex network environments.

[0006] The above objectives can be achieved through the following approach:

[0007] A multimodal identity verification method for internet companies includes: acquiring user historical behavior data, current transaction request data, device environmental parameters, and initial biometric data to generate a multidimensional risk feature vector; performing real-time analysis on the multidimensional risk feature vector to output an initial risk score; generating a verification path based on the initial risk score to obtain a personalized verification sequence instruction; executing the personalized verification sequence instruction to collect user facial dynamic video stream, real-time voice stream, and response action timing data to generate a multimodal biometric package; performing cross-modal correlation comparison between the multimodal biometric package and a pre-stored user baseline biometric template to output a biometric confidence matrix; acquiring user-submitted document image data, performing optical character recognition to extract document information fields, and generating a structured identity feature vector; performing feature-level correlation analysis on the biometric confidence matrix, the structured identity feature vector, and the initial risk score to generate a verification decision feature vector; and determining the verification result status based on the verification decision feature vector, generating an approval instruction, a rejection instruction, or a manual review request instruction.

[0008] Optionally, the personalized verification sequence instruction includes: when the initial risk score is lower than the first threshold, only the face fast comparison instruction is activated; when the initial risk score is between the first threshold and the second threshold, the face-voiceprint two-factor verification instruction is activated; when the initial risk score is higher than the second threshold, a dynamic action challenge instruction and a document information verification instruction are added to the face-voiceprint two-factor verification instruction.

[0009] Optionally, the method further includes: acquiring historical fraud case data packets and extracting the instruction response parameter set of attackers circumventing verification to establish an attack feature profile library, wherein the instruction response parameter set includes biometric simulation delay time, abnormal fluctuation range of sensor data, and action execution trajectory deviation angle parameters; calculating an adjustment coefficient based on the attack feature profile library and the initial risk score; using the adjustment coefficient to correct the initial risk score and update the personalized verification sequence instruction; and inserting a document information verification instruction when the current device environment parameters are detected to match the characteristics of a high-risk area.

[0010] Optionally, the document information verification instruction includes: comparing the consistency between the self-claimed name in the user's voiceprint feature segment and the name identified on the document in real time, and calculating a consistency score; detecting the optical response features of the document's anti-counterfeiting mark area under specific lighting conditions; when an abnormality is detected in the optical response features or the consistency score is lower than a preset score threshold, acquiring a fluorescence reaction image and extracting the proportion of the fluorescence area to perform similarity matching with a pre-stored template to obtain a verification matching result.

[0011] Optionally, the generation of the multimodal biometric package includes: based on the face fast comparison instruction, detecting the frequency of micro-expression changes and pupil reflection features in the facial dynamic video stream to generate a liveness detection flag; based on the face-voiceprint dual-factor verification instruction, sending a voice reading instruction containing a random number sequence to the user terminal and performing voiceprint matching to obtain an anti-spoofing consistency index and a voiceprint confidence score; based on the dynamic action challenge instruction, sending a specified head movement trajectory combination instruction to the user terminal to obtain a head posture change sequence; and timestamping and encapsulating the anti-spoofing consistency index, the voiceprint confidence score, and the liveness detection flag, and combining them with the verification matching result and the head posture change sequence to form a multimodal biometric package.

[0012] Optionally, obtaining the anti-counterfeiting consistency index and voiceprint confidence score includes: capturing lip movement features and speech spectrum features during the user's execution of the voice reading instruction; performing temporal alignment analysis on the lip movement features and the speech spectrum features to generate the anti-counterfeiting consistency index; extracting voiceprint feature segments from the speech spectrum features; performing similarity matching between the voiceprint feature segments and pre-stored voiceprint templates to generate the voiceprint confidence score.

[0013] Optionally, the generation of the verification decision feature vector includes: extracting a name field from the structured identity feature vector and extracting a self-proclaimed name from the voiceprint feature fragment; performing text similarity matching between the name field and the self-proclaimed name to obtain a matching value; and activating a high-risk handling sub-process when the matching value or the biometric confidence matrix is ​​detected to be lower than a preset threshold.

[0014] Optionally, the high-risk handling sub-process includes: sending a hand swiping command to the user terminal; capturing touch screen pressure sensing data and abnormal vibration characteristics of the gyroscope based on the hand swiping command to obtain gesture trajectory feature data; and performing spatiotemporal correlation analysis on the gesture trajectory feature data and the multimodal biometric package to generate a verification decision feature vector.

[0015] Optionally, the method further includes: extracting and parsing fraud case data packets that have been manually reviewed and confirmed to obtain biometric anomaly vectors; labeling the multidimensional risk feature vectors corresponding to the fraud case data packets based on the biometric anomaly vectors to obtain a risk labeling dataset; and updating the attack feature profile library using the risk labeling dataset.

[0016] Based on the same inventive concept, this invention also provides a multimodal identity verification system for internet enterprises. The system includes: a data acquisition module for acquiring user historical behavior data, current transaction request data, device environmental parameters, and initial biosignal data to generate a multidimensional risk feature vector; a risk perception module for real-time analysis of the multidimensional risk feature vector and outputting an initial risk score; a verification path generation module for generating a verification path based on the initial risk score to obtain personalized verification sequence instructions; and a multimodal acquisition module for executing the personalized verification sequence instructions, acquiring user facial dynamic video streams, real-time voice streams, and response action timing data to generate multimodal biosignal data. The system comprises: a multimodal biometric package; a modal comparison module, used to perform cross-modal correlation comparison between the multimodal biometric package and a pre-stored user baseline biometric template, and output a biometric confidence matrix; a structured data acquisition module, used to acquire user-submitted document image data, perform optical character recognition to extract document information fields, and generate a structured identity feature vector; an association decision module, used to perform feature-level correlation analysis on the biometric confidence matrix, the structured identity feature vector, and the initial risk score, and generate a verification decision feature vector; and a decision instruction generation module, used to determine the verification result status based on the verification decision feature vector, and generate an approval instruction, a rejection instruction, or a manual review request instruction.

[0017] Compared with the prior art, the present invention has the following advantages:

[0018] 1. This invention achieves adaptive dynamic verification intensity adjustment based on risk. Personalized verification sequences are generated through an initial risk score. In low-risk scenarios, the process is simplified to improve efficiency, while in high-risk scenarios, a multi-modal verification layer is overlaid, optimizing user experience while ensuring security.

[0019] 2. This invention overcomes the limitations of isolated verification of multimodal biometrics. It employs cross-modal correlation comparison technology to fuse facial dynamic video, real-time voice stream, and action timing data to construct a spatiotemporally synchronized biometric evidence chain, effectively resisting deepfakes and liveness attacks.

[0020] 3. This invention establishes a collaborative decision-making mechanism between structured identity and biometric behavior. It performs feature-level correlation analysis on document information fields, biometric confidence matrices, and risk scores to eliminate information silos and improve the accuracy of identifying identity theft.

[0021] 4. This invention forms a closed-loop, evolving fraud defense system. Fraud cases verified through manual review are used to feed back into the attack signature database, dynamically updating the risk model and enabling the system to continuously adapt to new attack methods.

[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating the networked enterprise multimodal identity verification method according to an embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram of cross-modal (lip movement-speech) temporal correlation analysis according to an embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram of the structure of the networked enterprise multimodal identity verification system according to an embodiment of the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Reference Figure 1 One embodiment of the present invention proposes a multimodal identity verification method for Internet enterprises, which adopts dynamic risk-driven multimodal verification path generation, cross-modal biometric association decision-making and incremental learning mechanism, and can achieve the optimal balance between security and user experience in complex network environments.

[0029] The multimodal identity verification method for internet enterprises described in this embodiment specifically includes:

[0030] Acquire user historical behavior data, current transaction request data, device environmental parameters, and initial biosignal data to generate a multidimensional risk feature vector;

[0031] The multidimensional risk feature vector is analyzed in real time, and an initial risk score is output.

[0032] Specifically, acquiring user historical behavior data includes extracting device fingerprint usage frequency distribution, historical transaction time distribution, and the number of abnormal operation markers from user behavior logs; current transaction request data is obtained by parsing real-time transaction messages to obtain transaction amount, transaction type code, and payee identification; device environment parameters are collected by calling the terminal system interface to obtain the International Mobile Equipment Identity (IMEI) string, GPS coordinates, and wireless network media access control address; initial biosignal data is captured by biosensors to obtain the mean grip pressure and standard deviation of heart rate variability. The above data is input into a feature standardization processor. For the device fingerprint usage frequency distribution, its information entropy is calculated as a feature value; for the transaction amount, maximum-minimum normalization is performed; for the GPS coordinates, the minimum Havesientic distance to the user's frequently used location set is calculated; and for the standard deviation of heart rate variability, Z-score standardization is used. The generated multidimensional risk feature vector is represented as V = [v1, v2, ..., v N ] T , where v N Let represent the standardized eigenvalue of the Nth dimension, where N is the total dimension of the features. When performing real-time analysis on a multidimensional risk feature vector, a pre-trained risk weight vector R = [r1, r2, ..., r] is loaded. N The initial risk score is calculated by linear weighted summation. Where r i Let be the weight coefficients of the i-th feature, obtained by training a logistic regression model on a historical fraud dataset and satisfying . v i These are the standardized values ​​for the corresponding feature dimensions. Ten-fold cross-validation is used during feature weight coefficient training to ensure generalization ability.

[0033] A verification path is generated based on the initial risk score, resulting in personalized verification sequence instructions.

[0034] Execute the personalized verification sequence instructions to collect the user's facial dynamic video stream, real-time voice stream, and response action timing data to generate a multimodal biometric package;

[0035] The multimodal biometric package is compared with the pre-stored user baseline biometric template across modalities, and a biometric confidence matrix is ​​output.

[0036] Specifically, when performing cross-modal association and comparison operations, the multimodal biometric package is first decomposed into synchronous data segments according to timestamps. The multimodal biometric package includes micro-expression frequency feature values ​​extracted from facial dynamic video streams, pupil reflection feature vectors, voiceprint feature segments extracted from real-time speech streams, anti-spoofing consistency indicators, and triaxial angle values ​​of head posture change sequences. The user baseline biometric template includes pre-stored baseline micro-expression frequency ranges, baseline pupil reflection vectors, baseline voiceprint template vectors, and baseline head posture sequences. For each time slice, the following steps are performed: the difference between the micro-expression frequency feature values ​​of the current time slice and the baseline micro-expression frequency range is calculated to obtain the expression matching degree; the cosine similarity between the pupil reflection feature vector and the baseline pupil reflection vector is calculated to obtain the pupil similarity; the voiceprint feature segments and the baseline voiceprint template vector are calculated using dynamic time-normalized distance calculation and then normalized to obtain the voiceprint confidence; the anti-spoofing consistency indicator is directly used as the lip-reading synchronization degree; the triaxial angle values ​​of the head posture change sequence are calculated using Euclidean distance calculation and then normalized to obtain the action matching degree. All five indicators are dimensionless values ​​ranging from zero to one. Weights are adjusted using a cross-modal coupling coefficient matrix, which is trained on a historical positive sample dataset and reflects the correlation strength between different modalities. Finally, a biometric confidence matrix is ​​generated.

[0037] Obtain the image data of the ID document submitted by the user, perform optical character recognition to extract the ID document information fields, and generate a structured identity feature vector;

[0038] The biometric confidence matrix, the structured identity feature vector, and the initial risk score are subjected to feature-level correlation analysis to generate a verification decision feature vector;

[0039] The verification result status is determined based on the verification decision feature vector, and an approval instruction, rejection instruction, or manual review request instruction is generated.

[0040] Specifically, based on a multi-source heterogeneous data fusion and dynamic risk assessment mechanism, this method constructs a multi-dimensional risk feature vector by acquiring user historical behavior, real-time transactions, device environment, and biometric signals to achieve initial risk quantification assessment. Personalized verification paths are generated based on risk scores, adaptively activating biometric collection modes of varying intensities. Cross-modal correlation comparison technology is employed to fuse dynamic facial video, real-time voice streams, and action time-series data, combined with structured identity information from optical character recognition. Through feature-level correlation analysis, a verification decision vector is constructed, ultimately achieving intelligent decision-making across three states. This method overcomes the limitations of traditional static verification, optimizing user experience through a risk-driven dynamic verification path. It simplifies processes and improves efficiency in low-risk scenarios, while enhancing multi-modal cross-verification strength in high-risk scenarios. Cross-modal biometric correlation comparison effectively resists deepfake attacks, and the collaborative analysis of structured identity and biometrics eliminates information silos. The overall approach forms a hierarchical defense system, ensuring enterprise transaction security while reducing the need for manual intervention and improving the system's adaptive protection capabilities.

[0041] Optionally, the instructions for generating personalized verification sequences include:

[0042] When the initial risk score is below the first threshold, only the fast face comparison command is activated;

[0043] When the initial risk score is between the first and second thresholds, activate the face-voiceprint two-factor verification command.

[0044] When the initial risk score is higher than the second threshold, a dynamic action challenge instruction and a document information verification instruction are added to the face-voiceprint dual-factor verification instruction.

[0045] Specifically, the system first receives an initial risk score output by real-time analysis of multidimensional risk feature vectors. This score is numerical data that represents a quantitative value of the user's trading risk. A first threshold and a second threshold are preset, where the first threshold is used to distinguish between low risk and medium risk, and the second threshold is used to distinguish between medium risk and high risk. These thresholds are preset based on historical trading data and fraud pattern analysis. The system compares the initial risk score with a first threshold and a second threshold. If the initial risk score is less than the first threshold, the system only activates the fast face comparison command, which triggers the device's camera to capture the user's facial image for rapid liveness detection and identity matching. If the initial risk score is between the first and second thresholds, the system activates the face-voiceprint dual-factor verification command, which simultaneously initiates facial video stream capture and voice command reading, requiring the user to read a random number sequence to simultaneously verify facial and voiceprint features. If the initial risk score is higher than the second threshold, the system, in addition to activating the face-voiceprint dual-factor verification command, adds a dynamic action challenge command and an ID verification command. The dynamic action challenge command captures the user's head movement trajectory through the device's sensors, while the ID verification command guides the user to upload an ID image for optical recognition and anti-counterfeiting detection. The entire process is based on the principle of risk adaptation, achieving a balance between security and user experience by dynamically adjusting the verification strength; simplifying the verification process in low-risk scenarios to improve user operation efficiency; introducing multi-factor verification in medium-risk scenarios to enhance the confirmation of identity authenticity; and superimposing action and document verification in high-risk scenarios to effectively resist complex fraud attacks, thereby improving the robustness and adaptability of the verification system as a whole.

[0046] Optionally, the method further includes:

[0047] Acquire historical fraud case data packages and extract the instruction response parameter set of attackers to circumvent verification to establish an attack feature profile library. The instruction response parameter set includes biometric simulation delay time, abnormal fluctuation range of sensor data, and deviation angle parameter of action execution trajectory.

[0048] Based on the attack feature profile database and the initial risk score, the adjustment coefficient is calculated;

[0049] The initial risk score is corrected using the adjustment coefficient, and the personalized verification sequence instruction is updated; when the current device environment parameters are detected to match the characteristics of a high-risk area, a document information verification instruction is inserted.

[0050] Specifically, the system first retrieves historical fraud case data packages from a pre-set fraud case database. These packages contain records of multiple confirmed fraud events. From these packages, it extracts a set of attack response parameters used by attackers to circumvent verification. These parameters include: biometric simulation delay time (the time difference between the attacker's response and the user's simulated biometric actions); abnormal fluctuation range of sensor data (the extent to which device sensor readings, such as accelerometers or gyroscopes, exceed normal thresholds); and action execution trajectory deviation angle parameters (the angle deviation between the actual trajectory and the standard trajectory when the user performs a specified action). Based on these parameter sets, an attack feature profile library is built, storing various attack feature vectors in a structured format. Next, the system combines this attack feature profile library with the initial risk score to calculate an adjustment coefficient. The adjustment coefficient equals the initial risk score multiplied by the attack feature similarity plus one. For the adjustment coefficient α, we have:

[0051] α=S risk Sim+1,

[0052] Where S risk The initial risk score is represented by the risk perception module's analysis of multi-dimensional risk feature vectors. Sim represents the similarity between the current trading parameters and the average features in the attack feature profile library. It is obtained by calculating the cosine similarity between the current biometric simulation delay time, the abnormal fluctuation range of sensor data, and the deviation angle of the action execution trajectory and the mean of the corresponding parameters in the library. This similarity value is normalized to between zero and one. An adjustment coefficient is used to correct the initial risk score. For the corrected risk score S... corrected ,have:

[0053] S corrected =S risk ·α,

[0054] The process involves regenerating personalized verification sequence instructions based on the revised risk score, employing the aforementioned threshold logic: if the revised risk score is below the first threshold, only the fast face comparison instruction is activated; if it is between the first and second thresholds, the face-voiceprint dual-factor verification instruction is activated; and if it is above the second threshold, dynamic action challenge instructions and document information verification instructions are added. Simultaneously, the system monitors device environmental parameters such as GPS coordinates or IP addresses in real time. When these parameters match high-risk area characteristics, such as known fraud-prone locations, a document information verification instruction is forcibly inserted. This instruction requires the user to upload a document image for optical recognition and anti-counterfeiting detection. The entire process is based on the principle of dynamic risk adaptation, learning from historical attack patterns and perceiving environmental risks to optimize verification strength in real time. It enhances the accuracy of risk scoring in fraud pattern recognition, avoiding misjudgments due to blind spots in historical attacks; proactively strengthens document verification during environmental risk detection, effectively intercepting location-related fraud; and comprehensively improves the system's defense capabilities against new attacks while maintaining the intelligent adaptability of the verification process, reducing user operation interruptions.

[0055] Optionally, the document information verification instruction includes:

[0056] The consistency score is calculated by comparing the user's self-proclaimed name in the voiceprint feature segment with the name identified on the ID card in real time.

[0057] The optical response characteristics of the anti-counterfeiting mark area on the document are detected under specific lighting conditions;

[0058] When an abnormality in the optical response characteristics is detected or the consistency score is lower than a preset score threshold, a fluorescence reaction image is acquired and the proportion of the fluorescence area is extracted and matched with a pre-stored template to obtain a verification matching result.

[0059] Specifically, when the document information verification process is activated according to the personalized verification sequence command, the system first acquires the voiceprint feature fragments from the user's real-time voice stream and extracts the user's self-proclaimed name text using speech recognition technology. Simultaneously, it reads the document recognition name text extracted via optical character recognition from the structured identity feature vector. The system then calculates the text similarity matching value between the self-proclaimed name and the document recognition name; the consistency score is equal to one minus the edit distance divided by the maximum name length. The consistency score is calculated as follows: con ,have:

[0060]

[0061] Where ED represents the minimum edit distance between two name strings, calculated using a dynamic programming algorithm, and L... max The score represents the character length of the longer of the two names; this score is normalized to the range of zero to one. Next, the system controls the device's flashlight to switch to a specific wavelength mode, such as ultraviolet light, to illuminate the predefined anti-counterfeiting area in the user-submitted document image. The image sensor captures the optical response characteristics of this area, extracting the number and distribution pattern of feature points. When an abnormality in the optical response characteristics is detected—that is, the number of feature points is lower than a preset threshold, the distribution pattern deviates from the standard template, or the consistency score is lower than a preset threshold, such as 0.8—the fluorescence reaction verification sub-process is activated: the ultraviolet light source is controlled to enhance irradiation, a fluorescence reaction image of the document is acquired, the pixel area of ​​the fluorescent region is extracted using an image segmentation algorithm, and its proportion to the total area of ​​the document is calculated. This proportion is then matched with the proportion of the pre-stored standard fluorescence template, and the absolute difference method is used to obtain the verification matching result. For the verification matching result, Match... result ,have:

[0062] Match result =1-|R current -R template |,

[0063] Where R current Represents the current percentage of fluorescent region area, RRtemplate This indicates the standard percentage value of the pre-stored template. This process is based on a multi-dimensional anti-counterfeiting cross-verification principle. Through dynamic consistency verification of voiceprint and name on the ID card, it effectively identifies identity theft; through specific spectral response detection, it accurately detects physical document forgery traces; and under abnormal conditions, it triggers a fluorescence verification layer to enhance the ability to identify highly realistic counterfeit documents, forming a progressive anti-counterfeiting verification mechanism that significantly improves the anti-fraud depth of the verification system.

[0064] Optionally, the generation of the multimodal biometric package includes:

[0065] Based on the fast face comparison instruction, the frequency of micro-expression changes and pupil reflection features in the dynamic facial video stream are detected to generate a liveness detection flag.

[0066] Based on the face-voiceprint dual-factor verification command, a voice reading command containing a random number sequence is sent to the user terminal and voiceprint matching is performed to obtain the anti-counterfeiting consistency index and voiceprint confidence score.

[0067] Based on the dynamic action challenge command, a specified head movement trajectory combination command is sent to the user terminal to obtain a head posture change sequence.

[0068] The anti-counterfeiting consistency index, the voiceprint confidence score, and the liveness detection flag are timestamped and encapsulated, and combined with the verification matching results and the head posture change sequence to form a multimodal biometric package.

[0069] Specifically, when the system activates the corresponding verification mode based on the personalized verification sequence command, it performs differentiated data acquisition and processing for different commands: If it is a fast face comparison command, the system captures the user's facial dynamic video stream through the camera, calculates the frequency of micro-expression changes in the specified facial area per unit time using optical flow, and extracts the displacement of the pupil's reflection feature points to the screen flash stimulus. When the micro-expression frequency is within the normal physiological range and the pupil reflection feature matches the liveness feature, a binary liveness detection flag is generated as true; otherwise, it is false. If it is a face-voiceprint dual-factor verification command, a voice reading command containing a random number sequence is sent to the user terminal and voiceprint matching is performed. The system obtains anti-counterfeiting consistency indicators and voiceprint confidence scores. For dynamic action challenge commands, the system sends a specified head movement trajectory combination command to the user terminal, such as "turn 30 degrees to the left and then nod." It captures three-axis angular velocity data using the device's gyroscope and calculates the head posture change sequence through integration. This sequence consists of time-series data of pitch, yaw, and roll angles. Finally, the system aligns and encapsulates all output data, including anti-counterfeiting consistency indicators, voiceprint confidence scores, and liveness detection flags, according to the acquisition timestamp at the millisecond level. If there is a document information verification command, the verification matching result is added; if there is a dynamic action challenge, the head posture change sequence is added, forming a structured multimodal biometric package. This process is based on the principle of multi-source biometric spatiotemporal fusion. It blocks static photo attacks through liveness detection flags, resists audio playback fraud through lip-reading synchronization indicators, verifies the authenticity of physical operations through action trajectory sequences, and ensures the integrity of the biological behavior evidence chain through multi-dimensional feature timestamp alignment and encapsulation, providing a highly reliable data foundation for subsequent cross-modal comparisons.

[0070] Optionally, obtaining the anti-counterfeiting consistency index and voiceprint confidence score includes:

[0071] Capture lip movement features and speech spectrum features during the user's execution of voice reading instructions;

[0072] The lip movement features and the speech spectrum features are subjected to time-series alignment analysis to generate anti-counterfeiting consistency indicators.

[0073] Extract the voiceprint feature segments from the speech spectrum features, perform similarity matching between the voiceprint feature segments and the pre-stored voiceprint template, and generate a voiceprint confidence score.

[0074] Specifically, such as Figure 2 As shown, when the system executes the face-voiceprint two-factor authentication command, it first sends a voice reading command containing a random number sequence to the user terminal. During the user's reading, two parallel acquisition threads are started simultaneously: the video acquisition thread captures a dynamic facial video stream of thirty frames per second through the camera, and uses a key point detection algorithm to extract the coordinate sequence of the lip center point as the lip movement feature, denoted as Lip. t={(x1,y1,t1),(x2,y2,t2),…,(x n ,y n ,t n )}, where x n y n The coordinates are normalized values, t n The timestamp is in milliseconds; the audio acquisition thread records real-time speech streams via microphone, performs short-time Fourier transform at a sampling rate of 16 kHz, and extracts the Mel-frequency cepstral coefficients of every 20 millisecond frames as speech spectral features, denoted as... Lip t With Spec t Alignment is performed at the millisecond level based on timestamps, and the Dynamic Time Warping Distance (DTW) is calculated. For the anti-counterfeiting consistency index, the following applies:

[0075]

[0076] Among them, DTW (Lip) t ,Spec t ) indicates Lip t Coordinate sequence and Spec t The minimum regularized path cumulative distance of the spectral sequence is calculated using a dynamic programming algorithm to obtain Max. DTW The maximum normalization distance threshold is determined statistically from historical positive sample data; Consistency index Normalized to the zero-to-one range as an indicator of anti-counterfeiting consistency. Meanwhile, from the Spec... t The fundamental frequency F0 and the first three formants F1-F3 are extracted to form a voiceprint feature segment, which is then matched with a pre-stored user voiceprint template. The voiceprint confidence score is calculated using cosine similarity. score :

[0077]

[0078] Among them, V candidate V represents the current voiceprint feature segment vector. template This represents the voiceprint template vector established during the user registration phase, and both have the same dimension. This process is based on the principle of spatiotemporal coupling of biometric features. Through rigorous temporal alignment analysis of lip movements and speech spectrum, it effectively identifies attacks such as audio playback or video synthesis. Professional extraction and matching of voiceprint features ensures the uniqueness of identity biometric verification. A two-factor cross-validation mechanism significantly enhances the defense against advanced deception techniques.

[0079] Optionally, the generation of the verification decision feature vector includes:

[0080] The name field is extracted from the structured identity feature vector, and the self-proclaimed name is extracted from the voiceprint feature fragment;

[0081] The name field is matched with the self-proclaimed name using text similarity to obtain a matching value;

[0082] When the detected matching value or the biometric confidence matrix is ​​lower than a preset threshold, the high-risk handling subprocess is activated.

[0083] Specifically, firstly, the name field is extracted from the structured identity feature vector, which is generated from the document image data after optical character recognition processing; simultaneously, the self-proclaimed name is extracted from the voiceprint feature segment, which comes from the real-time voice stream captured when the user executes a voice reading command, and is converted into text through speech recognition technology; the text similarity matching value between the name field and the self-proclaimed name is calculated, that is, the matching value is equal to one minus the edit distance divided by the maximum name length. For the matching value Match value ,have:

[0084]

[0085] The matching values ​​are normalized to the range of zero to one. A biometric confidence matrix is ​​obtained, output by the modal comparison module through cross-modal correlation comparison of multimodal biometric packages and the user's baseline biometric template. Preset name matching thresholds and biometric confidence thresholds are established based on historical verification data. When a matching value is detected to be less than the name matching threshold or any element in the biometric confidence matrix is ​​less than the biometric confidence threshold, a high-risk handling sub-process is activated. This process sends a hand swipe command to the user and captures touchscreen pressure sensing data and abnormal gyroscope vibration characteristics to generate gesture trajectory feature data. Finally, the gesture trajectory feature data is spatiotemporally correlated with the multimodal biometric package to output a verification decision feature vector. This process is based on the principle of multi-source information consistency verification. By matching the text of the name on the identification document with the self-proclaimed name via voiceprint, identity information conflicts are effectively detected. Combined with biometric confidence threshold judgment, a high-risk handling mechanism is triggered in a timely manner. Overall, this enhances the system's ability to identify identity theft and biometric forgery, improving the accuracy and robustness of verification decisions.

[0086] Optionally, the high-risk handling sub-process includes:

[0087] Send hand swipe commands to the user's device on the touchscreen;

[0088] Based on the sliding finger of the hand, touch screen pressure sensing data and abnormal vibration characteristics of the gyroscope are captured to obtain gesture trajectory feature data;

[0089] The gesture trajectory feature data and the multimodal biometric package are subjected to spatiotemporal correlation analysis to generate a verification decision feature vector.

[0090] Specifically, when the high-risk handling sub-process is activated, a hand swipe command containing a specific trajectory pattern is sent to the user's terminal. This command displays a visual swipe path on the touchscreen. During the user's swipe operation, the system simultaneously collects two types of sensor data: a touch pressure time sequence obtained through the touchscreen pressure sensor, denoted as... Where p k This is the pressure value, in Newtons (t). k The timestamps are in the millisecond range; three-axis angular velocity data are acquired through the device's gyroscope, and the magnitude and standard deviation of the angular velocity vector are calculated to obtain the abnormal vibration characteristics. level :

[0091]

[0092] Where ω x ,ω y ,ω z These represent the instantaneous values ​​of the three-axis angular velocities, and σ represents the standard deviation calculation function. The extracted gesture trajectory feature data includes three key dimensions: calculating the mean touch pressure as the pressure feature, calculating the Hausdorff distance between the actual swipe trajectory and the preset path as the trajectory deviation feature, and recording the Shake... level As vibration characteristics, the above characteristics are subjected to spatiotemporal correlation analysis with multimodal biometric packages: first, the gesture operation time window is aligned with the biometric acquisition time axis to establish a time correlation matrix; second, the biometric behavior coupling index Coupling is calculated. index :

[0093] Coupling index =w1·exp(-|T gesture -T bio |)+w2·cosθ,

[0094] Among them, T gesture T represents the duration of the gesture operation. bio The duration of voiceprint or facial verification is represented by θ, the spatial angle between the head posture change vector and the gesture trajectory vector, and w1 and w2 are preset weighting coefficients. The final result is a verification decision feature vector containing pressure features, trajectory deviation features, vibration features, and coupling degree indices. This process is based on the principle of behavioral biometric collaborative verification. It uses pressure pattern recognition to exclude automated robot operations, trajectory deviation detection to discover malicious software manipulation, and bio-behavioral spatiotemporal coupling analysis to confirm the consistency of the operating entity, forming a multi-layered defense mechanism for high-risk scenarios.

[0095] Optionally, the method further includes:

[0096] Extract and parse the data packets of fraud cases that have been manually reviewed and confirmed to obtain biometric anomaly vectors;

[0097] Based on the aforementioned biometric anomaly vectors, the multidimensional risk feature vectors corresponding to the fraud case data packets are labeled to obtain a risk labeling dataset;

[0098] The attack feature profile database is updated using the risk-labeled dataset.

[0099] Specifically, the process involves acquiring data packages of transaction cases ultimately confirmed as fraud through manual review. These packages contain complete verification process records. The data packages are then parsed to extract biometric anomaly vectors. These vectors are composed of differences between the multimodal biometric package and the user's baseline biometric template, specifically including four dimensions: abnormal status of liveness detection markers, deviation of anti-counterfeiting consistency indicators, voiceprint confidence score discrepancy, and standard deviation of head posture change sequences. Simultaneously, the original multidimensional risk feature vector corresponding to the fraud case is extracted. This vector is generated from user historical behavior data, current transaction request data, device environmental parameters, and initial biometric signal data. The biometric anomaly vectors are used as annotation labels and associated with the corresponding multidimensional risk feature vectors to form a risk-labeled dataset. A feature-weighted fusion algorithm is then used to update the attack feature profile library. For the new attack feature vector XT in the updated attack feature profile library, the following applies:

[0100] XT=λ·YT+(1-λ)·DFX,

[0101] Wherein, YT is the original attack feature vector, DFX is the current risk-labeled feature vector, and λ is the forgetting factor, ranging from zero to one, controlling the proportion of historical data retained. The current risk-labeled feature vector is obtained by averaging the risk-labeled dataset. The attack feature profile library continuously evolves through iterative updates. This process is based on the principle of incremental adversarial learning, using fraudulent samples confirmed by manual review to correct blind spots in the feature library and enhance the system's sensitivity to identifying new attack patterns. The dynamic update mechanism maintains the timeliness of the attack profile, forming a closed-loop system where defense capabilities and fraud methods are upgraded simultaneously.

[0102] Based on the same inventive concept, such as Figure 3 As shown, the present invention also provides a multimodal identity verification system for internet enterprises, the system comprising:

[0103] The data acquisition module is used to acquire user historical behavior data, current transaction request data, device environmental parameters and initial biosignal data, and generate a multidimensional risk feature vector;

[0104] The risk perception module is used to perform real-time analysis on the multidimensional risk feature vector and output an initial risk score;

[0105] The verification path generation module is used to generate a verification path based on the initial risk score to obtain personalized verification sequence instructions.

[0106] The multimodal acquisition module is used to execute the personalized verification sequence instructions, acquire the user's facial dynamic video stream, real-time voice stream and response action timing data, and generate a multimodal biometric package;

[0107] The modality comparison module is used to perform cross-modal correlation comparison between the multimodal biometric package and the pre-stored user baseline biometric template, and output a biometric confidence matrix;

[0108] The structured data acquisition module is used to acquire user-submitted document image data, perform optical character recognition to extract document information fields, and generate structured identity feature vectors.

[0109] The correlation decision module is used to perform feature-level correlation analysis on the biometric confidence matrix, the structured identity feature vector, and the initial risk score to generate a verification decision feature vector;

[0110] The decision instruction generation module is used to determine the verification result status based on the verification decision feature vector and generate an pass instruction, a rejection instruction, or a manual review request instruction.

[0111] Example 1

[0112] To verify the feasibility of this invention in practice, it was applied to the online business platform of a large commercial enterprise. This enterprise aims to effectively prevent the increasing risks of identity fraud, such as deepfakes and account theft, when handling high-risk transactions like user account opening, large-amount transfers, and password changes, while simultaneously optimizing the user experience for legitimate users and reducing unnecessary verification friction.

[0113] To verify the beneficial effects of this invention, the company selected 20,000 active users for a six-month pilot application. 10,000 users served as the experimental group, using the system described in this invention; the other 10,000 users with similar conditions served as the control group, using the bank's traditional verification scheme of "static facial recognition plus SMS verification code." During the experiment, the average verification time, fraud attack interception success rate, and the proportion requiring manual review were monitored and recorded under different risk scenarios.

[0114] In a low-risk transaction scenario, Mr. Zhang, a user in the experimental group, used his frequently used mobile phone and home Wi-Fi to log into a mobile banking app to check his account balance. The data acquisition module of this invention collected Mr. Zhang's historical behavioral data (frequently used devices, frequently used IP addresses), current transaction request data (balance inquiry, a non-sensitive operation), device environmental parameters (device model, system version normal), and initial biometric signals (slight shaking while holding the phone, normal). The risk perception module analyzed the multi-dimensional risk feature vector generated from this data, outputting an initial risk score of 12 (below the first threshold of 30), classifying it as a low-risk scenario. Based on the low-risk score, the verification path generation module only activated the "rapid face comparison instruction." The multimodal acquisition module activated the front-facing camera to capture a dynamic video stream of Mr. Zhang's face. The system detected the frequency of his natural micro-expression changes and the pupil's reflection characteristics in the video stream, generating a "liveness detection flag" indicating it was true. The modal comparison module compared this facial feature with a pre-stored benchmark template, with a confidence level of 99.5%. The associated decision-making module generates a verification decision feature vector based on high confidence and low risk scores, and the decision instruction generation module finally outputs "pass instruction". The entire process takes about 2.5 seconds, providing a smooth user experience.

[0115] In a medium-risk transaction involving a small transfer to a frequently used contact, Ms. Li, a user in the experimental group, initiated a transfer of 5,000 yuan to her client (a frequent recipient) using the company's office network. The system detected that the transaction amount was moderate, but the login environment (company network IP) was not frequently used. After comprehensive evaluation by the risk perception module, an initial risk score of 58 was output (between the first threshold of 30 and the second threshold of 70), classifying it as medium risk. The system activated the "face-voiceprint dual-factor verification command." The system prompted Ms. Li to read aloud the random number "8192" displayed on the screen. During the reading, the multimodal acquisition module simultaneously captured her facial dynamic video stream and real-time audio stream. The system analyzed the temporal correlation between her lip movement features and speech spectrum features, generating an anti-spoofing consistency index of 0.96 (high consistency), effectively eliminating the risk of recording attacks. At the same time, the system extracted voiceprint feature segments from the speech and compared them with pre-stored voiceprint templates, generating a voiceprint confidence score of 98.2%. Both biometric dimensions passed with high scores, and the system determined that the verification was successful. The entire process took approximately 9 seconds.

[0116] In a high-risk fraud attack involving Deepfake video and voice synthesis, an attacker attempted to use a Deepfake image of user Mr. Wang to forge a video and synthesize voice, initiating a large transfer of 200,000 yuan to an unknown account from a newly registered device. After obtaining the current transaction request data (large amount, first transfer to an unknown account) and device environment parameters (new device ID, out-of-town IP), the risk perception module output an initial risk score as high as 85 (above the second threshold of 70). The system then queried the attack feature profile database and found that several parameters of the current scenario (such as response latency and smooth, unfluctuating sensor data) were highly similar to known "emulator / script attack" characteristics. The system calculated an adjustment coefficient of 1.2, updating the corrected risk score to 102. Subsequently, the verification path generation module added "dynamic action challenge instructions" and "document information verification instructions" to the "face-voiceprint dual-factor verification." The system instructed the attacker to read aloud, "Please say your name and turn your head to the left." The attacker played the synthesized voice "Mr. Wang" and manipulated the forged video to make a left-turning motion. However, the system detected a consistency index of only 0.31 between lip movements and the speech spectrum, far below the threshold. Simultaneously, the voiceprint confidence score of the synthesized speech was only 45%, resulting in a failed comparison. The head posture change sequence captured by the system showed stiff movement trajectories and angular velocity changes that did not conform to human kinematics. The attacker uploaded a forged ID photo. The system performed OCR recognition through the structured data acquisition module and asked the attacker to read out their name. The correlation decision module comparison revealed an inconsistency between the identified "Mr. Wang" in the voiceprint and the name "Wang Moumou" on the ID, with an extremely low consistency score. Because the values ​​in multiple biometric confidence matrices were below the preset threshold, the system activated a high-risk handling sub-process, requiring the user to slide along a specific trajectory on the screen. The attacker could not simulate the touch pressure and subtle tremors of a real user, resulting in abnormal gesture trajectory feature data. Ultimately, the decision instruction generation module generated a "rejection instruction," the transaction was successfully intercepted, and a high-risk alert was sent to the backend security center and the real user, Mr. Wang. After manual verification, the abnormal biometric vectors of the intercepted fraud cases, such as low anti-counterfeiting consistency indicators and abnormal posture sequences, were used to update the attack feature profile database, enhancing the system's ability to identify similar attacks.

[0117] Table 1: Comparison of Verification Efficiency and User Experience Data

[0118]

[0119] Table 2 Comparison of interception success rates for different fraud types

[0120]

[0121]

[0122] As can be seen from the data in Tables 1-2 above, the experimental group, after applying the method of this invention, performed significantly better than the control group. In terms of efficiency and user experience, this invention, by dynamically adjusting the verification strength, reduced the verification time in low-risk scenarios from 15.2 seconds to 2.8 seconds, increasing user satisfaction by 17%. Regarding security, this invention achieved an interception rate of over 99.5% against advanced fraud methods such as Deepfake, far exceeding traditional solutions and significantly enhancing the system's anti-fraud capabilities. Furthermore, the proportion requiring manual review decreased from 4.8% to 0.5%, greatly saving on enterprise operating costs. These data fully demonstrate that the multimodal identity verification method for internet enterprises proposed in this invention, through dynamic risk-driven, cross-modal correlation decision-making, and incremental learning mechanisms, can achieve an excellent balance between ensuring enterprise security and optimizing user experience, possessing extremely high practical application value.

[0123] It should be noted that the formulas described above, through the principle of dimensional consistency and mathematical standardization methods (such as normalization, dimensionless parameter conversion, or unit system unification), can translate physical quantities with different properties into unitless standard values ​​or parameters that can be superimposed in the same dimension. This eliminates the interference of different dimensions on the computational logic, allowing the formulas to retain the original data distribution characteristics while possessing mathematical rationality and adaptability to objective laws. These are conventional technical methods and will not be elaborated further. The electrical connections between the various units described above do not necessarily represent direct or indirect connections; any indirect connection method is applicable to the embodiments of this invention as long as it achieves the purpose of this invention. The above descriptions are merely exemplary embodiments of this invention and should not be construed as limiting the scope of this invention.

[0124] All equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.

Claims

1. A multimodal identity verification method for internet enterprises, characterized in that, The method includes: Acquire user historical behavior data, current transaction request data, device environmental parameters, and initial biosignal data to generate a multidimensional risk feature vector; The multidimensional risk feature vector is analyzed in real time, and an initial risk score is output. A verification path is generated based on the initial risk score, resulting in personalized verification sequence instructions. Execute the personalized verification sequence instructions to collect the user's facial dynamic video stream, real-time voice stream, and response action timing data to generate a multimodal biometric package; The multimodal biometric package is compared with the pre-stored user baseline biometric template across modalities, and a biometric confidence matrix is ​​output. Obtain the image data of the ID document submitted by the user, perform optical character recognition to extract the ID document information fields, and generate a structured identity feature vector; The biometric confidence matrix, the structured identity feature vector, and the initial risk score are subjected to feature-level correlation analysis to generate a verification decision feature vector; The verification result status is determined based on the verification decision feature vector, and an approval instruction, rejection instruction, or manual review request instruction is generated.

2. The multimodal identity verification method for internet enterprises according to claim 1, characterized in that, The instructions for generating personalized verification sequences include: When the initial risk score is below the first threshold, only the fast face comparison command is activated; When the initial risk score is between the first and second thresholds, activate the face-voiceprint two-factor verification command. When the initial risk score is higher than the second threshold, a dynamic action challenge instruction and a document information verification instruction are added to the face-voiceprint dual-factor verification instruction.

3. The multimodal identity verification method for internet enterprises according to claim 2, characterized in that, The method further includes: Acquire historical fraud case data packages and extract the instruction response parameter set of attackers to circumvent verification to establish an attack feature profile library. The instruction response parameter set includes biometric simulation delay time, abnormal fluctuation range of sensor data, and deviation angle parameter of action execution trajectory. Based on the attack feature profile database and the initial risk score, the adjustment coefficient is calculated; The initial risk score is corrected using the adjustment coefficient, and the personalized verification sequence instruction is updated. When the current device environment parameters are detected to match the characteristics of a high-risk area, insert a document information verification instruction.

4. The multimodal identity verification method for internet enterprises according to claim 3, characterized in that, The document information verification instruction includes: The consistency score is calculated by comparing the user's self-proclaimed name in the voiceprint feature segment with the name identified on the ID card in real time. The optical response characteristics of the anti-counterfeiting mark area on the document are detected under specific lighting conditions; When an abnormality in the optical response characteristics is detected or the consistency score is lower than a preset score threshold, a fluorescence reaction image is acquired and the proportion of the fluorescence area is extracted and matched with a pre-stored template to obtain a verification matching result.

5. The multimodal identity verification method for Internet enterprises according to claim 4, characterized in that, The generated multimodal biometric package includes: Based on the fast face comparison instruction, the frequency of micro-expression changes and pupil reflection features in the dynamic facial video stream are detected to generate a liveness detection flag. Based on the face-voiceprint dual-factor verification command, a voice reading command containing a random number sequence is sent to the user terminal and voiceprint matching is performed to obtain the anti-counterfeiting consistency index and voiceprint confidence score. Based on the dynamic action challenge command, a specified head movement trajectory combination command is sent to the user terminal to obtain a head posture change sequence. The anti-counterfeiting consistency index, the voiceprint confidence score, and the liveness detection flag are timestamped and encapsulated, and combined with the verification matching results and the head posture change sequence to form a multimodal biometric package.

6. The multimodal identity verification method for Internet enterprises according to claim 5, characterized in that, The obtained anti-counterfeiting consistency index and voiceprint confidence score include: Capture lip movement features and speech spectrum features during the user's execution of voice reading instructions; The lip movement features and the speech spectrum features are subjected to time-series alignment analysis to generate anti-counterfeiting consistency indicators. Extract the voiceprint feature segments from the speech spectrum features, perform similarity matching between the voiceprint feature segments and the pre-stored voiceprint template, and generate a voiceprint confidence score.

7. The multimodal identity verification method for internet enterprises according to claim 1, characterized in that, The generated verification decision feature vector includes: The name field is extracted from the structured identity feature vector, and the self-proclaimed name is extracted from the voiceprint feature fragment; The name field is matched with the self-proclaimed name using text similarity to obtain a matching value; When the detected matching value or the biometric confidence matrix is ​​lower than a preset threshold, the high-risk handling subprocess is activated.

8. The multimodal identity verification method for Internet enterprises according to claim 7, characterized in that, The high-risk handling sub-process includes: Send hand swipe commands to the user's device on the touchscreen; Based on the sliding finger of the hand, touch screen pressure sensing data and abnormal vibration characteristics of the gyroscope are captured to obtain gesture trajectory feature data; The gesture trajectory feature data and the multimodal biometric package are subjected to spatiotemporal correlation analysis to generate a verification decision feature vector.

9. The multimodal identity verification method for Internet enterprises according to claim 3, characterized in that, The method further includes: Extract and parse the data packets of fraud cases that have been manually reviewed and confirmed to obtain biometric anomaly vectors; Based on the aforementioned biometric anomaly vectors, the multidimensional risk feature vectors corresponding to the fraud case data packets are labeled to obtain a risk labeling dataset; The attack feature profile database is updated using the risk-labeled dataset.

10. A multimodal identity verification system for internet enterprises, applied to the multimodal identity verification method for internet enterprises as described in any one of claims 1-9, characterized in that, The system includes: The data acquisition module is used to acquire user historical behavior data, current transaction request data, device environmental parameters and initial biosignal data, and generate a multidimensional risk feature vector; The risk perception module is used to perform real-time analysis on the multidimensional risk feature vector and output an initial risk score; The verification path generation module is used to generate a verification path based on the initial risk score to obtain personalized verification sequence instructions. The multimodal acquisition module is used to execute the personalized verification sequence instructions, acquire the user's facial dynamic video stream, real-time voice stream and response action timing data, and generate a multimodal biometric package; The modality comparison module is used to perform cross-modal correlation comparison between the multimodal biometric package and the pre-stored user baseline biometric template, and output a biometric confidence matrix; The structured data acquisition module is used to acquire user-submitted document image data, perform optical character recognition to extract document information fields, and generate structured identity feature vectors. The correlation decision module is used to perform feature-level correlation analysis on the biometric confidence matrix, the structured identity feature vector, and the initial risk score to generate a verification decision feature vector; The decision instruction generation module is used to determine the verification result status based on the verification decision feature vector and generate an pass instruction, a rejection instruction, or a manual review request instruction.

Citation Information

Patent Citations

  • Voiceprint identification, face identification and synchronous in-vivo detection-based identity authentication method and system

    CN105426723A

  • Living person identity authentication method based on voice pattern and image features

    CN106709402A

  • User adaptive verification method based on operator payment account system

    CN114548999A

  • Digital payment identity security verification method and system based on cloud platform

    CN118154194A

  • Right and interest exchange verification and confirmation system

    CN120337188A

Cited By

  • Man-machine interaction identity intelligent identification method and system based on multi-mode perception

    CN121808757A

  • Enterprise registration self-service terminal remote identity verification method and system

    CN121967094A