Network access identity auditing method and system based on multi-modal data fusion

By collecting multimodal data and performing feature extraction and fusion analysis, the problem of low accuracy in identifying minors in existing technologies has been solved, achieving efficient and accurate verification of minors' online identity.

CN122133085APending Publication Date: 2026-06-02CHINA UNICOM (JIANGXI) IND INTERNET CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNICOM (JIANGXI) IND INTERNET CO LTD
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for identifying minors suffer from problems such as low accuracy, susceptibility to environmental interference, and high requirements for user cooperation, making it difficult to meet the needs of internet platforms for large-scale and efficient interception of minors from the internet.

Method used

Collect data from multiple modalities (mouse movement trajectory, audio recording, facial video stream, and fingertip capacitive pressure sequence), and construct a multimodal fusion identity audit model through feature extraction, phase space reconstruction, chaotic mapping expansion, and spectral analysis. This model comprehensively analyzes differences in users' operating habits, vocalization status, visual attention, and keystroke mechanics.

Benefits of technology

It significantly improves the accuracy of identifying minors and adults, prevents identity theft, and provides a stable and robust identification system that can work effectively in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133085A_ABST
    Figure CN122133085A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for network access identity verification based on multimodal data fusion. The method includes: collecting multiple modal data from several known users during the network access process; extracting features from each modal data to obtain a basic feature set corresponding to each modal data; reconstructing the basic feature sets of each modality into a phase space trajectory matrix, and generating a four-dimensional feature matrix based on the phase space trajectory matrix; performing chaotic mapping expansion, slicing, and concatenation on the four-dimensional feature matrix to obtain a concatenated vector, and stacking all the concatenated vectors row-wise to obtain a two-dimensional fusion feature matrix; and training an initial multimodal fusion identity verification model based on the two-dimensional fusion feature matrix. This invention, through multimodal fusion, deeply extracts and fully amplifies the comprehensive differences between minors and adults in terms of operating habits, vocalization states, visual attention, and keystroke mechanics, significantly improving the recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of identity verification technology, and in particular to a method and system for network access identity verification based on multimodal data fusion. Background Technology

[0002] With the widespread adoption of internet services, problems such as internet addiction among minors, online fraud, and exposure to harmful information have become increasingly prominent, attracting widespread attention. To protect the physical and mental health of minors, relevant national laws and regulations explicitly require internet service providers to effectively verify user identities to prevent minors from registering for and using internet services such as online games, live streaming rewards, and social media platforms without the consent of their guardians. Therefore, accurately identifying whether a user is a minor has become a key technical means for internet platforms to fulfill their social responsibilities and implement regulatory requirements. How to efficiently and accurately complete identity verification during the user onboarding process is crucial not only for the protection of minors but also directly impacts user experience and platform operational efficiency.

[0003] Current technologies for identifying minors primarily rely on single-dimensional detection methods such as ID card verification, facial age estimation, and voice age determination. ID card verification requires users to actively upload photos of their documents, a cumbersome process that carries the risk of document misuse. Facial age estimation is significantly affected by factors such as lighting, angle, and makeup, making accuracy difficult to guarantee. Voice age determination is prone to failure when there is background noise or when users deliberately disguise their age. In practical applications, these technologies generally suffer from insufficient recognition accuracy, susceptibility to environmental interference, and high requirements for user cooperation, making it difficult to meet the needs of internet platforms for large-scale, high-efficiency interception of minors accessing the internet. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for network access identity verification based on multimodal data fusion, which aims to solve the problems of identity theft and low accuracy in traditional methods of minor identity verification.

[0005] In a first aspect, the present invention provides a network access identity verification method based on multimodal data fusion, the method comprising: Collect multiple modal data of multiple known users during the network access process. The multiple modal data includes mouse movement trajectory, audio recording, facial video stream in front of the screen, and fingertip capacitive pressure sequence when typing on the keyboard. Feature extraction is performed on each modality of data to obtain a basic feature set corresponding to each modality of data. Each basic feature set contains multiple basic features. The basic feature sets of each mode are reconstructed into a phase space trajectory matrix, and a four-dimensional feature matrix is ​​generated based on the phase space trajectory matrix; The four-dimensional feature matrix is ​​expanded by chaotic mapping to obtain a four-dimensional chaotic state tensor. The four-dimensional chaotic state tensor is then sliced ​​to obtain a three-dimensional sub-tensor. The spectral vectors of the three-dimensional sub-tensor in each dimension are calculated, and the spectral vectors are concatenated to obtain a concatenated vector. All the concatenated vectors are stacked row by row to obtain a two-dimensional fused feature matrix. The initial multimodal fusion identity audit model is trained based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model. The target fusion feature matrix of the unknown user is obtained and input into the trained multimodal fusion identity audit model to obtain the identity audit result.

[0006] Secondly, this invention provides an online identity verification system based on multimodal data fusion, the system comprising: The modal data acquisition module is used to collect multiple modal data of multiple known users during the network access process. The multiple modal data include mouse movement trajectory, audio recording, facial video stream in front of the screen, and fingertip capacitive pressure sequence when typing on the keyboard. The basic feature extraction module is used to extract features from each type of modality data to obtain a basic feature set corresponding to each type of modality data. Each basic feature set contains multiple basic features. The feature reconstruction module is used to reconstruct the basic feature set of each mode into a phase space trajectory matrix, and generate a four-dimensional feature matrix based on the phase space trajectory matrix; The feature fusion module is used to perform chaotic mapping expansion on the four-dimensional feature matrix to obtain a four-dimensional chaotic state tensor, slice the four-dimensional chaotic state tensor to obtain a three-dimensional sub-tensor, calculate the spectral vector of the three-dimensional sub-tensor in each dimension, concatenate the spectral vectors to obtain a concatenated vector, and stack all the concatenated vectors row by row to obtain a two-dimensional fused feature matrix. The training module is used to train the initial multimodal fusion identity audit model based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model, obtain the target fusion feature matrix of the unknown identity user, and input the target fusion feature matrix into the trained multimodal fusion identity audit model to obtain the identity audit result.

[0007] Thirdly, the present invention provides a storage medium that stores one or more programs, which, when executed by a processor, implement the above-described method for network access identity verification based on multimodal data fusion.

[0008] Fourthly, the present invention provides an electronic device, the electronic device comprising a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the above-mentioned network access identity verification method based on multimodal data fusion.

[0009] Compared with the prior art, the present invention has the following advantages: 1. This invention extracts multi-dimensional behavioral features from four modalities: mouse movement trajectory, audio recording, facial video stream, and fingertip capacitive pressure sequence. These features are then fused through multiple levels, including phase space reconstruction, chaotic mapping expansion, and spectral analysis, before being input into an identity verification model to output verification results. This method utilizes multi-modal behavioral data naturally generated during user registration to construct behavioral features. Because behavioral patterns are difficult to imitate and replicate, it effectively prevents identity theft. Simultaneously, through multi-modal fusion, it deeply extracts and fully amplifies the comprehensive differences between minors and adults in operating habits, vocalization states, visual attention, and keystroke mechanics, significantly improving recognition accuracy and solving the problems of low accuracy and susceptibility to bypassing traditional single-modal or static feature recognition methods.

[0010] 2. The mouse movement trajectory in this invention records the user's motion control ability during screen interaction; the audio recording reflects the physiological development of the vocal organs; the facial video stream reveals the distribution pattern of visual attention; and the fingertip capacitive pressure sequence reflects the fine motor control level of the hand muscles. These four modalities are independent yet complementary. A single modality may fail due to environmental interference or fluctuations in user state, but the comprehensive analysis of the four modalities can form stable behavioral characteristics. Based on this, the features extracted from each modality are designed specifically for the physiological immaturity unique to minors, transforming subjectively observed behavioral differences into calculable mathematical indicators. This effectively captures the comprehensive differences between adults and minors in terms of operating habits, vocal states, visual attention, and motor control, providing a highly discriminative and robust feature foundation for subsequent multimodal fusion.

[0011] 3. In the fusion stage, the static features of each modality are expanded into a phase space trajectory matrix containing a delay dimension using a chaotic modulation formula, allowing the originally isolated basic features to acquire evolutionary information on the time axis. Then, a four-dimensional feature matrix is ​​constructed through row vector outer products, enabling pairwise interactions between different features within the same modality. This process is not a simple feature combination, but rather couples the curvature and velocity of the mouse trajectory, the fundamental frequency and energy of speech, the displacement and time of the gaze point, and the force and rhythm of keystrokes along the same delay dimension through multiplication operations, forming high-order features that reflect complex behavioral patterns such as hand-eye coordination, audiovisual coordination, and motion control. Subsequently, the four-dimensional features are integrated using Logistic chaotic mapping. The feature matrix is ​​expanded into a four-dimensional chaotic tensor. By leveraging the sensitivity of chaotic systems to initial values, the weak correlations between different modal features are fully intertwined and fused during the iteration process, so that the behavioral information originally scattered in four independent spaces forms a unified representation in the same high-dimensional chaotic space. Finally, the spectral components of each dimension are extracted by Fourier transform, and the periodic structure in the chaotic tensor is transformed into a spectral vector. The vectors are then concatenated to obtain a two-dimensional fusion feature matrix. This allows the comprehensive differences between adults and minors in multiple dimensions such as operating habits, vocalization, visual attention, and keystroke mechanics to be presented in the same feature space. It also provides the identity verification model with high-quality input that contains both the original features and the high-order coupling relationships. Attached Figure Description

[0012] Figure 1 This is a flowchart of an online identity verification method based on multimodal data fusion proposed in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an online identity verification system based on multimodal data fusion proposed in an embodiment of the present invention.

[0013] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but does not exclude other elements or objects.

[0015] like Figure 1 As shown, an embodiment of the present invention proposes an access identity verification method based on multimodal data fusion. The method includes steps S101 to S105, wherein: Step S101: Collect multiple modal data of multiple known users during the network access process. The multiple modal data includes mouse movement trajectory, audio recording, facial video stream in front of the screen, and fingertip capacitive pressure sequence when typing on the keyboard. It should be noted that mouse movement trajectories are recorded with millisecond precision using front-end JavaScript or system-level hooks, tracking the user's mouse movements on the screen. During user onboarding, pre-defined click tasks are automatically displayed. For example, each click task requires clicking a pre-defined target. A new target appears after each successful click, and the verification process typically requires clicking three or more targets consecutively. The mouse movement trajectory includes the screen coordinates of the mouse at multiple consecutive sampling points, the timestamp of reaching each sampling point, and the timestamp of each screen click during each click task. The time interval between adjacent sampling points is equal. Audio recording is achieved using the microphone of a terminal device, such as a computer or mobile phone, while the user reads and aloud a specified text or answers random questions. Facial video streams are captured by a camera, recording continuous facial video of the user's actions. Fingertip capacitive pressure sequences are recorded at a high sampling rate using a multi-touch and pressure-sensitive keyboard, recording the capacitance and pressure values ​​of each finger when it contacts a keycap, forming pressure timing data describing the typing force, contact area, and release rhythm.

[0016] Furthermore, before feature extraction, the collected raw multimodal data needs to be cleaned and standardized to eliminate interference from noise and individual differences. Specifically, for mouse movement trajectories, drift points caused by system lag are removed and the sampling frequency is standardized; for audio recordings, endpoint detection algorithms are used to remove silent segments, and pre-emphasis and noise reduction are performed; for facial video streams, face alignment, scale normalization, and illumination compensation are performed frame by frame; for fingertip capacitive pressure sequences, sliding window filtering is used to remove high-frequency noise, and irregular key event sequences are aligned to a fixed time length through linear interpolation. After preprocessing, all modal data must be Z-score normalized to conform to a standard normal distribution, laying the foundation for subsequent feature extraction.

[0017] Step S102: Perform feature extraction on each modality of data to obtain a basic feature set corresponding to each modality of data. Each basic feature set contains multiple basic features. It should be noted that in the field of identity verification, traditional methods usually prioritize modalities with direct physiological identification, such as faces, voiceprints, or fingerprints. Mouse movement trajectories have long been considered merely peripheral input signals of the operating system, mainly used for human-computer interaction rather than identity representation, and are difficult to carry reliable biometric information. However, the applicant's research found that by incorporating mouse trajectories into the analysis, the researchers delved into the micro-dynamics level. By constructing basic indicators such as instantaneous velocity, acceleration, and curvature, and by uniquely extracting a series of fine-grained features such as average curvature, rate of change of curvature, proportion of abnormal curvature points, velocity fluctuation, reaction time, and click accuracy, these features can capture the subtle motor control traces revealed by users unconsciously, which stem from differences in neurodevelopment.

[0018] Specifically, in some embodiments, the base quantity is obtained according to the following formula: ; in, Let be the instantaneous velocities of the i-th and (i+1)-th sampling points, respectively. Let be the screen coordinates of the i-th and (i+1)-th sampling points, respectively. Let be the instantaneous acceleration and instantaneous curvature of the i-th sampling point, respectively. The basic feature set corresponding to the mouse movement trajectory is calculated using the following formula. : ; in, For the mean curvature, The rate of change of curvature, For maximum curvature, The proportion of abnormal curvature points, For velocity fluctuation, During the reaction, For mouse click precision, The total number of sampling points. Let be the instantaneous curvature of the (i+1)th sampling point. This represents the number of all sampling points whose instantaneous curvature is greater than a first preset threshold. The time interval between adjacent sampling points. This is the timestamp of the successful mouse click on the target during the j-th click of the task. Let j be the timestamp when the target appears during the j-th click on the task. The standard duration for the pre-set j-th click task, This represents the total number of click tasks. Each click task contains a circular target, and the next target only appears after each successful click. Let J represent the screen coordinates at which the mouse successfully clicks the target on the j-th click. Let the screen coordinates be the center of the target for the j-th click task. Let be the target radius of the j-th click task.

[0019] In summary, in the real-world scenario of a user joining a network, when faced with sequentially appearing clickable targets on the screen, the user's hand movement trajectory is not a simple displacement from point A to point B, but rather a complex behavioral signal containing neural developmental status. The seven features mentioned above are not merely a direct description of the mouse trajectory, but rather a micro-behavioral representation system with inherent connections constructed from the perspective of differences in neurophysiological development. Specifically, the average curvature captures the overall level of smooth hand tracking ability; the rate of change of curvature reflects the stability of neural regulation during movement; and the ratio of maximum curvature to abnormal curvature points jointly characterizes the transient jitters and corrective behaviors that are difficult to suppress due to incomplete cerebellar development. These micro-fluctuations are physiological traces that users cannot subjectively control. Speed ​​fluctuation measures the smoothness of acceleration and deceleration, reaction time reflects the neural transmission efficiency from visual stimulus to movement execution, and click accuracy comprehensively reflects the fine control ability of hand-eye coordination. These seven features, from different dimensions, jointly outline the user's neural developmental status throughout the entire process of movement planning, execution control, and feedback correction, and each feature is designed for a micro-level that users cannot deliberately imitate.

[0020] It's important to note that mouse movement trajectories, facial video streams, and fingertip capacitive pressure sequences primarily reflect a user's motor control and behavioral habits, while the audio recording modality provides a physiological development information dimension that is completely orthogonal to these. Specifically, even if a minor, through deliberate training, can mimic adult behavior patterns in mouse operation or control their gaze while looking at a screen, the nonlinear vibration characteristics of their vocal cords during vocalization—such as multi-frequency competition, formant jumps, and harmonic breaks—cannot be eliminated through conscious will. Conversely, even if an adult attempts to impersonate a minor and change their tone, their established vocal cord structure cannot reproduce the vocal cord vibration instability unique to the developmental period. Therefore, the introduction of the audio recording modality is not a simple supplement to existing modalities, but rather the construction of a physiological development detection dimension orthogonal to the behavioral modality. When other modalities fail due to user state fluctuations, environmental interference, or deliberate imitation, the audio recording modality can form a complete, multi-dimensional, mutually reinforcing, and complementary recognition system with other modalities, significantly improving the overall robustness and anti-spoofing ability of the solution in complex real-world scenarios.

[0021] Specifically, in some embodiments, the basic feature set corresponding to the recording is obtained according to the following formula. : ; in, For multimodal coefficients, The resonance peak jump rate, Harmonic fracture index, For the frequency of voice interruption; The recording is obtained by the user reading a preset text during the network access process. The recording is divided into several frames, and the frames are windowed and silent frames are removed to obtain a set of valid frames, which contains multiple valid frames. Calculate the autocorrelation function for each valid frame in the set of valid frames: ; in, For the t-th valid frame in time delay The autocorrelation value under the following conditions , They are the m-th and m-th frames in the t-th valid frame, respectively. One signal, Total number of sampling points for each valid frame; Within a preset fundamental frequency range, scan all autocorrelation values ​​for each valid frame. If... and and Then determine The first local peak value is calculated by summing all first local peak values ​​from all valid frames. ; in, The time delay of the t-th valid frame is respectively The autocorrelation value under the following conditions They are the 1st, 2nd, and 3rd respectively. The first local peak For the first local peak set, It is a constant. This represents the maximum value of the autocorrelation. The preset fundamental frequency range is divided into several fundamental frequency intervals, and all first local peaks are assigned to their corresponding fundamental frequency intervals according to the frequency of each first local peak. and and Then determine This is the second local peak value; in, These represent the number of first local peaks falling into the i-th, (i-1), and (i+1)-th intervals, respectively. It is a constant. The maximum number; If multiple second local peaks exist, the multipeak coefficient is calculated using the following formula: ; in, This represents the total number of the second local peaks. These are the average frequencies of the m-th and n-th fundamental frequency intervals, respectively. For the m-th and n-th second local peaks, This is the difference between the upper and lower limits of the preset baseband range.

[0022] In summary, during the identity verification process for online users, applicants discovered that when users read preset text, the physiological development of their vocal organs is implicitly present in the speech signal as a nonlinear acoustic phenomenon. However, conventional features such as MFCC and fundamental frequency mean can only describe the surface properties of sound and cannot touch the physical essence of the vocalization process. The reason this solution extracts the multi-peak coefficient is that during the vocal cord development process of minors, the laryngeal muscle control is not yet mature and the vocal cord mucosa fluctuations are unstable, resulting in the physical phenomenon that multiple competing frequencies often coexist in the vocal cord vibration pattern. This phenomenon cannot be eliminated through subjective imitation and is stable under different pronunciation content. In addition, because the autocorrelation peaks in real speech signals are mixed with various components such as fundamental frequency, harmonics, and environmental noise, simply counting the number or position of peaks cannot effectively identify true physiological multi-peak vibration. To address this, our scheme employs a dual screening mechanism: First, by scanning local peaks within the fundamental frequency range, candidate vibration modes are initially selected from the autocorrelation function of each frame. Then, through the division of a preset fundamental frequency interval and a second local peak screening, multiple dominant vibration modes stably existing across continuous frequency ranges are identified. This mechanism effectively eliminates accidental peaks caused by noise interference or single-frame anomalies, retaining only those vibration modes reflecting continuous physiological characteristics. Finally, by using a weighted distance formula, the frequency differences and intensity distribution among multiple dominant vibration modes are comprehensively considered, transforming the complex multi-peak phenomenon into a single numerical index. This preserves the physical essence of physiological characteristics while ensuring cross-sample comparability.

[0023] In addition, in some embodiments, the resonance peak jump rate needs to be calculated according to the following formula: ; in, The total number of valid frames. These are the first formant frequencies of the (t+1)th and tth valid frames, respectively. To preset the jump threshold, This is an indicator function; it returns 1 if true and 0 if false. The harmonic fracture index is calculated using the following formula: ; in, Let be the energies of the (p+1)th, (p-1)th, and (p)th harmonics of the t-th valid frame, respectively. The third preset threshold, For the FFT spectrum of the q-th valid frame, The fourth preset threshold, Let be the fundamental frequency of the t-th valid frame; The frequency of vocal interruptions can be calculated using the following formula: ; in, All are constants. These are the short-time energies of the t-th and t+1-th valid frames, respectively.

[0024] In summary, the formant jump rate captures the stability of tongue position and oral cavity shape during continuous articulation. The applicant found that minors, due to the immature differentiation of vocal tract muscle control, are prone to abnormal formant jumps during the transition between adjacent syllables. The harmonic breakage index reflects the synergy between vocal cord vibration and vocal tract tuning; irregular vibration of the vocal cords during development can lead to local distortions in the harmonic structure. The voice interruption frequency directly characterizes the stability of glottal closure; minors, due to insufficient laryngeal muscle endurance, are more prone to brief glottal leaks or vocal interruptions during articulation. These three features, from the dimensions of vocal tract regulation, harmonic synergy, and glottal closure, jointly outline the physiological development of the vocal organs during dynamic articulation. Each feature targets microscopic physiological phenomena that users cannot subjectively control. Even if minors consciously control their articulation, it is difficult to eliminate the instinctive tremors during vocal tract transitions, the physical distortions of the harmonic structure, and the muscle fatigue traces during glottal closure.

[0025] In addition, in some embodiments, the facial video stream includes consecutive frames of facial images, and each frame of the facial image is detected according to a trained eye detection model to identify the left eye region and the right eye region in each frame of the facial image. Obtain the center coordinates of the left and right eye regions respectively, and then obtain the fixation point coordinates based on the center coordinates of the left and right eye regions: ; The basic feature set corresponding to the facial video stream is extracted using the following formula. : ; in, The total distance moved by the fixation point. This represents the average velocity of the gaze point. For the coverage area of ​​the fixation point, This represents the total number of frames in the facial image. Let x and y be the x-coordinates of the gaze points corresponding to the face images in frame t+1 and frame t, respectively. Let be the ordinates of the gaze points corresponding to the face images in frame t+1 and frame t, respectively. The time interval between adjacent frames of face images. These represent the maximum and minimum values ​​of the x-coordinate of the fixation point, respectively. These represent the maximum and minimum values ​​of the ordinate of the fixation point, respectively. Let x and y be the x-coordinates of the centers of the right and left eye regions in the t-th frame of the face image, respectively. , y and y are the center ordinates of the right and left eye regions in the face image of frame t, respectively.

[0026] It should be noted that the eye detection model can employ a deep convolutional neural network-based object detection architecture, such as YOLOv5 or Faster R-CNN. Its training dataset is derived from a fusion of large-scale public face datasets and self-collected face images from multiple age groups. During data construction, the left and right eye regions in the face images are first manually labeled, with bounding boxes marking the positions of the left and right eyes respectively, generating corresponding label files. During training, the labeled face images are uniformly scaled to a fixed size, such as 416×416 pixels, and data augmentation operations are performed, including random rotation, brightness adjustment, and Gaussian noise addition, to improve the model's generalization ability and robustness. The network model takes the face image as input and outputs the bounding box coordinates and confidence scores of the detected left and right eye regions. The training process uses the SGD optimizer with an initial learning rate of 0.001 and a batch size of 32, iteratively training until the loss function converges. After training, the model is deployed in the identity verification system to detect the left and right eye regions in each frame of facial images in real time, providing input for the calculation of subsequent gaze trajectory features.

[0027] In online identity verification scenarios, the movement trajectory of a user's gaze point while operating in front of a screen is often considered a functional behavior related to the task. Traditional recognition methods typically focus only on facial recognition or expression analysis. Because the prefrontal cortex and attention control system of minors are not yet fully developed, they exhibit significantly different visual attention allocation patterns compared to adults when completing screen-clicking tasks. Specifically, they show longer gaze point search paths, faster scanning speeds with a lack of strategic pauses, wider visual coverage, and a lack of focusing efficiency. These characteristics are not visual behaviors that users can subjectively control, but rather instinctive reactions stemming from their neurodevelopmental level: minors tend to use broader visual searches to compensate for insufficient target localization abilities, lacking the path planning abilities of adults during movement, leading to increased ineffective displacements, and difficulty in stabilizing the gaze point within the effective operating area. Based on this, this embodiment detects the left and right eye regions and calculates the gaze point coordinates, then extracts three quantitative indicators: total movement distance, average speed, and coverage area, transforming the aforementioned subtle differences in visual attention development into calculable feature expressions. If conventional facial age estimation methods are used, they are easily affected by factors such as makeup, lighting, and angle; however, the gaze trajectory features extracted in this embodiment are derived from the unconscious visual behavior of users during operation, and have a natural anti-disguise capability.

[0028] Furthermore, in some embodiments, the basic feature set corresponding to the fingertip capacitive pressure sequence is obtained according to the following formula. : ; in, This is the peak pressure asymmetry index. For the press-release phase difference, For waveform steepness, The keystroke rhythm entropy, This is due to the pressure decay memory effect. This represents the total number of keystrokes. , These are the peak capacitances for the (k+1)th and kth keystrokes, respectively. These are the rise time and fall time for the k-th keystroke contact, respectively. This is the set of all sampling points within a 5ms window before and after the peak capacitance of the k-th keystroke. The capacitance value at the nth sampling point after the kth keystroke. Let be the window width of the set of all sampling points within a 5ms window before and after the peak capacitance of the k-th keystroke. Let be the normalized probability of the i-th keystroke. This is the average of the peak capacitance of all keystrokes; The peak capacitance is calculated using the following formula: ; Obtain the time interval between the peak capacitance of any adjacent keystrokes, and count the number of occurrences of each time interval. Use the ratio of the number of occurrences to the total number of time intervals as the normalized probability of the corresponding time interval. The rise time is the time from the start of the rise to the peak capacitance point, and the fall time is the time from the peak capacitance point to the end of the fall. ; in, Let the peak capacitance point, the start point of the rise, and the end point of the fall be the k-th keystroke. It is a constant.

[0029] It's important to note that the fingertip capacitive pressure sequence records the capacitance value change curve corresponding to each key press during a user's typing process using a multi-touch and pressure-sensitive keyboard or touchpad at a high sampling rate. When a user's fingertip contacts the keyboard surface, changes in the contact area cause corresponding changes in the capacitance value. As the fingertip presses down, the contact area gradually increases, causing the capacitance value to rise to its peak; as the fingertip releases, the contact area gradually decreases, causing the capacitance value to drop to the baseline. Therefore, each key press corresponds to a complete capacitance value timing curve, which includes microstructural information such as rise time, peak capacitance, fall time, and waveform steepness.

[0030] From the perspective of user behavior characteristics, this pressure sequence contains muscle memory and dynamic characteristics that are difficult to imitate, formed by users during long-term typing. Specifically, it includes the following five dimensions of information: First, the relative change pattern of peak pressure reflects the stability and fluctuation pattern of the force applied by the user during continuous keystrokes; second, the temporal phase relationship between pressing and releasing reflects the coordinated control ability of finger muscles in the two phases of pressing and releasing; third, the steepness of the waveform characterizes the speed of impact and the rate of increase in force at the moment the fingertips contact the keyboard; fourth, the regularity of the keystroke rhythm measures the rhythm stability of the user during typing through the entropy value of the time interval between adjacent keystrokes; fifth, the memory effect of pressure decay describes the autocorrelation between pressure values ​​during multiple consecutive keystrokes, reflecting the influence pattern of muscle fatigue or attention fluctuations on keystroke behavior.

[0031] Furthermore, because the bones and muscles of minors' hands are not yet fully developed, their fine motor control abilities differ fundamentally from those of adults. Specifically, this manifests as greater fluctuations in force during continuous keystrokes and a lack of stable adjustment mechanisms, reflecting their ability to regulate force; insufficient finger muscle coordination leading to a clumsy transition between pressing and lifting the keys, reflecting muscle coordination issues; difficulty maintaining a stable impact speed at the moment the fingertips touch the keyboard, reflecting precision in force control; poor rhythmic regularity and susceptibility to fluctuations in attention, reflecting rhythmic stability; and a different pattern of muscle fatigue affecting keystroke force compared to adults, reflecting muscle endurance characteristics. These five characteristics correspond to the complete keystroke dynamics chain from force preparation, contact impact, force adjustment, rhythm maintenance to fatigue response. They capture the microscopic behavioral traces etched into neuromuscular memory formed by users through millions of typing sessions. Even if minors deliberately imitate an adult's typing rhythm, they cannot replicate these muscle memory characteristics on a millisecond-level time scale and a micro-Newton-level force scale. Using only conventional macroscopic indicators such as key intervals or typing speed can easily mask the true differences due to brief periods of user focus or deliberate slowing down.

[0032] In summary, the above four modalities precisely cover the four physiological dimensions that users naturally exhibit during the network access process: neurodevelopment, vocal physiology, visual attention, and muscle memory. Moreover, the developmental timeline and physiological basis of each dimension are different. Motor control depends on the cerebellum, vocalization originates from the vocal cord structure, visual attention is controlled by the prefrontal cortex, and keystroke dynamics reflect the development of hand muscles. They cannot be synchronously deceived by a single dimension.

[0033] Step S103: Reconstruct the basic feature set of each mode into a phase space trajectory matrix, and generate a four-dimensional feature matrix based on the phase space trajectory matrix; It should be noted that the reconstruction is performed according to the following formula: ; in, The element in the i-th row and j-th column of the phase space trajectory matrix corresponding to the m-th mode. For the i-th basic feature of the basic feature set corresponding to the m-th modality, , , Let be the chaos intensity coefficient, development steepness coefficient, and attention decay coefficient of the i-th basic feature in the basic feature set corresponding to the m-th modality, respectively. Let be the standard deviation and skewness coefficient of the i-th basic feature of the basic feature set corresponding to the m-th modality in the known underage user group, respectively. These are the mean standard deviation and mean skewness coefficient of all basic features of the basic feature set corresponding to the m-th modality in the known underage user group, respectively; The element in the i-th row, j-th column, and k-th layer of the four-dimensional feature matrix is ​​obtained using the following formula. : When i = 1 to 7 ; When i = 8 to 11 ; When i = 12 to 14 ; When i = 14 to 19 ; in, , These are the elements in the i-th row and k-th column, and the j-th row and k-th column, respectively, of the phase space trajectory matrix corresponding to the mouse movement trajectory. , These are the elements in the (i-7)th row and kth column of the phase space trajectory matrix corresponding to the recording, respectively. , Let be the elements in the i-11th row and kth column of the phase space trajectory matrix corresponding to the facial video stream, and be the elements in the j-th row and kth column of the phase space trajectory matrix. , These are the elements in the i-14th row and kth column of the phase space trajectory matrix corresponding to the fingertip capacitive pressure sequence, respectively.

[0034] In summary, the designed chaotic modulation formula expands the static basic features of each mode into a phase space trajectory matrix containing a delay dimension. The chaotic intensity coefficient introduced in the formula controls the expansion amplitude of the feature in phase space, the steepness coefficient adjusts the nonlinear rate of feature evolution, and the attention decay coefficient simulates the decay of the feature's influence over time. The synergistic effect of these three coefficients allows the originally isolated basic features to acquire evolutionary information on the time axis, transforming the user's instantaneous behavior during network access into a trajectory with dynamic characteristics. Based on this, when constructing the four-dimensional feature matrix through row vector outer product, an interactive mechanism of pairwise multiplication of features within the same mode is adopted. For example, the average curvature of the mouse trajectory is coupled with the speed fluctuation on the same delay dimension, or the multi-peak coefficient of the recording is interacted with the formant jump rate. This multiplication operation captures the synergistic relationship between different aspects within the same behavioral dimension, revealing higher-order behavioral patterns such as how the user's speed stability changes when the curvature fluctuation is large. Compared to conventional feature splicing or weighted summation, this embodiment transforms the originally scattered basic features into a four-dimensional tensor containing temporal evolution information and feature interaction information through two levels: phase space expansion and feature interaction. This provides input rich in structural information for subsequent chaotic fusion, enabling the final identity verification model to simultaneously explore the essential differences between adults and minors from two dimensions: the dynamic evolution of behavioral trajectories and multi-feature collaboration.

[0035] Step S104: Perform chaotic mapping expansion on the four-dimensional feature matrix to obtain a four-dimensional chaotic state tensor, slice the four-dimensional chaotic state tensor to obtain a three-dimensional sub-tensor, calculate the spectral vector of the three-dimensional sub-tensor in each dimension, concatenate the spectral vectors to obtain a concatenated vector, and stack all the concatenated vectors by row to obtain a two-dimensional fused feature matrix. Specifically, in some embodiments, each element in the four-dimensional feature matrix is ​​first used as an initial value, and an M×M×M×M sub-block is generated iteratively through a Logistic mapping. All sub-blocks are combined to obtain a four-dimensional chaotic state tensor. Then, the four-dimensional chaotic state tensor is sliced ​​along the first dimension to obtain M three-dimensional sub-tensors. Fourier transforms are performed on the three dimensions of each three-dimensional sub-tensor, and the amplitudes of the first M / 2 frequency components are taken as spectral vectors. The three spectral vectors are concatenated into a 3M / 2-dimensional vector. Finally, the M 3M / 2-dimensional vectors are stacked row-wise to obtain a two-dimensional fusion feature matrix with M rows and 3M / 2 columns. For example, M can be 32.

[0036] In summary, in the context of online identity verification, complex interrelationships exist between mouse, voice, gaze points, and keystroke pressure signals generated during user operations, including hand-eye coordination and audiovisual collaboration. For example, when a minor clicks on a target, their gaze often lags behind the mouse, and their vocal cords may exhibit tension fluctuations at the moment of keystroke. These interrelationship patterns are sensitive indicators of the overall level of neurodevelopment. However, conventional linear fusion methods can only process information independent of each modality and cannot capture such cross-modal collaborative relationships. This embodiment iteratively expands each value in the four-dimensional feature matrix using Logistic chaotic mapping. By leveraging the sensitivity of chaotic systems to initial values, the weak interrelationships originally dispersed in the four independent modalities are fully interwoven in a high-dimensional chaotic space. This process is not a simple signal amplification but rather allows the features of different modalities to mutually influence and modulate each other during chaotic iteration. Ultimately, each element in the resulting four-dimensional chaotic tensor contains comprehensive information about multimodal collaboration. Then, the spectral components of each dimension are extracted by Fourier transform, and the complex nonlinear coupling mode in the chaotic tensor is transformed into a quantifiable spectral vector. The main frequency component in the spectrum corresponds to the most stable cooperative rhythm of the user during operation, while the high frequency component reflects the neural control precision during fine-tuning.

[0037] Step S105: Train the initial multimodal fusion identity audit model based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model, obtain the target fusion feature matrix of the unknown identity user, and input the target fusion feature matrix into the trained multimodal fusion identity audit model to obtain the identity audit result.

[0038] In this step, an initial multimodal fusion identity audit model is constructed based on the following formula: ; in, The probability is that it is an adult. , which are the elements in the i-th row and j-th column and the m-th row and n-th column of the two-dimensional fusion feature matrix, respectively, and z is the eigenvalue; Define a label for each known user: 1 for an adult and 0 for a minor. The training process uses gradient descent and cross-entropy loss function as the target. Iteratively optimize all parameters in the model. When the loss function converges, the trained multimodal fusion identity audit model is obtained. If the output probability of the trained multimodal fusion identity verification model is greater than the preset probability threshold, the identity verification result is an adult. If the output probability of the trained multimodal fusion identity verification model is less than or equal to the preset probability threshold, the identity verification result is a minor.

[0039] In summary, the two-dimensional fusion feature matrix generated through the preceding steps contains high-order coupled information for each element, formed by phase space reconstruction, chaotic mapping expansion, and spectral analysis of multiple modal features. The magnitude relationships and distribution patterns among the elements in the matrix inherently contain the essential differences between adult and minor groups. Therefore, without the need for complex multi-layer neural networks, the high-order discriminative information in the matrix can be aggregated into a single feature value simply through global mean removal and summation operations. When most elements in the fusion feature matrix are significantly higher or lower than the global average level, the aggregated z-value will exhibit a systematic shift, thereby outputting an accurate audit probability through the Sigmoid function.

[0040] Based on the aforementioned multimodal data fusion-based network access identity verification method, this invention extracts multi-dimensional behavioral features from four modalities: mouse movement trajectory, audio recording, facial video stream, and fingertip capacitive pressure sequence. These features are then fused through multiple levels, including phase space reconstruction, chaotic mapping expansion, and spectral analysis, before being input into the identity verification model to output verification results. This method utilizes multimodal behavioral data naturally generated during the user's network access process to construct behavioral features. Because behavioral patterns are difficult to imitate and replicate, it effectively prevents identity theft. Simultaneously, through multimodal fusion, it deeply extracts and fully amplifies the comprehensive differences between minors and adults in terms of operating habits, vocalization, visual attention, and keystroke mechanics, significantly improving recognition accuracy and solving the problems of low accuracy and susceptibility to bypassing traditional single-modal or static feature recognition methods.

[0041] like Figure 2 As shown, one embodiment of the present invention proposes an online identity verification system based on multimodal data fusion, the system comprising: The modal data acquisition module 10 is used to collect multiple modal data of multiple known users during the network access process. The multiple modal data includes mouse movement trajectory, audio recording, facial video stream in front of the screen, and fingertip capacitive pressure sequence of the contact surface when typing on the keyboard. The basic feature extraction module 20 is used to extract features for each type of modality data to obtain a basic feature set corresponding to each type of modality data. Each basic feature set contains multiple basic features. The feature reconstruction module 30 is used to reconstruct the basic feature set of each mode into a phase space trajectory matrix, and generate a four-dimensional feature matrix based on the phase space trajectory matrix; The feature fusion module 40 is used to perform chaotic mapping expansion on the four-dimensional feature matrix to obtain a four-dimensional chaotic state tensor, slice the four-dimensional chaotic state tensor to obtain a three-dimensional sub-tensor, calculate the spectral vector of the three-dimensional sub-tensor in each dimension, concatenate the spectral vectors to obtain a concatenated vector, and stack all the concatenated vectors by row to obtain a two-dimensional fusion feature matrix. The training module 50 is used to train the initial multimodal fusion identity audit model based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model, obtain the target fusion feature matrix of the unknown identity user, and input the target fusion feature matrix into the trained multimodal fusion identity audit model to obtain the identity audit result.

[0042] In another aspect, the present invention also proposes a storage medium on which one or more programs are stored, which, when executed by a processor, implement the above-described method for network access identity verification based on multimodal data fusion.

[0043] In another aspect, the present invention also proposes an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to realize the above-mentioned network access identity verification method based on multimodal data fusion.

[0044] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0045] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0046] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0047] While embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations fall within the scope and spirit of the invention as set forth in the claims. Furthermore, the invention described herein may have other embodiments and can be implemented or carried out in various ways.

Claims

1. A network access identity verification method based on multimodal data fusion, characterized in that, The method includes: Collect multiple modal data of multiple known users during the network access process. The multiple modal data includes mouse movement trajectory, audio recording, facial video stream in front of the screen, and fingertip capacitive pressure sequence when typing on the keyboard. Feature extraction is performed on each modality of data to obtain a basic feature set corresponding to each modality of data. Each basic feature set contains multiple basic features. The basic feature sets of each mode are reconstructed into a phase space trajectory matrix, and a four-dimensional feature matrix is ​​generated based on the phase space trajectory matrix; The four-dimensional feature matrix is ​​expanded by chaotic mapping to obtain a four-dimensional chaotic state tensor. The four-dimensional chaotic state tensor is then sliced ​​to obtain a three-dimensional sub-tensor. The spectral vectors of the three-dimensional sub-tensor in each dimension are calculated, and the spectral vectors are concatenated to obtain a concatenated vector. All the concatenated vectors are stacked row by row to obtain a two-dimensional fused feature matrix. The initial multimodal fusion identity audit model is trained based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model. The target fusion feature matrix of the unknown user is obtained and input into the trained multimodal fusion identity audit model to obtain the identity audit result.

2. The network access identity verification method based on multimodal data fusion according to claim 1, characterized in that, The steps for feature extraction of the mouse movement trajectory include: The mouse movement trajectory includes the screen coordinates of the mouse at multiple consecutive sampling points, the timestamp of reaching each sampling point, the timestamp of each screen click during each click task, and the time interval between adjacent sampling points is equal. The base quantity is obtained using the following formula: ; in, Let be the instantaneous velocities of the i-th and (i+1)-th sampling points, respectively. Let be the screen coordinates of the i-th and (i+1)-th sampling points, respectively. Let be the instantaneous acceleration and instantaneous curvature of the i-th sampling point, respectively. The basic feature set corresponding to the mouse movement trajectory is calculated using the following formula. : ; in, For the mean curvature, The rate of change of curvature, For maximum curvature, The proportion of abnormal curvature points, For velocity fluctuation, During the reaction, For mouse click precision, The total number of sampling points. Let be the instantaneous curvature of the (i+1)th sampling point. This represents the number of all sampling points whose instantaneous curvature is greater than a first preset threshold. The time interval between adjacent sampling points. This is the timestamp of the successful mouse click on the target during the j-th click of the task. Let j be the timestamp when the target appears during the j-th click on the task. The standard duration for the pre-set j-th click task, This represents the total number of click tasks. Each click task contains a circular target, and the next target only appears after each successful click. Let J represent the screen coordinates at which the mouse successfully clicks the target on the j-th click. Let the screen coordinates be the center of the target for the j-th click task. Let be the target radius of the j-th click task.

3. The network access identity verification method based on multimodal data fusion according to claim 2, characterized in that, The steps for feature extraction from the recording include: The basic feature set corresponding to the recording is obtained using the following formula. : ; in, For multimodal coefficients, The resonance peak jump rate, Harmonic fracture index, For the frequency of voice interruption; The recording is obtained by the user reading a preset text during the network access process. The recording is divided into several frames, and the frames are windowed and silent frames are removed to obtain an effective frame set. Calculate the autocorrelation function for each valid frame in the set of valid frames: ; in, For the t-th valid frame in time delay The autocorrelation value under the following conditions , They are the m-th and m-th frames in the t-th valid frame, respectively. One signal, Total number of sampling points for each valid frame; Within a preset fundamental frequency range, scan all autocorrelation values ​​for each valid frame. If... and and Then determine The first local peak value is calculated by summing all first local peak values ​​from all valid frames. ; in, The time delay of the t-th valid frame is respectively The autocorrelation value under the following conditions They are the 1st, 2nd, and 3rd respectively. The first local peak For the first local peak set, It is a constant. This represents the maximum value of the autocorrelation. The preset fundamental frequency range is divided into several fundamental frequency intervals, and all first local peaks are assigned to their corresponding fundamental frequency intervals according to the frequency of each first local peak. and and Then determine This is the second local peak value; in, These represent the number of first local peaks falling into the i-th, (i-1), and (i+1)-th intervals, respectively. It is a constant. The maximum number; If multiple second local peaks exist, the multipeak coefficient is calculated using the following formula: ; in, This represents the total number of the second local peaks. These are the average frequencies of the m-th and n-th fundamental frequency intervals, respectively. For the m-th and n-th second local peaks, This is the difference between the upper and lower limits of the preset baseband range.

4. The network access identity verification method based on multimodal data fusion according to claim 3, characterized in that, The step of feature extraction for the recording also includes: The resonance peak jump rate is calculated using the following formula: ; in, The total number of valid frames. These are the first formant frequencies of the (t+1)th and tth valid frames, respectively. To preset the jump threshold, This is an indicator function; it returns 1 if true and 0 if false. The harmonic fracture index is calculated using the following formula: ; in, Let be the energies of the (p+1)th, (p-1)th, and (p)th harmonics of the t-th valid frame, respectively. The third preset threshold, For the FFT spectrum of the q-th valid frame, The fourth preset threshold, Let be the fundamental frequency of the t-th valid frame; The frequency of vocal interruptions can be calculated using the following formula: ; in, All are constants. These are the short-time energies of the t-th and t+1-th valid frames, respectively.

5. The network access identity verification method based on multimodal data fusion according to claim 4, characterized in that, The steps for feature extraction of the facial video stream in front of the screen include: The facial video stream includes consecutive frames of facial images. Each frame of the facial image is detected based on a trained eye detection model to identify the left and right eye regions in each frame of the facial image. Obtain the center coordinates of the left and right eye regions respectively, and then obtain the fixation point coordinates based on the center coordinates of the left and right eye regions: ; The basic feature set corresponding to the facial video stream is extracted using the following formula. : ; in, The total distance moved by the fixation point. This represents the average velocity of the gaze point. For the coverage area of ​​the fixation point, This represents the total number of frames in the facial image. Let x and y be the x-coordinates of the gaze points corresponding to the face images in frame t+1 and frame t, respectively. Let be the ordinates of the gaze points corresponding to the face images in frame t+1 and frame t, respectively. The time interval between adjacent frames of face images. These represent the maximum and minimum values ​​of the x-coordinate of the fixation point, respectively. These represent the maximum and minimum values ​​of the ordinate of the fixation point, respectively. Let x and y be the x-coordinates of the centers of the right and left eye regions in the t-th frame of the face image, respectively. , y and y are the center ordinates of the right and left eye regions in the face image of frame t, respectively.

6. The network access identity verification method based on multimodal data fusion according to claim 5, characterized in that, The steps for feature extraction of the fingertip capacitive pressure sequence of the contact surface when the keyboard is tapped include: The basic feature set corresponding to the fingertip capacitive pressure sequence is obtained using the following formula. : ; in, This is the peak pressure asymmetry index. For the press-release phase difference, For waveform steepness, The keystroke rhythm entropy, This is due to the pressure decay memory effect. This represents the total number of keystrokes. , These are the peak capacitances for the (k+1)th and kth keystrokes, respectively. These are the rise time and fall time for the k-th keystroke contact, respectively. This is the set of all sampling points within a 5ms window before and after the peak capacitance of the k-th keystroke. The capacitance value at the nth sampling point after the kth keystroke. Let be the window width of the set of all sampling points within a 5ms window before and after the peak capacitance of the k-th keystroke. Let be the normalized probability of the i-th keystroke. This is the average of the peak capacitance of all keystrokes; The peak capacitance is calculated using the following formula: ; Obtain the time interval between the peak capacitance of any adjacent keystrokes, and count the number of occurrences of each time interval. Use the ratio of the number of occurrences to the total number of time intervals as the normalized probability of the corresponding time interval. The rise time is the time from the start of the rise to the peak capacitance point, and the fall time is the time from the peak capacitance point to the end of the fall. ; in, Let the peak capacitance point, the start point of the rise, and the end point of the fall be the k-th keystroke. It is a constant.

7. The network access identity verification method based on multimodal data fusion according to claim 6, characterized in that, The steps of reconstructing the basic feature sets of each mode into a phase space trajectory matrix and generating a four-dimensional feature matrix based on the phase space trajectory matrix include: Reconstruct according to the following formula: ; in, The element in the i-th row and j-th column of the phase space trajectory matrix corresponding to the m-th mode. For the i-th basic feature of the basic feature set corresponding to the m-th modality, , , Let be the chaos intensity coefficient, development steepness coefficient, and attention decay coefficient of the i-th basic feature in the basic feature set corresponding to the m-th modality, respectively. Let be the standard deviation and skewness coefficient of the i-th basic feature of the basic feature set corresponding to the m-th modality in the known underage user group, respectively. These are the mean standard deviation and mean skewness coefficient of all basic features of the basic feature set corresponding to the m-th modality in the known underage user group, respectively; The element in the i-th row, j-th column, and k-th layer of the four-dimensional feature matrix is ​​obtained using the following formula. : When i = 1 to 7 ; When i = 8 to 11 ; When i = 12 to 14 ; When i = 14 to 19 ; in, , These are the elements in the i-th row and k-th column, and the j-th row and k-th column, respectively, of the phase space trajectory matrix corresponding to the mouse movement trajectory. , These are the elements in the (i-7)th row and kth column of the phase space trajectory matrix corresponding to the recording, respectively. , Let be the elements in the i-11th row and kth column of the phase space trajectory matrix corresponding to the facial video stream, and be the elements in the j-th row and kth column of the phase space trajectory matrix. , These are the elements in the i-14th row and kth column of the phase space trajectory matrix corresponding to the fingertip capacitive pressure sequence, respectively.

8. The network access identity verification method based on multimodal data fusion according to claim 7, characterized in that, The steps of performing chaotic mapping expansion on the four-dimensional feature matrix to obtain a four-dimensional chaotic state tensor, slicing the four-dimensional chaotic state tensor to obtain a three-dimensional sub-tensor, calculating the spectral vectors of the three-dimensional sub-tensor in each dimension, concatenating the spectral vectors to obtain a concatenated vector, and stacking all the concatenated vectors row-wise to obtain a two-dimensional fused feature matrix include: Using each element in the four-dimensional feature matrix as an initial value, an M×M×M×M sub-block is generated iteratively through the Logistic mapping. All sub-blocks are combined to obtain a four-dimensional chaotic state tensor. Slice the four-dimensional chaotic tensor along the first dimension to obtain M three-dimensional subtensors. Perform Fourier transform on the three dimensions of each three-dimensional subtensor, take the amplitude of the first M / 2 frequency components as the spectral vector, and concatenate the three spectral vectors into a 3M / 2-dimensional vector. Stack the M 3M / 2 dimensional vectors by rows to obtain a two-dimensional fusion feature matrix with M rows and 3M / 2 columns.

9. The network access identity verification method based on multimodal data fusion according to claim 8, characterized in that, The steps of training the initial multimodal fusion identity audit model based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model, obtaining the target fusion feature matrix of the unknown identity user, and inputting the target fusion feature matrix into the trained multimodal fusion identity audit model to obtain the identity audit result include: The initial multimodal fusion identity audit model is constructed based on the following formula: ; in, The probability is that it is an adult. , which are the elements in the i-th row and j-th column and the m-th row and n-th column of the two-dimensional fusion feature matrix, respectively, and z is the eigenvalue; Define a label for each known user: 1 for an adult and 0 for a minor. The training process uses gradient descent and cross-entropy loss function as the target. Iteratively optimize all parameters in the model. When the loss function converges, the trained multimodal fusion identity audit model is obtained. If the output probability of the trained multimodal fusion identity verification model is greater than the preset probability threshold, the identity verification result is an adult. If the output probability of the trained multimodal fusion identity verification model is less than or equal to the preset probability threshold, the identity verification result is a minor.

10. A network access identity verification system based on multimodal data fusion, characterized in that, The system includes: The modal data acquisition module is used to collect multiple modal data of multiple known users during the network access process. The multiple modal data include mouse movement trajectory, audio recording, facial video stream in front of the screen, and fingertip capacitive pressure sequence when typing on the keyboard. The basic feature extraction module is used to extract features from each type of modality data to obtain a basic feature set corresponding to each type of modality data. Each basic feature set contains multiple basic features. The feature reconstruction module is used to reconstruct the basic feature set of each mode into a phase space trajectory matrix, and generate a four-dimensional feature matrix based on the phase space trajectory matrix; The feature fusion module is used to perform chaotic mapping expansion on the four-dimensional feature matrix to obtain a four-dimensional chaotic state tensor, slice the four-dimensional chaotic state tensor to obtain a three-dimensional sub-tensor, calculate the spectral vector of the three-dimensional sub-tensor in each dimension, concatenate the spectral vectors to obtain a concatenated vector, and stack all the concatenated vectors row by row to obtain a two-dimensional fused feature matrix. The training module is used to train the initial multimodal fusion identity audit model based on the two-dimensional fusion feature matrix to obtain the trained multimodal fusion identity audit model, obtain the target fusion feature matrix of the unknown identity user, and input the target fusion feature matrix into the trained multimodal fusion identity audit model to obtain the identity audit result.