Identity recognition method and system based on multi-modal data

By using multimodal data recognition methods, combining sensor motion trajectories, touch screen interactions, and voiceprint features, a non-biometric recognition system is constructed. This solves the problems of low recognition accuracy under lighting and occlusion conditions, as well as insufficient legal compliance, achieving improved accuracy and privacy protection across all scenarios.

CN121744293APending Publication Date: 2026-03-27LUZHOU VOCATIONAL & TECHN COLLEGE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing identity authentication systems suffer from low recognition accuracy under lighting conditions and when wearing obstructions, insufficient legal compliance, and poor security in biometric database storage, leading to high false recognition rates and privacy threats.

Method used

A multimodal data recognition method is adopted, which combines sensor motion trajectory, touch screen interaction and voiceprint features to construct a non-biological recognition system. Dynamic fusion recognition is achieved by capturing device motion inertia by sensors, modeling touch screen trajectory and abstracting voiceprint spectrum.

Benefits of technology

Reduce false recognition rates in low-light and noisy environments, avoid legal compliance risks, eliminate privacy litigation risks, and improve recognition accuracy across all scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744293A_ABST
    Figure CN121744293A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of identity recognition, in particular to an identity recognition method and system based on multi-modal data. The invention provides an identity recognition method based on multi-modal data, which adopts a user use dimension, a user interaction dimension and a user voiceprint dimension, thoroughly avoids legal compliance risks caused by a face recognition technology, constructs a non-biological recognition system based on sensor motion trail capture, touch screen interaction behavior modeling and voiceprint feature extraction, and improves the recognition efficiency. Therefore, all the feature data cannot restore the original biological attributes, and the hidden danger of privacy litigation is eliminated from the source. Meanwhile, depending on a three-dimensional dynamic fusion mechanism of the sensor, the touch screen and the voiceprint, full-scene precision jump, compensation of the dark light defect by the motion inertia of the sensor capture equipment, noise interference resistance by a touch screen track modeling fine operation habit and physiological uniqueness abstract representation of the voiceprint spectrum are realized, and the three forms a complementary enhancement effect through dynamic weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of identity recognition technology, specifically to an identity recognition method and system based on multimodal data. Background Technology

[0002] With the widespread application of artificial intelligence technology in the field of identity authentication, facial recognition-based solutions have become the mainstream technological approach. These systems collect users' facial biometric features and combine them with deep learning models to achieve high-precision identity verification, leading to their rapid adoption in scenarios such as financial payments and access control. However, existing technologies face a dual challenge: technically, facial recognition heavily relies on lighting conditions and shooting angles, with a false recognition rate (FRR) exceeding 35% in low-light environments and a performance drop by more than 50% when masks or hats are worn; legally, the implementation of the "Measures for the Security Management of Facial Recognition Technology Applications" has led to certain legal risks associated with using facial recognition data obtained in public places for identity verification.

[0003] Meanwhile, the current identity authentication system suffers from structural flaws: a single biometric recognition mode struggles to balance security and user experience. Voiceprint recognition performs well in quiet environments, but its false recognition rate deteriorates sharply under noise interference; while behavioral feature recognition avoids the risks associated with biometric data, it is limited by short-term fluctuations in user habits, requiring frequent calibration and increasing maintenance burden. More seriously, the storage security of biometric databases raises systemic concerns—once feature data is leaked, it cannot be reset, posing a permanent privacy threat.

[0004] The tightening policy environment is exacerbating the pressure for technological transformation. Regulators have explicitly mandated the use of non-biometric solutions, and mandatory facial recognition in public places is prohibited. The legality of existing systems in scenarios such as finance and security is questionable, and the industry urgently needs next-generation solutions. Summary of the Invention

[0005] The purpose of this invention is to provide an identity recognition method and system based on multimodal data to address the aforementioned problems.

[0006] The technical solution adopted in this invention is as follows: an identity recognition method based on multimodal data, suitable for contactless identity authentication scenarios, comprising the following steps:

[0007] During the seamless identity authentication process, multimodal data is collected, which includes three dimensions: user usage dimension, user interaction dimension, and user voiceprint dimension.

[0008] Preprocess the multimodal data;

[0009] Feature extraction is performed on multimodal data to obtain user usage dimension feature values, user interaction dimension feature values, and user voiceprint dimension feature values;

[0010] Multimodal feature values ​​are fused by linearly interpolating and synchronizing user usage dimension feature values ​​and user interaction dimension feature values ​​over time to obtain behavioral feature values. User voiceprint dimension feature values ​​are then aligned with behavioral feature values.

[0011] Build a seamless identity authentication and recognition decision.

[0012] Furthermore, the user-used data includes sensor data derived from microelectronic components built into the mobile device, including a three-axis accelerometer, gyroscope, and gravity sensor.

[0013] The triaxial accelerometer is used to acquire the linear acceleration of the user using the mobile device in the X, Y, and Z axis directions;

[0014] The gyroscope is used to obtain the rotational angular velocity of the user's mobile device in the X, Y, and Z axis directions;

[0015] The gravity sensor is used to obtain the direction vector of the user's mobile device relative to the direction of Earth's gravity.

[0016] Furthermore, the user interaction dimension data includes touch screen exchange data, which originates from microelectronic components built into the mobile device, including the capacitive touch screen module of the mobile device.

[0017] The capacitive touchscreen module of the mobile device is used to acquire the swiping trajectory, click events and pressure applied by the user when using the mobile device.

[0018] Furthermore, the user voiceprint dimension data includes voiceprint data derived from microelectronic components built into the mobile device, including the mobile device microphone;

[0019] The mobile device microphone is used for user voiceprint recognition after the user speaks a preset identity recognition activation word.

[0020] Furthermore, in the preprocessing of multimodal data, high-frequency noise suppression is performed on the data of the user-used dimension for subsequent acquisition of stable motion features;

[0021] Device size differences are eliminated from user interaction data to obtain standardized interaction features later.

[0022] Environmental noise is suppressed on the user's voiceprint data to preserve the user's voice biometric features in the future.

[0023] Furthermore, in the feature extraction of multimodal data, for the user usage dimension data, the long-term temporal dependency of device motion is captured from stable motion features, and the user's unique device holding habits are output.

[0024] Based on user interaction data, short-term behavioral patterns of user interaction with touchscreen are constructed from standardized interaction features;

[0025] For user voiceprint data, residual connection suppression training is performed on user voice biometrics.

[0026] Furthermore, in the process of fusing multimodal feature values, the user usage dimension features, user interaction dimension features, and user voiceprint dimension features output after feature extraction are fused to obtain a spliced ​​high-dimensional fused feature.

[0027] Furthermore, the process of fusing multimodal feature values ​​also includes dimensionality reduction of high-dimensional fused features and outputting dimensionality-reduced fused features.

[0028] Furthermore, in the process of constructing the seamless identity authentication and recognition decision, the dimensionality-reduced fusion features are compared with the user features stored in the database, and a judgment is made on whether the user is a legitimate user based on a set threshold.

[0029] The present invention also provides an identity recognition system based on multimodal data, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the aforementioned identity recognition method based on multimodal data.

[0030] The beneficial effects of the present invention include at least one of the following;

[0031] 1. A multimodal data-based identity recognition method is provided, which adopts user usage dimension, user interaction dimension, and user voiceprint dimension to completely avoid the legal compliance risks caused by facial recognition technology. Based on sensor motion trajectory capture, touch screen interaction behavior modeling, and voiceprint feature extraction, a non-biological recognition system is constructed, making it impossible to restore the original biological attributes of all feature data, thus escaping the regulatory scope of "facial information" in the "Measures for the Safety Management of Facial Recognition Technology Applications" and eliminating the risk of privacy litigation from the source.

[0032] 2. Simultaneously, relying on the three-dimensional dynamic fusion mechanism of sensors, touch screen and voiceprint, the false recognition rate is reduced by utilizing the device motion characteristics in low-light environments, and the false recognition rate is reduced by increasing the touch screen interaction weight in noisy scenes, thus achieving a leap in accuracy across all scenarios. The essence of the technology lies in the sensor capturing the device motion inertia to compensate for the defects in low light, the touch screen trajectory modeling to refine operating habits to resist noise interference, and the voiceprint spectrum abstractly representing physiological uniqueness. The three form a complementary and enhanced effect through dynamic weighting. Attached Figure Description

[0033] Figure 1 Here is a flowchart of an identity recognition method based on multimodal data;

[0034] Figure 2 This is a schematic diagram of the structure of an electronic device. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described in the accompanying drawings can generally be arranged and designed in various different configurations.

[0036] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0037] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.

[0038] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0039] like Figure 1 As shown, an identity recognition method based on multimodal data, suitable for contactless identity authentication scenarios, includes the following steps:

[0040] During the seamless identity authentication process, multimodal data is collected, which includes three dimensions: user usage dimension, user interaction dimension, and user voiceprint dimension.

[0041] Preprocess the multimodal data;

[0042] Feature extraction is performed on multimodal data to obtain user usage dimension feature values, user interaction dimension feature values, and user voiceprint dimension feature values;

[0043] Multimodal feature values ​​are fused by linearly interpolating and synchronizing user usage dimension feature values ​​and user interaction dimension feature values ​​over time to obtain behavioral feature values. User voiceprint dimension feature values ​​are then aligned with behavioral feature values.

[0044] Build a seamless identity authentication and recognition decision.

[0045] The purpose of this design is to provide an identity recognition method based on multimodal data. It employs user usage, user interaction, and user voiceprint dimensions to completely circumvent the legal compliance risks associated with facial recognition technology. By constructing a non-biological identification system based on sensor motion trajectory capture, touchscreen interaction behavior modeling, and voiceprint feature extraction, all feature data cannot be reconstructed to reveal original biological attributes, thus escaping the regulatory scope of "facial information" under the "Measures for the Safety Management of Facial Recognition Technology Applications," eliminating privacy litigation risks at the source. Simultaneously, relying on a three-dimensional dynamic fusion mechanism of sensors, touchscreen, and voiceprint, it reduces the false recognition rate in low-light environments by utilizing device motion characteristics and in noisy scenarios by increasing the weight of touchscreen interaction, achieving a leap in accuracy across all scenarios. The technical essence lies in the sensor capturing device motion inertia to compensate for low-light deficiencies, touchscreen trajectory modeling to refine operating habits and resist noise interference, and voiceprint spectrum abstraction to represent physiological uniqueness. These three elements form a complementary and enhancing effect through dynamic weighting.

[0046] In this embodiment, the data used by the user includes sensor data, which comes from microelectronic components built into the mobile device, including a three-axis accelerometer, a gyroscope, and a gravity sensor.

[0047] The triaxial accelerometer is used to acquire the linear acceleration of the user using the mobile device in the X, Y, and Z axis directions;

[0048] The gyroscope is used to obtain the rotational angular velocity of the user's mobile device in the X, Y, and Z axis directions;

[0049] The gravity sensor is used to obtain the direction vector of the user's mobile device relative to the direction of Earth's gravity.

[0050] Meanwhile, user interaction data includes touch screen exchange data, which comes from microelectronic components built into mobile devices, including capacitive touch screen modules of mobile devices.

[0051] The capacitive touchscreen module of the mobile device is used to acquire the swiping trajectory, click events and pressure applied by the user when using the mobile device.

[0052] Furthermore, the user voiceprint dimension data includes voiceprint data, which comes from the microelectronic components built into the mobile device, including the mobile device microphone;

[0053] The mobile device microphone is used for user voiceprint recognition after the user speaks a preset identity recognition activation word.

[0054] It should be noted that, in practice, sensor data comes directly from microelectronic components built into mobile devices, including:

[0055] Accelerometer: Measures the linear acceleration of a device along the X / Y / Z axes, capturing the dynamic behavior of users holding and moving devices (such as raising or lowering their hands).

[0056] Gyroscope: Detects the angular velocity of a device's rotation around the X / Y / Z axes, reflecting the user's habitual rotation of the device (such as tilting the screen).

[0057] Gravity sensor: Provides a direction vector relative to Earth's gravity, helping to identify the device's posture (such as landscape / portrait).

[0058] The trigger condition for data collection is that it is automatically activated when the user initiates an action (such as a payment request), and data collection continues for a set period of time, such as 2 seconds or 3 seconds.

[0059] Accelerometers at 100 Hz (one point every 10 ms) and gyroscopes at 50 Hz (one point every 20 ms) ensure high-precision time series coverage.

[0060] After data acquisition, high-frequency noise (such as equipment jitter) can be removed using low-pass filters to ensure smooth data. Normalization processing will standardize the original data and eliminate or reduce the influence of equipment differences and units.

[0061] As for touch screen interaction data, collection begins the moment the user's finger touches the screen and ends when the finger leaves the screen or when the collection continues for a preset period of time. The collected data mainly includes swipe trajectory, click events, and pressure intensity. The sampling frequency for swipe trajectory and pressure intensity is 100Hz, while click events need to be triggered by the event itself.

[0062] This approach collects user habits regarding touchscreen usage to create a unique identity verification method, especially useful in scenarios involving financial payments.

[0063] For voiceprint data, the user speaks a preset activation word, such as a payment verification or similar activation word, to trigger the collection. The voiceprint collection time is 2 seconds. Then, the existing VAD algorithm is used to segment the collected voiceprint segments, and then STFT conversion is performed to generate the spectrum.

[0064] Meanwhile, in this embodiment, during the preprocessing of multimodal data, high-frequency noise suppression is performed on the data of the user-used dimension for subsequent acquisition of stable motion features;

[0065] Device size differences are eliminated from user interaction data to obtain standardized interaction features later.

[0066] Environmental noise is suppressed on the user's voiceprint data to preserve the user's voice biometric features in the future.

[0067] Then, in the feature extraction of multimodal data, for the user usage dimension data, the long-term temporal dependency of device motion is captured from stable motion features, and the user's unique device holding habits are output.

[0068] Based on user interaction data, short-term behavioral patterns of user interaction with touchscreen are constructed from standardized interaction features;

[0069] For user voiceprint data, residual connection suppression training is performed on user voice biometrics.

[0070] Next, in the process of fusing multimodal feature values, the user usage dimension feature, user interaction dimension feature and user voiceprint dimension feature output after feature extraction are fused to obtain a spliced ​​high-dimensional fused feature.

[0071] Finally, the process of fusing multimodal feature values ​​also includes dimensionality reduction of high-dimensional fused features and outputting dimensionality-reduced fused features.

[0072] In practical applications, the following formula can be used for sensor data preprocessing:

[0073] out t =β⋅in t +(1−β)⋅out t−1

[0074] Among them in t This refers to the original input value at the current time step, i.e., the original sensor data.

[0075] out t−1 This is the filtered output value from the previous time step, i.e., the historical smoothing result;

[0076] β is the filter coefficient, used to control the weight of the new and old data;

[0077] out t This is the filtered output value at the current time step, i.e., the data after denoising.

[0078] In multimodal feature extraction, bidirectional LSTM is used for sensor data extraction. The input is a 128-dimensional feature vector, and the output is a 256-dimensional feature vector based on forward and backward propagation.

[0079] For touch screen data, a bidirectional GRU extraction method is used. The input is a 64-dimensional feature vector, which is processed by the update gate and the reset gate to output a 128-dimensional feature vector. The GRU gating mechanism efficiently captures short-term behavior and enhances the anti-counterfeiting capability of the pressure-sensitive layer.

[0080] The voiceprint data is extracted using ResNet-34. The input is a 256-dimensional feature vector, which is converted from a single channel to a dual channel by ResNet-34, and the output is a 512-dimensional feature vector.

[0081] Next, the data from the three dimensions—256 dimensions from the sensor data output, 128 dimensions from the touchscreen data output, and 512 dimensions from the voiceprint data output—is concatenated to obtain 896 dimensions. Then, dimensional compression is performed to improve overall computational efficiency, outputting 128-dimensional vector data using the following formula:

[0082] o=ReLU(W o ⋅c+b o )

[0083] Among them W o Let b be the weight matrix. o Here, is the bias vector, and ReLU is the activation function.

[0084] Finally, a seamless identity authentication and recognition decision is constructed. In specific implementation, the compressed output vector is compared with the user registration template data stored in the database by comparing the cosine similarity. When the similarity is greater than or equal to the upper limit of the set threshold, the user is determined to be a legitimate user. When the similarity is within the threshold, secondary verification is required. When the similarity is at the lower limit of the threshold, the user is determined to be an illegitimate user.

[0085] This application provides a schematic diagram of the structure of an electronic device. The electronic device includes a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions stored in the memory to implement the multimodal data-based identity recognition method described in any of the above embodiments.

[0086] In one embodiment of this application, the electronic device further includes a bus and a computer program stored in the memory and executable on the processor, such as an identity recognition program based on multimodal data.

[0087] Figure 2 Only an electronic device with memory and processor is shown. Those skilled in the art will understand that the structure shown does not constitute a limitation on the electronic device and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0088] Combination Figure 2 The memory in the electronic device stores a plurality of computer-readable instructions to implement an identity recognition method based on multimodal data, and the processor can execute the plurality of instructions to implement the method.

[0089] Specifically, the processor's implementation method for the above instructions can be found in the description of the relevant steps in the corresponding embodiment of the figure, and will not be repeated here.

[0090] Those skilled in the art will understand that the schematic diagram is merely an example of an electronic device and does not constitute a limitation on the electronic device. The electronic device may be a bus-type structure or a star-type structure. The electronic device may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, the electronic device may also include input / output devices, network access devices, etc.

[0091] It should be noted that electronic devices are merely examples. Other existing or future electronic products that are suitable for this application should also be included within the scope of protection of this application and are incorporated herein by reference.

[0092] The memory includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, portable hard drives, multimedia cards, card-type memory (e.g., SD or DX memory), magnetic storage, magnetic disks, optical disks, etc. In some embodiments, the memory can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory can be an external storage device of the electronic device, such as a plug-in portable hard drive, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory can be used not only to store application software and various types of data installed on the electronic device, such as identification codes based on multimodal data, but also to temporarily store data that has been output or will be output.

[0093] In some embodiments, a processor can be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions. This includes combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor is the control unit of an electronic device, connecting various components of the device through various interfaces and lines. It performs various functions and processes data by running or executing programs or modules stored in the memory (e.g., executing a multimodal data-based identification program) and calling data stored in the memory.

[0094] The processor executes the operating system of the electronic device and various installed applications. The processor executes the applications to implement the steps in each of the above embodiments of the identity recognition method based on multimodal data, such as the steps shown in the figure.

[0095] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in an electronic device. For example, the computer program may be divided into a receiving module, a preprocessing module, a projection module, and a determining module.

[0096] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the multimodal data-based identity recognition method described in the various embodiments of this application.

[0097] When modules / units integrated into an electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0098] This application provides an identity recognition method based on multimodal data, which can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0099] Electronic devices can be any electronic product that allows human-computer interaction with a customer, such as personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, interactive network television (IPTV), smart wearable devices, etc.

[0100] Electronic devices may also include network devices and / or client devices. The network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0101] The networks in which electronic devices are located include, but are not limited to, the Internet, wide area networks, metropolitan area networks, local area networks, and virtual private networks (VPNs).

[0102] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, and other memory.

[0103] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0104] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus. The bus is configured to implement the connection and communication between the memory and at least one processor, etc.

[0105] This application also provides a computer-readable storage medium (not shown), which stores computer-readable instructions. These computer-readable instructions are executed by a processor in an electronic device to implement a multimodal data-based identity recognition method as described in any of the above embodiments.

[0106] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0107] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0108] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0109] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the specification may also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0110] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An identity recognition method based on multimodal data, suitable for contactless identity authentication scenarios, characterized in that, Includes the following steps: During the seamless identity authentication process, multimodal data is collected, which includes three dimensions: user usage dimension, user interaction dimension, and user voiceprint dimension. Preprocess the multimodal data; Feature extraction is performed on multimodal data to obtain user usage dimension feature values, user interaction dimension feature values, and user voiceprint dimension feature values; Multimodal feature values ​​are fused by linearly interpolating and synchronizing user usage dimension feature values ​​and user interaction dimension feature values ​​over time to obtain behavioral feature values. User voiceprint dimension feature values ​​are then aligned with behavioral feature values. Build a seamless identity authentication and recognition decision.

2. The identity recognition method based on multimodal data according to claim 1, characterized in that, The user-used data includes sensor data, which comes from microelectronic components built into mobile devices, including triaxial accelerometers, gyroscopes, and gravity sensors. The triaxial accelerometer is used to acquire the linear acceleration of the user using the mobile device in the X, Y, and Z axis directions; The gyroscope is used to obtain the rotational angular velocity of the user's mobile device in the X, Y, and Z axis directions; The gravity sensor is used to obtain the direction vector of the user's mobile device relative to the direction of Earth's gravity.

3. The identity recognition method based on multimodal data according to claim 1, characterized in that, The user interaction dimension data includes touch screen exchange data, which comes from microelectronic components built into mobile devices, including the capacitive touch screen module of the mobile device. The capacitive touchscreen module of the mobile device is used to acquire the swiping trajectory, click events and pressure applied by the user when using the mobile device.

4. The identity recognition method based on multimodal data according to claim 1, characterized in that, The user voiceprint dimension data includes voiceprint data, which comes from microelectronic components built into the mobile device, including the mobile device microphone; The mobile device microphone is used for user voiceprint recognition after the user speaks a preset identity recognition activation word.

5. The identity recognition method based on multimodal data according to claim 1, characterized in that, In the preprocessing of multimodal data, high-frequency noise suppression is performed on the data of the user-used dimension for subsequent acquisition of stable motion features; Device size differences are eliminated from user interaction data to obtain standardized interaction features later. Environmental noise is suppressed on the user's voiceprint data to preserve the user's voice biometric features in the future.

6. The identity recognition method based on multimodal data according to claim 5, characterized in that, In the feature extraction of multimodal data, for the user usage dimension data, the long-term temporal dependency of device motion is captured from stable motion features, and the user's unique device holding habits are output. Based on user interaction data, short-term behavioral patterns of user interaction with touchscreen are constructed from standardized interaction features; For user voiceprint data, residual connection suppression training is performed on user voice biometrics.

7. The identity recognition method based on multimodal data according to claim 6, characterized in that, In the process of fusing multimodal feature values, the user usage dimension feature, user interaction dimension feature and user voiceprint dimension feature output after feature extraction are fused to obtain a spliced ​​high-dimensional fused feature.

8. The identity recognition method based on multimodal data according to claim 7, characterized in that, The process of fusing multimodal feature values ​​also includes dimensionality reduction of high-dimensional fused features and outputting dimensionality-reduced fused features.

9. The identity recognition method based on multimodal data according to claim 8, characterized in that, In the process of constructing a seamless identity authentication and recognition decision, the dimensionality-reduced fusion features are compared with the user features stored in the database, and a judgment is made on whether the user is legitimate based on a set threshold.

10. An identity recognition system based on multimodal data, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements an identity recognition method based on multimodal data as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Identity recognition method based on comparative learning and multi-modal biological characteristics

    CN118349985A

  • Credit identity verification enhancement technology based on voiceprint recognition and behavior analysis

    CN121331142A

Cited By

  • An intelligent interactive system

    CN122174843A