Identity information identification method and system based on multi-modal large model
Through the identity information identification method based on the multimodal large model, the multimodal data is integrated for feature extraction and matching, which solves the shortcomings of identification accuracy and robustness in the traditional methods, and achieves more efficient identity information identification.
Patent Information
- Application Number
- CN202411886879.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional identity information recognition methods rely on single modal data and cannot integrate multimodal information such as text, images, and voice, resulting in insufficient robustness and accuracy of recognition, poor generalization ability, and difficulty in rapid expansion.
The identity information recognition method based on multimodal large model is adopted, multimodal data is collected through cameras and microphones, multimodal databases and data sets are built, multimodal large model is trained to extract multimodal fusion features, and the features are stored in the vector library, and the multimodal data is matched in real time for identity information recognition.
It improves the accuracy, robustness, generalization and scalability of identity information identification, can process multimodal data more effectively, reduce manual intervention, and improve the degree of automation and user experience.
Smart Images

Figure CN120014720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identity information recognition, and in particular to an identity information recognition method and system based on a multimodal large model. Background Art
[0002] When a user conducts company business, it is usually necessary to identify (verify) the user's identity information. When the identity information verification passes, the user will be provided with relevant business processing services, otherwise the business will not be processed.
[0003] Traditionally, a single identity verification method is usually used for identity information recognition, such as manual, OCR recognition, voiceprint recognition technology, or a combination of multiple technologies, such as voiceprint recognition and OCR recognition. However, traditional methods have the following disadvantages:
[0004] 1. The use of a single or a combination of technologies to verify identity information has some limitations and often relies on single-modal data such as pictures, and cannot integrate multi-modal information such as text, images, and voice, which limits the robustness and accuracy of identity information recognition; 2. The test results on the training set are usually good, but when facing actual scenes, factors such as lighting and shooting angle affect the quality of the pictures taken, and the identity documents are oily and worn, which will lead to errors in identity information recognition in actual scenes, resulting in the inability to recognize identity information, that is, the generalization ability of identity information recognition is poor; 3. It is often difficult to understand the relationship between the information on the certificate and other related information. When multiple technologies are used for identity information recognition, there is usually a series relationship. If one technology is incorrectly recognized, the final recognition result will be wrong, and the information correlation is relatively poor; 4. It needs to be customized and developed for different application scenarios, and its versatility is relatively low. With the development of information technology and changes in application needs, traditional methods are difficult to quickly upgrade or expand functions due to the constraints of software and hardware.
[0005] Therefore, how to provide an identity information recognition method and system based on a multimodal large model to improve the accuracy, robustness, generalization and scalability of identity information recognition has become a technical problem that needs to be solved urgently. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide an identity information recognition method and system based on a multimodal large model, so as to improve the accuracy, robustness, generalization and extensibility of identity information recognition.
[0007] In a first aspect, the present invention provides an identity information recognition method based on a multimodal large model, comprising the following steps:
[0008] Step S10: Collect a large amount of historical multimodal data including person images, document images and audio data through a camera and a microphone, and store each of the historical multimodal data in a pre-created multimodal database;
[0009] Step S20: uploading each of the historical multimodal data to a server for preprocessing to construct a data set;
[0010] Step S30: the server creates a multimodal large model for extracting multimodal fusion features, and trains the multimodal large model using the data set;
[0011] Step S40: extracting the registered multimodal fusion features of each user from the user information database through the trained multimodal large model, and storing each registered multimodal fusion feature in a pre-created vector library;
[0012] Step S50: pre-process the input real-time multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, and match the real-time modal fusion features with the registered multimodal fusion features stored in the vector library to perform identity information recognition.
[0013] Furthermore, the step S20 specifically includes:
[0014] Step S21, uploading each of the historical multimodal data to the server, the server performs image quality assessment on the person image and the document image in each of the historical multimodal data through the multi-size structural similarity index, and determines whether the image quality is greater than a preset quality threshold, if not, proceeding to step S22; if yes, proceeding to step S23;
[0015] Step S22, performing image enhancement operation on each of the person images and the document image by using an image enhancement algorithm;
[0016] Step S23, after performing data cleaning operations including deleting duplicate data and noise data on each of the historical multimodal data, rotating and scaling the person image and the document image in each of the historical multimodal data, and adding background noise to the audio data to perform a sample expansion operation;
[0017] Step S24, normalize the size of each of the character images and the document image, unify the sampling rate of each of the audio data, complete the preprocessing of each of the historical multimodal data, and construct a data set based on the preprocessed historical multimodal data.
[0018] Furthermore, the step S30 specifically includes:
[0019] Step S31: The server creates a multimodal large model for extracting multimodal fusion features;
[0020] The multimodal large model locates the face area from the character image by the Yolo algorithm, extracts face features from the face area by the Resnet algorithm, extracts text features from the document image by the OCR technology and the BERT algorithm, and extracts audio features from the audio data by the MFCC algorithm, thereby obtaining multimodal fusion features including the face features, text features and audio features;
[0021] Step S32, dividing the data set into a training set, a test set and a validation set in a ratio of 7:2:1;
[0022] Step S33, training the multimodal large model through the training set, performing a Concat operation on the extracted facial features, text features, and audio features during the training process as input Embedding vectors of the multimodal large model, and fine-tuning the multimodal large model through Adapter Tuning to optimize hyperparameters including at least a learning rate, an adapter dimension, and a number of adapter layers;
[0023] Step S34: testing the trained multimodal large model using the test set, and verifying the tested multimodal large model using the verification set.
[0024] Furthermore, the step S40 is specifically as follows:
[0025] The server deploys the trained multimodal large model, and through the deployed multimodal large model, extracts the registered multimodal fusion features of each user from the user information database storing registered multimodal data, and stores each registered multimodal fusion feature in real time to a pre-created Milvus vector library.
[0026] Furthermore, the step S50 is specifically as follows:
[0027] Collect real-time multimodal data, detect user behavior during the collection process, and perform security analysis on the user behavior. Pre-process the multimodal data and input it into the multimodal big model to obtain real-time modal fusion features. Calculate the similarity between the real-time modal fusion features and each registered multimodal fusion feature stored in the vector library, and perform identity information recognition based on the similarity.
[0028] In a second aspect, the present invention provides an identity information recognition system based on a multimodal large model, comprising the following modules:
[0029] A historical multimodal data acquisition module is used to collect a large amount of historical multimodal data including person images, document images and audio data through a camera and a microphone, and store each of the historical multimodal data in a pre-created multimodal database;
[0030] A data set construction module, used for uploading each of the historical multimodal data to a server for preprocessing to construct a data set;
[0031] A multimodal large model training module is used for the server to create a multimodal large model for extracting multimodal fusion features, and train the multimodal large model through the data set;
[0032] A registered multimodal fusion feature extraction module is used to extract the registered multimodal fusion features of each user from the user information database through the trained multimodal large model, and store each registered multimodal fusion feature in a pre-created vector library;
[0033] The identity information recognition module is used to pre-process the input real-time multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, and match the real-time modal fusion features with the registered multimodal fusion features stored in the vector library to perform identity information recognition.
[0034] Furthermore, the data set construction module specifically includes:
[0035] A data quality assessment unit, used to upload each of the historical multimodal data to a server, and the server performs image quality assessment on the person images and document images in each of the historical multimodal data through a multi-size structural similarity index to determine whether the image quality is greater than a preset quality threshold. If not, the image enhancement unit is entered; if yes, the sample expansion unit is entered;
[0036] An image enhancement unit, used to perform image enhancement operations on each of the person images and the document image by using an image enhancement algorithm;
[0037] A sample expansion unit, configured to perform a data cleaning operation including deleting duplicate data and noise data on each of the historical multimodal data, rotate and scale the person image and the document image in each of the historical multimodal data, and add background noise to the audio data to perform a sample expansion operation;
[0038] The unified processing unit is used to normalize the size of each of the character images and the document image, unify the sampling rate of each of the audio data, complete the preprocessing of each of the historical multimodal data, and construct a data set based on the preprocessed historical multimodal data.
[0039] Furthermore, the multimodal large model training module specifically includes:
[0040] A multimodal large model creation unit, used for the server to create a multimodal large model for extracting multimodal fusion features;
[0041] The multimodal large model locates the face area from the character image by the Yolo algorithm, extracts face features from the face area by the Resnet algorithm, extracts text features from the document image by the OCR technology and the BERT algorithm, and extracts audio features from the audio data by the MFCC algorithm, thereby obtaining multimodal fusion features including the face features, text features and audio features;
[0042] A data set division unit, used for dividing the data set into a training set, a test set and a validation set in a ratio of 7:2:1;
[0043] A training unit, used to train the multimodal large model through the training set, perform a Concat operation on the extracted facial features, text features, and audio features during the training process as input Embedding vectors of the multimodal large model, and fine-tune the multimodal large model through Adapter Tuning to optimize hyperparameters including at least a learning rate, an adapter dimension, and a number of adapter layers;
[0044] A testing and verification unit is used to test the trained multimodal large model through the test set, and to verify the tested multimodal large model through the verification set.
[0045] Furthermore, the registered multimodal fusion feature extraction module is specifically used for:
[0046] The server deploys the trained multimodal large model, and through the deployed multimodal large model, extracts the registered multimodal fusion features of each user from the user information database storing registered multimodal data, and stores each registered multimodal fusion feature in real time to a pre-created Milvus vector library.
[0047] Furthermore, the identity information recognition module is specifically used for:
[0048] Collect real-time multimodal data, detect user behavior during the collection process, and perform security analysis on the user behavior. Pre-process the multimodal data and input it into the multimodal big model to obtain real-time modal fusion features. Calculate the similarity between the real-time modal fusion features and each registered multimodal fusion feature stored in the vector library, and perform identity information recognition based on the similarity.
[0049] The advantages of the present invention are:
[0050] 1. Through the camera and microphone, a large amount of historical multimodal data including character images, document images and audio data is collected and uploaded to the server for preprocessing to build a data set; then the server creates a multimodal large model for extracting multimodal fusion features, and trains the multimodal large model through the data set; then, through the trained multimodal large model, the registered multimodal fusion features of each user are extracted from the user information database and stored in the vector library; finally, the input real-time multimodal data is preprocessed and input into the multimodal large model to obtain real-time modal fusion features, and the real-time modal fusion features are matched with the registered multimodal fusion features stored in the vector library to perform identity information recognition; that is, the comprehensive character image (image), document image (text) and audio data The multimodal data of (voice) is used to characterize the characteristics of user identity, so that the multimodal large model can extract rich features from multiple data sources, thereby improving the accuracy of identity information recognition; by fusing multimodal data, even when the quality of a certain modal data is poor, it can maintain a good recognition effect to improve the robustness of identity information recognition; by fusing multiple modal features (face features, text features, audio features) through the multimodal large model, the ability to construct user portraits is improved, and unknown data can be better identified, thereby improving the generalization ability of the multimodal large model; and compared with traditional methods, it is easier to integrate new modal data, which improves the scalability of identity information recognition, and ultimately greatly improves the accuracy, robustness, generalization, and scalability of identity information recognition.
[0051] 2. By combining security analysis of user behavior (behavior verification) during the identity identification process, fraud can be effectively prevented, thereby improving the security of identity identification.
[0052] 3. Use the MS-SSIM algorithm to evaluate image quality. If the image quality is lower than the threshold, use the MSR algorithm to enhance the image, thereby improving the image quality and enhancing the image analysis capability.
[0053] 4. Use Milvus to store registered multimodal fusion features. When identity information recognition is required, you can quickly query the registered multimodal fusion features corresponding to the user based on Milvus, and calculate the cosine similarity with the real-time multimodal data to improve the speed of identity information recognition.
[0054] 5. By outputting and visualizing the identity recognition results in a structured manner, it is convenient for subsequent troubleshooting of identity recognition issues and for users to understand and view them.
[0055] 6. Using Adapter Tuning training improves the ability of multimodal feature extraction and fusion feature representation, while also improving the training speed.
[0056] 7. Perform identity information recognition based on a multimodal large model to reduce manual intervention, improve the degree of automation of identity information recognition, and enhance user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present invention will be further described below in conjunction with embodiments with reference to the accompanying drawings.
[0058] Figure 1 It is a flow chart of an identity information recognition method based on a multimodal large model of the present invention.
[0059] Figure 2 It is a structural schematic diagram of an identity information recognition system based on a multimodal large model of the present invention. DETAILED DESCRIPTION
[0060] The technical solution in the embodiment of the present application has the following overall idea: the multimodal data of character images, ID images and audio data are integrated to characterize the characteristics of the user identity, so that the multimodal large model can extract rich features from multiple data sources, thereby improving the accuracy of identity information recognition; by fusing multimodal data, even when the quality of a certain modal data is poor, a good recognition effect can be maintained to improve the robustness of identity information recognition; by fusing multiple modal features through the multimodal large model, the ability to construct user portraits is improved, and unknown data can be better identified, thereby improving the generalization ability of the multimodal large model; and compared with traditional methods, it is easier to integrate new modal data, thereby improving the scalability of identity information recognition, and thus improving the accuracy, robustness, generalization and scalability of identity information recognition.
[0061] Please refer to Figure 1 to Figure 2 As shown, a preferred embodiment of the identity information recognition method based on a multimodal large model of the present invention includes the following steps:
[0062] Step S10: Collect a large amount of historical multimodal data including person images, document images and audio data through a camera and a microphone, and store each of the historical multimodal data in a pre-created multimodal database;
[0063] Step S20: uploading each of the historical multimodal data to a server for preprocessing to construct a data set;
[0064] Step S30: the server creates a multimodal large model for extracting multimodal fusion features, and trains the multimodal large model using the data set;
[0065] Step S40: extracting the registered multimodal fusion features of each user from the user information database through the trained multimodal large model, and storing each registered multimodal fusion feature in a pre-created vector library;
[0066] Step S50: pre-process the input real-time multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, and match the real-time modal fusion features with the registered multimodal fusion features stored in the vector library to perform identity information recognition.
[0067] Compared with the traditional single method, the present invention integrates multimodal data of text, image and voice to characterize the characteristics of user identity, thereby improving the accuracy of user identity information recognition; when collecting customer images, it also performs behavior anomaly detection (security analysis of user behavior). If the customer behavior is abnormal, manual review is added, otherwise no manual review verification is performed, thereby improving the efficiency of user identity information recognition; the MS-SSIM algorithm is used for image quality assessment. If the image quality is lower than the threshold, the MSR algorithm is used for image enhancement, thereby improving the image quality and thus enhancing the image analysis capability; the Milvus storage registration is used Multimodal fusion features: When identity information recognition is required, the registered multimodal fusion features corresponding to the user can be quickly queried based on Milvus, and the cosine similarity with the real-time multimodal data can be calculated to improve the speed of identity information recognition; by structured output and visualization of identity recognition results, it is convenient for subsequent troubleshooting of identity recognition problems and for users to understand and view; the user's audio data is collected through a microphone, the user's person image and ID image are collected through a camera, and OCR technology is used to obtain text information from the ID image, thereby collecting more modal data of the user to improve the robustness of user identity information recognition; using Adapter Tuning training, the multimodal feature extraction and fusion feature representation capabilities are improved, while the training speed is also improved; identity information recognition is performed based on a large multimodal model to reduce manual intervention, improve the degree of automation of identity information recognition, and improve user experience.
[0068] The step S20 specifically includes:
[0069] Step S21, uploading each of the historical multimodal data to the server, and the server performs image quality assessment on the person images and document images in each of the historical multimodal data through the multi-scale structural similarity index (MS-SSIM), and determines whether the image quality is greater than a preset quality threshold. If not, proceed to step S22; if yes, proceed to step S23;
[0070] The calculation formula of the multi-size structural similarity index is:
[0071]
[0072] Where x and y represent the horizontal and vertical coordinates of the pixel in the image respectively; SSIM(x,y) represents the multi-scale structural similarity index of the pixel (x,y); lM (x,y) represents the brightness measurement; C j (x,y) represents the contrast measure; S j (x,y) represents the structural measure; α M , β j and γ j All represent importance adjustment weight coefficients; j represents the pixel number; M represents the total number of pixels;
[0073] Step S22, performing image enhancement operation on each of the person images and the document image by using an image enhancement algorithm (MSR algorithm);
[0074] The calculation formula of the image enhancement algorithm is:
[0075]
[0076] Among them, R(x,y) represents the pixel point after image enhancement; S(x,y) represents the original pixel point; F k () represents the Gaussian center surround function, c represents the standard deviation of the Gaussian function, σ represents the normalization coefficient; K represents the number of Gaussian center surround functions. When K = 1, MSR degenerates into SSR; w k represents the scale coefficient;
[0077] Step S23, after performing data cleaning operations including deleting duplicate data and noise data on each of the historical multimodal data, rotating and scaling the person image and the document image in each of the historical multimodal data, and adding background noise to the audio data to perform a sample expansion operation;
[0078] Step S24, normalize the size of each of the character images and the document image, unify the sampling rate of each of the audio data, complete the preprocessing of each of the historical multimodal data, and construct a data set based on the preprocessed historical multimodal data.
[0079] The step S30 specifically includes:
[0080] Step S31: The server creates a multimodal large model for extracting multimodal fusion features;
[0081] The multimodal large model locates the face area from the character image by the Yolo algorithm, extracts the face features from the face area by the Resnet algorithm, extracts the text features from the document image by the OCR technology and the BERT algorithm, and extracts the audio features from the audio data by the MFCC algorithm, thereby obtaining the multimodal fusion features including the face features, text features and audio features; that is, the face features, text features and audio features are fused to obtain the multimodal fusion features, so as to establish the intrinsic connection between the different modal data;
[0082] Extracting the text features specifically includes: first extracting text information from the document image by using OCR technology, and preprocessing the text information by removing special characters, converting full-width to half-width, removing blank characters, segmenting words, and removing stop words, and then extracting text features from the preprocessed text information by using the BERT algorithm;
[0083] Step S32, dividing the data set into a training set, a test set and a validation set in a ratio of 7:2:1;
[0084] Step S33, training the multimodal large model through the training set, performing a Concat operation on the extracted facial features, text features, and audio features during the training process as input Embedding vectors of the multimodal large model, and fine-tuning the multimodal large model through Adapter Tuning to optimize hyperparameters including at least a learning rate, an adapter dimension, and a number of adapter layers;
[0085] Step S34: testing the trained multimodal large model using the test set, and verifying the tested multimodal large model using the verification set.
[0086] The step S40 is specifically as follows:
[0087] The server deploys the trained multimodal large model, and through the deployed multimodal large model, extracts the registered multimodal fusion features of each user from the user information database storing registered multimodal data, and stores each registered multimodal fusion feature in real time to a pre-created Milvus vector library.
[0088] The step S50 is specifically as follows:
[0089] Collect real-time multimodal data, detect user behavior during the collection process, and perform security analysis on the user behavior, such as determining whether there is abnormal behavior such as deliberately covering the head portrait or using other people's identification documents. Then pre-process the multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, calculate the similarity between the real-time modal fusion features and each registered multimodal fusion feature stored in the vector library, and perform identity information recognition based on the similarity.
[0090] That is, the user's person image, ID image and audio data are collected through the camera and microphone; then OCR recognition is performed based on the multimodal large model to identify the text information of the ID image; based on the facial features of the person image, the text features of the text information, and the audio features of the audio data, the multimodal large model is used to fuse the multimodal features and extract the real-time modal fusion features; the cosine similarity of the extracted real-time modal fusion features and the registered multimodal fusion features stored in the vector library is calculated. When the similarity exceeds the preset similarity threshold, the user identity information is successfully recognized, and relevant business processing is carried out. The recognition result is structured and output and visualized; otherwise, the user identity information is unsuccessful and the relevant business processing is not carried out.
[0091] A preferred embodiment of an identity information recognition system based on a multimodal large model of the present invention includes the following modules:
[0092] A historical multimodal data acquisition module is used to collect a large amount of historical multimodal data including person images, document images and audio data through a camera and a microphone, and store each of the historical multimodal data in a pre-created multimodal database;
[0093] A data set construction module, used for uploading each of the historical multimodal data to a server for preprocessing to construct a data set;
[0094] A multimodal large model training module is used for the server to create a multimodal large model for extracting multimodal fusion features, and train the multimodal large model through the data set;
[0095] A registered multimodal fusion feature extraction module is used to extract the registered multimodal fusion features of each user from the user information database through the trained multimodal large model, and store each registered multimodal fusion feature in a pre-created vector library;
[0096] The identity information recognition module is used to pre-process the input real-time multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, and match the real-time modal fusion features with the registered multimodal fusion features stored in the vector library to perform identity information recognition.
[0097] Compared with the traditional single method, the present invention integrates multimodal data of text, image and voice to characterize the characteristics of user identity, thereby improving the accuracy of user identity information recognition; when collecting customer images, it also performs behavior anomaly detection (security analysis of user behavior). If the customer behavior is abnormal, manual review is added, otherwise no manual review verification is performed, thereby improving the efficiency of user identity information recognition; the MS-SSIM algorithm is used for image quality assessment. If the image quality is lower than the threshold, the MSR algorithm is used for image enhancement, thereby improving the image quality and thus enhancing the image analysis capability; the Milvus storage registration is used Multimodal fusion features: When identity information recognition is required, the registered multimodal fusion features corresponding to the user can be quickly queried based on Milvus, and the cosine similarity with the real-time multimodal data can be calculated to improve the speed of identity information recognition; by structured output and visualization of identity recognition results, it is convenient for subsequent troubleshooting of identity recognition problems and for users to understand and view; the user's audio data is collected through a microphone, the user's person image and ID image are collected through a camera, and OCR technology is used to obtain text information from the ID image, thereby collecting more modal data of the user to improve the robustness of user identity information recognition; using Adapter Tuning training, the multimodal feature extraction and fusion feature representation capabilities are improved, while the training speed is also improved; identity information recognition is performed based on a large multimodal model to reduce manual intervention, improve the degree of automation of identity information recognition, and improve user experience.
[0098] The data set construction module specifically includes:
[0099] A data quality assessment unit, used to upload each of the historical multimodal data to a server, and the server performs image quality assessment on the person images and document images in each of the historical multimodal data through a multi-scale structural similarity index (MS-SSIM), and determines whether the image quality is greater than a preset quality threshold. If not, the image enhancement unit is entered; if yes, the sample expansion unit is entered;
[0100] The calculation formula of the multi-size structural similarity index is:
[0101]
[0102] Where x and y represent the horizontal and vertical coordinates of the pixel in the image respectively; SSIM(x,y) represents the multi-scale structural similarity index of the pixel (x,y); l M (x,y) represents the brightness measurement; C j (x,y) represents the contrast measure; S j (x,y) represents the structural measure; α M , β j and γ jAll represent importance adjustment weight coefficients; j represents the pixel number; M represents the total number of pixels;
[0103] An image enhancement unit, used to perform image enhancement operations on each of the person images and the document image by using an image enhancement algorithm (MSR algorithm);
[0104] The calculation formula of the image enhancement algorithm is:
[0105]
[0106] Among them, R(x,y) represents the pixel point after image enhancement; S(x,y) represents the original pixel point; F k () represents the Gaussian center surround function, c represents the standard deviation of the Gaussian function, σ represents the normalization coefficient; K represents the number of Gaussian center surround functions. When K = 1, MSR degenerates into SSR; w k represents the scale coefficient;
[0107] A sample expansion unit, configured to perform a data cleaning operation including deleting duplicate data and noise data on each of the historical multimodal data, rotate and scale the person image and the document image in each of the historical multimodal data, and add background noise to the audio data to perform a sample expansion operation;
[0108] The unified processing unit is used to normalize the size of each of the character images and the document image, unify the sampling rate of each of the audio data, complete the preprocessing of each of the historical multimodal data, and construct a data set based on the preprocessed historical multimodal data.
[0109] The multimodal large model training module specifically includes:
[0110] A multimodal large model creation unit, used for the server to create a multimodal large model for extracting multimodal fusion features;
[0111] The multimodal large model locates the face area from the character image by the Yolo algorithm, extracts the face features from the face area by the Resnet algorithm, extracts the text features from the document image by the OCR technology and the BERT algorithm, and extracts the audio features from the audio data by the MFCC algorithm, thereby obtaining the multimodal fusion features including the face features, text features and audio features; that is, the face features, text features and audio features are fused to obtain the multimodal fusion features, so as to establish the intrinsic connection between the different modal data;
[0112] Extracting the text features specifically includes: first extracting text information from the document image by using OCR technology, and preprocessing the text information by removing special characters, converting full-width to half-width, removing blank characters, segmenting words, and removing stop words, and then extracting text features from the preprocessed text information by using the BERT algorithm;
[0113] A data set division unit, used for dividing the data set into a training set, a test set and a validation set in a ratio of 7:2:1;
[0114] A training unit, used to train the multimodal large model through the training set, perform a Concat operation on the extracted facial features, text features, and audio features during the training process as input Embedding vectors of the multimodal large model, and fine-tune the multimodal large model through Adapter Tuning to optimize hyperparameters including at least a learning rate, an adapter dimension, and a number of adapter layers;
[0115] A testing and verification unit is used to test the trained multimodal large model through the test set, and to verify the tested multimodal large model through the verification set.
[0116] The registered multimodal fusion feature extraction module is specifically used for:
[0117] The server deploys the trained multimodal large model, and through the deployed multimodal large model, extracts the registered multimodal fusion features of each user from the user information database storing registered multimodal data, and stores each registered multimodal fusion feature in real time to a pre-created Milvus vector library.
[0118] The identity information recognition module is specifically used for:
[0119] Collect real-time multimodal data, detect user behavior during the collection process, and perform security analysis on the user behavior, such as determining whether there is abnormal behavior such as deliberately covering the head portrait or using other people's identification documents. Then pre-process the multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, calculate the similarity between the real-time modal fusion features and each registered multimodal fusion feature stored in the vector library, and perform identity information recognition based on the similarity.
[0120] That is, the user's person image, ID image and audio data are collected through the camera and microphone; then OCR recognition is performed based on the multimodal large model to identify the text information of the ID image; based on the facial features of the person image, the text features of the text information, and the audio features of the audio data, the multimodal large model is used to fuse the multimodal features and extract the real-time modal fusion features; the cosine similarity of the extracted real-time modal fusion features and the registered multimodal fusion features stored in the vector library is calculated. When the similarity exceeds the preset similarity threshold, the user identity information is successfully recognized, and relevant business processing is carried out. The recognition result is structured and output and visualized; otherwise, the user identity information is unsuccessful and the relevant business processing is not carried out.
[0121] In summary, the advantages of the present invention are:
[0122] 1. Through the camera and microphone, a large amount of historical multimodal data including character images, document images and audio data is collected and uploaded to the server for preprocessing to build a data set; then the server creates a multimodal large model for extracting multimodal fusion features, and trains the multimodal large model through the data set; then, through the trained multimodal large model, the registered multimodal fusion features of each user are extracted from the user information database and stored in the vector library; finally, the input real-time multimodal data is preprocessed and input into the multimodal large model to obtain real-time modal fusion features, and the real-time modal fusion features are matched with the registered multimodal fusion features stored in the vector library to perform identity information recognition; that is, the comprehensive character image (image), document image (text) and audio data The multimodal data of (voice) is used to characterize the characteristics of user identity, so that the multimodal large model can extract rich features from multiple data sources, thereby improving the accuracy of identity information recognition; by fusing multimodal data, even when the quality of a certain modal data is poor, it can maintain a good recognition effect to improve the robustness of identity information recognition; by fusing multiple modal features (face features, text features, audio features) through the multimodal large model, the ability to construct user portraits is improved, and unknown data can be better identified, thereby improving the generalization ability of the multimodal large model; and compared with traditional methods, it is easier to integrate new modal data, which improves the scalability of identity information recognition, and ultimately greatly improves the accuracy, robustness, generalization, and scalability of identity information recognition.
[0123] 2. By combining security analysis of user behavior (behavior verification) during the identity identification process, fraud can be effectively prevented, thereby improving the security of identity identification.
[0124] 3. Use the MS-SSIM algorithm to evaluate image quality. If the image quality is lower than the threshold, use the MSR algorithm to enhance the image, thereby improving the image quality and enhancing the image analysis capability.
[0125] 4. Use Milvus to store registered multimodal fusion features. When identity information recognition is required, you can quickly query the registered multimodal fusion features corresponding to the user based on Milvus, and calculate the cosine similarity with the real-time multimodal data to improve the speed of identity information recognition.
[0126] 5. By outputting and visualizing the identity recognition results in a structured manner, it is convenient for subsequent troubleshooting of identity recognition issues and for users to understand and view them.
[0127] 6. Using Adapter Tuning training improves the ability of multimodal feature extraction and fusion feature representation, while also improving the training speed.
[0128] 7. Perform identity information recognition based on a multimodal large model to reduce manual intervention, improve the degree of automation of identity information recognition, and enhance user experience.
[0129] Although the specific implementation modes of the present invention are described above, those skilled in the art should understand that the specific implementation modes described are only illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. A method for identifying identity information based on a multimodal large model, characterized in that: The steps include: Step S10: Collect a large amount of historical multimodal data including person images, document images and audio data through a camera and a microphone, and store each of the historical multimodal data in a pre-created multimodal database; Step S20: uploading each of the historical multimodal data to a server for preprocessing to construct a data set; Step S30: the server creates a multimodal large model for extracting multimodal fusion features, and trains the multimodal large model using the data set; Step S40: extracting the registered multimodal fusion features of each user from the user information database through the trained multimodal large model, and storing each registered multimodal fusion feature in a pre-created vector library; Step S50: pre-process the input real-time multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, and match the real-time modal fusion features with the registered multimodal fusion features stored in the vector library to perform identity information recognition.
2. The identity information recognition method based on a multimodal large model as claimed in claim 1, characterized in that: The step S20 specifically includes: Step S21, uploading each of the historical multimodal data to the server, the server performs image quality assessment on the person image and the document image in each of the historical multimodal data through the multi-size structural similarity index, and determines whether the image quality is greater than a preset quality threshold, if not, proceeding to step S22; if yes, proceeding to step S23; Step S22, performing image enhancement operation on each of the person images and the document image by using an image enhancement algorithm; Step S23, after performing data cleaning operations including deleting duplicate data and noise data on each of the historical multimodal data, rotating and scaling the person image and the document image in each of the historical multimodal data, and adding background noise to the audio data to perform a sample expansion operation; Step S24, normalize the size of each of the character images and the document image, unify the sampling rate of each of the audio data, complete the preprocessing of each of the historical multimodal data, and construct a data set based on the preprocessed historical multimodal data.
3. The identity information recognition method based on a multimodal large model as claimed in claim 1, characterized in that: The step S30 specifically includes: Step S31: The server creates a multimodal large model for extracting multimodal fusion features; The multimodal large model locates the face area from the character image by the Yolo algorithm, extracts face features from the face area by the Resnet algorithm, extracts text features from the document image by the OCR technology and the BERT algorithm, and extracts audio features from the audio data by the MFCC algorithm, thereby obtaining multimodal fusion features including the face features, text features and audio features; Step S32, dividing the data set into a training set, a test set and a validation set in a ratio of 7:2:1; Step S33, training the multimodal large model through the training set, performing a Concat operation on the extracted facial features, text features, and audio features during the training process as input Embedding vectors of the multimodal large model, and fine-tuning the multimodal large model through Adapter Tuning to optimize hyperparameters including at least a learning rate, an adapter dimension, and a number of adapter layers; Step S34: testing the trained multimodal large model using the test set, and verifying the tested multimodal large model using the verification set.
4. The identity information recognition method based on a multimodal large model as claimed in claim 1, characterized in that: The step S40 is specifically as follows: The server deploys the trained multimodal large model, and through the deployed multimodal large model, extracts the registered multimodal fusion features of each user from the user information database storing registered multimodal data, and stores each registered multimodal fusion feature in real time to a pre-created Milvus vector library.
5. The identity information recognition method based on a multimodal large model as claimed in claim 1, characterized in that: The step S50 is specifically as follows: Collect real-time multimodal data, detect user behavior during the collection process, and perform security analysis on the user behavior. Pre-process the multimodal data and input it into the multimodal big model to obtain real-time modal fusion features. Calculate the similarity between the real-time modal fusion features and each registered multimodal fusion feature stored in the vector library, and perform identity information recognition based on the similarity.
6. An identity information recognition system based on a multimodal large model, characterized in that: Includes the following modules: A historical multimodal data acquisition module is used to collect a large amount of historical multimodal data including person images, document images and audio data through a camera and a microphone, and store each of the historical multimodal data in a pre-created multimodal database; A data set construction module, used for uploading each of the historical multimodal data to a server for preprocessing to construct a data set; A multimodal large model training module is used for the server to create a multimodal large model for extracting multimodal fusion features, and train the multimodal large model through the data set; A registered multimodal fusion feature extraction module is used to extract the registered multimodal fusion features of each user from the user information database through the trained multimodal large model, and store each registered multimodal fusion feature in a pre-created vector library; The identity information recognition module is used to pre-process the input real-time multimodal data and input it into the multimodal large model to obtain real-time modal fusion features, and match the real-time modal fusion features with the registered multimodal fusion features stored in the vector library to perform identity information recognition.
7. The identity information recognition system based on a multimodal large model as claimed in claim 6, characterized in that: The data set construction module specifically includes: A data quality assessment unit, used to upload each of the historical multimodal data to a server, and the server performs image quality assessment on the person images and document images in each of the historical multimodal data through a multi-size structural similarity index to determine whether the image quality is greater than a preset quality threshold. If not, the image enhancement unit is entered; if yes, the sample expansion unit is entered; An image enhancement unit, used to perform image enhancement operations on each of the person images and the document image by using an image enhancement algorithm; A sample expansion unit, configured to perform a data cleaning operation including deleting duplicate data and noise data on each of the historical multimodal data, rotate and scale the person image and the document image in each of the historical multimodal data, and add background noise to the audio data to perform a sample expansion operation; The unified processing unit is used to normalize the size of each of the character images and the document image, unify the sampling rate of each of the audio data, complete the preprocessing of each of the historical multimodal data, and construct a data set based on the preprocessed historical multimodal data.
8. The identity information recognition system based on a multimodal large model as claimed in claim 6, characterized in that: The multimodal large model training module specifically includes: A multimodal large model creation unit, used for the server to create a multimodal large model for extracting multimodal fusion features; The multimodal large model locates the face area from the character image by the Yolo algorithm, extracts face features from the face area by the Resnet algorithm, extracts text features from the document image by the OCR technology and the BERT algorithm, and extracts audio features from the audio data by the MFCC algorithm, thereby obtaining multimodal fusion features including the face features, text features and audio features; A data set division unit, used for dividing the data set into a training set, a test set and a validation set in a ratio of 7:2:1; A training unit, used to train the multimodal large model through the training set, perform a Concat operation on the extracted facial features, text features, and audio features during the training process as input Embedding vectors of the multimodal large model, and fine-tune the multimodal large model through Adapter Tuning to optimize hyperparameters including at least a learning rate, an adapter dimension, and a number of adapter layers; A testing and verification unit is used to test the trained multimodal large model through the test set, and to verify the tested multimodal large model through the verification set.
9. The identity information recognition system based on a multimodal large model as claimed in claim 6, characterized in that: The registered multimodal fusion feature extraction module is specifically used for: The server deploys the trained multimodal large model, and through the deployed multimodal large model, extracts the registered multimodal fusion features of each user from the user information database storing registered multimodal data, and stores each registered multimodal fusion feature in real time to a pre-created Milvus vector library.
10. The identity information recognition system based on a multimodal large model as claimed in claim 6, characterized in that: The identity information recognition module is specifically used for: Collect real-time multimodal data, detect user behavior during the collection process, and perform security analysis on the user behavior. Pre-process the multimodal data and input it into the multimodal big model to obtain real-time modal fusion features. Calculate the similarity between the real-time modal fusion features and each registered multimodal fusion feature stored in the vector library, and perform identity information recognition based on the similarity.