Information self-verification
By detecting fraud risk in user information and fusion processing of multimodal large models, users' self-certification results are generated, and the problem of user information being easily tampered with and fraud in the existing self-certification system is solved, and a high-accuracy self-certification system is achieved.
Patent Information
- Application Number
- PCT/CN2024/127256
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-19
AI Technical Summary
Since user information is easily tampered with, the existing self-certification system has the risk of using fraudulent means to bypass the self-certification system, resulting in a reduction in the accuracy and effectiveness of the system.
By obtaining user information, fraud risk detection, text information, image features and user behavior auxiliary information are extracted, and information fusion is used to process it using multimodal large models to generate user self-certified results.
It realizes the completeness and practical feasibility of the information self-certification function, reduces the risk of fraud, improves the accuracy and effectiveness of the self-certification system, and improves the user experience.
Smart Images

Figure CN2024127256_19062025_PF_FP_ABST
Abstract
Description
Information self-verification Technical Field
[0001] One or more embodiments of this specification relate to the field of information technology, and specifically to an information self-certification method, system, electronic device, and medium. Background Art
[0002] In the information technology field, self-certification generally refers to a method in which users provide specific documents, records, or information to prove the validity of their information. By self-certifying user information, auditing agencies can more effectively conduct data audits and risk assessments, while also providing more convenient and secure services. However, due to the frequent tampering of user information, there is a risk of fraudulent means being used to circumvent self-certification systems, which reduces their accuracy and effectiveness.
[0003] Summary of the Invention
[0004] The embodiments of this specification provide a method, system, electronic device and medium for information self-certification, and the technical solutions are as follows.
[0005] On the first aspect, the embodiments of this specification provide a method for information self-certification, including: obtaining user information; performing fraud risk detection on the user information to obtain the fraud risk level corresponding to the user; based on the fraud risk level, extracting text information, image features and user behavior auxiliary information corresponding to the user information; based on a multimodal large model, fusing the text information, image features and user behavior auxiliary information to generate a self-certification result corresponding to the user.
[0006] On the second aspect, the embodiments of this specification provide an information self-certification device, including: an information acquisition module for acquiring user information; a risk detection module for performing fraud risk detection on user information to obtain the fraud risk level corresponding to the user; an information extraction module for extracting text information, image features and user behavior auxiliary information corresponding to the user information based on the fraud risk level; a fusion module for fusing text information, image features and user behavior auxiliary information based on a multimodal large model to generate a self-certification result corresponding to the user.
[0007] In a third aspect, an embodiment of this specification provides an electronic device comprising a processor and a memory; the processor is connected to the memory; the memory is used to store executable program code; the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the steps of the information self-certification method of the first aspect of the above embodiment.
[0008] In a fourth aspect, an embodiment of this specification provides a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the steps of the information self-certification method of the first aspect of the above embodiment.
[0009] The beneficial effects brought about by the technical solutions provided by some embodiments of this specification include at least the following: being able to first obtain user information; then performing fraud risk detection on the user information to obtain the corresponding fraud risk level of the user; then extracting the text information, image features and user behavior auxiliary information corresponding to the user information based on the fraud risk level; then, based on the multimodal large model, fusing the text information, image features and user behavior auxiliary information to generate the corresponding self-certification result for the user. The embodiments of this specification make full use of the obtained user information and combine it with the multimodal large model for information processing to achieve the completeness and practical feasibility of the information self-certification function; moreover, the embodiments of this specification perform fraud risk detection on user information to obtain the corresponding fraud risk level of the user, further reducing the possibility of financial loss; the embodiments of this specification can fully automatically evaluate the credibility of user information, improve user experience, improve the accuracy and effectiveness of the self-certification system based on fraud risk detection and the multimodal large model, and provide basic information guarantee for building the self-certification system. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG1 is a schematic diagram of an application scenario of an information self-certification system provided in this specification.
[0011] FIG2 is a flow chart of an information self-certification method provided in this specification.
[0012] FIG3 is a flow chart of another information self-certification method provided in this specification.
[0013] FIG4 is a flow chart of another information self-certification method provided in this specification.
[0014] FIG5 is a schematic diagram of the multimodal large model data processing flow provided in this specification.
[0015] FIG6 is a flow chart of another information self-certification method provided in this specification.
[0016] FIG7 is a timing flow chart of the information self-certification system provided in this specification.
[0017] FIG8 is a schematic structural diagram of an information self-certification device provided in this specification.
[0018] FIG9 is a schematic structural diagram of an electronic device provided in this specification. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of this specification will be described clearly and completely below in conjunction with the drawings in the embodiments of this specification.
[0020] Throughout this specification, the claims, and the accompanying drawings, the terms "first," "second," and so forth are used to distinguish between different items, not to describe a particular order. Furthermore, the term "comprises" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may include other steps or elements inherent to the process, method, product, or apparatus.
[0021] The information self-certification method provided in multiple embodiments of this specification can be executed by the information self-certification device provided in the embodiment of the present invention, or a server integrated with the information self-certification device, wherein the information self-certification device can be implemented in hardware or software.
[0022] Before describing the technical solution of the present invention, a brief explanation of related technical terms is given first.
[0023] Multimodality: In the field of artificial intelligence and machine learning, multimodality refers to the ability to process multiple different types of data. These data types can include text, images, audio, video, etc. Multimodal methods aim to integrate and process these heterogeneous data to obtain more comprehensive information and better performance in various complex tasks. In the field of deep learning, multimodal methods typically involve using multiple input modalities (such as text and images) to train models so that the model can understand and process multiple data types simultaneously. This approach can be applied to tasks such as sentiment analysis, visual question answering, and video description generation, thereby improving the model's ability to understand complex real-world data.
[0024] Large models: In machine learning and deep learning, these are models with a very large number of parameters and complexity. These models typically consist of millions to billions of parameters and contain multiple layers of neural network structures or other complex learning structures.
[0025] Self-certification system: Before using identity-related benefits, users need to upload relevant supporting documents to verify their identity. The self-certification system verifies whether the uploaded information is consistent with the user's true information.
[0026] A multimodal large model is a deep learning model composed of multiple sub-models, each specialized for a specific data type. Joint training is then used to integrate information from different modalities to improve the model's overall performance. Multimodal large models are capable of processing a variety of different types of data, such as text, images, and audio. They combine information from multiple input data types for joint training and inference to solve complex cross-modal tasks. By integrating information from different modalities, multimodal large models demonstrate powerful capabilities in tasks such as visual question answering, video description generation, and multimedia retrieval.
[0027] Before describing the information self-certification method in detail in conjunction with one or more embodiments, this specification first introduces the application scenarios of the information self-certification method.
[0028] Please refer to Figure 1, which is a scenario diagram of the information self-certification system 100 provided in an embodiment of the present invention. The information self-certification system 100 may include an information self-certification device 110, a user terminal 120, and a storage terminal 130, etc. Among them, the user terminal 120 can be a terminal based on the Android system or a terminal based on the IOS system, or a PC based on the Windows system or MAC system, etc. The user terminal 120 and the information self-certification device 110 can be connected through a communication network, and the communication network includes a wireless network and a wired network, wherein the wireless network includes a combination of one or more of a wireless wide area network, a wireless local area network, a wireless metropolitan area network, and a wireless personal network. The network includes network entities such as routers and gateways, which are not shown in the figure. The user terminal 120 can exchange information with the information self-certification device 110 through the communication network. For example, the user can upload user information to the information self-certification device 110 through the user terminal 120.
[0029] The storage terminal 130 can be connected to the information self-certification device 110 and the user terminal 120. The user can store information in the storage terminal 130 through the user terminal 120, and the information self-certification device 110 can store relevant information such as user information, self-certification results, mining information, and image type recognition results in the storage terminal 130.
[0030] Specifically, the information self-certification device can be integrated into an electronic device, which can be a terminal, a server, or other device. A terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC); a server can be a single server or a server cluster consisting of multiple servers. In some embodiments, the information self-certification device can also be integrated into multiple electronic devices. For example, the information self-certification device can be integrated into multiple servers, and the information self-certification method of the present application can be implemented by multiple servers.
[0031] In the embodiments of this specification, the information self-certification device 110 can be used to obtain user information; perform fraud risk detection on user information to obtain the fraud risk level corresponding to the user; based on the fraud risk level, extract text information, image features and user behavior auxiliary information corresponding to the user information; based on a multimodal large model, fuse the text information, image features and user behavior auxiliary information to generate a self-certification result corresponding to the user, etc.
[0032] It should be noted that the scenario diagram of the information self-certification system shown in Figure 1 is only an example. The information self-certification system and scenario described in the embodiment of the present invention are for the purpose of more clearly illustrating the technical solution of the embodiment of the present invention, and do not constitute a limitation on the technical solution provided by the embodiment of the present invention. Ordinary technicians in this field can know that with the evolution of the information self-certification system and the emergence of new scenarios, the technical solution provided by the embodiment of the present invention is also applicable to similar technical problems.
[0033] Please refer to FIG2 , which is a flow chart of an information self-certification method provided by an embodiment of the present invention. The information self-certification method can be executed by the information self-certification device 110 shown in FIG1 . The information self-certification method may include at least the following steps:
[0034] 200. Get user information.
[0035] In this embodiment, user information may include profile information uploaded by the user when self-certifying, and may also include statistical information related to the user obtained by the self-certifying device. User information may include text information, image information, and user behavior information. User behavior information may include preference and interest information, behavioral data information, geographic location information, and social information.
[0036] For example, text information can include basic information, contact information, identity information, and account information. Basic information can include information such as name and date of birth. Contact information can include information such as email address, phone number, and mailing address. Identity information can include information used for identity verification, such as an ID. Account information can include information used for login and account management, such as username, password, and answers to security questions.
[0037] For example, the image information may be a user image, ID card photo, social security photo, and other information.
[0038] For example, user behavior information may include preference and interest information, behavior data information, geographic location information, and social information. Preference and interest information may include user preferences, subscription information, browsing history, etc., and preference and interest information may be used by the self-verification device to provide personalized recommendations and customized experiences for users. Behavioral data information may include click records, purchase history, browsing behavior, etc., and behavioral data information may be used by the self-verification device to analyze user behavior and trends. Geographic location information may include information about the user's geographic location, and geographic location information may be used by the self-verification device to provide positioning services and geo-related personalized content. Social information may include social media accounts, social relationship networks, etc., and social information may be used for social interaction with users and personalized recommendations.
[0039] 210. Perform fraud risk detection on the user information to obtain the fraud risk level corresponding to the user.
[0040] In this embodiment, the fraud risk level may be an evaluation value of the degree of fraud risk that may exist in user information, etc.
[0041] In this embodiment, since the self-certification is initiated by the user, the source and authenticity of the information uploaded by the user cannot be determined. Therefore, this embodiment performs fraud risk detection on the user information to predict the risk of uploading the information, so as to obtain the corresponding fraud risk level of the user, thereby further increasing the accuracy and effectiveness of the self-certification.
[0042] In some embodiments, fraud risk detection is performed on user information to obtain the fraud risk level corresponding to the user, including: obtaining the image fraud risk level, behavioral fraud risk level, and blacklist fraud risk level corresponding to the user based on the user information; and determining the fraud risk level corresponding to the user based on the image fraud risk level, behavioral fraud risk level, and blacklist fraud risk level.
[0043] In this embodiment, the image fraud risk score can be a fraud risk score derived based on the image information in the user information. The behavior fraud risk score can be a fraud risk score derived based on the user behavior information in the user information. The blacklist fraud risk score can be a fraud risk score derived by comparing the user information with a blacklist database.
[0044] This embodiment can obtain the user's corresponding image fraud risk, behavioral fraud risk, and blacklist fraud risk based on user information, and then determine the user's corresponding fraud risk based on the image fraud risk, behavioral fraud risk, and blacklist fraud risk. This embodiment utilizes user information to detect user fraud risk from multiple different perspectives, thereby improving the accuracy of fraud risk detection.
[0045] For example, this embodiment can compare the image fraud risk, behavioral fraud risk, and blacklist fraud risk, and select the risk with the highest value as the user's corresponding fraud risk. This embodiment can also perform a weighted summation of the image fraud risk, behavioral fraud risk, and blacklist fraud risk to obtain the user's corresponding fraud risk.
[0046] In some embodiments, based on user information, the user's corresponding picture fraud risk, behavior fraud risk and blacklist fraud risk are obtained, including: obtaining picture information and user behavior information in the user information; performing image detection on the user based on the picture information to obtain the user's corresponding picture fraud risk; performing behavior detection on the user based on preset behavior conditions and user behavior information to obtain the user's corresponding behavior fraud risk; performing blacklist detection on the user based on a blacklist database to obtain the user's corresponding blacklist fraud risk.
[0047] In this embodiment, the preset behavior condition may include whether users upload information simultaneously in different locations within a short period of time, whether users upload information in a concentrated manner, etc. User behavior information may include the address where the user uploads information, the number of times the user uploads information within a preset period of time, etc.
[0048] This embodiment can perform image detection on users based on their image information to determine their corresponding image fraud risk. For example, this embodiment can detect whether a self-certified image uploaded by a user has been tampered with or photoshopped. If such manipulations are present, the user may have modified the original image, indicating a high fraud risk, resulting in a higher image fraud risk.
[0049] This embodiment can also perform behavioral detection on users based on preset behavioral conditions and user behavior information to determine the corresponding behavioral fraud risk level. For example, if a user submitted more than 100 profile information in Shanghai, Beijing, and Guangzhou within five minutes, it can be determined that the user has uploaded profile information simultaneously in different locations within a short period of time and has uploaded profile information in a concentrated manner, indicating a high fraud risk and thus a relatively high behavioral fraud risk level.
[0050] This embodiment can also perform blacklist checks on users based on a blacklist database to obtain the user's corresponding blacklist fraud risk. For example, this embodiment includes a historical database, which includes a whitelist database and a blacklist database, and the blacklist database contains a list of fraud risk users. This embodiment can use user information to compare the user with the list of fraud risk users in the blacklist database. When the user matches certain users on the fraud risk user list, it indicates that the user has a high fraud risk, thereby obtaining a relatively high blacklist fraud risk.
[0051] 220. Based on the fraud risk, extract the text information, image features and user behavior auxiliary information corresponding to the user information.
[0052] In this embodiment, if the fraud risk level is not lower than the preset risk threshold, the user's fraud risk level can be determined to be high, and a notification message indicating that the fraud risk detection failed can be sent to the user terminal. If the fraud risk level is lower than the preset risk threshold, the user's fraud risk level can be determined to be low, and further text information, image features, and user behavior auxiliary information corresponding to the user information can be extracted to perform self-verification processing on the user information.
[0053] In some embodiments, based on the fraud risk, text information, image features and user behavior auxiliary information corresponding to the user information are extracted, including: when the fraud risk is lower than a preset risk threshold, the text information, image features and user behavior auxiliary information corresponding to the user information are extracted through a feature extraction model.
[0054] In this embodiment, the feature extraction model may be a model that extracts relevant text, image, and other features from user information. Feature extraction models may include text recognition models, image processing models, and feature splicing models. User behavior auxiliary information may be user behavior information used to assist in verifying the validity of user self-certification.
[0055] For example, in this embodiment, the preset risk threshold can be set to thirty percent, and the obtained fraud risk corresponding to the user is ten percent. It can be seen that the fraud risk corresponding to the user is lower than the preset risk threshold, and the user's fraud risk level is determined to be relatively low. The text information, image features and user behavior auxiliary information corresponding to the user information can be extracted through the feature extraction model, and a prompt message indicating that the fraud risk detection has passed can be sent to the user end.
[0056] In some embodiments, text information, image features, and user behavior auxiliary information corresponding to the user information are extracted through a feature extraction model, including: performing text recognition on the user information based on a text recognition model in the feature extraction model to obtain text information corresponding to the user information; extracting image features corresponding to the user information based on an image processing model in the feature extraction model; and splicing the user behavior information in the user information through a feature splicing model in the feature extraction model to obtain the spliced user behavior auxiliary information.
[0057] In this embodiment, the text recognition model can be a machine learning model or a deep learning model for recognizing, understanding, and processing text data in user information. The text recognition model can include an optical character recognition model, a natural language processing model, a named entity recognition model, a text classification model, etc.
[0058] For example, this embodiment can extract printed text and handwritten text from images or scanned documents in user information through an optical character recognition model. This embodiment can also process and understand natural language text in user information through a natural language processing model, including models such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and transformers. This embodiment can also identify and extract named entities, such as names of people, places, and organizations, from the text of user information through a named entity recognition model. This embodiment can also classify the text of user information through a text classification model.
[0059] In this embodiment, the image processing model can be based on a machine learning model or deep learning model using a convolutional neural network (CNN) or an attention mechanism called a Transformer. This embodiment can use the image processing model to extract image features corresponding to user information, which are used to identify the type of images uploaded by the user. This image-text feature pairing can then be combined with the text information to form image-text feature pairs. These image-text feature pairs can then be used to mine user-related information in a large multimodal model through image-text question-and-answer and other modes.
[0060] In this embodiment, the feature splicing model may be a model that maps user behavior information in the user information into vector representations and then performs splicing processing. In this embodiment, the user behavior information in the user information is first mapped into vector representations using the feature splicing model, and then the user behavior information mapped into vector representations is spliced together to obtain spliced user behavior auxiliary information.
[0061] 230. Based on a multimodal large model, text information, image features and user behavior auxiliary information are integrated and processed to generate the user's corresponding self-verification results.
[0062] In this embodiment, feature extraction is performed on user information to obtain text information, image features, and auxiliary information about user behavior. This information is then integrated with a multimodal large model to generate a self-verification result corresponding to the user. The corresponding self-verification result can be either valid (passed), or invalid (failed).
[0063] In some embodiments, based on a multimodal large model, text information, image features, and user behavior auxiliary information are fused and processed to generate a self-verification result corresponding to the user, including: obtaining a fusion feature item based on text information, image features, and user behavior auxiliary information; obtaining a first model response of the multimodal large model to the fusion feature item and the user behavior auxiliary information to generate a self-verification result corresponding to the user.
[0064] In this embodiment, the fused feature item may be a feature item obtained by fusing the preset feature fusion tag vector with the triple feature item after position encoding and feature embedding based on the preset feature fusion tag vector. The triple feature item may be a feature item composed of text information, image features, and user behavior auxiliary information.
[0065] In this embodiment, the first model response may be a model response of the multimodal large model to the fused feature item and the user behavior auxiliary information.
[0066] In some embodiments, based on text information, image features and user behavior auxiliary information, fused feature items are obtained, including: splicing text information, image features and user behavior auxiliary information into triple feature items; performing feature embedding processing on the triple feature items to obtain triple embedding vectors corresponding to the triple feature items; performing position encoding processing on the triple embedding vectors to obtain position-encoded triple embedding vectors; based on a preset feature fusion mark vector, fusing the preset feature fusion mark vector with the position-encoded triple embedding vector to obtain the fused feature items after fusion processing.
[0067] In this embodiment, the preset feature fusion marker vector may be a vector used to fuse all feature items (corresponding to the triplet embedding vector after position encoding).
[0068] This embodiment uses a preset feature fusion marker vector to fuse the position-encoded triple embedding vectors into an overall representation, so that the multimodal large model can comprehensively consider and process them. This fusion method can help the multimodal large model better understand and utilize different types of feature information, thereby improving the performance of the multimodal large model. In this embodiment, the preset feature fusion marker vector is set to be used in the training process of the model, and its better representation is learned through the back propagation algorithm. The representation of the preset feature fusion marker vector will be dynamically adjusted according to the optimization goal of the model and the characteristics of the input data to maximize the performance of the model.
[0069] The feature item embedding process in this embodiment is used to embed all feature items in a triplet feature item, thereby converting all feature items in the triplet feature item into a low-dimensional representation. For example, text information can be converted into a word embedding vector, image features can be converted into an image embedding vector, and user behavior auxiliary information can be converted into a behavior information embedding vector. This embodiment can also perform position encoding on the triplet embedding vector to preserve the position information of the feature item in the sequence.
[0070] This embodiment fuses the preset feature fusion tag vector with all other feature items to obtain a fused feature item after fusion processing. This can be achieved through weighted summation, splicing, or other fusion methods. Next, this embodiment can input the fused feature item after fusion processing together with other feature items (such as user behavior auxiliary information) into the multimodal large model for processing to obtain the user's corresponding self-verification result.
[0071] Please refer to FIG. 3 , which shows a flow chart of an information self-certification method provided in another embodiment of this specification. The method can be executed by the information self-certification device 110 shown in FIG. 1 .
[0072] As shown in FIG3 , the information self-certification method may include at least the following steps:
[0073] 300. Obtain user information;
[0074] 310. Perform fraud risk detection on user information to obtain the user's corresponding fraud risk level;
[0075] 320. Based on the fraud risk, extract the text information, image features and user behavior auxiliary information corresponding to the user information;
[0076] 330. Based on a multimodal large model, text information, image features and user behavior auxiliary information are integrated to generate the user's corresponding self-verification results;
[0077] 340. Based on the text information, obtain a second model response of the multimodal large model to the text information.
[0078] In this embodiment, the second model response is the mined information corresponding to the text information.
[0079] In this embodiment, the text information can be used as the input of the multimodal large model, and the multimodal large model can output the mining information corresponding to the text information. The second model response is the model response of the multimodal large model to the text information.
[0080] This embodiment can not only fuse text information, image features, and user behavior auxiliary information through a multimodal large model to generate a self-verification result corresponding to the user. This embodiment can also extract user features and use a multimodal large model (such as a natural language large model GPT, a picture-text pairing model Clip, a picture-text understanding large model VisualGLM, etc.) to perform information mining on the text information corresponding to the user, thereby mining out valid information related to the user. The embodiments of this specification can use the mined information corresponding to the text information in the training of the multimodal large model to improve the accuracy of the multimodal large model.
[0081] For example, taking the verification of teacher identity information through social security information uploaded by users as an example, when mining user information through a multimodal large model, relevant query information can be set to obtain information such as the payment base, payment ratio and payment amount in the social security uploaded by the user, and the user's salary situation can be further inferred. Then, the user's salary situation can be used as the mining information corresponding to the text information.
[0082] Please refer to FIG. 4 , which shows a flow chart of an information self-certification method provided in another embodiment of this specification. The method can be executed by the information self-certification device 110 shown in FIG. 1 .
[0083] As shown in FIG4 , the information self-certification method may include at least the following steps:
[0084] 400. Get user information;
[0085] 410. Perform fraud risk detection on the user information to obtain the user's corresponding fraud risk level;
[0086] 420. Based on the fraud risk, extract the text information, image features and user behavior auxiliary information corresponding to the user information;
[0087] 430. Based on a multimodal large model, text information, image features, and user behavior auxiliary information are integrated to generate the user's corresponding self-verification results;
[0088] 440. Based on the text information, obtain a second model response of the multimodal large model to the text information;
[0089] 450. Based on the image features, obtain a third model response of the multimodal large model to the image features.
[0090] In this embodiment, the third model response is an image type recognition result corresponding to the image feature.
[0091] As shown in FIG5 , in this embodiment, text information can be input as a multimodal large model, and after being processed by the multimodal large model, the text information f is outputted through the first multilayer perceptron (MLP). text The second model response is the model response of the multimodal large model to the text information. imageAs the input of the multimodal large model, after being processed by the multimodal large model, the second multilayer perceptron MLP outputs the image type recognition result corresponding to the image feature. This embodiment can also fuse the preset feature fusion tag vector token based on the preset feature fusion tag vector token and the position-encoded triple embedding vector to obtain a fused feature item after fusion processing; then the fused feature item after fusion processing is combined with other feature items (such as user behavior auxiliary information f action ) are input together into the multimodal large model for processing, and the third multi-layer perceptron MLP outputs the user's corresponding self-certification result.
[0092] This embodiment can not only fuse text information, image features and user behavior auxiliary information through a multimodal large model to generate a self-certification result corresponding to the user. This embodiment can also perform information mining on the text information corresponding to the user, thereby mining out valid information related to the user. This embodiment can also use a multimodal large model (such as a natural language large model GPT, a picture-text pairing model Clip, a picture-text understanding large model VisualGLM, etc.) to distinguish the type of pictures uploaded by users, thereby extracting user-related information from the image and obtaining image type recognition results corresponding to image features. The embodiment of this specification can use the image type recognition results corresponding to image features in the training of a multimodal large model to improve the accuracy of the multimodal large model.
[0093] For example, taking the example of verifying a teacher's identity information by uploading social security information, after the user uploads a social security screenshot, the image features corresponding to the social security screenshot are processed through a multimodal large model to further determine whether it belongs to the social security screenshot information. It can also determine whether the identity ID and name match the personal information for identity verification.
[0094] This embodiment can send relevant self-certification prompt information to the user terminal based on the user's corresponding self-certification result, for example, sending a prompt message to the user terminal indicating whether the self-certification passed or failed. This embodiment can also store relevant information such as user information, self-certification results, mining information, and image type recognition results in a database for archiving. This embodiment of the specification can use relevant information such as user information, self-certification results, mining information, and image type recognition results in the training of a multimodal large model to improve the accuracy of the multimodal large model.
[0095] The embodiments of this specification fully utilize the acquired user information and combine it with a multimodal big model for information processing to achieve the completeness and practical feasibility of the information self-certification function; in addition, the embodiments of this specification perform fraud risk detection on user information to obtain the user's corresponding fraud risk level, further reducing the possibility of financial loss; the embodiments of this specification can fully automatically evaluate the credibility of user information, improve user experience, and improve the accuracy and effectiveness of the self-certification system based on fraud risk detection and a multimodal big model, providing basic information guarantee for building a self-certification system.
[0096] Please refer to Figure 6, which shows a flowchart of an information self-certification method provided by another embodiment of this specification. In this embodiment, the information self-certification device is specifically integrated into a server and applied to a scenario where a user performs self-certification when applying for a loan.
[0097] As shown in FIG6 , the information self-certification method may include at least the following steps:
[0098] 600. Based on the user's loan application request, obtain user information, including self-certified application information and user-related statistical information;
[0099] 610. Perform fraud risk detection on the user information to obtain the user's corresponding fraud risk level;
[0100] 620. Based on the fraud risk, extract the text information, image features, and user behavior auxiliary information corresponding to the user information;
[0101] 630. Based on a multimodal large model, text information, image features, and user behavior auxiliary information are integrated to generate the user's corresponding loan application self-certification result;
[0102] 640. Based on the text information, obtain mining information corresponding to the text information output by the multimodal large model;
[0103] 650. Based on the image features, obtain the image type recognition result output by the multimodal large model for the image features.
[0104] In this embodiment, the output of the modal large model can be refined into the output of various attention factors (corresponding to the various features corresponding to the user information extracted based on the fraud risk), or it can directly output the self-certification results and evidence items to form an end-to-end self-certification verification scheme.
[0105] Please refer to Figure 7, which is a timing flow chart of the information self-certification system of this embodiment. As shown in Figure 7, the user uploads self-certification information to the information self-certification device through the client. The information self-certification device uses the anti-fraud model to detect fraud on the user information containing the information. When a fraud risk is detected, a message indicating that the self-certification verification failed is sent to the client page. When no fraud risk is detected, the text information, image features and user behavior auxiliary information corresponding to the user information are extracted. Based on the multimodal large model, the text information, image features and user behavior auxiliary information are fused and processed to generate the user's corresponding application self-certification result. The image type can be verified, the person-image real-name matching can be performed, and the user information can be mined. Then, it is determined whether the self-certification is passed. When the self-certification fails, a message indicating that the self-certification verification fails is sent to the client page. When the self-certification passes, a message indicating that the self-certification verification passes is sent to the client page.
[0106] In this embodiment, to obtain authentic user information and prevent financial losses from user fraud, user-uploaded information must be verified. This embodiment not only prevents users from using fraudulent means to circumvent self-certification risks, but also parses user-uploaded information, verifies user identity through a multimodal macro model, and mines valid information from the information. Building this information self-certification system will improve the performance of online self-certification systems, mine more user asset information, reduce user complaint rates, and mitigate the risk of financial losses associated with self-certification products. In this embodiment, the validity of user-uploaded information can be determined through a multimodal macro model, achieving a fully automated process. The macro model can also output the evidence items upon which judgments are based, allowing attribution and backtracking of the output results. Using an anti-fraud model (i.e., performing fraud risk detection on user information) can intercept potential fraud risk points, ensuring the effectiveness of the entire system. This embodiment also integrates the anti-fraud model and the multimodal macro model to verify and extract value from user-uploaded information, providing effective evidence factors for user asset verification, identity verification, credit assessment, and borrowing capacity.
[0107] The embodiments of this specification fully utilize the acquired user information and combine it with a multimodal big model for information processing to achieve the completeness and practical feasibility of the information self-certification function; in addition, the embodiments of this specification perform fraud risk detection on user information to obtain the user's corresponding fraud risk level, further reducing the possibility of financial loss; the embodiments of this specification can fully automatically evaluate the credibility of user information, improve user experience, and improve the accuracy and effectiveness of the self-certification system based on fraud risk detection and a multimodal big model, providing basic information guarantee for building a self-certification system.
[0108] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0109] Please refer to FIG8 , which is a schematic diagram of the structure of an information self-certification device provided in an embodiment of this specification.
[0110] As shown in FIG8 , the information self-certification device may at least include an information acquisition module 800 , a risk detection module 810 , an information extraction module 820 and a fusion module 830 .
[0111] The information acquisition module 800 is used to obtain user information; the risk detection module 810 is used to perform fraud risk detection on user information to obtain the fraud risk level corresponding to the user; the information extraction module 820 is used to extract the text information, image features and user behavior auxiliary information corresponding to the user information based on the fraud risk level; the fusion module 830 is used to fuse the text information, image features and user behavior auxiliary information based on a multimodal large model to generate a self-certification result corresponding to the user.
[0112] In some embodiments, the risk detection module includes: a risk sub-module, which is used to obtain the user's corresponding image fraud risk, behavioral fraud risk and blacklist fraud risk based on user information; and a risk determination module, which is used to determine the user's corresponding fraud risk based on the image fraud risk, behavioral fraud risk and blacklist fraud risk.
[0113] In some embodiments, the risk submodule includes: an information acquisition submodule, which is used to obtain picture information and user behavior information in user information; an image detection module, which is used to perform image detection on the user based on the picture information to obtain the user's corresponding picture fraud risk; a behavior detection module, which is used to perform behavior detection on the user based on preset behavior conditions and user behavior information to obtain the user's corresponding behavior fraud risk; a blacklist detection module, which is used to perform blacklist detection on the user based on the blacklist database to obtain the user's corresponding blacklist fraud risk.
[0114] In some embodiments, the information extraction module includes: a risk determination module, which is used to extract text information, image features and user behavior auxiliary information corresponding to the user information through a feature extraction model when the fraud risk is lower than a preset risk threshold.
[0115] In some embodiments, the risk determination module includes a feature extraction module, and the feature extraction module includes: a text recognition module, which is used to perform text recognition on user information based on the text recognition model in the feature extraction model to obtain text information corresponding to the user information; an image feature module, which is used to extract image features corresponding to the user information based on the image processing model in the feature extraction model; and a behavior splicing module, which is used to splice user behavior information in the user information through the feature splicing model in the feature extraction model to obtain user behavior auxiliary information after splicing.
[0116] In some embodiments, the fusion module includes: a fusion feature item module, which is used to obtain fusion feature items based on text information, image features and user behavior auxiliary information; a first response module, which is used to obtain the first model response of the multimodal large model to the fusion feature items and user behavior auxiliary information to generate a self-certification result corresponding to the user.
[0117] In some embodiments, the fusion feature item module includes: an information splicing module for splicing text information, image features and user behavior auxiliary information into triple feature items; a feature embedding module for performing feature embedding processing on the triple feature items to obtain a triple embedding vector corresponding to the triple feature item; a position encoding module for performing position encoding processing on the triple embedding vector to obtain a triple embedding vector after position encoding; a fusion processing module for fusing the preset feature fusion mark vector with the position encoded triple embedding vector based on a preset feature fusion mark vector to obtain a fused feature item after fusion processing.
[0118] In some embodiments, the information self-certification device also includes: a second response module, which is used to obtain a second model response of the multimodal large model to the text information based on the text information, and the second model response is the mining information corresponding to the text information.
[0119] In some embodiments, the information self-certification device also includes: a third response module, which is used to obtain a third model response of the multimodal large model to the image feature based on the image feature, and the third model response is an image type recognition result corresponding to the image feature.
[0120] Based on the content of the information self-certification system in multiple embodiments of this specification, it can be seen that the embodiments of this specification make full use of the acquired user information, combine it with a multimodal big model for information processing, and realize the completeness and practical feasibility of the information self-certification function; and, the embodiments of this specification perform fraud risk detection on user information to obtain the corresponding fraud risk level of the user, further reducing the possibility of capital loss; the embodiments of this specification can fully automatically evaluate the credibility of user information, improve user experience, and improve the accuracy and effectiveness of the self-certification system based on fraud risk detection and a multimodal big model, providing basic information guarantee for building a self-certification system.
[0121] Each embodiment in this specification is described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the description of the information self-verification system embodiment is relatively simple, as it is fundamentally similar to the information self-verification method embodiment. For relevant portions, refer to the description of the method embodiment.
[0122] Please refer to FIG9 , which shows a schematic structural diagram of an electronic device provided in an embodiment of this specification.
[0123] As shown in FIG. 9 , the electronic device 900 may include: at least one processor 901 , at least one network interface 904 , a user interface 903 , a memory 905 , and at least one communication bus 902 .
[0124] The communication bus 902 can be used to realize the connection and communication of the above components. The user interface 903 can include buttons, and the optional user interface can also include a standard wired interface or a wireless interface.
[0125] The network interface 904 may include, but is not limited to, a Bluetooth module, an NFC module, a Wi-Fi module, and the like.
[0126] Among them, the processor 901 may include one or more processing cores. The processor 901 uses various interfaces and lines to connect the various parts of the entire electronic device 900, and executes various functions of the electronic device 900 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 905, and calling data stored in the memory 905. Optionally, the processor 901 can be implemented in at least one hardware form of DSP, FPGA, and PLA. The processor 901 can integrate one or a combination of CPU, GPU, modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may not be integrated into the processor 901, but may be implemented separately through a chip.
[0127] Among them, the memory 905 may include RAM or ROM. Optionally, the memory 905 includes a non-transitory computer-readable medium. The memory 905 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 905 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 905 may also be at least one storage device located away from the aforementioned processor 901. The memory 905 as a computer storage medium may include an operating system, a network communication module, a user interface module and an information self-certification application. The processor 901 can be used to call the information self-certification application stored in the memory 905 and execute the information self-certification and formulation steps mentioned in the above-mentioned embodiments.
[0128] The embodiments of this specification also provide a computer-readable storage medium having instructions stored therein that, when executed on a computer or processor, cause the computer or processor to perform one or more steps in the embodiments shown in Figures 2 to 6 above. If the various component modules of the electronic device described above are implemented as software functional units and sold or used as independent products, they may be stored in the computer-readable storage medium.
[0129] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiments of this specification is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0130] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. The technical features of this embodiment and the implementation scheme can be combined in any manner unless they conflict.
[0131] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims of this specification.
Claims
1. An information self-certification method, comprising: Get user information; Performing fraud risk detection on the user information to obtain a fraud risk level corresponding to the user; Based on the fraud risk, extracting text information, image features, and user behavior auxiliary information corresponding to the user information; Based on a multimodal large model, the text information, the image features and the user behavior auxiliary information are fused to generate a self-certification result corresponding to the user.
2. According to the method of claim 1, the step of performing fraud risk detection on the user information to obtain the fraud risk level corresponding to the user comprises: Based on the user information, obtain the image fraud risk, behavior fraud risk and blacklist fraud risk corresponding to the user; The fraud risk level corresponding to the user is determined according to the image fraud risk level, the behavior fraud risk level and the blacklist fraud risk level.
3. The method according to claim 2, wherein obtaining the image fraud risk, behavior fraud risk, and blacklist fraud risk corresponding to the user based on the user information comprises: Obtaining picture information and user behavior information from the user information; Performing image detection on the user based on the picture information to obtain the picture fraud risk level corresponding to the user; Based on the preset behavior conditions and the user behavior information, the user's behavior detection is performed to obtain the behavior fraud risk corresponding to the user; Based on the blacklist database, a blacklist check is performed on the user to obtain the blacklist fraud risk corresponding to the user.
4. The method according to claim 1, wherein extracting text information, image features and user behavior auxiliary information corresponding to the user information based on the fraud risk level comprises: When the fraud risk is lower than a preset risk threshold, the text information, image features and user behavior auxiliary information corresponding to the user information are extracted through a feature extraction model.
5. The method according to claim 4, wherein extracting text information, image features and user behavior auxiliary information corresponding to the user information through a feature extraction model comprises: Based on the text recognition model in the feature extraction model, performing text recognition on the user information to obtain text information corresponding to the user information; Extracting image features corresponding to the user information based on the image processing model in the feature extraction model; The user behavior information in the user information is spliced through the feature splicing model in the feature extraction model to obtain the spliced user behavior auxiliary information.
6. The method according to claim 1, wherein the text information, the image features and the user behavior auxiliary information are fused based on a multimodal large model to generate a self-certification result corresponding to the user, comprising: Based on the text information, the image features and the user behavior auxiliary information, obtaining a fusion feature item; Obtain a first model response of the multimodal large model to the fused feature item and the user behavior auxiliary information to generate a self-certification result corresponding to the user.
7. The method according to claim 6, wherein the acquiring of fusion feature items based on the text information, the image features and the user behavior auxiliary information comprises: splicing the text information, the image features and the user behavior auxiliary information into a triple feature item; Performing feature embedding processing on the triple feature item to obtain a triple embedding vector corresponding to the triple feature item; Performing position encoding processing on the triple embedding vector to obtain a triple embedding vector after position encoding; Based on a preset feature fusion mark vector, the preset feature fusion mark vector is fused with the position-encoded triple embedding vector to obtain a fused feature item after fusion processing.
8. The method according to claim 1, further comprising: Based on the text information, a second model response of the multimodal large model to the text information is obtained, where the second model response is mined information corresponding to the text information.
9. The method according to claim 1, further comprising: Based on the image feature, a third model response of the multimodal large model to the image feature is obtained, where the third model response is an image type recognition result corresponding to the image feature.
10. An information self-certification device, comprising: Information acquisition module, used to obtain user information; A risk detection module, used to perform fraud risk detection on the user information to obtain the fraud risk level corresponding to the user; An information extraction module, used to extract text information, image features and user behavior auxiliary information corresponding to the user information based on the fraud risk level; A fusion module is used to fuse the text information, the image features and the user behavior auxiliary information based on a multimodal large model to generate a self-certification result corresponding to the user.
11. The device according to claim 10, wherein the risk detection module comprises: A risk submodule, used to obtain the image fraud risk, behavior fraud risk and blacklist fraud risk corresponding to the user based on the user information; The risk determination module is used to determine the fraud risk corresponding to the user according to the image fraud risk, the behavior fraud risk and the blacklist fraud risk.
12. The device according to claim 11, wherein the risk submodule comprises: An information acquisition submodule, used to acquire image information and user behavior information in the user information; An image detection module, used to perform image detection on the user based on the image information to obtain the image fraud risk level corresponding to the user; A behavior detection module, used to perform behavior detection on the user based on preset behavior conditions and the user behavior information, and obtain the behavior fraud risk corresponding to the user; The blacklist detection module is used to perform blacklist detection on the user based on the blacklist database to obtain the blacklist fraud risk corresponding to the user.
13. The device according to claim 10, wherein the information extraction module comprises: The risk determination module is used to extract text information, image features and user behavior auxiliary information corresponding to the user information through a feature extraction model when the fraud risk is lower than a preset risk threshold.
14. The device according to claim 13, wherein the risk determination module comprises a feature extraction module, wherein the feature extraction module comprises: A text recognition module, used to perform text recognition on the user information based on the text recognition model in the feature extraction model to obtain text information corresponding to the user information; An image feature module, used to extract image features corresponding to the user information based on an image processing model in the feature extraction model; The behavior splicing module is used to splice the user behavior information in the user information through the feature splicing model in the feature extraction model to obtain the spliced user behavior auxiliary information.
15. The device according to claim 10, wherein the fusion module comprises: A fusion feature item module, used for obtaining a fusion feature item based on the text information, the image feature and the user behavior auxiliary information; The first response module is used to obtain the first model response of the multimodal large model to the fusion feature item and the user behavior auxiliary information to generate a self-certification result corresponding to the user.
16. The device according to claim 15, wherein the fusion feature item module comprises: An information splicing module, used for splicing the text information, the image features and the user behavior auxiliary information into a triple feature item; A feature embedding module, used to perform feature embedding processing on the triple feature item to obtain a triple embedding vector corresponding to the triple feature item; A position encoding module, used for performing position encoding processing on the triple embedding vector to obtain a triple embedding vector after position encoding; The fusion processing module is used to fuse the preset feature fusion label vector based on the preset feature, and fuse the preset feature fusion label vector with the triplet embedding vector after the position encoding to obtain a fused feature item after the fusion processing.
17. The apparatus according to claim 10, further comprising: The second response module is used to obtain a second model response of the multimodal large model to the text information based on the text information, where the second model response is the mined information corresponding to the text information.
18. The apparatus according to claim 10, further comprising: The third response module is used to obtain a third model response of the multimodal large model to the image feature based on the image feature, and the third model response is an image type recognition result corresponding to the image feature.
19. An electronic device comprising a processor and a memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method according to any one of claims 1 to 9.
20. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Internet financial gang fraud behavior detection method based on knowledge graph
CN112053221A
Financial fraud risk identification method and device, computer equipment and storage medium
CN112561684A
Fraud risk prediction method and device, equipment and storage medium
CN115713407A
Information self-certification method and device, electronic equipment and medium
CN117541379A
Fraud detection
US20180365687A1