User fraud risk identification method and apparatus, and electronic device
Through the combination of natural language processing and machine learning model and parameter server, the high cost and time-consuming problem of user fraud risk identification in the credit field is solved, low-cost and low-time-consuming user fraud risk identification is achieved, and identification efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510332895.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
The existing user fraud risk identification methods in the credit field have problems such as high cost, low intelligence and long-term consumption, resulting in poor user experience.
The natural language processing model and machine learning model are combined with parameter servers. By obtaining the user's basic information and multi-dimensional unstructured data, the feature data is generated and the fraud risk level is output, and the model parameters are trained using the training data set to achieve rapid decision-making.
Identify potential user fraud risks in a low-cost and low-time consuming framework, reduce manual intervention, and improve identification efficiency and user experience.
Smart Images

Figure CN120338939A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and in particular, to a method, device, and electronic device for identifying user fraud risks. Background Art
[0002] In the credit field, it is common to encounter the situation where users default on their loans. In order to strengthen supervision, it is usually necessary to evaluate the fraud risks of users in order to take corresponding measures in a timely manner to avoid bad debts.
[0003] Currently, the common methods for judging user fraud risks in the credit field include: credit scoring, behavior analysis, social network analysis, rule engines, blacklist systems, real-time monitoring and early warning, manual review, etc. These methods above are usually selectively used when identifying fraud risks.
[0004] This leads to the following three defects in the prior art:
[0005] 1. High cost: The systems to be built are complex and intricate, and there are dependencies between some systems, which requires continuous repeated investment, especially the server costs involved;
[0006] 2. Low intelligence: In traditional credit applications, especially on the bank side, most of the process ends require manual intervention, which does not conform to the current trend of AI participation in decision-making and greatly affects work efficiency;
[0007] 3. Serious time-consuming problem: For the individual decision-making links of each system, whether in series or in parallel, the time-consuming problem needs to be considered. Only relying on the way of stacking applications cannot meet the efficient credit decision-making process. If the decision is made according to the traditional system process, the final response time cannot reach a reasonable level (such as a few hundred milliseconds), which ultimately brings poor user experience. Summary of the Invention
[0008] The purpose of this application is to provide a method, device, and electronic device for identifying user fraud risks, which can identify the potential fraud risks of users in a low-cost, low-time-consuming, and relatively simple framework, without the need for additional manual intervention, quickly classify the fraud risks for users, and complete the final decision.
[0009] In a first aspect, the present application provides a method for identifying user fraud risks. The method includes: obtaining the first basic information of a target user; the first basic information includes at least one of the following: name, ID number, mobile phone number, device number; based on the first basic information, obtaining the first multi-dimensional unstructured data of the target user; the first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, and e-commerce transaction records; processing the first multi-dimensional unstructured data according to multiple preset time windows to generate first feature data under multiple time windows; inputting the first feature data under multiple time windows into a preset user fraud risk identification model; wherein, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; the user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set; processing the first feature data into a semantic feature vector through the natural language processing model; outputting the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data through the parameter server and the machine learning model.
[0010] Further, the training process of the above user fraud risk identification model is as follows: obtaining a training data set; the data in the training data set includes: second feature data under multiple time windows obtained by processing the second multi-dimensional unstructured data of multiple sample users; performing word segmentation, encoding, multi-layer Transformer processing, and feature extraction on the second feature data of multiple sample users through the natural language processing model to obtain semantic feature vectors corresponding to the second multi-dimensional unstructured data of multiple sample users respectively, and obtaining key parameters involved in generating the semantic feature vectors; the key parameters include at least one of the following: Transformer layer parameters, word embedding parameters, position encoding parameters; wherein, the semantic feature vector includes lexical semantics and context information; taking the semantic feature vectors of multiple sample users and the fraud risk level labels respectively corresponding to multiple sample users as a training sample set, and inputting the key parameters into the parameter server; training the model parameters of the machine learning model with the training sample set, and storing and updating the key parameters through the parameter server during the training process, and obtaining the user fraud risk identification model after the training is completed; the model parameters include: tree depth, number of leaf nodes, learning rate, and L1 regularization term.
[0011] Further, the step of obtaining the training data set includes: for each sample user, obtaining the second basic information of the sample user; based on the second basic information, obtaining the second multi-dimensional unstructured data of the sample user; processing the second multi-dimensional unstructured data according to multiple preset time windows to generate second feature data under multiple time windows, and obtaining the training data set.
[0012] Further, the above-mentioned parameter server includes: a server side, a client side, and a scheduler; the server side is used to store model parameters, receive gradients uploaded by the client side, and update local parameters, responsible for the storage and update of parameters to ensure the consistency of model parameters; the client side is used to obtain the latest parameters from the server side, calculate gradients using local data, and upload the gradients to the server side, responsible for calculating gradients and synchronizing model parameters; the scheduler is used to manage server and client nodes and complete functions such as data synchronization and node addition / removal between nodes.
[0013] Further, the above-mentioned natural language processing model includes: a GPT model or a BERT model; the machine learning model includes: an XGB model or an LGB model.
[0014] Further, the step of outputting the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data by the parameter server and the machine learning model includes: processing the semantic feature vector corresponding to the first feature data by the parameter server and the machine learning model to obtain a probability value; if the probability value is greater than the threshold, determining that the fraud risk level of the target user is a high risk level; if the probability value is less than or equal to the threshold, determining that the fraud risk level of the target user is a low risk level.
[0015] In a second aspect, the present application further provides a device for identifying user fraud risks. The device includes: a user information registration module for obtaining the first basic information of the target user; the first basic information includes at least one of the following: name, ID number, mobile phone number, device number; a user information collection module for obtaining the first multi-dimensional unstructured data of the target user based on the first basic information; the first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, e-commerce transaction records; a user information processing module for processing the first multi-dimensional unstructured data according to multiple preset time windows to generate first feature data under multiple time windows; a model input module for inputting the first feature data under multiple time windows into a preset user fraud risk identification model; wherein, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; the user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set; a model prediction module for processing the first feature data into a semantic feature vector by the natural language processing model; outputting the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data by the parameter server and the machine learning model.
[0016] In a third aspect, the present application further provides an electronic device, including a processor and a memory. The memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method described in the first aspect above.
[0017] In a fourth aspect, the present application further provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the method described in the first aspect above.
[0018] In the user fraud risk identification method, device, and electronic device provided by the present application, first, first basic information of a target user is obtained; the first basic information includes at least one of the following: name, ID number, mobile phone number, and device number; then, based on the first basic information, first multi-dimensional unstructured data of the target user is obtained; the first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, and e-commerce transaction records; then, the first multi-dimensional unstructured data is processed according to multiple preset time windows to generate first feature data under multiple time windows; finally, the first feature data under multiple time windows is input into a preset user fraud risk identification model; wherein, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; the user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set; the first feature data is processed into a semantic feature vector by the natural language processing model; and the fraud risk level of the target user is output by the parameter server and the machine learning model based on the semantic feature vector corresponding to the first feature data. The present application can identify potential fraud risks of users in a low-cost, low-time-consuming, and relatively simple framework without additional manual intervention, quickly classify the fraud risks of users, and complete the final decision. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a flowchart of a method for identifying user fraud risk provided by an embodiment of the present application;
[0021] Figure 2 It is a flowchart of a model training process provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic framework diagram of a model training process provided by an embodiment of the present application;
[0023] Figure 4Structural block diagram of a user fraud risk identification device provided by an embodiment of the present application;
[0024] Figure 5 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0025] Next, the technical solutions of the present application will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are some of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0026] In the credit field, common methods for judging user fraud risks include:
[0027] 1. Credit scoring model: Using data such as the user's credit history and repayment records to generate a credit score. The lower the score, the higher the fraud risk.
[0028] 2. Machine learning model: Using algorithms such as logistic regression, random forest, gradient boosting machine (GBM), neural network, etc., training the model through historical data to predict the fraud probability.
[0029] 3. Behavior analysis: Analyzing the user's behavior patterns, such as login frequency, device information, IP address, etc. Abnormal behaviors may indicate fraud risks.
[0030] 4. Social network analysis: By analyzing the user's social relationships, identifying fraud gangs or associated risks.
[0031] 5. Rule engine: Setting rules (such as applying for loans multiple times within a short period), users who trigger the rules are regarded as high-risk.
[0032] 6. Blacklist system: Listing known fraud users in the blacklist and rejecting their loan applications.
[0033] 7. Real-time monitoring and early warning: Real-time monitoring of transactions and behaviors, and giving early warnings in a timely manner when abnormalities are found.
[0034] 8. Multi-dimensional data verification: Verifying user information through third-party data (such as credit investigation agencies, operators, etc.), and giving risk warnings when there are inconsistencies.
[0035] 9. Manual review: Conducting manual reviews on high-risk users to further confirm the possibility of fraud.
[0036] 10. Fraud pattern recognition: By analyzing historical fraud cases, identifying common patterns and applying them to the risk assessment of new users.
[0037] These methods mentioned above are usually selectively used when identifying fraud risks, which are prone to cause the following three defects: 1. High cost: The systems to be built are huge and intricate, and there are dependencies between some systems, requiring continuous repeated investment, especially the server costs involved; 2. Low intelligence: In traditional credit application processes, especially on the bank side, most of the end processes require manual intervention, which does not conform to the current trend of AI participation in decision-making and greatly affects work efficiency; 3. Serious time-consuming problem: For the individual decision-making links of each system, whether in series or in parallel, the time-consuming problem needs to be considered. Relying solely on the way of stacking applications cannot meet the efficient credit decision-making process. If decisions are made according to the traditional system process, the final response time cannot reach a reasonable level (such as in the order of hundreds of milliseconds), ultimately resulting in poor user experience.
[0038] Based on this, the embodiments of the present application provide a method, device and electronic device for identifying user fraud risks, which can identify potential fraud risks of users in a framework with low cost, low time consumption and relatively simple structure, without the need for additional manual intervention, quickly classify the fraud risks for users and complete the final decision.
[0039] To facilitate the understanding of this embodiment, first, a method for identifying user fraud risks disclosed in the embodiments of the present application will be introduced in detail.
[0040] In the credit field, the definition of fraud is broad. In this embodiment, the behavior that the user fails to repay the money for 30 consecutive days in the first billing period after borrowing is considered, that is, the first overdue for 30 days (fpd30). Figure 1 The following is a flowchart of a method for identifying user fraud risks provided by the embodiments of the present application. The method specifically includes the following steps:
[0041] Step S102, obtain the first basic information of the target user; the first basic information includes at least one of the following: name, ID number, mobile phone number, device number;
[0042] Step S104, based on the first basic information, obtain the first multi-dimensional unstructured data of the target user; the first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, and e-commerce transaction records; these multi-dimensional unstructured data are the information obtained with the basic information of the user as the primary key.
[0043] Step S106, process the first multi-dimensional unstructured data according to multiple preset time windows to generate first feature data under multiple time windows;
[0044] This step processes data according to time windows to obtain corresponding feature data. Taking e-commerce transaction records as an example, the time windows here can be determined daily, weekly, or monthly. For example, process data within time windows such as the last 1 day, the last 1 week, 3 weeks, 5 weeks, 7 weeks, etc., and the last 3 months, 6 months, 12 months, etc. of some users, which is convenient for characterizing unstructured data in dimensions such as short / medium / long term of users, and storing it for subsequent vectorization processing.
[0045] Step S108: Input the first feature data under multiple time windows into a preset user fraud risk identification model; among them, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; the above natural language processing model can include: a GPT model or a BERT model; the machine learning model can include: an XGB model or an LGB model. The user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set.
[0046] Step S110: Process the first feature data into semantic feature vectors through the natural language processing model; output the fraud risk level of the target user based on the semantic feature vectors corresponding to the first feature data through the parameter server and the machine learning model.
[0047] Specifically, when implemented, the parameter server and the machine learning model process the semantic feature vectors corresponding to the first feature data to obtain a probability value; if the probability value is greater than the threshold, it is determined that the fraud risk level of the target user is a high risk level; if the probability value is less than or equal to the threshold, it is determined that the fraud risk level of the target user is a low risk level.
[0048] For example, the probability value determined by the parameter server and the machine learning model is 0.7, which is greater than the threshold 0.5, so 1 is output, indicating a high risk level.
[0049] The user fraud risk identification method provided by the embodiments of this application can identify potential fraud risks of users in a low-cost, low-time-consuming, and relatively simple framework through the above multiple consecutive processes of basic information acquisition, unstructured data acquisition, semantic feature vector determination, and model prediction, without the need for additional manual intervention, quickly classifying the fraud risks of users and making a final decision.
[0050] The following elaborates in detail on the training process of the above user fraud risk identification model:
[0051] See Figure 2 As shown, the training process includes the following steps:
[0052] Step S202: Obtain a training data set; the data in the training data set includes: second feature data under multiple time windows obtained by processing the second multi-dimensional unstructured data for multiple sample users.
[0053] In specific implementation, for each sample user, the second basic information of the sample user can be obtained; based on the second basic information, the second multi-dimensional unstructured data of the sample user can be obtained; the second multi-dimensional unstructured data is processed according to multiple preset time windows to generate second feature data under multiple time windows, thereby obtaining the training data set.
[0054] The process of obtaining this training data set is similar to the processing process for the target user described above and will not be elaborated here.
[0055] Step S204: Perform word segmentation, encoding, multi-layer Transformer processing, and feature extraction on the second feature data of multiple sample users through a natural language processing model to obtain semantic feature vectors corresponding to the second multi-dimensional unstructured data of multiple sample users respectively, and obtain the key parameters involved when generating the semantic feature vectors; the key parameters include at least one of the following: Transformer layer parameters, word embedding parameters, position encoding parameters; among them, the semantic feature vectors include lexical semantics and context information and are applicable to the processing of various downstream tasks or modules.
[0056] Step S206: Use the semantic feature vectors of multiple sample users and the fraud risk level labels respectively marked for multiple sample users as a training sample set, and input the key parameters into the parameter server.
[0057] Here, the semantic feature vector can be used as x of the user, and whether the user is finally fraudulent (label 1 or 0) can be used as y, and y can be determined based on the risk performance of the user (mainly fpd30) or determined through clustering. In practical applications, the above key parameters can be adjusted or not adjusted. If adjustment is required, for example, the number of layers of the transformer can be 12 layers or 24 layers, and the adjustment is based on the training effect of the model. The main function of the parameter server is still storage, and these parameters are recorded first.
[0058] Parameter Server: It is a programming framework mainly used to support the distributed storage and collaboration of large-scale parameters. In machine learning, training a model often generates a large number of model parameters, which need to be saved and shared for parallel computing and optimization among multiple computing nodes. The Parameter Server is used to manage and distribute these model parameters. For example, tree depth, subsample, learning rate, L1 regularization term, L2 regularization term, etc. The Parameter Server supports distributed training and real-time prediction and is the operating brain of machine learning.
[0059] In step S208, use the training sample set to train the model parameters of the machine learning model. During the training process, store and update the key parameters through the Parameter Server. After the training is completed, obtain the user fraud risk identification model; the model parameters include: tree depth, number of leaf nodes, learning rate, and L1 regularization term.
[0060] In this embodiment, the above-mentioned Parameter Server includes: a server side, a client side, and a scheduler; the server side is used to store model parameters, receive the gradients uploaded by the client side, and update the local parameters, responsible for the storage and update of parameters to ensure the consistency of model parameters; the client side is used to obtain the latest parameters from the server side, calculate the gradients using local data, and upload the gradients to the server side, responsible for calculating gradients and synchronizing model parameters; the scheduler is used to manage the server and client nodes, and complete functions such as data synchronization between nodes and node addition / deletion.
[0061] Figure 3 Shows the entire training framework diagram of the user fraud risk identification model, where the training and scoring module and the decision-making module actually correspond to the functions of the above-mentioned machine learning model. The decision-making module can convert the probability value output by the model into a score, usually between 1 and 100 points. By the level of the score, judge the fraud risk level of the user. The final trained result can be judged by referring to the threshold method. For example, if the finally trained score exceeds 90 points, it can be judged that the user fraud risk is relatively low, and it is recommended to judge the user as a low-risk user; if the trained score is lower than 60 points, judge that the user's fraud risk is relatively high, and it is recommended to judge the user as a high fraud risk user.
[0062] The user fraud risk identification method provided by the embodiments of the present application is a solution with high performance, perfect framework, and user-friendly. It can identify the potential fraud risk of users under a low-cost, low-time-consuming, and relatively simple framework without additional manual intervention, quickly classify the fraud risk of users and complete the final decision.
[0063] Based on the above method embodiments, the embodiments of the present application further provide a user fraud risk identification device. SeeFigure 4 As shown in Figure 4 , the device includes: a user information registration module 402, configured to obtain first basic information of a target user; the first basic information includes at least one of the following: name, ID number, mobile phone number, device number; a user information collection module 404, configured to obtain first multi-dimensional unstructured data of the target user based on the first basic information; the first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, e-commerce transaction record; a user information processing module 406, configured to process the first multi-dimensional unstructured data according to multiple preset time windows to generate first feature data under multiple time windows; a model input module 408, configured to input the first feature data under multiple time windows into a preset user fraud risk identification model; wherein, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; the user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set; a model prediction module 410, configured to process the first feature data into a semantic feature vector through the natural language processing model; output the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data through the parameter server and the machine learning model.
[0064] Further, the above device further includes: a model training module, configured to execute the following training process of the user fraud risk identification model: obtain a training data set; the data in the training data set includes: second feature data under multiple time windows obtained by processing second multi-dimensional unstructured data of multiple sample users; perform word segmentation, encoding, multi-layer Transformer processing, and feature extraction on the second feature data of multiple sample users through the natural language processing model to obtain semantic feature vectors respectively corresponding to the second multi-dimensional unstructured data of multiple sample users, and obtain key parameters involved when generating the semantic feature vectors; the key parameters include at least one of the following: Transformer layer parameters, word embedding parameters, position encoding parameters; wherein, the semantic feature vector includes lexical semantics and context information; use the semantic feature vectors of multiple sample users and the fraud risk level labels respectively corresponding to multiple sample users as a training sample set, and input the key parameters into the parameter server; use the training sample set to train the model parameters of the machine learning model, and during the training process, store and update the key parameters through the parameter server, and obtain the user fraud risk identification model after the training is completed; the model parameters include: tree depth, number of leaf nodes, learning rate, and L1 regularization term.
[0065] Further, the above-mentioned model training module is used to obtain the second basic information of each sample user; based on the second basic information, obtain the second multi-dimensional unstructured data of the sample user; process the second multi-dimensional unstructured data according to multiple preset time windows to generate second feature data under multiple time windows, and obtain a training dataset.
[0066] Further, the above-mentioned parameter server includes: a server side, a client side, and a scheduler; the server side is used to store model parameters, receive gradients uploaded by the client side, and update local parameters, responsible for parameter storage and update to ensure the consistency of model parameters; the client side is used to obtain the latest parameters from the server side, calculate gradients using local data, and upload the gradients to the server side, responsible for calculating gradients and synchronizing model parameters; the scheduler is used to manage server and client nodes, and complete functions such as data synchronization and node addition / deletion between nodes.
[0067] Further, the above-mentioned natural language processing model includes: GPT model or BERT model; the machine learning model includes: XGB model or LGB model.
[0068] Further, the above-mentioned model prediction module 410 is used to process the semantic feature vector corresponding to the first feature data through a parameter server and a machine learning model to obtain a probability value; if the probability value is greater than a threshold, determine that the fraud risk level of the target user is a high risk level; if the probability value is less than or equal to the threshold, determine that the fraud risk level of the target user is a low risk level.
[0069] The device provided in the embodiments of the present application has the same implementation principle and the same technical effects as those in the foregoing method embodiments. For a brief description, for the parts not mentioned in the embodiments of the device, reference may be made to the corresponding content in the foregoing method embodiments.
[0070] The embodiments of the present application also provide an electronic device, as Figure 5 shown, which is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 51 and a memory 50. The memory 50 stores computer executable instructions that can be executed by the processor 51, and the processor 51 executes the computer executable instructions to implement the above method.
[0071] In Figure 5 the illustrated embodiment, the electronic device further includes a bus 52 and a communication interface 53. Among them, the processor 51, the communication interface 53, and the memory 50 are connected through the bus 52.
[0072] Among them, the memory 50 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 53 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 52 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 52 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0073] The processor 51 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 51 or instructions in the form of software. The above-mentioned processor 51 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor 51 reads the information in the memory and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0074] The embodiments of the present application also provide a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above method. For specific implementation, reference can be made to the foregoing method embodiments and will not be elaborated herein.
[0075] The computer program products of the method, apparatus, and electronic device provided by the embodiments of the present application include a computer-readable storage medium storing program codes. The instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated herein.
[0076] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0077] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0078] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0079] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for identifying user fraud risks, characterized in that, The method includes: Obtaining the first basic information of the target user; the first basic information includes at least one of the following: name, ID number, mobile phone number, device number; Based on the first basic information, obtaining the first multi-dimensional unstructured data of the target user; the first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, and e-commerce transaction records; Processing the first multi-dimensional unstructured data according to multiple preset time windows to generate first feature data under multiple time windows; Inputting the first feature data under the multiple time windows into a preset user fraud risk identification model; wherein, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; the user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set; Processing the first feature data into a semantic feature vector through the natural language processing model; outputting the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data through the parameter server and the machine learning model.
2. The method according to claim 1, wherein The training process of the user fraud risk identification model is as follows: Obtaining a training data set; the data in the training data set includes: second feature data under multiple time windows obtained by processing the second multi-dimensional unstructured data of multiple sample users; Performing word segmentation, encoding, multi-layer Transformer processing, and feature extraction on the second feature data of multiple sample users through a natural language processing model to obtain semantic feature vectors corresponding to the second multi-dimensional unstructured data of multiple sample users respectively, and obtaining key parameters involved in generating the semantic feature vectors; wherein, the semantic feature vector includes lexical semantics and context information; Using the semantic feature vectors of multiple sample users and the fraud risk level labels respectively marked for multiple sample users as a training sample set, and inputting the key parameters into the parameter server; Training the model parameters of the machine learning model using the training sample set, and storing and updating the key parameters through the parameter server during the training process. After the training is completed, a user fraud risk identification model is obtained; the model parameters include: tree depth, number of leaf nodes, learning rate, and L1 regularization term.
3. The method according to claim 2, wherein The key parameters include at least one of the following: Transformer layer parameters, word embedding parameters, position encoding parameters.
4. The method according to claim 2, wherein The step of obtaining the training data set includes: For each sample user, obtaining the second basic information of the sample user; based on the second basic information, obtaining the second multi-dimensional unstructured data of the sample user; processing the second multi-dimensional unstructured data according to multiple preset time windows to generate second feature data under multiple time windows, thereby obtaining the training data set.
5. The method according to claim 2, characterized in that The parameter server includes: a server side, a client side, and a scheduler; The server side is used to store model parameters, receive gradients uploaded by the client side, and update local parameters, responsible for the storage and update of parameters to ensure the consistency of model parameters; The client is used to obtain the latest parameters from the server, calculate gradients using local data, and upload the gradients to the server, and is responsible for calculating gradients and synchronizing model parameters; The scheduler is used to manage the server and client nodes, and complete functions such as data synchronization and node addition / removal between nodes.
6. The method according to claim 1, wherein The natural language processing model includes: GPT model or BERT model; The machine learning model includes: XGB model or LGB model.
7. The method according to claim 1, wherein The step of outputting the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data by the parameter server and the machine learning model includes: Processing the semantic feature vector corresponding to the first feature data by the parameter server and the machine learning model to obtain a probability value; If the probability value is greater than the threshold, determine that the fraud risk level of the target user is a high risk level; If the probability value is less than or equal to the threshold, determine that the fraud risk level of the target user is a low risk level.
8. An identification device for user fraud risk, characterized in that, The device includes: A user information registration module, configured to obtain the first basic information of the target user; The first basic information includes at least one of the following: name, ID number, mobile phone number, device number; A user information collection module, configured to obtain the first multi-dimensional unstructured data of the target user based on the first basic information; The first multi-dimensional unstructured data includes: user browsing behavior, address book marking situation, e-commerce transaction record; A user information processing module, configured to process the first multi-dimensional unstructured data according to multiple preset time windows to generate first feature data under multiple time windows; A model input module, configured to input the first feature data under multiple time windows into a preset user fraud risk identification model; Among them, the user fraud risk identification model includes: a natural language processing model, a parameter server, and a machine learning model; The user fraud risk identification model is obtained by training the parameter server and the machine learning model with a training data set; A model prediction module, configured to process the first feature data into a semantic feature vector through the natural language processing model; Output the fraud risk level of the target user based on the semantic feature vector corresponding to the first feature data by the parameter server and the machine learning model.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by the processor, the computer executable instructions cause the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Financial field consumption fraud detection method based on big data
CN121258522A