Credit risk prediction method and device, electronic equipment and storage medium
This credit risk prediction method, which optimizes scoring weights using large language models and genetic algorithms, solves the problem of insufficient utilization of unstructured data by traditional risk control models. It improves the accuracy and adaptability of credit risk prediction and meets the complex risk control needs of internet finance.
Patent Information
- Application Number
- CN202511183225.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-04
AI Technical Summary
Traditional risk control models struggle to effectively utilize unstructured data, resulting in one-sided risk identification, poor model adaptability, and an inability to meet the complex and diverse risk control needs of internet finance.
A large language model is introduced to perform semantic understanding and feature extraction on unstructured data, generating multiple risk sub-scores. The score weights are optimized through a genetic algorithm, and the model is self-updated by a feedback learning mechanism.
It improves the accuracy and adaptability of credit risk prediction, ensures the scientific nature of credit decisions, and meets the rapidly changing needs of internet finance risk control.
Smart Images

Figure CN120894129A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a credit risk prediction method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the rapid development of Internet financial business, the data forms of users in online lending are increasingly diversified, including structured numerical data and unstructured text, behavior logs, etc. Traditional risk control models are usually based on statistics and rule engines, mainly using classic methods such as logistic regression, decision tree, and score card to predict user loan risks. These methods generally rely on explicit structured data features, such as user income status, credit records, historical repayment situations, etc., to combine risk scores with manually set weights. However, with the continuous enrichment of user data dimensions, especially the emergence of a large amount of unstructured data such as text information and social behavior data, traditional risk control methods gradually show limitations such as insufficient data utilization, one-sided risk identification, poor model adaptability, and are difficult to meet the complex and diverse needs of Internet financial risk control. SUMMARY
[0003] The present application provides a credit risk prediction method, device, electronic device, and storage medium, which can improve the accuracy of credit risk prediction. The technical solution is as follows: According to one aspect of the present application, a credit risk prediction method is provided, the method comprising: obtaining target credit data of a target object; inputting the target credit data into a trained large language model to obtain a plurality of risk sub-scores output by the trained large language model, different risk sub-scores being risk values of the target object in different risk dimensions, the risk dimensions at least including a behavior dimension, a consumption dimension, and a trust dimension; determining a credit overdue risk value of the target object based on the plurality of risk sub-scores and a score weight corresponding to each risk sub-score.
[0004] According to another aspect of the present application, a credit risk prediction device is provided, the device comprising: a first obtaining module configured to obtain target credit data of a target object; a first prediction module configured to input the target credit data into a trained large language model to obtain a plurality of risk sub-scores output by the trained large language model, different risk sub-scores being risk values of the target object in different risk dimensions, the risk dimensions at least including a behavior dimension, a consumption dimension, and a trust dimension; The determination module is used to determine the credit delinquency risk value of the target object based on the multiple risk sub-scores and the scoring weights corresponding to each risk sub-score.
[0005] According to one aspect of this application, an electronic device is provided, comprising: a processor and a memory storing a program, the program including instructions that, when executed by the processor, cause the processor to perform the credit risk prediction method as described above.
[0006] According to another aspect of this application, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the credit risk prediction method as described above.
[0007] According to another aspect of this application, a computer program product is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned credit risk prediction method.
[0008] The beneficial effects of the technical solutions provided in this application include at least the following: By incorporating the understanding capabilities of a large language model for unstructured data, rich multi-dimensional risk signals are extracted from target credit data to obtain multiple quantified risk sub-scores. Furthermore, an intelligent optimization algorithm generates multiple score weights to integrate the multi-dimensional risk sub-scores, ensuring the accuracy of the final credit delinquency risk value and thus guaranteeing the scientific and adaptable nature of credit decisions. Finally, a feedback learning mechanism enables the large language model to self-update, meeting the rapidly changing risk control needs of internet finance. Attached Figure Description
[0009] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 A flowchart of a credit risk prediction method according to an exemplary embodiment of this application is shown; Figure 2 This is a flowchart of a method for generating scoring weights provided in an exemplary embodiment of this application; Figure 3 This is a schematic diagram of the structure of a credit risk prediction device provided in an embodiment of this application; Figure 4 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of this application is shown. Detailed Implementation
[0010] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0011] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0012] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications "a" and "a plurality" mentioned in this application are illustrative and not restrictive, and those skilled in the art should understand that unless explicitly indicated in the context, they should be understood as "one or more". The names of messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0013] The following description of the present application's solution, with reference to the accompanying drawings, provides a detailed explanation of the technical solutions provided by the embodiments of the present application through specific examples and application scenarios.
[0014] Please refer to Figure 1 The document illustrates a flowchart of a credit risk prediction method according to an exemplary embodiment of this application. An example of this method applied to an electronic device is provided for illustrative purposes. Figure 1 As shown, the method includes: Step 101: Obtain the target credit data of the target object.
[0015] Targeted credit data includes various information and behavioral data provided by the target applicant when applying for a loan, such as: a description of the loan purpose, past transaction and repayment records, user spending records, social media reviews, etc. For example, the target applicant could be a target user.
[0016] In one possible implementation, when assessing whether to allow a target individual to borrow money, target credit data such as a description of the loan purpose, past transactions, repayment records, and consumption records can be obtained, provided the target individual grants permission. It should be noted that target credit data must be obtained with user authorization.
[0017] Step 102: Input the target credit data into the trained large language model to obtain multiple risk sub-scores output by the trained large language model. Different risk sub-scores are the risk values of the target object in different risk dimensions. The risk dimensions include at least the behavioral dimension, consumption dimension and trust dimension.
[0018] The trained Large Language Model (LLM) is used for semantic understanding and feature extraction of target credit data. On the one hand, LLM can transform textual descriptions into numerical representations, capturing implicit information such as sentiment, behavioral patterns, and credit quality. On the other hand, LLM can also combine structured historical behavioral data to analyze patterns in user behavior sequences. For example, the model might focus on whether a user's wording when describing the purpose of a loan is cautious and honest, and combine this with their historical consumption structure to judge their financial habits. Through this analysis, the LLM module outputs multiple risk sub-scores, each representing an evaluation on a specific risk dimension.
[0019] One possible implementation involves inputting the target credit data into a trained large language model. The model performs semantic understanding and feature analysis, outputting multiple risk sub-scores, where each sub-score represents the target object's risk value across different risk dimensions. These risk dimensions can include behavioral, consumption, trust, and other dimensions. Correspondingly, the risk sub-scores can include: a behavioral pattern score (a risk sub-score on the behavioral dimension: measuring the stability and reliability of a user's borrowing and repayment patterns), a consumption structure score (a risk sub-score on the consumption dimension: measuring the impact of a user's income and expenditure structure and consumption habits on their debt repayment ability), a trust feature score (a risk sub-score on the trust dimension: assessing the user's creditworthiness based on the credibility of their text descriptions, historical credit records, etc.), and other feature scores (risk sub-scores on other dimensions: user risks that can be considered in addition to the factors mentioned above).
[0020] Optionally, the values of each risk sub-score can be standardized to a certain range (e.g., 0-1), with higher values indicating higher risk in that dimension. Generating multiple sub-scores instead of a single score helps maintain a fine-grained analysis of risk and improves model interpretability.
[0021] Before applying the trained large language model to predict risk sub-scores, further training of the large language model may be necessary. The training data used in this process can include collecting raw user lending information, such as loan purpose descriptions, customer-sales dialogues, and supplementary materials uploaded by the user. Simultaneously, time series data such as user transaction history, monthly income and expenditure structure, repayment schedule, and asset-liability snapshots should be collected. This information constitutes the input for training the LLM. This information collectively depicts the user's "linguistic features + behavioral patterns" in the lending scenario. The analysis conclusions of advanced models (such as DeepSeek-R1) are sampled, and existing rule-based models or expert experience are combined with the advanced model analysis to provide four initial sub-scores (behavioral pattern score, consumption structure score, trust characteristic score, and other characteristic score) for each sample as labels. These sub-scores are then combined with the analysis conclusions to form the output of the trained LLM. The overall scale is approximately 100,000 samples.
[0022] Training Methodology: A supervised fine-tuning process was adopted, using an open-source LLM model, such as Qwen3-32B, that aligns with the Chinese financial context. The model makes predictions on given input data and compares these predictions with standard output data, calculating word-level cross-entropy to ensure the model generates four complete sub-segments while maintaining template format consistency. Each sub-segment is then trained directly as an independent continuous value, enabling the LLM to initially accurately rate users. Dimensions with larger errors are given greater weight to ensure effective model training. Finally, the model's training status is assessed by monitoring its loss curves on the training and testing sets, combined with the Spearman correlation coefficients of each sub-segment and the output format integrity rate, to select the optimal model.
[0023] Step 103: Determine the credit delinquency risk value of the target object based on multiple risk sub-scores and the corresponding score weights of each risk sub-score.
[0024] After obtaining risk sub-scores across multiple risk dimensions, the impact of different risk sub-scores on the final credit delinquency risk may vary. Therefore, it is necessary to synthesize multiple risk sub-scores to obtain the overall credit delinquency risk value for the target entity. Corresponding scoring weights are provided for each risk sub-score. The sum of the products of each risk sub-score and its scoring weight is used to determine the credit delinquency risk value for the target entity.
[0025] For example, the method for determining the credit delinquency risk value can be as shown in formula (1): (1) in, This represents the i-th risk sub-score. The combined credit delinquency risk value. , ,... A set of scoring weights.
[0026] Unlike related technologies where the weights of each risk sub-score are manually set, this application's embodiments use a genetic algorithm (GA) to automatically search for and obtain the optimal set of score weights. Please refer to... Figure 2 This is a flowchart of a method for generating scoring weights provided in an exemplary embodiment of this application. The method includes the following steps: Step 201: Obtain historical credit data and credit delinquency tags for multiple sample objects.
[0027] Step 202: For each sample object, input the historical credit data into the trained large language model to obtain multiple sample sub-scores output by the trained large language model.
[0028] Step 203: Based on multiple sample sub-scores, credit delinquency labels, and multiple sets of candidate score weights, a genetic algorithm is used to generate the score weights for each risk sub-score.
[0029] Specifically, step 203 may also include steps 203A to 203C.
[0030] Step 203A: Calculate the sample delinquency risk value corresponding to each sample object based on multiple sample sub-scores and candidate score weights; Step 203B: Based on the sample delinquency risk value and credit delinquency label, determine the candidate fitness value corresponding to the weights of multiple candidate scores; Step 203C: Based on the candidate fitness values, the score weights of each risk sub-score are generated through iterative genetic algorithm.
[0031] The process of using a genetic algorithm to search for weights and obtain a set of optimal rating weights can be described as follows: (1) Initialize the population: randomly generate several sets of candidate weights w1, w2, ... w N (2) Evaluation of fitness: For each group of candidate weights, calculate the final score S of all training samples. final , and evaluate the prediction error or accuracy as the fitness value; (3) Iterative cycle: select a weighted group with high fitness from the current population as the parent generation; perform cross-combination on the selected weighted group to generate a new weighted group; randomly fine-tune (mutate) the new weighted group to increase diversity; calculate the fitness of each weighted group in the new population, and select the best performers to enter the next generation; cycle for several generations until convergence or the iteration limit is reached; output the optimal weighted combination as the final model parameters; after optimization, the module outputs a final numerical risk score as a quantitative assessment of the user's overall default risk.
[0032] In one possible implementation, several training data pairs are first acquired, each pair including historical credit data and a credit delinquency label corresponding to the sample object. Then, the historical credit data is input into the trained large language model to obtain multiple sample sub-scores for each sample object. For the initialized multiple sets of candidate score weights, the sample delinquency risk value of the sample object under each set of candidate score weights is calculated based on the multiple sample sub-scores and each set of candidate score weights. Next, for each set of candidate score weights, a candidate fitness value is determined based on the sample delinquency risk value and the credit delinquency label. This candidate fitness value can be the prediction accuracy or the loss value between the two. Then, based on the candidate fitness value, a weight set with high fitness is selected from the multiple sets of candidate score weights as the parent set. Based on this parent set, subsequent iterative processes are executed until the optimal weight combination is output as the final score weight.
[0033] Once the credit delinquency risk value of the target entity is determined, corresponding risk control decisions can be made based on this value. One possible implementation involves setting thresholds. For example, if the credit delinquency risk value is below a first threshold, it indicates that the risk is controllable, and a first credit decision can be executed, allowing the target entity to engage in credit activity. Conversely, if the credit delinquency risk value is above a second threshold, it indicates that the user's delinquency risk is high, and a second credit decision can be executed, disallowing the target entity from engaging in credit activity.
[0034] Optionally, the credit delinquency risk value and each risk sub-score can also be provided to risk control personnel for reference, enabling human-machine collaborative decision-making. For example, if a user's final risk score is near the critical value, risk control personnel can review the composition of each risk sub-score and find that the user may have a good "behavioral pattern score" but a high "trust characteristic score," which means that the user's written description may be unreliable. This will assist humans in more accurately assessing the risk. This module ensures the interpretability of the decision-making process.
[0035] After the initial credit decision is made, the user's actual repayment performance will be fed back as a real result after a period of time. Specifically, if the user obtains a loan, it is observed whether they repay on time, whether there are any overdue payments, and the severity of the overdue payments (e.g., number of overdue days or whether it ultimately becomes a bad debt). This feedback module compares the actual repayment results with the risk score predicted by the previous model, generating an error signal or evaluation metric. Specifically, the error can be calculated based on whether there was an actual default (0 or 1) and the default probability predicted by the model. If the model prediction deviates significantly from the actual result, the model needs to be adjusted accordingly. This feedback information is compiled into training samples (features are the user's contextual data, and labels are the actual results) and sent to the next module for model updates.
[0036] The model is updated and trained using error backpropagation, forming a closed-loop optimization. The update process mainly involves updating the LLM parameters. For LLM model training, a supervised fine-tuning + preference fine-tuning approach is adopted. Newly obtained samples are added to the LLM training set to reduce the error between the predicted risk score and the actual risk. Specifically, subsets and model responses that are correctly judged in the final result are selected as chosen (correct) responses, and subsets and model responses that are incorrectly judged in the final result are selected as rejected (incorrect) responses. Preference pairs are constructed, and the Direct Preference Optimization (DPO) algorithm is used for optimization training. At the same time, a second round of imitation training is performed on the chosen responses using supervised fine-tuning methods to directly improve the model's ability to make correct decisions. With the continuous feedback of new data, LLM can gradually correct its judgment biases on various risk dimensions and improve the accuracy of sub-scores. After multiple iterations, the entire system forms a closed loop of "data-model-decision-feedback", and its risk discrimination ability evolves on its own.
[0037] In one possible implementation, after executing the first credit decision, the actual credit outcome of the target entity can be obtained. This actual credit outcome could include whether repayment was made on time, whether delinquency occurred, and the severity of the delinquency (e.g., number of delinquent days or whether it ultimately resulted in a bad debt). Then, by comparing the actual credit outcome with the credit delinquency risk value, it can be determined whether the large language model needs optimization. Specifically, an error threshold can be set; if the error between the actual credit outcome and the credit delinquency risk value exceeds the error threshold, the process of optimizing the large language model can be executed.
[0038] Specifically, the process of optimizing a large language model after training can include the following steps: First, obtain target credit data and multiple risk sub-scores that match actual credit results with credit delinquency risk values, and determine them as positive training samples.
[0039] The actual credit outcome matches the credit delinquency risk value when the actual credit outcome indicates no delinquency has occurred and the credit delinquency risk value is also below the first threshold. In this matching scenario, the target credit data from which the predicted credit delinquency risk value is obtained, along with multiple risk sub-scores, can be used as positive training samples for the large language model. These samples can then be used for subsequent imitation training to directly improve the large language model's ability to make correct decisions.
[0040] Second, obtain target credit data and multiple risk sub-scores that do not match the actual credit results and credit delinquency risk values, and identify them as negative training samples.
[0041] A mismatch occurs when the actual credit outcome indicates an overdue payment, but the overdue risk value is below the first threshold. In such cases, the target credit data predicting the overdue risk value and multiple risk sub-scores can be used as negative training samples for the large language model. Preference pairs are then constructed, and the Direct Preference Optimization (DPO) algorithm is used for optimization training.
[0042] Third, optimize the trained large language model based on positive and negative training samples.
[0043] By continuously optimizing the large language model through positive and negative training samples, LLM can gradually correct its judgment bias on various risk dimensions and improve the prediction accuracy of risk sub-scores.
[0044] In summary, the embodiments of this application provide a credit risk prediction method: by introducing the understanding ability of a large language model for unstructured data, rich multi-dimensional risk signals are extracted from the target credit data to obtain multiple quantified risk sub-scores; furthermore, multiple score weights are generated through intelligent optimization algorithms to integrate the multi-dimensional risk sub-scores, ensuring the accuracy of the final generated credit delinquency risk value, thereby guaranteeing the scientific nature and adaptability of credit decisions; finally, by leveraging a feedback learning mechanism, the large language model is self-updated to meet the rapidly changing risk control needs of internet finance.
[0045] Please refer to Figure 3 This is a schematic diagram of the structure of a credit risk prediction device provided in an embodiment of this application. For example, as shown... Figure 3 As shown, the device 300 includes...
[0046] The first acquisition module 301 is used to acquire the target credit data of the target object; The first prediction module 302 is used to input the target credit data into a trained large language model to obtain multiple risk sub-scores output by the trained large language model. The different risk sub-scores are the risk values of the target object in different risk dimensions. The risk dimensions include at least behavioral dimension, consumption dimension and trust dimension. The determination module 303 is used to determine the credit delinquency risk value of the target object based on the multiple risk sub-scores and the scoring weights corresponding to each risk sub-score.
[0047] Optionally, the device further includes: The second acquisition module is used to acquire historical credit data and credit delinquency tags of multiple sample objects; The second prediction module is used to input the historical credit data into the trained large language model for each sample object to obtain multiple sample sub-scores output by the trained large language model. The generation module is used to generate the score weights of each risk sub-score based on the multiple sample sub-scores, the credit delinquency label, and multiple sets of candidate score weights using a genetic algorithm.
[0048] Optionally, the generation module is further configured to: Based on the multiple sample sub-scores and the candidate score weights, calculate the sample delinquency risk value corresponding to each sample object; Based on the sample delinquency risk value and the credit delinquency label, determine the candidate fitness value corresponding to the multiple sets of candidate scoring weights; Based on the candidate fitness values, the score weights of each risk sub-score are generated iteratively using the genetic algorithm.
[0049] Optionally, the device further includes: A first execution module is configured to execute a first credit decision when the credit delinquency risk value is lower than a first threshold, wherein the first credit decision indicates that the target object is permitted to engage in credit activities; The second execution module is used to execute a second credit decision when the credit delinquency risk value is higher than a second threshold, the second credit decision indicating that the target object is not allowed to engage in credit activities.
[0050] Optionally, the device further includes: The third acquisition module is used to acquire the actual credit result of the target object after executing the first credit decision; The training module is used to optimize the trained large language model when the error between the actual credit result and the credit delinquency risk value exceeds an error threshold.
[0051] Optionally, the training module is further configured to: Obtain target credit data and multiple risk sub-scores that match the actual credit results with the credit delinquency risk values, and determine them as positive training samples; Obtain target credit data and multiple risk sub-scores that do not match the actual credit results and the credit delinquency risk values, and determine them as negative training samples; Based on the positive training samples and the negative training samples, the trained large language model is optimized. In summary, the embodiments of this application provide a credit risk prediction method: by introducing the understanding ability of a large language model for unstructured data, rich multi-dimensional risk signals are extracted from the target credit data to obtain multiple quantified risk sub-scores; furthermore, multiple score weights are generated through intelligent optimization algorithms to integrate the multi-dimensional risk sub-scores, ensuring the accuracy of the final generated credit delinquency risk value, thereby guaranteeing the scientific nature and adaptability of credit decisions; finally, by leveraging a feedback learning mechanism, the large language model is self-updated to meet the rapidly changing risk control needs of internet finance.
[0052] An exemplary embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, causing the electronic device to perform a credit risk prediction method according to an embodiment of this application.
[0053] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a credit risk prediction method according to an embodiment of this application.
[0054] An exemplary embodiment of this application also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a credit risk prediction method according to an embodiment of this application.
[0055] refer to Figure 4 The present invention describes a structural block diagram of an electronic device 400 that can serve as a server or client of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein.
[0056] like Figure 4As shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0057] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information to electronic device 400. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0058] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, Figure 1 , Figure 2 The method shown can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 400 via ROM 402 and / or communication unit 409. In some embodiments, computing unit 401 can be configured to execute by any other suitable means (e.g., by means of firmware). Figure 1 , Figure 2 The method shown.
[0059] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0060] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0061] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0062] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0063] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0064] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. A credit risk prediction method, characterized in that, The method includes: Obtain the target credit data of the target object; The target credit data is input into the trained large language model to obtain multiple risk sub-scores output by the trained large language model. Different risk sub-scores are the risk values of the target object in different risk dimensions. The risk dimensions include at least behavioral dimension, consumption dimension and trust dimension. Based on the multiple risk sub-scores and the corresponding score weights for each risk sub-score, the credit delinquency risk value of the target object is determined.
2. The method according to claim 1, characterized in that, The method further includes: Obtain historical credit data and credit delinquency tags for multiple sample objects; For each of the sample objects, the historical credit data is input into the trained large language model to obtain multiple sample sub-scores output by the trained large language model; Based on the multiple sample sub-scores, the credit delinquency label, and multiple sets of candidate score weights, the score weights for each of the risk sub-scores are generated using a genetic algorithm.
3. The method according to claim 2, characterized in that, The step of generating the score weights for each risk sub-score using a genetic algorithm based on the multiple sample sub-scores, the credit delinquency label, and multiple sets of candidate score weights includes: Based on the multiple sample sub-scores and the candidate score weights, calculate the sample delinquency risk value corresponding to each sample object; Based on the sample delinquency risk value and the credit delinquency label, determine the candidate fitness value corresponding to the multiple sets of candidate scoring weights; Based on the candidate fitness values, the score weights of each risk sub-score are generated iteratively using the genetic algorithm.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: If the credit delinquency risk value is lower than a first threshold, a first credit decision is executed, which instructs the target to engage in credit activities. If the credit delinquency risk value is higher than the second threshold, a second credit decision is executed, which indicates that the target object is not allowed to engage in credit activities.
5. The method according to claim 4, characterized in that, The method further includes: After executing the first credit decision, the actual credit result of the target object is obtained; If the error between the actual credit result and the credit delinquency risk value exceeds an error threshold, the trained large language model is optimized.
6. The method according to claim 5, characterized in that, The optimization of the trained large language model also includes: Obtain target credit data and multiple risk sub-scores that match the actual credit results with the credit delinquency risk values, and determine them as positive training samples; Obtain target credit data and multiple risk sub-scores that do not match the actual credit results and the credit delinquency risk, and determine them as negative training samples; Based on the positive training samples and the negative training samples, the trained large language model is optimized.
7. A credit risk prediction device, characterized in that, The device includes: The first acquisition module is used to acquire the target credit data of the target object; The first prediction module is used to input the target credit data into a trained large language model to obtain multiple risk sub-scores output by the trained large language model. The different risk sub-scores are the risk values of the target object in different risk dimensions. The risk dimensions include at least behavioral dimension, consumption dimension and trust dimension. The determination module is used to determine the credit delinquency risk value of the target object based on the multiple risk sub-scores and the scoring weights corresponding to each risk sub-score.
8. The apparatus according to claim 7, characterized in that, The device further includes: The second acquisition module is used to acquire historical credit data and credit delinquency tags of multiple sample objects; The second prediction module is used to input the historical credit data into the trained large language model for each sample object to obtain multiple sample sub-scores output by the trained large language model. The generation module is used to generate the score weights of each risk sub-score based on the multiple sample sub-scores, the credit delinquency label, and multiple sets of candidate score weights using a genetic algorithm.
9. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the credit risk prediction method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the credit risk prediction method according to any one of claims 1-6.