Financial risk control collection behavior score card construction method and device, equipment, medium and product
By analyzing the semantics of collection texts using large language models and retrieval-enhanced generation techniques, and combining prompt words and historical cases, the problem of unstructured feature modeling and knowledge reuse in financial risk control collection behavior scoring cards was solved. This enabled the construction of efficient and interpretable scoring cards, improving their accuracy and applicability.
Patent Information
- Application Number
- CN202511091652.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Existing financial risk control and debt collection behavior scoring cards have deficiencies in unstructured feature modeling, are unable to analyze implicit risk signals in debt collection texts, lack knowledge reuse and reasoning capabilities, have lagging dynamic optimization mechanisms, and complex models cannot provide interpretability.
Unstructured text is parsed using a large language model and prompt word templates. Similar historical cases are retrieved from the knowledge base using retrieval enhancement generation technology. Scores are calculated by weighting structured and semantic features and case reference scores to construct a financial risk control and debt collection behavior scorecard.
It improved feature extraction accuracy by 40%, shortened the rule iteration cycle to 24 hours, reduced the development cost of cross-product scorecards, met regulatory interpretability requirements, and improved the accuracy and applicability of scorecards.
Smart Images

Figure CN120996925A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of financial risk control management, and particularly relates to a financial risk control collection behavior scorecard construction method, device, equipment, medium and product. BACKGROUND
[0002] The financial risk control collection behavior scorecard is a core tool for providing decision basis for differentiated collection strategies by quantifying the risk characteristics (such as repayment willingness, performance ability and disconnection probability) of debtors in collection interaction; its technical essence is a risk characteristic quantification mapping system, and is widely applied to post-loan management links of credit products such as consumer loans and business loans, and directly affects the collection efficiency and bad debt rate control.
[0003] At present, the existing financial risk control collection behavior scorecards mainly have the following two types: one is a rule engine driven scorecard: specifically, a score rule is constructed by using IF-THEN logic, for example: if the overdue days > 90 days and the last 3 times of collection have no response → score ≤ 400 points (high risk); the score depends on artificial preset dimensions (such as overdue days and historical repayment times, etc.) and weights, and a typical scheme is a 15-dimensional fixed weight model of a certain bank, and the artificial update is performed every quarter; the other is a machine learning prediction type scorecard: specifically, a score model is trained based on algorithms such as logistic regression and XGBoost, for example, in a certain patent disclosed technical scheme, the input features are structured data (such as credit inquiry times and monthly income stability, etc.), the model output is a repayment probability mapping value of 0-1000 points, and the weight is determined by feature importance evaluation (such as "the number of overdue times in the last 6 months" weight 0.23).
[0004] However, the construction process of the above two types of financial risk control collection behavior scorecards has the following defects:
[0005] (1) Non-structured feature modeling is missing: the existing technology cannot analyze the implicit risk signals in the collection text, for example: the text "I was laid off recently, and I can find a job next month" contains "unemployment" event and "explicit repayment time", but the traditional model can only extract the keyword feature of "no repayment commitment", resulting in score deviation ≥ 18%;
[0006] (2) Lack of knowledge reuse and reasoning ability: there is no structured sedimentation of historical cases, and the score rules of similar scenarios need to be repeatedly developed, and the cross-case score consistency is < 70%; the model can only make statistical prediction based on the input features, and cannot combine domain knowledge (such as "business loan customers are sensitive to policy subsidies") for logical reasoning;
[0007] (3) Dynamic optimization mechanism lags behind: rule iteration relies on manual feature engineering, and the response to market changes (such as economic policy adjustments) is more than 30 days; no real-time feedback loop is established between "score output and collection effect", and the feature weight update lags behind the actual risk fluctuations;
[0008] (4) Contradiction between explainability and complexity: complex models (such as deep learning models) can improve prediction accuracy, but cannot output "the specific reason why a certain customer's score is low", which violates the explainability requirements of relevant laws. SUMMARY
[0009] The purpose of the present application is to provide a financial risk control collection behavior scorecard construction method, device, computer equipment, computer readable storage medium and computer program product, to solve the problems of missing non-structured feature modeling and insufficient knowledge reuse and reasoning ability in the existing financial risk control collection behavior scorecard construction method.
[0010] In order to achieve the above purpose, the present application adopts the following technical solutions:
[0011] In a first aspect, a financial risk control collection behavior scorecard construction method is provided, comprising:
[0012] Obtaining multi-source heterogeneous data generated by a target debtor and in the process of collection interaction for a target credit product;
[0013] According to the structured field content in the multi-source heterogeneous data, a structured feature score for quantifying the financial risk control collection behavior score from the structured feature dimension is processed;
[0014] According to the unstructured text in the multi-source heterogeneous data, a structured semantic feature is output by applying a semantic feature analysis method based on a large language model and a prompt word template, and a semantic feature score for quantifying the financial risk control collection behavior score from the semantic feature dimension is processed by combining a preset semantic feature quantization rule;
[0015] According to the unstructured text in the multi-source heterogeneous data, a similar historical collection case is retrieved from a knowledge base by applying a retrieval enhancement generation technology, and a case reference score for quantifying the financial risk control collection behavior score from the similar case reference dimension is processed based on the similar historical collection case;
[0016] According to the structured feature score, the semantic feature score and the case reference score, the financial risk control collection behavior score P of the target debtor on the target credit product is calculated according to the following formula:
[0017] P = α × p structured + β × p semantic + γ × p case
[0018] wherein p structured denotes the structured feature score, p semantic denotes the semantic feature score, p case denotes the case reference score, and a, b and g respectively denote preset weight coefficients and have a+b+g=1;
[0019] According to the financial risk control collection behavior score P, a financial risk control collection behavior score card is constructed for the target debtor and the target credit product and outputted.
[0020] Based on the above invention content, a new scheme for constructing a financial risk control collection behavior score card based on a large language model and combining prompt word guidance and retrieval enhancement generation technology is provided, that is, after obtaining the target debtor and the multi-source heterogeneous data generated in the collection interaction process for the target credit product, on the one hand, the structured feature score is obtained according to the structured field content therein, and on the other hand, the structured semantic feature is obtained based on the large language model and the prompt word template according to the unstructured text therein, and the semantic feature score is obtained by processing according to the semantic feature quantization rule, and the case reference score is obtained by retrieving similar historical collection cases by using the retrieval enhancement generation technology, and finally the final score is calculated based on the weighting of the three scores, and the score card construction is completed based on the calculation result, so that the LLM can deeply analyze the collection text semantics, and the implicit risk features can be accurately extracted combined with the prompt words, the limitation of structured data can be broken through, the feature extraction accuracy of the traditional keyword matching can be improved by 40%, the historical cases can be retrieved by using the RAG technology to enhance the knowledge reuse and reasoning ability, and then the accuracy of the score card construction result is improved, which is beneficial to the post-loan collection strategy decision and management, and is convenient for practical application and popularization.
[0021] In one possible design, the multi-source heterogeneous data obtained for the target debtor and generated in the collection interaction process for the target credit product includes:
[0022] The structured field content and the unstructured text obtained for the target debtor and generated in the collection interaction process for the target credit product, wherein the structured field content contains the overdue days, the historical repayment rate and / or the credit limit utilization rate, and the unstructured text contains the collection call text converted by using the automatic speech recognition technology and / or the collection chat record based on the social tool communication;
[0023] The data cleaning processing is performed on the unstructured text to obtain new unstructured text, wherein the data cleaning processing includes the greeting filtering processing and / or the semantic fuzzy expression correction processing;
[0024] aggregate the structured field content and the new unstructured text to obtain multi-source heterogeneous data generated for the target debtor and in a collection interaction process for the target credit product.
[0025] In one possible design, the prompt word template includes a basic prompt word and a dynamic constraint prompt word appended according to a type of the target credit product, where the basic prompt word is used to prompt output of risk features such as a key event, a repayment intention, and an emotional tendency.
[0026] In one possible design, according to unstructured text in the multi-source heterogeneous data, a retrieval enhancement generation technique is applied to retrieve a similar historical collection case from a knowledge base, and a case reference score for quantifying a financial risk control collection behavior score from a similar case reference dimension is processed based on the similar historical collection case, including:
[0027] According to unstructured text in the multi-source heterogeneous data, a Sentence-BERT model is used to generate a current text vector corresponding to the unstructured text;
[0028] According to the current text vector, at least one similar historical collection case is retrieved from a knowledge base based on a pre-established FAISS index, where the knowledge base includes a case sub-database and a rule sub-database, the case sub-database is used to store record information of all historical collection cases, the record information includes a text vector and a score adjustment record of a corresponding case, and the rule sub-database is used to store risk control rules and characteristics of different credit products;
[0029] According to the score adjustment record in the record information of the at least one similar historical collection case, the risk control rules are adjusted to obtain adjusted risk control rules;
[0030] According to characteristics corresponding to the target credit product, the adjusted risk control rules are corrected to obtain corrected risk control rules;
[0031] According to the structured semantic features, the corrected risk control rules are combined to process a case reference score for quantifying a financial risk control collection behavior score from a similar case reference dimension.
[0032] In one possible design, when the risk control rules are adjusted according to the score adjustment record in the record information of the at least one similar historical collection case, the method further includes: according to the score adjustment record in the record information of the at least one similar historical collection case, a natural language explanation for expressing an adjustment reason is extracted.
[0033] When constructing the financial risk control collection behavior scorecard, the method further comprises: adding the natural language explanation for expressing the adjustment reason to the financial risk control collection behavior scorecard.
[0034] In one possible design, after constructing and outputting the financial risk control collection behavior scorecard, the method further comprises:
[0035] Collecting actual feedback data for the financial risk control collection behavior scorecard;
[0036] According to the actual feedback data, new historical collection cases are sorted out and added to the knowledge base, and if it is found that the absolute value of the deviation between the predicted repayment rate and the actual repayment rate of the target debtor on the target credit product exceeds a preset threshold, the PPO algorithm is used to update and adjust the semantic feature quantization rule and / or all weight coefficients for calculating the financial risk control collection behavior score P.
[0037] In a second aspect, a financial risk control collection behavior scorecard construction device is provided, comprising a heterogeneous data acquisition unit, a structured scoring unit, a semantic analysis scoring unit, a case retrieval scoring unit, a risk control scoring calculation unit, and a scorecard construction unit;
[0038] The heterogeneous data acquisition unit is configured to acquire multi-source heterogeneous data generated by a target debtor in a collection interaction process for a target credit product.
[0039] The structured scoring unit is in communication connection with the heterogeneous data acquisition unit and is configured to process structured feature scores for quantifying financial risk control collection behavior scores from a structured feature dimension according to the content of structured fields in the multi-source heterogeneous data.
[0040] The semantic analysis scoring unit is in communication connection with the heterogeneous data acquisition unit and is configured to output structured semantic features by applying a semantic feature analysis method based on a large language model and a prompt word template according to unstructured text in the multi-source heterogeneous data, and process semantic feature scores for quantifying financial risk control collection behavior scores from a semantic feature dimension in combination with a preset semantic feature quantization rule.
[0041] The case retrieval scoring unit is in communication connection with the heterogeneous data acquisition unit and is configured to retrieve similar historical collection cases from a knowledge base by applying retrieval enhancement generation technology according to unstructured text in the multi-source heterogeneous data, and process case reference scores for quantifying financial risk control collection behavior scores from a similar case reference dimension based on the similar historical collection cases.
[0042] The risk control score calculation unit is respectively communicatively connected with the structured score unit, the semantic analysis score unit and the case search score unit, and is configured to calculate a financial risk control and collection behavior score P of the target debtor on the target credit product according to the structured feature score, the semantic feature score and the case reference score according to the following formula:
[0043] P = α × p structured + β × p semantic + γ × p case
[0044] In the formula, p structured represents the structured feature score, p semantic represents the semantic feature score, and p case represents the case reference score, and α, β and γ respectively represent preset weight coefficients and have α + β + γ = 1.
[0045] The score card construction unit is communicatively connected with the risk control score calculation unit, and is configured to construct a financial risk control and collection behavior score card for the target debtor and the target credit product according to the financial risk control and collection behavior score P and output the score card.
[0046] In a third aspect, the present application provides a computer device, comprising a storage module, a processing module and a transceiver module which are communicatively connected in sequence, wherein the storage module is configured to store a computer program, the transceiver module is configured to transceive messages, and the processing module is configured to read the computer program and execute the financial risk control and collection behavior score card construction method according to any possible design of the first aspect.
[0047] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores instructions, and when the instructions are executed on a computer, the financial risk control and collection behavior score card construction method according to any possible design of the first aspect is executed.
[0048] In a fifth aspect, the present application provides a computer program product, comprising a computer program or instructions, and when the computer program or the instructions are executed on a computer, the financial risk control and collection behavior score card construction method according to any possible design of the first aspect is implemented.
[0049] The above-mentioned scheme has the following beneficial effects:
[0050] (1) The application creatively provides a new scheme for constructing a financial risk control collection behavior scorecard based on a large language model and combining prompt word guidance and retrieval enhancement generation technology, that is, after obtaining target debtors and multi-source heterogeneous data generated in the collection interaction process for a target credit product, on the one hand, a structured feature score is obtained by processing the structured field content therein, on the other hand, structured semantic features are obtained based on a large language model and a prompt word template according to the unstructured text therein, and a semantic feature score is obtained by processing the semantic features in combination with a semantic feature quantization rule, and a case reference score is obtained by retrieving similar historical collection cases by using a retrieval enhancement generation technology, and finally a final score is calculated based on the weighting of the three scores, and the scorecard construction is completed based on the calculation result, so that the LLM can deeply analyze the collection text semantics, and the implicit risk features can be accurately extracted in combination with the prompt words, the limitation of structured data is broken through, the feature extraction accuracy is improved by 40% compared with the traditional keyword matching, the RAG technology can be used to retrieve historical cases to enhance the knowledge reuse and reasoning ability, and then the accuracy of the scorecard construction result is improved, which is beneficial to the post-loan collection strategy decision and management;
[0051] (2) By using the RAG technology to retrieve historical cases and risk control rules, an interpretable logic chain of "semantic feature-score mapping" and a complete logic chain of "text-semantic feature-similar case-score adjustment" can be constructed, meeting the regulatory interpretability requirements;
[0052] (3) By forming a "perception-decision-feedback" closed loop, the feature weight and reasoning rule can be dynamically optimized based on the collection effect, and the iteration cycle can be shortened to 24 hours;
[0053] (4) It can be adapted to multiple credit product scenarios, and the development cost of cross-product scorecards can be reduced by ≥60% through RAG knowledge transfer, facilitating practical application and promotion. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0055] Figure 1 The flowchart of the financial risk control collection behavior scorecard construction method provided by the embodiments of the application.
[0056] Figure 2 The structure diagram of the financial risk control collection behavior scorecard construction device provided by the embodiments of the application.
[0057] Figure 3A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.
[0059] It should be understood that although the terms "first" and "second", etc., may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the invention.
[0060] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, or A and B exist simultaneously. Another example is A, B and / or C, which can mean that any one of A, B, and C or any combination thereof exists. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone or A and B exist simultaneously. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.
[0061] Example
[0062] like Figure 1 As shown, the financial risk control and debt collection behavior scoring card construction method provided in the first aspect of this embodiment can be executed, but is not limited to, by computer devices with certain computing resources, such as servers, personal computers (PCs, referring to a type of multi-purpose computer suitable for personal use in terms of size, price, and performance; desktops, laptops, mini-laptops, tablets, and ultrabooks are all considered personal computers), smartphones, personal digital assistants (PDAs), or wearable devices. Figure 1 As shown, the method for constructing a financial risk control and debt collection behavior scoring card includes, but is not limited to, the following steps S1 to S6.
[0063] S1. Obtain multi-source heterogeneous data of a target debtor and generated in a collection interaction process for a target credit product.
[0064] In the step S1, the multi-source heterogeneous data specifically includes but is not limited to structured field content and unstructured text, wherein the structured field content includes but is not limited to overdue days, historical repayment rate and / or credit limit usage rate, and the unstructured text includes but is not limited to collection call text converted by applying automatic speech recognition technology and / or collection chat record based on social tool communication. The foregoing social tool can include but is not limited to short message, email and / or WeChat, etc. The multi-source heterogeneous data can be regularly collected in the collection interaction process. In order to improve the usability of the unstructured text and the multi-source heterogeneous data, preferably, the multi-source heterogeneous data of the target debtor and generated in the collection interaction process for the target credit product includes but is not limited to the following steps S11-S13.
[0065] S11. Obtain structured field content and unstructured text of a target debtor and generated in a collection interaction process for a target credit product, wherein the structured field content includes but is not limited to overdue days, historical repayment rate and / or credit limit usage rate, and the unstructured text includes but is not limited to collection call text converted by applying automatic speech recognition technology and / or collection chat record based on social tool communication, etc.
[0066] S12. Perform data cleaning processing on the unstructured text to obtain new unstructured text, wherein the data cleaning processing includes but is not limited to greeting filtering processing and / or semantic fuzzy expression correction processing, etc.
[0067] In the step S12, the greeting can be exemplified but is not limited to "Hello", etc., which is meaningless for scoring, and thus needs to be filtered out. The specific process of the semantic fuzzy expression correction processing is a prior art means, for example, marking "may repay" as "fuzzy commitment", etc.
[0068] S13. Aggregate the structured field content and the new unstructured text to obtain the multi-source heterogeneous data of the target debtor and generated in the collection interaction process for the target credit product.
[0069] S2. Process structured feature scores for quantifying financial risk control collection behavior scores from structured feature dimensions according to the structured field content in the multi-source heterogeneous data.
[0070] In the step S2, the specific processing manner of the structured feature score can be derived from the existing rule engine driven scoring manner and / or machine learning prediction type scoring manner, or the min-max normalization manner can be used to map the overdue days, the historical repayment rate and / or the credit limit utilization rate to [0, 10] points (for example, overdue days D = 5 → mapping formula: corresponding structured feature score = 10 - (D / 180) * 10 = 9.72 points).
[0071] S3. According to the unstructured text in the multi-source heterogeneous data, a semantic feature analysis method based on a large language model and a prompt word template is applied to output structured semantic features, and a preset semantic feature quantification rule is combined to process a semantic feature score for quantifying the financial risk control collection behavior score from the semantic feature dimension.
[0072] In the step S3, the large language model (Large Language Model, LLM) refers to a deep learning model trained using a large amount of text data, so that the model can generate natural language text or understand the meaning of language text, thereby realizing semantic feature analysis of the unstructured text through the combination of the large language model and the prompt word template applicable in the credit scenario. Specifically, the prompt word template includes but is not limited to basic prompt words and dynamic constraint prompt words added according to the type of the target credit product (for example, if the target credit product is a business loan type, the prompt "focus on identifying 'policy subsidies''supply chain funds' and other industry-related events" is added). The basic prompt words are used to prompt the output of the following risk features: key events (such as unemployment or investment failure, etc.), repayment intention (such as explicit commitment, vague commitment or refusal, etc.), and emotional tendency (such as positive, neutral or negative, etc.). In addition, the prompt word template can also define the output format, such as the prompt "output format: event: [], intention: [], emotion: [], confidence: []"; for example, for the text "the customer said that the store closed down and cannot repay temporarily", the large language model can output the following semantic structured features through the combination of the prompt word template: "event: [store closure (unemployment)]", intention: [refuse], emotion: [negative], confidence: [0.96]". At this time, combined with the preset semantic feature quantification rule, "refuse" intention can be mapped to -15 points, and confidence 0.96 can be weighted by a weight coefficient 1.0. The semantic feature score can be obtained by weighted accumulation.
[0073] S4. According to the unstructured text in the multi-source heterogeneous data, a semantic feature analysis method based on a large language model and a prompt word template is applied to output structured semantic features, and a preset semantic feature quantification rule is combined to process a semantic feature score for quantifying the financial risk control collection behavior score from the semantic feature dimension.
[0074] In the step S4, the retrieval-augmented generation (RAG) is a way to combine large language models with real-time data retrieval to generate more accurate, up-to-date and contextually relevant responses, which generates answers or content by referencing information from external knowledge bases, has strong explainability and customization ability, and is suitable for multiple natural language processing tasks such as question and answer systems, document generation and intelligent assistants, so it can be modified to apply to the present embodiment to retrieve similar historical collection cases from the knowledge base. Specifically, according to the unstructured text in the multi-source heterogeneous data, the retrieval-augmented generation technology is applied to retrieve similar historical collection cases from the knowledge base, and based on the similar historical collection cases, the case reference score for quantifying the financial risk control collection behavior score from the similar case reference dimension is processed, including but not limited to the following steps S41-S45.
[0075] S41. According to the unstructured text in the multi-source heterogeneous data, a Sentence-BERT model is used to generate a current text vector corresponding to the unstructured text.
[0076] In the step S41, the Sentence-BERT model is an existing sentence embedding model (i.e. an existing large language model) improved based on the BERT (Bidirectional Encoder Representations from Transformers, pre-trained language model based on Transformer architecture) model, mainly used to improve the efficiency and accuracy of text semantic similarity matching; it is optimized through Siamese network architecture, which significantly shortens the time for processing large-scale text data while maintaining high semantic representation quality. Therefore, by importing the unstructured text into the Sentence-BERT model, the current text vector can be output.
[0077] S42. According to the current text vector, at least one similar historical collection case is retrieved from the knowledge base based on the pre-established FAISS index, wherein the knowledge base includes but is not limited to a case sub-library and a rule sub-library, etc., the case sub-library is used to store the record information of all historical collection cases, the record information includes but is not limited to the text vector and the score adjustment record of the corresponding case, etc., and the rule sub-library is used to store the risk control rules and the characteristics of different credit products.
[0078] In the step S42, the historical collection case can be exemplified as "the customer mentioned 'unemployment' but eventually repaid"; the specific format of the record information can be exemplified as {text vector, score adjustment record, repayment result}, wherein the text vector can be generated in advance based on the unstructured text of the corresponding case by using the Sentence-BERT model. The FAISS (Facebook AI Similarity Search) is an open source similarity search tool developed by Facebook Artificial Intelligence Laboratory, mainly used to solve the problem of fast retrieval of large-scale high-dimensional vector data, so that at least one similar historical collection case (for example, five similar historical collection cases are retrieved when TopK=5) with the most similar text vector to the current text vector can be retrieved based on the current text vector and the text vector in the record information by supporting cosine similarity and the like. The risk control rules and the semantic feature quantification rules can be the same or different, for example, the risk control rules are: "'malicious default' tactics → score deduction 30 points", and the like. In addition, the characteristics are exemplified as: "consumption loan customers are highly sensitive to 'credit influence'".
[0079] S43. Adjust the risk control rules according to the score adjustment record in the record information of the at least one similar historical collection case, to obtain adjusted risk control rules.
[0080] In the step S43, for example, for the "store closure" event, if it is found according to the score adjustment record in the record information of the at least one similar historical collection case that the average score of similar events is -20 points, the specific content "'store closure' deducts 25 points" in the risk control rules can be adjusted to "'store closure' deducts 20 points". In addition, in order to make the rule adjustment interpretable, preferably, when the risk control rules are adjusted according to the score adjustment record in the record information of the at least one similar historical collection case, the method further comprises: extracting a natural language explanation for expressing the adjustment reason according to the score adjustment record in the record information of the at least one similar historical collection case; for example, based on the foregoing example, the natural language explanation "'store closure' deducts 20 points (reference case ID 202305)" can be extracted.
[0081] S44. Modify the adjusted risk control rules according to the characteristics corresponding to the target credit product, to obtain modified risk control rules.
[0082] In the step S44, for example, based on the above example, if the target credit product is a business loan type, the "store closure" is modified from "store closure deducts 20 points" to "store closure deducts 15 points" (because the business risk recovery probability is higher). In addition, in the rule modification, the natural language explanation for expressing the modification reason can also be extracted.
[0083] S45. According to the structured semantic features, combined with the modified risk control rules, the case reference score for quantifying the financial risk control and collection behavior score from similar case reference dimensions is processed.
[0084] S5. According to the structured feature score, the semantic feature score and the case reference score, the financial risk control and collection behavior score P of the target debtor on the target credit product is calculated according to the following formula:
[0085] P = a x p structured + b x p semantic + g x p case
[0086] In the formula, p structured represents the structured feature score, p semantic represents the semantic feature score, and p case represents the case reference score, a, b and g respectively represent the preset weight coefficients and a+b+g=1.
[0087] In the step S5, the weight coefficients a, b and g can be respectively preset as follows: a=0.4, b=0.4 and g=0.2, which can be dynamically adjusted by reinforcement learning.
[0088] S6. According to the financial risk control and collection behavior score P, the financial risk control and collection behavior score card is constructed for the target debtor and the target credit product and output.
[0089] In the step S6, the specific content of the financial risk control and collection behavior score card can include other contents in addition to the financial risk control and collection behavior score P, such as the name and / or customer ID of the target debtor, the name and / or product ID of the target credit product, the structured feature score, the semantic feature score and the case reference score, etc. In order to enhance the interpretability of the financial risk control and collection behavior score card, preferably, when constructing the financial risk control and collection behavior score card, the method further comprises: adding the natural language explanation for expressing the adjustment reason to the financial risk control and collection behavior score card. In addition, other extracted natural language explanations can also be added to the financial risk control and collection behavior score card.
[0090] After the step S6, in order to realize the dynamic optimization of the knowledge base, rules and important parameters, preferably, after the financial risk control collection behavior scorecard is built and output, the method further comprises: collecting actual feedback data for the financial risk control collection behavior scorecard; according to the actual feedback data, new historical collection cases are sorted out and added to the knowledge base, and if it is found that the absolute value of the deviation between the predicted repayment rate and the actual repayment rate of the target debtor on the target credit product exceeds a preset threshold, the semantic feature quantization rules and / or all weight coefficients for calculating the financial risk control collection behavior score P are updated and adjusted by using the PPO algorithm. The specific format of the actual feedback data can be exemplified as: [customer ID, score, actual repayment rate, manual correction record]; the preset threshold can be exemplified as 10%. The PPO (Proximal Policy Optimization) algorithm is a reinforcement learning algorithm proposed by OpenAI in 2017, which aims to balance training stability and efficiency by limiting the update amplitude of the policy, and is widely used in complex tasks such as continuous control, so it can be used in the embodiment to update and adjust the semantic feature quantization rules and / or all weight coefficients (i.e. the weight coefficients a, b and g) for calculating the financial risk control collection behavior score P; in the process of reinforcement learning dynamic adjustment based on the PPO algorithm, the specific reward function can be designed as: reward value Reward = 0.8 x repayment rate - 0.2 x complaint rate. In addition, the PPO algorithm can be used to update and adjust the risk control rules, and the optimized risk control rules and the new historical collection cases are added to the knowledge base, so as to ensure the timeliness of retrieval.
[0091] The embodiment also compares the financial risk control collection behavior scorecard construction method with the traditional machine learning scorecard, and obtains the technical index comparison results shown in Table 1 as follows:
[0092] Table 1. Comparison of technical indexes between the method of the embodiment and the traditional machine learning scorecard
[0093] Technical indicators Traditional machine learning scorecard The method of the present embodiment Unstructured feature utilization <30% >90% Scoring consistency (across cases) 70% 92% Rule iteration cycle ≥ 30 days ≤ 24 hours Cross-product development cost 100% (repeated development) 40% (RAG knowledge transfer) Interpretability Low (black box model) High (reasoning chain visualized)
[0094] In combination with Table 1 above, the financial risk control collection behavior scorecard construction method of the embodiment has the following specific effects:
[0095] (1) The risk prediction accuracy is improved: the pilot data of a certain city commercial bank shows that the recognition rate of the embodiment method for "high-risk customers" (final bad debts) reaches 89%, which is 23.6% higher than that of the traditional model (the recognition rate is 72%), and the bad debt rate is reduced by 1.2 percentage points;
[0096] (2) Collection resource optimization: Through accurate scoring, the proportion of manual collection resource investment for high-risk customers is reduced from 60% to 45%, and the collection rate is increased by 18%;
[0097] (3) Enhanced regulatory compliance: The output "semantic feature-case reference-score adjustment" reasoning chain passes the national financial supervision and management bureau risk control model compliance review, and the review time is shortened from 5 days to 1 day;
[0098] (4) Cross-product adaptation efficiency: When expanding from consumer loans to business loan scenarios, only 300 industry cases need to be supplemented to the RAG knowledge base, and the development cycle is shortened from 45 days to 7 days.
[0099] Thus, based on the financial risk control collection behavior scorecard construction method described in the preceding steps S1-S6, a new scheme is provided for constructing a financial risk control collection behavior scorecard based on a large language model and combining prompt word guidance and retrieval enhancement generation technology, that is, after obtaining multi-source heterogeneous data generated in the collection interaction process for a target credit product, on the one hand, the structured feature score for quantifying the financial risk control collection behavior from the structured feature dimension is processed according to the structured field content therein, on the other hand, the structured semantic feature is obtained based on the large language model and the prompt word template according to the unstructured text therein, and the semantic feature score is processed according to the semantic feature quantification rule, and the case reference score is obtained by retrieving similar historical collection cases using the retrieval enhancement generation technology, and finally the final score is calculated based on the weighting of the three scores, and the scorecard construction is completed based on the calculation result, so that the LLM can deeply analyze the collection text semantics, and the implicit risk features can be accurately extracted by combining the prompt words, breaking through the limitations of structured data, and improving the accuracy of feature extraction by 40% compared with traditional keyword matching. The RAG technology is used to retrieve historical cases to enhance the knowledge reuse and reasoning ability, thereby improving the accuracy of the scorecard construction result, which is beneficial to the post-loan collection strategy decision and management, and facilitates practical application and promotion.
[0100] As shown in Figure 2 the second aspect of the present embodiment provides a virtual device for implementing the financial risk control collection behavior scorecard construction method of the first aspect, comprising a heterogeneous data acquisition unit, a structured scoring unit, a semantic analysis scoring unit, a case retrieval scoring unit, a risk control scoring calculation unit and a scorecard construction unit;
[0101] The heterogeneous data acquisition unit is configured to acquire multi-source heterogeneous data of a target debtor and generated in a collection interaction process for a target credit product;
[0102] The structured scoring unit is in communication connection with the heterogeneous data acquisition unit, and is configured to process a structured feature score for quantifying a financial risk control collection behavior from a structured feature dimension according to the structured field content in the multi-source heterogeneous data.
[0103] The semantic analysis scoring unit is in communication connection with the heterogeneous data acquisition unit, configured to apply a semantic feature analysis mode based on a large language model and a prompt word template to output structured semantic features according to unstructured texts in the multi-source heterogeneous data, and process to obtain semantic feature scores for quantifying financial risk control collection behavior scores from a semantic feature dimension in combination with preset semantic feature quantization rules;
[0104] The case retrieval scoring unit is in communication connection with the heterogeneous data acquisition unit, configured to apply retrieval enhancement generation technology to retrieve similar historical collection cases from a knowledge base according to unstructured texts in the multi-source heterogeneous data, and process to obtain case reference scores for quantifying financial risk control collection behavior scores from a similar case reference dimension based on the similar historical collection cases;
[0105] The risk control score calculation unit is in communication connection with the structured scoring unit, the semantic analysis scoring unit and the case retrieval scoring unit respectively, configured to calculate the financial risk control collection behavior score P of the target debtor on the target credit product according to the structured feature score, the semantic feature score and the case reference score according to the following formula:
[0106] P = a x p structured + b x p semantic + g x p case
[0107] In the formula, p structured represents the structured feature score, p semantic represents the semantic feature score, and p case represents the case reference score, a, b and g respectively represent preset weight coefficients and have a + b + g = 1;
[0108] The score card construction unit is in communication connection with the risk control score calculation unit, configured to construct a financial risk control collection behavior score card for the target debtor and the target credit product according to the financial risk control collection behavior score P and output the same.
[0109] The working process, working details and technical effects of the foregoing device provided by the second aspect of the embodiment can be referred to the financial risk control collection behavior score card construction method described in the first aspect, which will not be repeated here.
[0110] As Figure 3As shown, the third aspect of the present embodiment provides a computer device for performing the financial risk control collection behavior scorecard construction method of the first aspect, comprising a storage module, a processing module and a transceiver module connected in sequence, wherein the storage module is configured to store a computer program, the transceiver module is configured to transceive messages, and the processing module is configured to read the computer program and perform the financial risk control collection behavior scorecard construction method of the first aspect. Specifically, the storage module can include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, a first-in first-out memory (FIFO) and / or a first-in last-out memory (FILO), etc.; and the processing module can use a microprocessor of the STM32F105 series, but is not limited thereto. In addition, the computer device can further include a power module, a display screen and other necessary components, but is not limited thereto.
[0111] The working process, working details and technical effects of the aforementioned computer device provided by the third aspect of the present embodiment can be referred to the financial risk control collection behavior scorecard construction method of the first aspect, which will not be described herein again.
[0112] The fourth aspect of the present embodiment provides a computer readable storage medium storing instructions of the financial risk control collection behavior scorecard construction method of the first aspect, i.e., the computer readable storage medium stores instructions, and when the instructions are run on a computer, the financial risk control collection behavior scorecard construction method of the first aspect is performed. The computer readable storage medium refers to a carrier for storing data, which can include, but is not limited to, floppy disks, optical disks, hard disks, flash memories, USB flash drives and / or memory sticks, etc. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.
[0113] The working process, working details and technical effects of the aforementioned computer readable storage medium provided by the fourth aspect of the present embodiment can be referred to the financial risk control collection behavior scorecard construction method of the first aspect, which will not be described herein again.
[0114] The fifth aspect of the present embodiment provides a computer program product comprising a computer program or instructions, which, when executed by a computer, implement the financial risk control collection behavior scorecard construction method of the first aspect. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.
[0115] It should be pointed out finally that the above description is only for preferred embodiments of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for constructing a financial risk control collection behavior scorecard, characterized in that, The method comprises the following steps: obtaining multi-source heterogeneous data of a target debtor and generated in a collection interaction process for a target credit product; processing structured feature points for quantifying a financial risk control collection behavior score from a structured feature dimension according to structured field content in the multi-source heterogeneous data; applying a semantic feature analysis method based on a large language model and a prompt word template to output structured semantic features, and processing semantic feature points for quantifying a financial risk control collection behavior score from a semantic feature dimension according to the structured semantic features and a preset semantic feature quantization rule; applying a retrieval enhancement generation technology to retrieve similar historical collection cases from a knowledge base according to unstructured text in the multi-source heterogeneous data, and processing case reference points for quantifying a financial risk control collection behavior score from a similar case reference dimension based on the similar historical collection cases; calculating a financial risk control collection behavior score P of the target debtor on the target credit product according to the structured feature points, the semantic feature points and the case reference points according to the following formula: P = a x p structured + b x p semantic + g x p case where p structured denotes the structured feature score, p semantic denotes the semantic feature score, p case denotes the case reference score, and a, b and g denote preset weight coefficients and have a+b+g=1. constructing a financial risk control collection behavior score card for the target debtor and the target credit product according to the financial risk control collection behavior score P and outputting the financial risk control collection behavior score card.
2. The method of claim 1, wherein the financial risk control collection behavior scorecard is constructed by, obtaining multi-source heterogeneous data of a target debtor and generated in a collection interaction process for a target credit product, comprising: obtaining structured field content and unstructured text of a target debtor and generated in a collection interaction process for a target credit product, wherein the structured field content contains overdue days, historical repayment rate and / or credit limit utilization rate, and the unstructured text contains collection call text converted by applying automatic speech recognition technology and / or collection chat records based on social tool communication; performing data cleaning processing on the unstructured text to obtain new unstructured text, wherein the data cleaning processing includes greeting filtering processing and / or semantic fuzzy expression correction processing; summarizing the structured field content and the new unstructured text to obtain multi-source heterogeneous data of the target debtor and generated in the collection interaction process for the target credit product.
3. The method of claim 1, wherein the financial risk control collection behavior scorecard is constructed by, The prompt word template includes a basic prompt word and a dynamic constraint prompt word added according to the type of the target credit product, wherein the basic prompt word is used to prompt the output of the following risk features: key events, repayment intention and emotional tendency.
4. The method of claim 1, wherein the financial risk control collection behavior scorecard is constructed by, applying a retrieval enhancement generation technology to retrieve similar historical collection cases from a knowledge base according to unstructured text in the multi-source heterogeneous data, and processing case reference points for quantifying a financial risk control collection behavior score from a similar case reference dimension based on the similar historical collection cases, comprising: generating a current text vector corresponding to the unstructured text in the multi-source heterogeneous data by using a Sentence-BERT model; According to the current text vector, at least one similar historical collection case is retrieved from a knowledge base based on a pre-established FAISS index, wherein the knowledge base includes a case sub-library and a rule sub-library, the case sub-library is used to store record information of all historical collection cases, the record information includes a text vector and a score adjustment record of a corresponding case, and the rule sub-library is used to store risk control rules and characteristics of different credit products; According to the score adjustment record in the record information of the at least one similar historical collection case, the risk control rules are adjusted to obtain adjusted risk control rules; According to the characteristics corresponding to the target credit product, the adjusted risk control rules are modified to obtain modified risk control rules; According to the structured semantic features, the modified risk control rules are combined to obtain a case reference score for quantifying the financial risk control collection behavior score from the similar case reference dimension.
5. The method of claim 4, wherein, When the risk control rules are adjusted according to the score adjustment record in the record information of the at least one similar historical collection case, the method further comprises: extracting a natural language explanation for expressing the adjustment reason according to the score adjustment record in the record information of the at least one similar historical collection case; When the financial risk control collection behavior scorecard is constructed, the method further comprises: adding the natural language explanation for expressing the adjustment reason to the financial risk control collection behavior scorecard.
6. The method for constructing a financial risk control and debt collection behavior scoring card according to claim 1, characterized in that, After the financial risk control collection behavior scorecard is constructed and output, the method further comprises: Collecting actual feedback data for the financial risk control collection behavior scorecard; According to the actual feedback data, new historical collection cases are sorted out and added to the knowledge base, and if it is found that the absolute value of the deviation between the predicted repayment rate and the actual repayment rate of the target debtor on the target credit product exceeds a preset threshold, the PPO algorithm is used to update the semantic feature quantization rules and / or all weight coefficients for calculating the financial risk control collection behavior score P.
7. A financial risk control collection behavior scorecard construction device characterized by, The system comprises a heterogeneous data acquisition unit, a structured scoring unit, a semantic analysis scoring unit, a case retrieval scoring unit, a risk control scoring calculation unit, and a scorecard construction unit. The heterogeneous data acquisition unit is configured to acquire multi-source heterogeneous data generated by a target debtor in a collection interaction process for a target credit product. The structured scoring unit is communicatively connected to the heterogeneous data acquisition unit and is configured to process structured feature scores for quantifying a financial risk control collection behavior score from a structured feature dimension according to structured field content in the multi-source heterogeneous data. The semantic analysis scoring unit is communicatively connected to the heterogeneous data acquisition unit and is configured to output structured semantic features by applying a semantic feature analysis method based on a large language model and a prompt word template according to unstructured text in the multi-source heterogeneous data, and process semantic feature scores for quantifying a financial risk control collection behavior score from a semantic feature dimension in combination with preset semantic feature quantization rules. The case search scoring unit is communicatively connected to the heterogeneous data acquisition unit, configured to search similar historical collection cases from a knowledge base according to unstructured texts in the multi-source heterogeneous data, and process the similar historical collection cases to obtain a case reference score for quantifying the financial risk control collection behavior score from similar case reference dimensions; The risk control score calculation unit is respectively communicatively connected to the structured scoring unit, the semantic analysis scoring unit and the case search scoring unit, configured to calculate the financial risk control collection behavior score P of the target debtor on the target credit product according to the structured feature score, the semantic feature score and the case reference score according to the following formula: P = a x p structured + β x p semantic + γ x p case where p structured denotes the structured feature score, p semantic denotes the semantic feature score, p case denotes the case reference score, and a, b and g denote preset weight coefficients and have a+b+g=1. The score card construction unit is communicatively connected to the risk control score calculation unit, configured to construct a financial risk control collection behavior score card for the target debtor and the target credit product according to the financial risk control collection behavior score P and output the score card.
8. A computer device, comprising: The computer readable storage medium has instructions stored thereon, and when the instructions are executed on the computer, the financial risk control collection behavior score card construction method according to any one of claims 1-6 is executed.
9. A computer-readable storage medium, characterized in that The computer program or the instructions realize the financial risk control collection behavior score card construction method according to any one of claims 1-6 when executed on the computer.
10. A computer program product comprising computer programs or instructions, characterized in that, The computer program or the instructions realize the financial risk control collection behavior score card construction method according to any one of claims 1-6 when executed on the computer.