Debtor portrait construction method and system based on repayment willingness mining

By collecting and processing multiple types of data to form a fusion dataset, feature mining and causal screening are performed. Feature weights are optimized by combining collection business feedback, and a lightweight debtor profiling model is constructed. This solves the problems of insufficient repayment willingness mining and poor dynamic adaptability in existing technologies, and achieves accurate debtor profiling and collection strategy matching, thereby improving the efficiency of credit collection business.

CN121998752APending Publication Date: 2026-05-08SHUSHE (SHENZHEN) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHUSHE (SHENZHEN) TECHNOLOGY CO LTD
Filing Date
2026-01-16
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing debtor profiling technologies fail to fully uncover the implicit characteristics of repayment willingness, resulting in profiling that cannot accurately reflect the debtor's subjective repayment tendency. Furthermore, the models fail to dynamically iterate and optimize, making it difficult to adapt to changes in the debtor's repayment willingness over time and with external events.

Method used

After collecting and preprocessing multiple types of data to form a fused dataset, feature mining and cross-domain alignment are performed to remove features without causal relationships, a causal feature library is constructed, and weight calibration and counterfactual reasoning are performed using collection business feedback data to form a core feature library. Finally, a lightweight debtor profile model is constructed.

Benefits of technology

It enables real-time generation of debtor profiles and precise matching of collection strategies, improving the efficiency of credit collection and repayment performance, and solving the problems of insufficient accuracy and poor dynamic adaptability in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998752A_ABST
    Figure CN121998752A_ABST
Patent Text Reader

Abstract

The invention provides a debtor portrait construction method and system based on repayment willingness mining. The method comprises the following steps: collecting multiple types of data of a debtor to form a fusion data set; performing feature mining based on the fused data set to obtain a multi-dimensional repayment willingness feature set; performing cross-domain alignment and causal screening on the multi-dimensional repayment willingness feature set to form a causal feature library; performing weight calibration on the features in the causal feature library based on intervention effect analysis; adjusting parameters of feature mining in a targeted manner according to a weight calibration result, and verifying the effectiveness of a causal feature library through anti-factual reasoning to obtain a core feature library; and constructing a debtor portrait model based on the core feature library. According to the method, the related repayment attributes of the debtor can be described from multiple dimensions of static willingness, dynamic trend and associated influence, real-time generation of the portrait and accurate matching of the collection strategy are realized, and the overall efficiency and the repayment fulfillment rate of the credit collection service are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of debtor profiling, specifically to a method and system for constructing debtor profiling based on repayment willingness mining. Background Technology

[0002] In the risk management and collection processes of financial lending, debtor profiling is a core technological support for achieving precise collection and improving repayment efficiency. Existing debtor profiling technologies mostly focus on the debtor's basic business data (such as credit limit, overdue period, historical performance records, etc.), generating profile labels through conventional feature extraction and model training to assist in collection business decisions.

[0003] However, existing debtor profiling technologies have significant shortcomings. For example, the analysis of debtors' repayment intentions is superficial, failing to fully capture the implicit characteristics related to repayment intentions, resulting in profiling that cannot truly reflect the debtor's subjective repayment tendency. Furthermore, profiling models are mostly statically constructed, without dynamic iterative optimization based on collection business feedback data, making it difficult to adapt to the dynamic characteristics of debtors' repayment intentions changing over time and with external events.

[0004] Therefore, there is an urgent need for a debtor profiling method that can deeply explore repayment intentions in order to solve the problems in existing technologies and improve the efficiency and effectiveness of credit collection. Summary of the Invention

[0005] To address the problems existing in current technologies, this application provides a method and system for constructing debtor profiles based on repayment willingness mining. The specific solution is as follows: A method for constructing a debtor profile based on repayment willingness mining, characterized by comprising: Collect various types of data from debtors, and map the heterogeneous data formed by preprocessing these various types of data to a unified feature space to form a fused dataset; Feature mining is performed on the fused dataset to obtain a multidimensional repayment intention feature set; cross-domain alignment and causal filtering are performed on the multidimensional repayment intention feature set to remove features that are not directly causally related to repayment behavior, thus forming a causal feature library; Using collection business feedback data as incremental input, the features in the causal feature library are weighted and calibrated based on intervention effect analysis; the parameters of the feature mining are adjusted in a targeted manner according to the weight calibration results, and the effectiveness of the causal feature library is verified by counterfactual reasoning to obtain the core feature library; Based on the core feature library, a debtor profile model is constructed. After lightweight processing, the model is encapsulated into a standardized interface and embedded into the existing debt collection system to achieve real-time generation of debtor profiles and accurate matching of collection strategies.

[0006] In some specific embodiments, the feature mining includes semantic-driven latent state decoding of repayment intention, time-series-driven modeling of repayment intention trends, and related person-driven mining of the transmission effect of repayment intention.

[0007] In some specific embodiments, the multiple types of data include interactive semantic data that characterizes semantic signals of repayment willingness, time-series correlation data that reflects dynamic changes in willingness, correlation data that reflects the influence of correlation, and business tag data that supports basic judgment of the profile. The preprocessing includes noise removal and keyword extraction of interactive semantic data, timestamp sorting and three-dimensional sequence integration of time-series related data, label standardization of related person data, and compliance desensitization and redundant data filtering of all data.

[0008] In some specific embodiments, semantic-driven latent state decoding of repayment intention includes: calculating the initial latent state of repayment intention and state transition probability based on business tag data, extracting semantic tags and polarity scores from interactive semantic data, and adjusting the initial state transition probability using the semantic tags and polarity scores as correction factors to obtain static repayment intention features. The time-driven repayment intention trend modeling includes: aligning the static repayment intention latent state with the collection strategy execution time sequence and external event time sequence to form a three-dimensional data stream; statistically analyzing the intention state transition probability in segments according to preset key time intervals; generating a repayment intention trend curve and marking the high intention window period to obtain dynamic repayment intention characteristics. The mining of the transmission effect of repayment intention driven by related parties includes: calculating the transmission effect value based on the relationship type, performance status and response behavior in the related party data, adjusting the implicit state of the debtor's static repayment intention according to the transmission effect value, and supplementing the implicit association influence characteristics; By integrating static repayment willingness characteristics, dynamic repayment willingness characteristics, and implicit correlation influence characteristics, a multi-dimensional repayment willingness feature set is formed, which includes static willingness, dynamic trends, and correlation influence.

[0009] In some specific embodiments, the feature range for cross-domain alignment is determined, which includes not only the multidimensional repayment willingness feature set, but also business-related features related to the construction of debtor profiles; The multidimensional repayment willingness feature set and the business-related features are mapped to the same latent space, and a causal relationship graph between the features in the latent space and the debtor's repayment behavior is constructed to identify the direct causal relationship between the features and the repayment behavior in the causal relationship graph. The pseudo-correlation features that have no direct causal relationship with the repayment behavior in the causal relationship graph are removed, and the core features with direct causal relationship are retained and integrated to form a causal feature library.

[0010] In some specific embodiments, the type of the collection business feedback data is determined, and the collection business feedback data includes the repayment achievement rate after the collection strategy is implemented, the performance timeliness, the second overdue situation, and the content of communication objections; Define intervention variables, which are variables related to the execution of collection strategies, including at least one of collection timing selection, collection strategy type, and intervention method of related parties; A sample matching algorithm was used to match the intervention-treated debtor samples with similar characteristics to the control group samples that had not been intervened, thereby eliminating sample selection bias and obtaining sample data. Based on the sample data, the intervention effect value of each feature in the causal feature library on repayment behavior is calculated; weights are assigned to the corresponding features according to the magnitude of the intervention effect value, with the higher the intervention effect value, the greater the weight of the feature, thus completing the weight calibration of the features in the causal feature library.

[0011] In some specific embodiments, the process of acquiring the core feature library specifically includes: Extract target features from the weight calibration results whose weight fluctuation exceeds a preset threshold, and adjust the parameters of the feature mining step corresponding to the target features; Construct a causal relationship network between features in the causal feature library and debtor repayment behavior, and clarify the correlation between features and repayment behavior; perform counterfactual hypothesis deduction on the core features in the causal relationship network, and calculate the deviation between predicted repayment behavior and actual repayment behavior under the assumed conditions; The deviation values ​​are compared with a preset validity threshold. Valid features with deviation values ​​exceeding the threshold are retained, while invalid features with deviation values ​​below the threshold are removed. The valid features are then integrated to form a core feature library.

[0012] In some specific embodiments, parameter adjustments are made for the feature mining stage corresponding to the target feature, specifically including: If the target feature originates from time-driven repayment intention trend modeling, then the key time intervals are redefined and the calculation weights of the repayment intention state transition probability within the intervals are adjusted. If the target feature originates from the mining of the repayment willingness transmission effect driven by related parties, then the calculation coefficient of the related party transmission effect should be recalibrated. If the target feature originates from semantically driven latent state decoding of repayment intention, then optimize the matching rules of semantic labels and adjust the correction magnitude of polarity scores to the initial state transition probability.

[0013] In some specific embodiments, a model architecture adapted to the debt collection business scenario is selected, and a profile model is constructed based on the model architecture; the model architecture includes gradient boosting tree model, random forest model, single-layer perceptron, and shallow neural network.

[0014] A debtor profiling system based on repayment willingness mining includes: The data acquisition unit is used to collect various types of data from debtors, and to map the heterogeneous data formed by preprocessing these various types of data to a unified feature space to form a fused dataset. The causal feature unit is used to perform feature mining based on the fused dataset to obtain a multidimensional repayment intention feature set; cross-domain alignment and causal filtering are performed on the multidimensional repayment intention feature set to remove features that have no direct causal relationship with repayment behavior, forming a causal feature library; The weight calibration unit is used to take collection business feedback data as incremental input, perform weight calibration on the features in the causal feature library based on intervention effect analysis, adjust the parameters of the feature mining according to the weight calibration results, and verify the effectiveness of the causal feature library through counterfactual reasoning to obtain the core feature library. The profile building unit is used to build a debtor profile model based on the core feature library. After the model is lightweighted, it is encapsulated into a standardized interface and embedded into the existing collection system to realize the real-time generation of debtor profiles and accurate matching of collection strategies.

[0015] Beneficial Effects: This application proposes a method and system for constructing debtor profiles based on repayment willingness mining. By collecting multiple types of data and comprehensively capturing repayment willingness characteristics, it provides a data foundation for accurate profiling, enabling the profile to depict the debtor's repayment-related attributes from multiple dimensions such as static willingness, dynamic trends, and related influences. This achieves real-time profile generation and precise matching with collection strategies, effectively improving the overall efficiency of credit collection business and repayment performance rate. It solves the problems of insufficient accuracy, poor dynamic adaptability, and poor business implementation in existing debtor profile construction technologies, significantly improving the scientific nature and practicality of profile construction and optimizing the effectiveness of credit collection business.

[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the debtor profile construction method in this application; Figure 2 This is a schematic diagram illustrating the principle of the debtor profile construction method in this application; Figure 3 This is a schematic diagram of the various data processing flows in this application; Figure 4 This is a schematic diagram of the core feature library processing flow of this application; Figure 5 These are real-world demonstration images of model training and optimization. Figure 6 These are actual deployment demonstration images of the real machine; Figure 7 This is a schematic diagram of the debtor profile construction system module in this application.

[0019] Figure labels: 1-Data acquisition unit; 2-Causal feature unit; 3-Weight calibration unit; 4-Profile construction unit. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0021] This application proposes a method for constructing debtor profiles based on repayment willingness mining, effectively addressing issues such as insufficient accuracy, poor dynamic adaptability, and inefficient business implementation in existing debtor profile construction methods, significantly improving the efficiency and effectiveness of credit collection. A flowchart illustrating the debtor profile construction process based on repayment willingness mining is attached. Figure 1 As shown in the attached diagram, the principle is as follows: Figure 2 As shown, the specific solution is as follows: A method for constructing debtor profiles based on repayment willingness mining includes: 101. Collect various types of data from debtors, and map the heterogeneous data formed by preprocessing these various types of data to a unified feature space to form a fused dataset; 102. Based on the fused dataset, feature mining is performed to obtain a multidimensional repayment intention feature set; cross-domain alignment and causal filtering are performed on the multidimensional repayment intention feature set to remove features that have no direct causal relationship with repayment behavior, forming a causal feature library; 103. Using collection business feedback data as incremental input, the features in the causal feature library are weighted and calibrated based on intervention effect analysis; the parameters of feature mining are adjusted in a targeted manner according to the weight calibration results, and the effectiveness of the causal feature library is verified through counterfactual reasoning to obtain the core feature library; 104. Construct a debtor profile model based on the core feature library, and after lightweight processing, encapsulate the model into a standardized interface, embed it into the existing collection system, and realize the real-time generation of debtor profiles and accurate matching of collection strategies.

[0022] The core purpose of step 101 is to provide a unified, high-quality data foundation for subsequent repayment willingness mining and profile construction. Its implementation logic revolves around data acquisition, data governance, and data integration.

[0023] Among these, "multi-category data" refers to a collection of various information related to debtor repayment, not limited to a single data type. Collecting this data aims to comprehensively cover potential factors influencing repayment behavior, avoiding incomplete feature mining due to missing data dimensions. Preprocessing involves basic governance of the collected raw data, with the core function of eliminating data noise, standardizing data formats, and ensuring data processability. This includes removing invalid data, correcting data errors, and standardizing data time formats or numerical units. The principle is to filter redundant and abnormal information using basic data cleaning rules, clearing obstacles for subsequent data fusion. Heterogeneous data refers to data from different sources, in different formats, or with different dimensions. Due to the inherent differences in this type of data, it cannot be directly used for joint analysis. Therefore, it needs to be mapped to a unified feature space. The principle of this process is to transform different types of data into feature vectors of the same dimension and distribution through common feature standardization or feature transformation methods, achieving "homogenization" at the data level. The resulting fused dataset provides structurally unified and informationally complete data support for subsequent feature mining, ensuring the accuracy and comprehensiveness of feature mining.

[0024] Step 102 is a key step in achieving accurate discovery of repayment willingness and effective feature screening. The core logic is feature extraction, feature optimization, and feature purification.

[0025] Feature mining based on fused datasets aims to extract core information that characterizes repayment willingness from unified data, forming a multi-dimensional repayment willingness feature set. Essentially, feature mining uses data association analysis and feature extraction algorithms to identify and integrate feature dimensions directly related to repayment willingness from the fused data. Its role is to transform raw data into feature information with practical business significance, providing core materials for subsequent profile building.

[0026] The core function of cross-domain alignment is to eliminate domain differences between different feature sources. Since the features in the multidimensional repayment willingness feature set may come from different data scenarios or data types and have distribution differences, directly using them for analysis will affect the accuracy of the results. Therefore, cross-domain alignment methods are used to adjust these features to the same distribution space to ensure comparability between features. The principle is to eliminate feature distribution offset through domain adaptation technology.

[0027] Causal screening is a core step in improving feature quality. Its purpose is to eliminate pseudo-correlated features that have no direct causal relationship with repayment behavior, thus preventing such features from interfering with the judgment of subsequent models. The principle is based on the logic of causal relationship identification, analyzing the intrinsic relationship between features and repayment behavior, distinguishing between "correlation" and "causation", and finally retaining features with direct causal relationship to form a causal feature library. The core value of this feature library is to provide a high-quality and highly relevant feature foundation for subsequent weight calibration and model construction, thereby improving the accuracy of profile construction from the source.

[0028] The core objective of step 103 is to achieve dynamic optimization and validity verification of the feature library, ensuring that the final core feature library accurately adapts to the needs of repayment behavior prediction. The logic involves feedback iteration, weight optimization, parameter adjustment, and validity verification. Using collection business feedback data as incremental input essentially introduces real-world results from actual business scenarios. This data reflects the effectiveness of previous features and models in practical applications, providing real business feedback for feature optimization. The principle is to use incremental feedback data to correct the perceived correlation between features and repayment behavior, preventing the model from becoming detached from real-world business scenarios.

[0029] The core function of weight calibration based on intervention effect analysis is to adjust the feature weights according to the actual impact of features on repayment behavior, so that more decisive features can be given higher weights. The principle is to set intervention variables related to debt collection business, compare the differences in repayment behavior under different intervention conditions, quantify the impact of each feature on the repayment result, and then allocate weights according to the quantification results to achieve accurate ranking of feature importance.

[0030] Adjusting feature mining parameters based on weight calibration results is a process of reverse optimization of the previous feature mining process based on weight feedback. For example, if the weight of a certain type of feature is too low, it means that there may be a deviation in the previous mining parameters. By adjusting the parameters, the extraction efficiency of such effective features can be improved. The principle is to use the deviation information of weight feedback to correct the core parameters of feature mining and ensure the targeting of subsequent feature mining.

[0031] The core function of counterfactual reasoning verification is to test the effectiveness of the causal feature library. The principle is to construct a scenario of "hypothetical feature changes", simulate the predicted results of repayment behavior after feature adjustment, and then compare it with the actual repayment results to determine whether the correlation between features and repayment behavior is real and reliable. The core feature library formed after eliminating invalid features can maximize the accuracy and reliability of subsequent model construction.

[0032] Step 104 is the core step in realizing the practical application of debt profiling. Building a debtor profiling model based on a core feature library essentially establishes a mapping relationship between core features and debtor repayment-related attributes. Its function is to transform abstract feature data into concrete profiling results. The principle is to learn the correlation between core features and repayment behavior and willingness through model training algorithms, enabling the model to output profiling information that characterizes the debtor's repayment attributes based on the input core features. The core purpose of lightweight processing is to adapt to the deployment environment of existing collection systems and reduce the model's consumption of system resources. This is achieved by removing redundant structures in the model, simplifying model parameters, and optimizing the calculation process, thereby improving the model's operating efficiency and resource adaptability without significantly affecting model performance. Examples include trimming network structures with minimal impact and converting high-precision parameters to low-precision parameters. Standardized interface encapsulation is crucial for achieving compatibility and integration between the model and existing collection systems. The principle is to unify the model's input and output formats according to industry-standard interface specifications, ensuring smooth data interaction between the model and existing systems and avoiding deployment obstacles caused by interface incompatibility. The principle of embedding the encapsulated model into the existing debt collection system to achieve real-time profile generation is that after the model receives real-time data of the debtor from the system, it quickly processes and outputs the profile results through a lightweight calculation process. The principle of accurate matching of collection strategies is that the profile results can accurately represent the debtor's willingness and ability to repay, providing a clear basis for the system's strategy recommendation module. This enables the system to match appropriate collection strategies for debtors with different profiles, ultimately achieving the business goal of improving collection efficiency and optimizing collection results.

[0033] In some specific embodiments, feature mining includes semantic-driven decoding of the latent state of repayment intention, time-series-driven modeling of repayment intention trends, and identification of the transmission effect of repayment intention driven by related parties. By clarifying the specific implementation path of feature mining, it targets and mines repayment intention-related features from three core dimensions: semantics, time series, and related parties. This ensures that the multi-dimensional repayment intention feature set can comprehensively and accurately capture the core characteristics of debtors' repayment intentions, providing a high-quality foundation for subsequent feature selection and profile construction.

[0034] The semantic-driven decoding of latent repayment intention is based on analyzing semantic data from interactions between debtors and debt collectors. This semantic data includes communication texts and speech-to-text formats that convey subjective attitudes. The principle is to extract semantic tags related to repayment intention using natural language processing technology, such as willingness to repay but requiring an extension, or refusal to repay. These tags are then scored to quantify subjective inclination. This is combined with the initial latent repayment intention state and state transition probabilities calculated from basic data. The polarity score is used as a correction factor to adjust the initial state, ultimately yielding a static repayment intention feature that truly reflects the debtor's subjective repayment tendency. This static repayment intention feature overcomes the limitations of traditional methods that rely solely on objective data, accurately capturing the debtor's unexpressed potential repayment attitude.

[0035] The core of time-driven repayment intention trend modeling is to uncover the dynamic changes in repayment intention over time. Its principle involves aligning the previously obtained static repayment intention features with the time series of collection strategy execution and external events, forming a three-dimensional data stream encompassing time, static intention, and external events. Then, based on business experience, pre-defined key time intervals are defined, such as 24 hours after collection or 3 days after payday. The probability of transition in repayment intention status within each interval is calculated, generating a repayment intention trend curve and marking high-intention windows, ultimately yielding dynamic repayment intention features. The purpose of these features is to adapt to the dynamic characteristics of repayment intention fluctuating with time and external events, avoiding the limitation of static features failing to reflect changes in intention.

[0036] The core of the linked-person-driven repayment willingness transmission effect mining is to consider the influence of entities related to the debtor on their repayment willingness. The principle is to extract key information related to repayment from linked-person data, including the relationship type between the linked person and the debtor, the linked person's own performance status, and their response to collection efforts. This is then quantitatively calculated to obtain the transmission effect value of linked persons on the debtor's repayment willingness. This value is then used to correct the static repayment willingness characteristics and supplement implicit linked influence characteristics. The role of this feature is to overcome the limitation of focusing only on the debtor's own data and capture the indirect influence of external linked factors on repayment willingness. After feature mining is completed in these three dimensions, the integrated multidimensional repayment willingness feature set can comprehensively represent the debtor's repayment willingness from three levels: subjective static, dynamic change, and external linkage, providing comprehensive feature support for subsequent causal screening and accurate profile construction.

[0037] In some specific embodiments, the multiple data types include interactive semantic data representing semantic signals of repayment willingness, time-series correlation data reflecting dynamic changes in willingness, correlation data of related individuals demonstrating the influence of correlation, and business tag data supporting basic judgments in the profiling process. Preprocessing includes noise removal and keyword extraction for interactive semantic data, timestamp sorting and three-dimensional sequence integration for time-series correlation data, tag standardization for related individual data, and compliance-compliant desensitization and redundant data filtering for all data. By clearly defining the specific scope and preprocessing methods of the multiple data types, the comprehensiveness and accuracy of subsequent feature mining are ensured from the data source, laying a core foundation for building a high-quality fused dataset. The processing flow for the multiple data types is attached. Figure 3 As shown.

[0038] The classification of data into multiple categories revolves around the core need of mining repayment willingness. Interactive semantic data, representing semantic signals of repayment willingness, encompasses information containing subjective attitudes, such as communication texts and voice-to-text records between debtors and debt collectors. Its core function is to provide semantic evidence for mining the debtor's subjective repayment tendency. Time-series correlation data, reflecting dynamic changes in willingness, includes records of the debtor's repayment-related behaviors at different time points, time sequences of collection strategy execution, and time information of external events (such as paydays and holidays), used to capture the dynamic patterns of repayment willingness changing with time and the external environment. Related party data, reflecting the influence of connections, refers to information such as the performance status and collection response behavior of entities related to the debtor, such as relatives and colleagues, which can explore the transmission effect of external factors on the debtor's repayment willingness. Business tag data, supporting basic profiling judgments, covers objective data such as the debtor's basic credit information, overdue duration, and judicial default labels, providing a basic reference for mining repayment willingness.

[0039] The preprocessing stage adopts a "classification and standardization" approach. Noise removal for interactive semantic data aims to eliminate redundant information such as irrelevant interjections and interfering characters in communication records. Keyword extraction uses natural language processing technology to capture core semantic information such as "willing to repay" and "refusal to fulfill obligations," ensuring the validity of the semantic data. Timestamp sorting of time-series related data helps to organize the temporal logic of the data. Three-dimensional sequence integration structurally integrates time information, willingness-related data, and external event data to form a data format that meets the needs of subsequent time-series modeling. Standardization of labels for related person data addresses the heterogeneity of related person data from different sources by unifying the expression format of labels such as relationship type and performance status level. Compliance anonymization is performed on all data to comply with relevant regulations on data security and privacy protection. Sensitive information such as ID card numbers and mobile phone numbers is encrypted or anonymized to avoid compliance risks. Redundant data filtering removes duplicate, invalid, and missing data, reducing the computational cost of subsequent data processing. Through the above targeted preprocessing, various heterogeneous data are transformed into a data form with a unified format and effective information. After being mapped to a unified feature space, a fusion dataset covering all dimensions of repayment willingness mining can be formed, ensuring that subsequent feature mining can accurately capture the core information related to repayment willingness.

[0040] Semantic-driven latent state decoding of repayment intention includes: calculating the initial latent state of repayment intention and state transition probability based on business tag data; extracting semantic tags and polarity scores from interactive semantic data; adjusting the initial state transition probability using semantic tags and polarity scores as correction factors to obtain static repayment intention features.

[0041] Semantic-driven latent state decoding of repayment intention is based on objective business data and supplemented by subjective semantic information to accurately characterize the static characteristics of repayment intention. Its implementation logic involves first using business tag data (such as overdue duration, credit limit, and repayment history) to calculate the initial latent state of repayment intention and the probability of state transition through probability statistics or a basic model. The initial latent state is the debtor's repayment intention level (e.g., high, medium, low) initially determined based on objective data. The state transition probability is the possibility of transformation between different intention levels. This step provides an objective framework for judging repayment intention. For example, debtor A's business tag data is "45 days overdue, 3 good repayment records, no judicial default label." By statistically analyzing the distribution of intentions among similar debtors, their initial latent state is determined to be "medium intention." Simultaneously, the initial state transition probabilities are calculated as follows: medium → high: 0.2, medium → low: 0.3, medium → medium: 0.5. Subsequently, semantic tags are extracted from interactive semantic data (such as collection communication texts and voice-to-text records) using natural language processing technology. Semantic tags are key information that directly reflects repayment attitude, such as agreeing to full repayment, refusing to communicate, and applying for an extension. Simultaneously, the semantic tags are quantified with polarity scoring to reflect subjective tendencies; positive scores correspond to a positive repayment attitude, and negative scores correspond to a negative attitude. The absolute value of the score reflects the degree of determination of the attitude. For example, debtor A's interactive semantic data is "I will repay the full amount next month when I get my salary. I am indeed short of funds during this period, but don't worry, I definitely won't default." The extracted semantic tags are "promise to repay + financial difficulties," which is comprehensively judged as a positive attitude, and assigned a polarity score of +0.7. Finally, the semantic tags and polarity scores are used as correction factors to adjust the initial state transition probability. The reason for this adjustment is that the initial latent state relies solely on objective data and cannot reflect the debtor's true subjective attitude. By correcting subjective semantic information, the accuracy of the repayment intention judgment can be significantly improved. The final static repayment intention feature can accurately represent the debtor's stable subjective repayment tendency at a certain point in time. Semantic labels and polarity scores are used as correction factors and substituted into a preset correction formula to adjust the initial state transition probability. The correction formula must reflect the logic that "the higher the polarity score, the greater the increase in the positive transition probability." After adjustment, the willingness level corresponding to the transition direction with the highest probability is selected as the final static repayment willingness feature. For example, substituting debtor A's polarity score +0.7 into the correction formula, the initial transition probability "medium → high" increases from 0.2 to 0.8, "medium → low" decreases from 0.3 to 0.1, and "medium → medium" decreases from 0.5 to 0.1. At this time, the transition probability of "medium → high" is the highest, so its static repayment willingness feature is determined to be "medium to high willingness".

[0042] Time-series-driven repayment intention trend modeling involves aligning the static latent state of repayment intention with the execution timeline of collection strategies and the timeline of external events to form a three-dimensional data stream. It then statistically analyzes the probability of intention state transitions across preset key time intervals, generating a repayment intention trend curve and marking high-intention windows to obtain dynamic repayment intention characteristics. This time-series-driven repayment intention trend modeling focuses on the dynamic changes in repayment intention. Its core principle is to combine the time dimension with external influencing factors to uncover patterns in intention changes. Specifically, it first aligns the aforementioned static latent state of repayment intention with the execution timeline of collection strategies (such as the time, method, and content of each collection) and the timeline of external events (such as the debtor's payday, holidays, and the date of delivery of judicial notices, etc., events that may affect repayment intention). This ensures the correspondence between intention state and influencing factors within the same time dimension, forming a three-dimensional data stream of static repayment intention, collection strategies, and external events. The value of this data stream lies in establishing a temporal correlation between intention state and influencing factors, providing a structured data foundation for subsequent pattern discovery. For example, the key nodes of debtor A's three-dimensional data flow are: March 1st (telephone collection) - medium to high willingness; March 8th (payday) - no collection action; March 10th (SMS collection) - medium to high willingness. Then, based on practical experience in credit collection, pre-defined key time intervals are established (such as 24 hours after the execution of the collection strategy, 1-3 days after the debtor's payday, 24 hours after the delivery of judicial notice, etc.). These intervals are key periods where repayment willingness is prone to change, as summarized in the business. Subsequently, based on practical experience in credit collection, pre-defined key time intervals are established. These intervals are core periods where repayment willingness is prone to fluctuation, common intervals include "1-3 days after the execution of the collection strategy," "3-5 days before payday," "1-3 days after payday," and "24 hours after the delivery of judicial notice," etc. Then, the transition probability of the implicit state of static repayment willingness is statistically analyzed segment by interval, and the changing trend of willingness level within different intervals is analyzed. For example, consider debtor A, dividing the timeframe into three key periods: March 1st-3rd (1-3 days after collection), March 5th-7th (1-3 days before payday), and March 8th-10th (1-3 days after payday). The transition probabilities for each period are statistically analyzed: 0.1 for the transition from medium-high to high in the 1-3 days after collection, 0.4 for the transition from medium-high to high in the 1-3 days before payday, and 0.9 for the transition from medium-high to high in the 1-3 days after payday. A repayment willingness trend curve is plotted with time on the horizontal axis and the corresponding quantitative value of the willingness level on the vertical axis. Based on the statistical results of the transition probabilities, the periods of significant increase in willingness level are marked, i.e., the high willingness window. Finally, the "trend curve change pattern + high willingness window" are integrated to form a dynamic repayment willingness characteristic.For example, Debtor A's trend curve shows a pattern of "slow rise after collection, accelerated rise before payday, and sharp rise after payday". The high willingness window is from March 8th to March 10th (1-3 days after payday). The corresponding dynamic repayment willingness characteristics are "significant increase in willingness after payday, with the high willingness window concentrated within 3 days after payday".

[0043] The mining of the transmission effect of repayment willingness driven by related parties includes: calculating the transmission effect value based on the relationship type, performance status, and response behavior in related party data; adjusting the implicit state of the debtor's static repayment willingness based on the transmission effect value; and supplementing the implicit related influence characteristics. The mining of the transmission effect of repayment willingness driven by related parties focuses on the indirect impact of external related factors on repayment willingness. Its implementation logic is to quantify the transmission effect of related parties on the debtor's repayment willingness. First, three core types of information are extracted from the related party data: the relationship type between the related party and the debtor (such as relatives, colleagues, guarantors, etc., the influence of different relationship types varies), the performance status of the related party (such as the related party's own credit performance; related parties with good performance may have a positive influence on the debtor), and the collection response behavior of the related party (such as the degree of cooperation of the related party with collection efforts, whether they assist in urging the debtor to repay). A weight range (such as 0-1) is set for each dimension, and the weight value is determined based on business experience or statistical models. For example, debtor A's related party is their spouse (relationship type weight 0.8), the spouse has a good credit repayment record (repayment status weight 0.9), and proactively contacted collection personnel to indicate they would urge the debtor to repay (response behavior weight 0.9). Based on these three types of information, a quantitative model is used to calculate the transmission effect value. The magnitude of the transmission effect value represents the strength of the related party's influence; positive or negative indicates the direction of the influence (positive for promoting repayment, negative for hindering repayment). A weighted product formula is used to calculate the transmission effect value: Transmission effect value = Relationship type weight × Repayment status weight × Response behavior weight. The transmission effect value ranges from 0 to 1; a higher value indicates a stronger positive transmission influence from the related party to the debtor. If the related party has a poor repayment status or refuses to cooperate, the corresponding dimension weight is negative, and the transmission effect value is negative, representing a negative transmission influence. For example, the transmission effect value of debtor A's spouse = 0.8 × 0.9 × 0.9 = 0.648, which is positive and relatively high, indicating a strong positive transmission influence. Subsequently, the transmission effect value is used as the adjustment basis to correct the debtor's implicit static repayment intention state. Simultaneously, implicit correlation influence characteristics are supplemented (i.e., the fluctuation characteristics of intention brought about by the influence of related parties, which cannot be directly obtained through the debtor's own data). The value of this step lies in breaking through the limitation of only focusing on the debtor's own data, capturing the implicit influence of external related factors on repayment intention, and making the characterization of repayment intention more comprehensive. The transmission effect value is used as an adjustment factor to correct the debtor's implicit static repayment intention state. The adjustment logic is "a high positive transmission effect value increases the intention level, and a high negative transmission effect value decreases the intention level." At the same time, "related party relationship type, transmission effect direction and intensity" are integrated to form implicit correlation influence characteristics. For example, debtor A's static repayment intention characteristic is "medium-high intention." Combined with the spouse's positive transmission effect value of 0.648, it is adjusted to "high intention." The supplemented implicit correlation influence characteristic is "strong positive transmission from spouse (direct relative), increasing repayment intention."

[0044] By integrating static repayment willingness characteristics, dynamic repayment willingness characteristics, and implicit correlation influence characteristics, a multi-dimensional repayment willingness feature set is formed, encompassing static willingness, dynamic trends, and correlation influences. These three types of features comprehensively cover the core characteristics of repayment willingness from three core dimensions: static subjective tendency, dynamic change patterns, and external correlation influences. These three elements complement each other and are indispensable, ensuring that the feature set can comprehensively and accurately represent the debtor's repayment willingness. This lays a solid feature foundation for subsequent cross-domain alignment, causal screening, and the construction of a high-precision debtor profile model.

[0045] In some specific embodiments, the feature range for cross-domain alignment is determined. This range includes not only the multidimensional repayment willingness feature set but also business-related features relevant to debtor profiling. The multidimensional repayment willingness feature set and business-related features are mapped to the same latent space, and a causal relationship graph between features and debtor repayment behavior is constructed within this latent space. The direct causal relationship between features and repayment behavior in the causal relationship graph is identified. Falsely related features without a direct causal relationship to repayment behavior are removed from the causal relationship graph, while core features with a direct causal relationship are retained and integrated to form a causal feature library. By eliminating feature heterogeneity and refining causal relationship features, the problem of "falsely related feature interference" in traditional feature selection is solved, providing a high-quality causal feature library for subsequent weight calibration and model construction. Its implementation logic follows a progressive path of defining the feature range, unifying the feature space, mining causal relationships, and refining core features.

[0046] First, it is necessary to determine the feature scope for cross-domain alignment. This feature scope is not limited to the multidimensional repayment intention feature set, but includes business-related features related to the construction of the debtor profile. The multidimensional repayment intention feature set is the core feature that represents the debtor's subjective tendency to repay, dynamic changes, and external influences. The business-related features are an effective supplement to the core features, specifically including objective data such as the debtor's asset status certificate, income flow records, debt structure information, and occupational stability labels. These features can provide support for the judgment of repayment intention from the perspective of repayment ability. The combination of the two can achieve dual feature coverage of "repayment intention + repayment ability", avoiding the one-sidedness of the profile caused by a single feature dimension. The purpose of determining the feature scope is to clarify the processing objects of subsequent cross-domain alignment and ensure the comprehensiveness of feature coverage.

[0047] The second step is to map the multidimensional repayment willingness feature set and business-related features to the same latent space. The core of this step is to solve the heterogeneity problem between different features. The semantic features in the multidimensional repayment willingness feature set belong to text-derived features, the temporal features belong to sequence features, and the related person features belong to relational features. However, the business-related features are mostly numerical or categorical features. The expression format, scale, and dimension of different types of features are different. If causal analysis is performed directly, the results will be distorted due to the incomparability of features. The latent space is a standardized high-dimensional space constructed based on feature embedding technology. Its principle is to transform different types of features into low-dimensional dense vectors of the same dimension through feature encoding models (such as embedding layers and normalized mapping functions), eliminate the differences in scale and format between different features, and realize the comparability and fusion of features. For example, the semantic features of text are transformed into numerical vectors, and the occupational tags of categories are transformed into one-hot encoded vectors. Then, the mapping function projects all vectors into the same latent space to ensure that all types of features are subjected to subsequent causal analysis under the same standard.

[0048] Next, a causal relationship graph is constructed between features and debtor repayment behavior in the latent space, and direct causal relationships are identified. The causal relationship graph is a directed graph with "features" and "repayment behavior" as nodes and "causal relationship strength" as edge weights. Its construction principle is based on the analysis of the inherent logical relationship between each feature vector in the latent space and repayment behavior labels (such as "on-time repayment" and "overdue repayment") by the causal discovery algorithm, rather than simple statistical correlation. In specific implementation, by calculating the conditional independence between features and repayment behavior, it is determined whether the influence of features on repayment behavior is direct or indirect. For example, the feature "high semantic positive score" can directly affect repayment behavior, while the feature "occupation type" needs to affect repayment behavior indirectly by affecting "income stability". In this case, only the direct causal edge between "high semantic positive score" and repayment behavior is retained, and the indirect causal edge of "occupation type" is removed. The role of identifying direct causal relationships is to distinguish between "causal relationship" and "correlation relationship", and to provide a basis for subsequent removal of pseudo-correlated features.

[0049] Finally, spurious features are eliminated and integrated to form a causal feature library. Spurious features refer to features that are statistically related to repayment behavior but have no direct causal relationship. The existence of such features will seriously interfere with the accuracy of the model's judgment. For example, features such as "the debtor's zodiac sign" and "the geographical location of the bank" may have a weak correlation with repayment behavior in terms of data statistics, but they do not have an inherent logic to affect repayment behavior. Based on the causal relationship diagram constructed above, all spurious and indirect causal features that have no direct causal relationship with repayment behavior are eliminated, and only core features with direct causal relationship are retained. These core features include static intention features and dynamic window period features in the multidimensional repayment intention feature set, as well as income flow features and asset status features in business-related features. These core features are structured and integrated to form a causal feature library. The core value of this feature library is that it achieves precise purification of features, ensuring that all features included in the library are effective features that have a direct impact on repayment behavior, laying a solid foundation for subsequent weight calibration based on intervention effect analysis.

[0050] In some specific embodiments, the type of collection business feedback data is determined. This data includes repayment achievement rate, performance timeliness, second-time delinquency cases, and communication objections after the implementation of collection strategies. Intervention variables are set, which are variables related to the implementation of collection strategies, including at least one of collection timing selection, collection strategy type, and related party intervention methods. A sample matching algorithm is used to match the intervention sample with a control group sample that has similar characteristics but has not been intervened, eliminating sample selection bias and obtaining sample data. Based on the sample data, the intervention effect value of each feature in the causal feature library on repayment behavior is calculated. Weights are assigned to corresponding features according to the magnitude of the intervention effect value; the higher the intervention effect value, the greater the weight of the feature, thus completing the weight calibration of the features in the causal feature library. By quantifying the importance of features through feedback information from real business scenarios, the problem of traditional feature weight settings being detached from actual business is solved, ensuring that the core features in the causal feature library can accurately match the needs of collection business. Its implementation logic follows a closed-loop path: determining the type of feedback data, setting intervention variables, sample matching to eliminate bias, calculating the intervention effect, and weight calibration and assignment.

[0051] First, the types of feedback data from the collection process are determined. The four selected data categories—repayment achievement rate, performance timeliness, second delinquency, and communication objections—are all direct results generated after the implementation of collection strategies, comprehensively reflecting the effectiveness of the business. Repayment achievement rate refers to the proportion of debtors who successfully repay after the implementation of a specific collection strategy, serving as a core quantitative indicator for measuring collection effectiveness. Performance timeliness refers to the time interval between the execution of collection and the completion of repayment, used to assess the timeliness of repayment. Second delinquency refers to the probability of a debtor defaulting again after repayment, reflecting the stability of repayment performance. Communication objections refer to subjective feedback information raised by debtors during the collection process, such as obstacles to repayment and requests, supplementing the quantitative data. These four data categories work together to provide comprehensive and accurate business feedback from four dimensions—repayment results, repayment efficiency, performance stability, and subjective requests—for subsequent intervention effect analysis, avoiding weighting bias caused by a single data dimension.

[0052] Secondly, intervention variables are set. These variables are directly related to the execution of the collection strategy, including the timing of collection, the type of collection strategy, and the intervention method of related parties. These variables are the core basis for distinguishing between the "intervention group" and the "control group," and directly affect the collection effect and the performance of the characteristics. The timing of collection includes different time nodes such as collection on payday, collection on holidays, and collection on weekdays; the type of collection strategy includes different collection methods such as telephone collection, SMS collection, and door-to-door collection; and the intervention method of related parties includes different intervention paths such as directly contacting related parties to urge repayment and conveying collection information through related parties. The purpose of setting intervention variables is to clarify the "intervention actions" for subsequent analysis, and to ensure that the actual impact of each characteristic in the causal feature database on repayment behavior can be accurately quantified by comparing the differences between the intervention group and the control group.

[0053] Next, a sample matching algorithm is used to eliminate sample selection bias. Sample selection bias refers to the difference in characteristics between the directly selected intervention group and the control group, which leads to the comparison results failing to truly reflect the effects of the intervention variables and features. For example, directly comparing debtors who are contacted on payday (intervention group) with randomly selected debtors who are not contacted (control group) may result in distorted results because the debtors in the intervention group have a higher willingness to repay. The core principle of the sample matching algorithm is to match one or more control group samples that are highly similar to the sample in terms of causal feature dimensions but have not been intervened, for each debtor sample that has been intervened. The matching dimensions include core features such as static repayment willingness level, overdue duration, and asset status. For example, for a debtor who is "overdue for 30 days, has a medium to high repayment willingness, a monthly income of 5,000 yuan" and is contacted on payday, a debtor with exactly the same characteristics but who has not been contacted is matched as a control group. In this way, the interference caused by the differences in the characteristics of the samples can be eliminated, ensuring that the intervention effect value calculated later only reflects the true impact of the intervention variables and features on repayment behavior, and finally obtaining sample data with a balanced feature distribution.

[0054] Next, the intervention effect value is calculated based on the matched sample data. The intervention effect value is the core indicator for quantifying the influence of each feature in the causal feature library on repayment behavior. Its calculation principle is to compare the difference in repayment behavior between the intervention group sample and the control group sample under the same feature dimension. The specific calculation formula can adopt the average intervention effect formula, that is, intervention effect value = (average repayment achievement rate of intervention group sample - average repayment achievement rate of control group sample) / average repayment achievement rate of control group sample. The higher the value, the stronger the positive impact of the feature on repayment behavior. For example, for the feature of "high willingness window period", the repayment achievement rate of the intervention group (collection during the window period) is 60%, and the repayment achievement rate of the control group (collection outside the window period) is 30%. Then the intervention effect value of this feature is 100%, which shows that this feature has a significant effect on improving the repayment achievement rate. In the calculation process, it is necessary to combine the feedback data of multiple dimensions such as repayment achievement rate, performance timeliness, and second delinquency to conduct a comprehensive evaluation to ensure that the intervention effect value can fully reflect the business value of the feature.

[0055] Finally, features are weighted according to the magnitude of the intervention effect value to complete weight calibration. The core logic of weight assignment is that "the higher the intervention effect value, the greater the feature weight." That is, the more significant the influence of a feature on repayment behavior, the higher its weight ratio in subsequent model construction. For example, the intervention effect value of the "high willingness window" feature is 100%, and it is assigned a weight of 0.3; the intervention effect value of the "stable income flow" feature is 50%, and it is assigned a weight of 0.15. In this way, the role of core features can be highlighted and the interference of secondary features can be weakened. After the weight calibration is completed, the features in the causal feature library not only have direct causal relationship, but are also assigned importance weights that match the actual business, which effectively improves the accuracy of subsequent core feature library construction and profile model training, laying a key foundation for finally achieving high-precision debtor profiles and accurate collection strategy matching.

[0056] In some specific embodiments, the process of acquiring the core feature library includes: extracting target features whose weight fluctuation exceeds a preset threshold from the weight calibration results; adjusting parameters for the feature mining steps corresponding to the target features; constructing a causal relationship network between features in the causal feature library and debtor repayment behavior, clarifying the association between features and repayment behavior; performing counterfactual hypothesis deduction on the core features in the causal relationship network, calculating the deviation between predicted repayment behavior and actual repayment behavior under the assumed conditions; comparing the deviation value with a preset validity threshold, retaining valid features with deviation values ​​exceeding the threshold, eliminating invalid features with deviation values ​​below the threshold, and integrating valid features to form the core feature library. The process is attached. Figure 4 As shown, through a progressive operation of parameter iterative optimization, causal relationship verification, and counterfactual feature filtering, a secondary purification of the causal feature library is achieved, ensuring that the final core feature library retains only key features that significantly affect repayment behavior.

[0057] First, target features whose weight fluctuation exceeds a preset threshold are extracted from the weight calibration results, and parameters are adjusted for the corresponding feature mining stages. Here, "weight fluctuation" refers to the percentage change in feature weight after calibration relative to before calibration. The preset threshold is a critical value set based on business experience or statistical models, such as 15%. The core purpose of setting the threshold is to filter out features whose importance has changed significantly. Weight fluctuations in these features often stem from unreasonable parameter settings in the early stages of feature mining, resulting in the features failing to fully capture information related to repayment willingness. "Target features" are those whose weight fluctuation exceeds this threshold, and the corresponding feature mining stages are semantic-driven decoding, time-series-driven modeling, or association-driven mining stages that generate these features. The core principle of parameter adjustment is to optimize key parameters at the source of feature mining to ensure that the adjusted features can more accurately represent repayment willingness. For example, if the target feature is a time-driven dynamic willingness feature and the weight fluctuation range reaches 22%, then the key time interval division rules in the time-series modeling stage should be adjusted accordingly, changing the original window period of "3 days after payday" to "2 days after payday", or optimizing the calculation weight of the willingness state transition probability. If the target feature is a semantically driven static willingness feature, then the matching threshold of semantic labels or the correction coefficient of polarity score on the initial state transition probability should be adjusted. This operation can improve the quality of the target feature and avoid feature failure caused by mining parameter deviation.

[0058] Secondly, a causal relationship network is constructed between features in the causal feature library and debtors' repayment behavior to clarify the correlation between features and repayment behavior. This "causal relationship network" is a structured topological network with "features in the causal feature library" and "repayment behavior" as nodes, "direct causal associations" as directed edges, and "intervention effect weights" as edge weights. Its construction principle is based on the direct causal association results obtained from previous causal screening, transforming abstract causal relationships into a visualized and quantifiable network structure. During construction, each feature node establishes a direct directed edge only with the repayment behavior node, with the edge weight being the intervention effect weight of that feature. This clarifies "which features directly affect repayment behavior" and "to what extent." For example, the "high willingness window" feature node points to the "increased repayment achievement rate" behavior node, with an edge weight of 0.8, while features without a direct causal relationship, such as "occupation type," are not included in the network. The core function of this causal relationship network is to define a clear analytical scope for subsequent counterfactual reasoning, avoiding irrelevant features from interfering with the verification process and ensuring that counterfactual inference is only conducted on core features directly related to repayment behavior.

[0059] Then, counterfactual hypothesis deduction is performed on the core features in the causal relationship network to calculate the deviation between predicted repayment behavior and actual repayment behavior under the assumed conditions. The core principle of counterfactual hypothesis deduction is "how would the repayment behavior change if this feature were not present?" Specifically, counterfactual assumptions are artificially set for each core feature in the causal relationship network. For example, "reduce the weight of this feature by 50%", "remove this feature", or "adjust the value of this feature to the opposite state". Then, based on a pre-set repayment behavior prediction model, the assumed conditions are substituted to calculate predicted repayment behavior data. The predicted data is then compared with actual repayment behavior data in real business scenarios to calculate the deviation. The dimensions of the deviation calculation must be consistent with the core business indicators, including repayment achievement rate deviation, performance timeliness deviation, and secondary delinquency rate deviation. The larger the absolute value of the deviation, the more significant the impact of the feature on repayment behavior. For example, regarding the "high willingness window" feature, if removing this feature reduces the predicted repayment achievement rate from the actual 60% to 28%, the repayment achievement rate deviation is 32%. This result directly reflects the key role of this feature in improving the repayment achievement rate.

[0060] Finally, the deviation values ​​are compared with the preset validity threshold, and valid features are retained and integrated to form a core feature library. The "preset validity threshold" here is a critical deviation value set based on the actual needs of debt collection operations, for example, 20%. The logic is that "if the deviation value exceeds the threshold, it indicates that the feature has a significant impact on repayment behavior and is a valid feature; if the deviation value does not reach the threshold, it indicates that the feature has a weak impact on repayment behavior and is an invalid feature." During the comparison process, core features in the causal relationship network are screened one by one. Features with deviation values ​​exceeding the validity threshold are retained, while invalid features with deviation values ​​below the threshold are removed. For example, the "high willingness window" feature has a deviation value of 32%, exceeding the 20% threshold, and is retained; the "debtor's bank" feature has a deviation value of only 5%, and is removed. Finally, all retained valid features are structurally integrated to form the core feature library. This core feature library is the final feature set after multiple rounds of purification through "causal screening, weight calibration, parameter optimization, and counterfactual verification." It contains only the key features that have the most significant impact on repayment behavior, providing the most core and effective feature support for subsequently building a high-precision debtor profile model.

[0061] In some specific embodiments, parameter adjustments are made to the feature mining stage corresponding to the target feature. Specifically, if the target feature originates from time-series driven repayment intention trend modeling, the key time intervals are redefined, and the calculation weights of the repayment intention state transition probability within the intervals are adjusted. If the target feature originates from the mining of the repayment intention transmission effect driven by related parties, the calculation coefficients of the related party transmission effect are recalibrated. If the target feature originates from semantic-driven latent state decoding of repayment intention, the matching rules of semantic labels are optimized, and the correction magnitude of polarity scores to the initial state transition probability is adjusted. The core logic of this differentiated parameter adjustment strategy for target features from different sources is "feature tracing - parameter matching - precise optimization." By adjusting the parameters at the source of feature mining, the problem of excessive fluctuations in target feature weights is solved, improving the accuracy of feature representation of repayment behavior and providing a high-quality feature foundation for the subsequent purification of the core feature library.

[0062] If the target feature originates from time-series-driven repayment intention trend modeling, the corresponding parameter adjustments focus on two core dimensions: the division of key time intervals and the weighting of the calculation of repayment intention state transition probabilities within each interval. The division of key time intervals directly determines whether the time-series feature can capture the core change nodes of repayment intention. Large fluctuations in weights indicate that the original interval division rules do not match the actual business scenario. For example, the original rule divides "5 days after payday" as the key interval, but the actual repayment peak is concentrated within 2 days after payday. Re-dividing the intervals adjusts the key period to a range that better reflects the actual business situation, ensuring that the time-series feature can accurately anchor the high-intention window. Adjusting the calculation weight of the state transition probability within the interval is based on the principle that different sub-intervals contribute differently to changes in repayment intention. For example, increasing the calculation weight of the transition probability 1-2 days after payday from 0.3 to 0.6 reduces the weight of off-peak periods, thereby strengthening the representation ability of the intention change pattern in the core period and improving the effectiveness of the time-series-driven target feature.

[0063] If the target feature originates from the mining of the repayment willingness transmission effect driven by related parties, the core of parameter adjustment is to recalibrate the calculation coefficients of the related party transmission effect. The calculation of the related party transmission effect value depends on the weight coefficients of three core indicators: relationship type, performance status, and response behavior. Excessive fluctuations in weights indicate that the original coefficient settings cannot truly reflect the degree of influence of related parties on the debtor's repayment willingness. For example, in the original rule, the coefficient for the relationship type "spouse" was 0.7, but in actual business, the spouse's influence on the debtor is significantly higher. The calibration coefficient is to adjust the weight of this type of relationship to 0.9. Another example is that in the original rule, the coefficient for "related party performs well" was 0.5, which was increased to 0.8 after calibration. By optimizing the coefficients, the calculation of the transmission effect value is made to better reflect the actual strength of the related party influence, correcting the matching deviation between related party-driven features and repayment behavior.

[0064] If the target feature originates from semantically driven latent state decoding of repayment intention, parameter adjustments revolve around the semantic label matching rules and the correction magnitude of polarity scores on the initial state transition probability. Semantic label matching rules determine the accuracy of extracting core attitude information from interactive semantic data. Large weight fluctuations indicate that the original rules have ambiguous label definitions. For example, the original rules uniformly categorize "considering repayment" and "promising repayment" as positive labels, leading to an overly high polarity score. Optimizing the matching rules involves refining the label classification standards, clearly distinguishing the boundaries between different semantics such as "promising repayment," "considering repayment," and "refusing repayment," thereby improving the accuracy of label extraction. Adjusting the correction magnitude of polarity scores is based on the principle that the original correction coefficient cannot reasonably balance the weight ratio of objective business data and subjective semantic data. For example, if the original polarity score correction coefficient is 0.5, the subjective semantic correction effect on the initial latent state is weak. Increasing the correction coefficient to 0.7 strengthens the correction effect of the true subjective repayment attitude on the initial latent state, allowing semantically driven static features to more accurately reflect the debtor's actual repayment tendency.

[0065] Differentiated parameter adjustment strategies ensure that the mining parameters for each type of target feature can be adapted to its own technical logic and business scenario, effectively reducing the fluctuation range of target feature weights, improving feature quality, and laying a solid foundation for subsequent causal relationship network construction and counterfactual reasoning verification.

[0066] In some specific embodiments, a model architecture suitable for the debt collection business scenario is selected, and a profiling model is built based on this architecture. The model architectures include gradient boosting tree models, random forest models, single-layer perceptrons, and shallow neural networks. The selection criteria for the model architecture are anchored to the needs of the debt collection business scenario. The core requirements of the profiling model for the debt collection business are threefold: first, real-time performance, requiring the generation of debtor profiles within milliseconds to support the immediate matching of collection strategies; second, interpretability, requiring clear understanding of the impact weights of core output features on the profiling results, facilitating debt collectors' understanding of the profiling basis and the development of targeted strategies; and third, adaptability, requiring compatibility with the structured, moderately dimensional feature set (static intention, dynamic trend, and correlation influence features) in the core feature library, without needing to process high-dimensional unstructured data. Based on these requirements, gradient boosting tree models, random forest models, single-layer perceptrons, and shallow neural networks are preferentially selected. These four models all possess the common advantages of being "lightweight and efficient, highly interpretable, and adaptable to small-to-medium-scale structured features," which are highly compatible with the debt collection business scenario.

[0067] Gradient boosting tree models, commonly implemented as GBDT, XGBoost, and LightGBM, work by iteratively generating multiple weak decision trees, successively fitting the prediction errors of the previous model, and finally weighting and integrating the predictions of all weak decision trees to obtain the final output. During model construction, feature vectors from a core feature library are used as the input layer, and debtor repayment behavior labels (such as "high willingness to repay," "medium willingness to repay," and "low willingness to repay") are used as the output layer. The model is optimized by adjusting key parameters such as the learning rate, tree depth, and number of leaf nodes. The core advantage of this model is its strong interpretability, directly outputting the contribution of each feature to the profile result (e.g., the "high willingness window" feature contributes up to 35%), and its high prediction accuracy, making it suitable for accurately identifying and constructing core profiles of debtors with high repayment willingness.

[0068] The core principle of the Random Forest model is to construct multiple independent decision trees. By using "randomly sampled features + random sampled samples," the overfitting risk of a single decision tree is reduced. Finally, a voting mechanism is used to integrate the prediction results of all decision trees. When building the model, the input and output are consistent with the gradient boosting tree model. During training, the model's generalization ability is improved by controlling parameters such as the number of decision trees and the feature sampling ratio. The advantage of this model is its strong anti-interference ability, which can effectively handle the small amount of noisy data in the core feature library (such as the polarity score bias of some semantic labels). It is suitable for building general debtor profiles with high stability requirements.

[0069] The single-layer perceptron, as the simplest neural network model, is based on the principle of linear weighted summation plus activation function mapping. The model structure contains only an input layer and an output layer, with no hidden layers. When building the model, the dimension of the input layer is perfectly matched with the feature dimension of the core feature library (e.g., if the core feature library has 12 features, then the input layer has 12 neurons). The input features are linearly weighted by a weight matrix, and then the weighted result is mapped to the probability value corresponding to the profile label by an activation function (such as the Sigmoid function or the ReLU function). The advantages of this model are its extremely simple structure and extremely fast operation speed. It can generate profiles in milliseconds, which fully meets the real-time requirements of debt collection systems and is suitable for high-frequency, large-scale scenarios of rapid generation of debtor profiles.

[0070] Shallow neural networks, in their core principle, add 1-2 hidden layers to a single-layer perceptron. The hidden layers capture implicit correlations between features through nonlinear transformations. The model is constructed with an "input layer-hidden layer-output layer" structure. The number of neurons in the hidden layer is typically 2-3 times the feature dimension of the input layer. During training, the weights of each layer are optimized using the backpropagation algorithm. The advantage of this model is that it combines the efficiency of linear models with the nonlinear fitting ability of deep learning models. It can capture the interactive influence between core features (such as the superposition effect of "static high willingness" and "strong positive transmission from spouse"), achieving a better balance between accuracy and efficiency. It is suitable for complex debt collection scenarios with high requirements for profile accuracy.

[0071] Regardless of the chosen model architecture, a unified construction process must be followed: First, preprocess the features in the core feature library by normalizing or standardizing to eliminate dimensional differences between features. Second, divide the model into training, validation, and test sets. The training set uses core feature data of historical debtors and corresponding repayment behavior labels; the validation set is used to adjust model parameters; and the test set is used to evaluate the final model performance. Third, train the model based on the selected architecture, optimizing model parameters by monitoring metrics such as prediction accuracy and AUC value on the validation set. Fourth, after completing model training, verify the model's generalization ability using the test set to ensure stable prediction accuracy on unseen new data, ultimately forming a debtor profile model adapted to debt collection scenarios. Figure 5 It is a live demonstration diagram of model training and optimization, which includes the unique identifier of the corresponding training model, the associated training task number, the algorithm type or structure of the model, the number of data samples used for training, the current training status of the model, the training completion time of the model, the number of optimization rounds the model has gone through, and the model's performance evaluation metrics.

[0072] The first step is to preprocess the core feature library, which contains various types of features such as static repayment willingness features, dynamic willingness trend features, and implicit correlation influence features. It contains heterogeneous data such as numerical (e.g., transmission effect value), categorical (e.g., willingness level label), and sequence (e.g., window period time series). It is necessary to eliminate the dimensional differences of numerical features through normalization, convert categorical features into a vector format that the model can recognize through one-hot encoding, and compress sequence features into fixed-length vectors through sequence sampling to ensure that all features are in a structured and standardized input format. The second step is to select a model architecture based on the core requirements of the debt collection business. If high interpretability is desired, a gradient boosting tree or random forest model is chosen, using the core feature vector as the input layer and "repayment willingness level, performance potential score, and high willingness window" as the output layer. The model is optimized by adjusting parameters such as tree depth, number, and learning rate. During training, the core feature data of historical debtors and corresponding repayment behavior labels are used as the dataset, divided into training, validation, and test sets in a 7:2:1 ratio. The validation set is used to monitor the model's prediction accuracy, AUC value, recall, and other indicators to avoid overfitting. If extreme real-time performance is desired, a single-layer perceptron or shallow neural network is chosen, with the input layer dimension perfectly matching the core feature dimension. The number of neurons in the hidden layer is 2-3 times that of the input layer. The weights are optimized using the backpropagation algorithm. An early stopping strategy is adopted during training, terminating training when the validation set loss no longer decreases after several consecutive rounds to ensure the model's generalization ability. After completing the model training in the third step, the performance is verified using a test set. The core indicators of the collection business must be met: prediction accuracy of no less than 85% and prediction time of a single sample of no more than 5ms. After the verification is passed, the basic profile model is obtained.

[0073] The model of this application is deployed to a specific application scenario, and a real-world demonstration is attached. Figure 6 As shown. Users can click the "Deployment Scheme Configuration" button to enter the deployment scheme configuration interface or module and set relevant parameters. Before deployment, the model can be lightweighted.

[0074] The core of lightweight processing is to reduce model size and computational consumption without significantly sacrificing model accuracy, adapting to the hardware resources and real-time response requirements of existing debt collection systems. This is achieved through four targeted methods: First, model pruning: for tree models, redundant decision trees and tree nodes with low contribution are pruned; for neural network models, redundant connections with absolute weights below a preset threshold are pruned, reducing the number of model parameters. Second, parameter quantization: the storage format of model parameters is compressed from 32-bit floating-point to 16-bit floating-point or 8-bit integer, significantly reducing the model's memory footprint. For example, the FP16 format can compress the model size by 50% with less than 3% impact on model accuracy in debt collection scenarios. Third, model distillation: if a complex model such as a gradient boosting tree is selected, distillation techniques can be used to guide the training of a simpler model with the prediction probability of the complex model, improving the model inference speed by 2-3 times while retaining over 95% accuracy. Fourth, structural optimization: for shallow neural networks, redundant hidden layers are merged and the number of neurons is adjusted to the optimal scale; for tree models, similar decision rules are merged to further simplify the model structure. After lightweighting, the model performance needs to be verified again to ensure that the single-sample prediction time is compressed to within 2ms, the model size is compressed to 30%-50% of the original size, and the prediction accuracy decreases by no more than 5%, so as to meet the real-time requirements of the collection system.

[0075] In some embodiments, a dynamic mapping rule between profile tags and collection strategies is constructed. The specific implementation logic is as follows: First, a pre-defined strategy mapping rule library is established. Based on collection business experience, a correspondence between profile tags and collection strategies is established. For example, "high willingness + payday window" is mapped to "SMS reminder + low-frequency phone collection," "medium willingness + strong positive transmission from spouse" is mapped to "intervention by related parties + recommendation of installment repayment plan," and "low willingness + no transmission from related parties" is mapped to "legal notification + in-person collection." Priorities are also set for different strategies. Second, real-time profile generation is implemented. When the collection system obtains the latest core characteristic data of the debtor, the model interface is automatically triggered, generating a debtor profile within 10ms and returning it to the system. Third, precise strategy matching is achieved. The system reads the core tags from the profile results, retrieves the optimal matching collection strategy from the mapping rule library, and dynamically adjusts the strategy based on the debtor's historical collection records. Fourth, establish an effect monitoring and iteration mechanism to track business indicators such as repayment achievement rate and performance timeliness after each strategy matching, regularly analyze the matching effect of profile tags and strategies, optimize the mapping rule base and model parameters, form a closed loop, and continuously improve the accuracy and efficiency of collection business.

[0076] A debtor profiling system based on repayment willingness mining is shown in the attached diagram. Figure 7 As shown, the system includes: Data acquisition unit 1 is used to collect various types of data from debtors, and to map the heterogeneous data formed by the preprocessing of various types of data to a unified feature space to form a fused dataset; Causal feature unit 2 is used to perform feature mining based on the fused dataset to obtain a multidimensional repayment intention feature set; cross-domain alignment and causal filtering are performed on the multidimensional repayment intention feature set to remove features that are not directly causally related to repayment behavior, forming a causal feature library; The weight calibration unit 3 is used to take the collection business feedback data as incremental input, perform weight calibration on the features in the causal feature library based on intervention effect analysis, adjust the parameters of feature mining according to the weight calibration results, verify the effectiveness of the causal feature library through counterfactual reasoning, and obtain the core feature library. The profile building unit 4 is used to build a debtor profile model based on the core feature library. After the model is lightweighted, it is encapsulated into a standardized interface and embedded into the existing collection system to realize the real-time generation of debtor profiles and accurate matching of collection strategies.

[0077] Those skilled in the art will understand that the components of this application described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage system for execution by the computing system. Alternatively, they can be fabricated as separate integrated circuit components, or multiple components or steps can be fabricated as a single integrated circuit component. Thus, this application is not limited to any particular combination of hardware and software.

[0078] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

[0079] The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for constructing a debtor profile based on repayment willingness mining, characterized in that, include: Collect various types of data from debtors, and map the heterogeneous data formed by preprocessing these various types of data to a unified feature space to form a fused dataset; Feature mining is performed on the fused dataset to obtain a multidimensional repayment willingness feature set; Cross-domain alignment and causal filtering are performed on the multidimensional repayment intention feature set to remove features that have no direct causal relationship with repayment behavior, thus forming a causal feature library. Using collection business feedback data as incremental input, the features in the causal feature library are weighted and calibrated based on intervention effect analysis. Based on the weight calibration results, the parameters of the feature mining are adjusted in a targeted manner, and the effectiveness of the causal feature library is verified by counterfactual reasoning to obtain the core feature library; Based on the core feature library, a debtor profile model is constructed. After lightweight processing, the model is encapsulated into a standardized interface and embedded into the existing debt collection system to achieve real-time generation of debtor profiles and accurate matching of collection strategies.

2. The debtor profiling construction method according to claim 1, characterized in that, The feature mining includes semantic-driven latent state decoding of repayment intention, time-series-driven trend modeling of repayment intention, and related person-driven mining of the transmission effect of repayment intention.

3. The debtor profile construction method according to claim 2, characterized in that, The various types of data include interactive semantic data that characterizes semantic signals of repayment willingness, time-series correlation data that reflects dynamic changes in willingness, correlation data that reflects the influence of correlation, and business tag data that supports basic judgment of the profile. The preprocessing includes noise removal and keyword extraction of interactive semantic data, timestamp sorting and three-dimensional sequence integration of time-series related data, label standardization of related person data, and compliance desensitization and redundant data filtering of all data.

4. The debtor profile construction method according to claim 2, characterized in that, Semantic-driven latent state decoding of repayment intention includes: calculating the initial latent state of repayment intention and state transition probability based on business tag data, extracting semantic tags and polarity scores from interactive semantic data, and using the semantic tags and polarity scores as correction factors to adjust the initial state transition probability to obtain static repayment intention features; The time-driven repayment intention trend modeling includes: aligning the static repayment intention latent state with the collection strategy execution time sequence and external event time sequence to form a three-dimensional data stream; statistically analyzing the intention state transition probability in segments according to preset key time intervals; generating a repayment intention trend curve and marking the high intention window period to obtain dynamic repayment intention characteristics. The mining of the transmission effect of repayment intention driven by related parties includes: calculating the transmission effect value based on the relationship type, performance status and response behavior in the related party data, adjusting the implicit state of the debtor's static repayment intention according to the transmission effect value, and supplementing the implicit association influence characteristics; By integrating static repayment willingness characteristics, dynamic repayment willingness characteristics, and implicit correlation influence characteristics, a multi-dimensional repayment willingness feature set is formed, which includes static willingness, dynamic trends, and correlation influence.

5. The debtor profiling construction method according to claim 1, characterized in that, The feature range for cross-domain alignment is determined, which includes not only the multidimensional repayment willingness feature set, but also business-related features related to the construction of debtor profiles; The multidimensional repayment willingness feature set and the business-related features are mapped to the same latent space, and a causal relationship graph between the features in the latent space and the debtor's repayment behavior is constructed to identify the direct causal relationship between the features and the repayment behavior in the causal relationship graph. The pseudo-correlation features that have no direct causal relationship with the repayment behavior in the causal relationship graph are removed, and the core features with direct causal relationship are retained and integrated to form a causal feature library.

6. The debtor profile construction method according to claim 1, characterized in that, Determine the type of the collection business feedback data, which includes the repayment achievement rate, performance timeliness, second delinquency cases, and content of communication objections after the collection strategy is implemented; Define intervention variables, which are variables related to the execution of collection strategies, including at least one of collection timing selection, collection strategy type, and intervention method of related parties; A sample matching algorithm was used to match the intervention-treated debtor samples with similar characteristics to the control group samples that had not been intervened, thereby eliminating sample selection bias and obtaining sample data. Based on the sample data, the intervention effect value of each feature in the causal feature library on repayment behavior is calculated; weights are assigned to the corresponding features according to the magnitude of the intervention effect value, with the higher the intervention effect value, the greater the weight of the feature, thus completing the weight calibration of the features in the causal feature library.

7. The debtor profile construction method according to claim 2, characterized in that, The process of acquiring the core feature library specifically includes: Extract target features from the weight calibration results whose weight fluctuation exceeds a preset threshold, and adjust the parameters of the feature mining step corresponding to the target features; Construct a causal relationship network between features in the causal feature library and debtor repayment behavior, and clarify the correlation between features and repayment behavior; perform counterfactual hypothesis deduction on the core features in the causal relationship network, and calculate the deviation between predicted repayment behavior and actual repayment behavior under the assumed conditions; The deviation values ​​are compared with a preset validity threshold. Valid features with deviation values ​​exceeding the threshold are retained, while invalid features with deviation values ​​below the threshold are removed. The valid features are then integrated to form a core feature library.

8. The debtor profile construction method according to claim 2, characterized in that, Parameter adjustments are made to the feature mining process corresponding to the target features, specifically including: If the target feature originates from time-driven repayment intention trend modeling, then the key time intervals are redefined and the calculation weights of the repayment intention state transition probability within the intervals are adjusted. If the target feature originates from the mining of the repayment willingness transmission effect driven by related parties, then the calculation coefficient of the related party transmission effect should be recalibrated. If the target feature originates from semantically driven latent state decoding of repayment intention, then optimize the matching rules of semantic labels and adjust the correction magnitude of polarity scores to the initial state transition probability.

9. The debtor profiling construction method according to claim 1, characterized in that, A model architecture suitable for debt collection business scenarios is selected, and a profile model is constructed based on the model architecture; the model architecture includes gradient boosting tree model, random forest model, single-layer perceptron, and shallow neural network.

10. A debtor profiling system based on repayment willingness mining, characterized in that, include: The data acquisition unit is used to collect various types of data from debtors, and to map the heterogeneous data formed by the preprocessing of these various types of data to a unified feature space to form a fused dataset. The causal feature unit is used to perform feature mining based on the fused dataset to obtain a multidimensional repayment willingness feature set. Cross-domain alignment and causal filtering are performed on the multidimensional repayment intention feature set to remove features that have no direct causal relationship with repayment behavior, thus forming a causal feature library. The weight calibration unit is used to take the collection business feedback data as incremental input and perform weight calibration on the features in the causal feature library based on intervention effect analysis. Based on the weight calibration results, the parameters of the feature mining are adjusted in a targeted manner, and the effectiveness of the causal feature library is verified by counterfactual reasoning to obtain the core feature library; The profile building unit is used to build a debtor profile model based on the core feature library. After the model is lightweighted, it is encapsulated into a standardized interface and embedded into the existing collection system to realize the real-time generation of debtor profiles and accurate matching of collection strategies.