A time series map-based collection implicit fraud identification method and system

By constructing a conversational temporal semantic graph centered on the target users of the collection, the problem of the inability to identify hidden fraud in collection in existing technologies has been solved. This has enabled accurate identification and comprehensive improvement of hidden fraud, provided accurate risk basis, and improved the efficiency of intelligent collection.

CN122155825APending Publication Date: 2026-06-05SHENZHEN HUOCHUANG ZHIHUI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN HUOCHUANG ZHIHUI TECHNOLOGY CO LTD
Filing Date
2026-02-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing intelligent collection systems are unable to effectively identify in-depth mining and multi-dimensional correlation analysis of the entire collection cycle of conversation data, making it difficult to identify hidden fraudulent behavior and unable to provide accurate risk basis for the formulation of collection strategies.

Method used

Construct a conversation temporal semantic graph centered on the target users of debt collection. By collecting the associated data of the entire debt collection conversation, extract skeleton attributes, mine fine-grained attributes, integrate exclusive fraud features, and perform multi-dimensional reasoning and quantification to form a comprehensive risk feature matrix. Input the matrix into a pre-trained risk judgment model to determine implicit fraud.

Benefits of technology

It has achieved accurate identification of hidden fraud, reduced the false negative rate, improved the accuracy and comprehensiveness of debt collection fraud identification, provided financial institutions with accurate risk information, and improved the efficiency of intelligent debt collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122155825A_ABST
    Figure CN122155825A_ABST
Patent Text Reader

Abstract

The application provides a collection implicit fraud identification method and system based on a timing graph, including: collecting full-cycle data and preprocessing to form an original data set; constructing a basic graph with a collection target user as the core; mining fine-grained attributes and extracting exclusive fraud features to form a feature-rich conversation timing semantic graph; performing multi-dimensional reasoning, extracting core risk features and quantifying to form a comprehensive risk feature matrix; inputting a risk judgment model to perform implicit fraud judgment and outputting a judgment result. The application constructs a conversation timing semantic basic graph with a collection target user as the core, not only deeply fuses the implicit fraud features and the collection interaction basic features, but also multi-dimensionally and deeply mines and quantitatively represents the implicit fraud features, realizes the graphed and timed unified representation of the collection full-cycle interaction data, and improves the accuracy and comprehensiveness of the collection fraud identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent debt collection, and more specifically, to a method and system for identifying hidden fraud in debt collection based on time-series graphs. Background Technology

[0002] In intelligent debt collection scenarios for post-credit management in the financial sector, fraud detection is a core element in improving collection efficiency and reducing bad debt risk. Existing intelligent debt collection systems' fraud detection solutions are mostly based on single performance records, semantic keyword matching, or simple behavioral feature analysis to build identification models. These models can only identify explicit fraudulent behaviors with clear fraudulent characteristics, lacking the ability to deeply mine and multi-dimensionally correlate data across the entire debt collection cycle. This makes it difficult to effectively identify implicit fraudulent behaviors without clear fraud indicators, such as semantic evasion, false cooperation, or collusion among multiple parties, and thus fails to provide accurate risk basis for debt collection strategy formulation. Summary of the Invention

[0003] To address the problems existing in current technologies, this application provides a method and system for identifying hidden fraud in debt collection based on time-series graphs. The specific solution is as follows: A method for identifying hidden fraud in debt collection based on time-series graphs, comprising: Collect and preprocess data related to the entire collection cycle of the target debt collection session to form the raw dataset; Extract skeleton attributes for building a time series graph from the original dataset, and construct a basic graph centered on the target users of debt collection based on the skeleton attributes; Based on the original dataset, fine-grained attributes of the skeleton attributes are mined and specific fraud features for debt collection scenarios are extracted. These specific fraud features are then fused into the base graph to form a feature-enriched conversational temporal semantic graph. Based on the aforementioned conversational temporal semantic graph, multi-dimensional reasoning is performed to extract the core risk characteristics of three types of hidden fraud: deflection fraud, false cooperation fraud, and collusion fraud. A comprehensive risk characteristic matrix is ​​then formed based on these core risk characteristics. The comprehensive risk feature matrix is ​​input into the pre-trained risk judgment model to determine implicit fraud, and the judgment result includes implicit fraud type, fraud probability, and comprehensive risk level.

[0004] In some specific embodiments, the full-cycle collection session associated data includes semantic interaction data, non-semantic behavior data, time-series benchmark data, entity association data, and performance record data.

[0005] In some specific embodiments, the acquisition of skeleton attributes includes: extracting semantic class attributes representing the core semantic logic of the session from semantic interaction data, extracting behavioral class attributes representing user collection interaction behavior from non-semantic behavior data, extracting time-series class attributes representing session time correlation features from time-series benchmark data, extracting entity class attributes representing subject correlation relationships from entity association data, and extracting performance class attributes representing user performance behavior status from performance record data; performing fraud identification correlation verification on the extracted attributes, eliminating redundant attributes, and obtaining skeleton attributes.

[0006] In some specific embodiments, the process of obtaining the exclusive fraud features includes: performing feature preprocessing on the fine-grained attributes; based on the behavioral and semantic features of three types of implicit fraud in the debt collection scenario—evasive fraud, feigned cooperation fraud, and collusion fraud—feature clustering is performed on the preprocessed fine-grained attributes, and the core fine-grained attributes corresponding to each type of implicit fraud are extracted as exclusive fraud features; the exclusive fraud features all possess quantifiable, temporally sequential, and graphically mapped feature attributes.

[0007] In some specific embodiments, feature mapping associations are established with the node set and edge set of the basic graph based on the type and representation dimension of the exclusive fraud features; quantitative attribute fields of exclusive fraud features are added to the nodes, and fraud feature correlation degree and temporal coupling feature attributes are added to the edges between nodes; triggering time sequence identifiers and time sequence trajectories of exclusive fraud features are added at the time sequence dimension layer to achieve deep integration and graph representation of exclusive fraud features and the original skeleton attributes of the basic graph.

[0008] In some specific embodiments, multi-dimensional reasoning includes: Anomaly detection and correlation mining are performed on the exclusive fraud features added to each node to identify fraud feature aggregation anomalies at the node level and realize node feature reasoning. Based on the correlation and temporal coupling characteristics of fraud features between nodes, we can mine the collaborative anomalies and transmission patterns of fraud features between nodes to achieve edge feature reasoning. Tracing the temporal markers and evolutionary trajectories of fraud features, identifying temporal abrupt changes and continuous evolution anomalies of fraud features, and realizing temporal feature processing; By cross-validating the inference results from node feature inference, edge feature inference, and temporal feature processing, a multi-dimensional inference result is obtained.

[0009] In some specific embodiments, based on multi-dimensional reasoning results, the core risk characteristics of three types of hidden fraud are extracted respectively: the core risk characteristics of evasive fraud focus on semantic ambiguity, behavioral evasion, and temporal inconsistency; the core risk characteristics of false cooperation fraud focus on contradictory statements, false behavior, and breach of promise; and the core risk characteristics of collusive fraud focus on entity relevance, behavioral coordination, and feature synchronization.

[0010] In some specific embodiments, the core risk characteristics of the three types of hidden fraud are quantified, and the quantitative score of each type of hidden fraud is calculated by combining the importance weight of each core risk characteristic. In the comprehensive risk feature matrix, each element value corresponds to the quantitative score of the collection target and the corresponding core risk feature at the corresponding time node. The last column of the matrix adds the summary score of single-type hidden fraud risk and the time series cumulative risk score.

[0011] In some specific embodiments, the comprehensive risk feature matrix is ​​input into the pre-trained risk judgment model. Through feature weight matching and multi-dimensional feature fusion calculation, the three types of hidden fraud are judged independently. At the same time, the time-series cumulative risk score is combined for cross-validation to complete the multi-dimensional judgment of hidden fraud of the collection target.

[0012] A debt collection fraud detection system based on time-series graphs includes: The input unit is used to collect and preprocess data related to the entire collection cycle of the target collection session to form the raw dataset. The basic graph unit is used to extract skeleton attributes from the original dataset for building a time series graph, and to construct a basic graph with the debt collection target user as the core based on the skeleton attributes; The temporal graph unit is used to mine fine-grained attributes about the skeleton attributes based on the original dataset and extract exclusive fraud features in the collection scenario. The exclusive fraud features are then fused into the base graph to form a feature-enriched conversational temporal semantic graph. The feature matrix unit is used to perform multi-dimensional reasoning based on the conversation temporal semantic graph, extract the core risk features of three types of hidden fraud: evasive fraud, false cooperation fraud, and collusion fraud, and quantify the core risk features to form a comprehensive risk feature matrix. The output unit is used to input the comprehensive risk feature matrix into the pre-trained risk judgment model for implicit fraud judgment, and output the judgment result including implicit fraud type, fraud probability, and comprehensive risk level.

[0013] Beneficial Effects: This application proposes a method and system for identifying latent fraud in debt collection based on temporal graphs. By constructing a conversational temporal semantic foundation graph centered on the target user of the debt collection, it not only deeply integrates latent fraud features with the basic features of debt collection interactions, but also performs multi-dimensional and in-depth mining and quantitative representation of latent fraud features. This achieves a unified graph-based and temporal representation of the entire debt collection cycle interaction data, significantly reducing the false negative rate of latent fraud identification, improving the accuracy and comprehensiveness of debt collection fraud identification, providing financial institutions with accurate and quantitative risk basis for formulating differentiated debt collection strategies, effectively improving the efficiency of intelligent debt collection, and reducing the risk of bad debts.

[0014] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the method for identifying hidden fraud in debt collection as described in this application; Figure 2 This is a schematic diagram illustrating the principle of the debt collection hidden fraud identification method of this application; Figure 3 This is a schematic diagram of the generation process of the temporal semantic graph in this application; Figure 4 This is a schematic diagram illustrating the construction process of the feature matrix of this application; Figure 5 This is a schematic diagram of the debt collection hidden fraud identification system module of this application.

[0017] Figure labels: 1-Input unit; 2-Basic graph unit; 3-Time series graph unit; 4-Feature matrix unit; 5-Output unit. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0019] This application proposes a time-series graph-based method for identifying hidden fraud in debt collection, addressing the core problem that existing intelligent debt collection technologies cannot effectively identify hidden fraud. A flowchart illustrating the hidden fraud identification process is attached. Figure 1 As shown in the attached diagram, the principle is as follows. Figure 2 As shown, the specific solution is as follows: A method for identifying hidden fraud in debt collection based on time-series graphs, comprising: 101. Collect and preprocess data related to the entire collection cycle of the target debt collection session to form the raw dataset; 102. Extract skeleton attributes from the original dataset to build a time series graph, and construct a basic graph with the target users of the collection based on the skeleton attributes; 103. Based on the original dataset, fine-grained attributes about skeleton attributes are mined and exclusive fraud features in the debt collection scenario are extracted. The exclusive fraud features are then fused into the basic graph to form a feature-enriched conversational temporal semantic graph. 104. Based on the conversational temporal semantic graph, multi-dimensional reasoning is carried out to extract the core risk characteristics of three types of hidden fraud: evasive fraud, false cooperation fraud, and collusion fraud. A comprehensive risk characteristic matrix is ​​formed based on the quantification of the core risk characteristics. 105. Input the comprehensive risk feature matrix into the pre-trained risk judgment model to determine implicit fraud, and output the judgment result including implicit fraud type, fraud probability, and comprehensive risk level.

[0020] Step 101 is the foundational data acquisition stage of the entire method for identifying hidden fraud in debt collection. The collected data on the entire debt collection cycle covers all time dimensions of the debt collection business, from the initial contact to the termination of the collection process. All session interaction data are included in the collection scope. The data types include multi-dimensional basic data that support the identification of debt collection fraud, ensuring the integrity and continuity of the data and avoiding deviations in subsequent feature extraction and fraud determination due to missing data.

[0021] The data preprocessing stage performs a series of standardized operations on the collected raw, unprocessed data, including removing useless or missing data generated during the collection process, standardizing the format, and de-identifying the data. Finally, the processed multi-dimensional data is integrated into a structured raw dataset to ensure the dataset's standardization, consistency, and usability, laying a data foundation for the orderly implementation of subsequent steps.

[0022] Step 102 involves extracting skeleton attributes from the original dataset to build a time series graph, and constructing a basic graph centered on the target users for debt collection based on the skeleton attributes.

[0023] Skeleton attributes are the fundamental attributes that support the construction of time-series graphs and represent the core logic of collection interactions. The extraction process is not a simple attribute screening of the original dataset, but rather the mining of key attributes from the original data across the entire lifecycle and all dimensions that reflect core elements such as target users, collection interaction behaviors, and time-series relationships. The selected skeleton attributes possess strong correlation and core characteristics, and can fully support the subsequent construction of nodes, edges, and time-series dimensions in the graph.

[0024] This application abandons the existing representation methods that lack a clear data association structure, and innovatively constructs a conversational temporal semantic foundation graph with the target user as the sole core. The target user is used as the root node of the graph. Around this root node, based on extracted skeleton attributes, node sets, edge sets, and temporal dimension layers are built, forming a foundation graph with logical node associations, strong edge association representations, and temporal traceability capabilities, rather than a simple set of attributes or a data association table. This foundation graph achieves a structured, graph-based, and temporally unified representation of the entire collection cycle's related data, transforming previously scattered and unrelated collection data into a logical, temporally ordered, and strongly correlated graph structure. It can intuitively and accurately reflect the interaction logic and temporal relationships throughout the entire collection cycle, providing a core carrier for subsequent fusion of specific fraud features and multi-dimensional reasoning.

[0025] Step 103 involves mining fine-grained attributes related to skeleton attributes based on the original dataset and extracting specific fraud features for the debt collection scenario. These specific fraud features are then fused into the basic graph to form a feature-enriched conversational temporal semantic graph. This step is the core step in achieving accurate capture of latent fraud features in this application, directly addressing the core pain point of existing technologies that lack targeted fraud features and cannot identify latent fraud.

[0026] Fine-grained attributes are derived from the skeleton attributes extracted in step 102 and then deeply mined from the original dataset. This process involves disassembling and mining more detailed and representative fine-grained attributes from the underlying dimensions of the skeleton attributes. This mining process is not simply a breakdown of the skeleton attributes, but rather a deep analysis of the skeleton attributes from the perspectives of behavioral and semantic features of implicit fraud, taking into account the interactive characteristics of the debt collection scenario. This process uncovers fine-grained attributes that reflect subtle behavioral and semantic changes in the debt collection target. These fine-grained attributes are crucial for capturing implicit fraud characteristics and represent attribute dimensions that have never been addressed or explored in existing technologies.

[0027] From the mined fine-grained attributes, and combining the behavioral patterns and characteristics of three types of covert fraud in debt collection scenarios—evasiveness, feigned cooperation, and collusion—targeted, exclusive fraud features are extracted. These exclusive fraud features are specifically adapted for identifying covert fraud in debt collection scenarios, differing from commonly used fraud features in existing technologies. They possess strong scenario adaptability, accurately characterize the potential features of the three types of covert fraud, and the extracted exclusive fraud features naturally possess quantifiable and graph-mapping attributes, laying the foundation for subsequent integration into the basic graph.

[0028] The feature fusion and enrichment graph formation involves selectively integrating the extracted, specific fraud features into the base graph constructed in step 102. This is not a simple feature overlay, but rather a deep fusion of the specific fraud features with the nodes, edges, and temporal dimensions of the base graph. This allows the base graph to represent fraud features, ultimately forming a feature-enriched conversational temporal semantic graph. This enriched graph retains the full-cycle temporal association logic of the base graph while adding specific feature dimensions for latent fraud. It achieves an organic combination of basic collection interaction features and specific latent fraud features, allowing previously undetectable latent fraud features to be intuitively represented and traced through the graph.

[0029] Step 104 is a key step in achieving in-depth mining and quantitative characterization of implicit fraud features in this application, breaking through the technical limitations of existing technologies that rely on single-dimensional feature analysis and cannot effectively mine and quantify implicit fraud features.

[0030] Multi-dimensional reasoning relies on feature-enriched conversational temporal semantic graphs for inference analysis. This reasoning differs from existing single-dimensional feature analysis; instead, it involves interconnected reasoning based on the graph's nodes, edges, and temporal sequence. Leveraging the graph's strong correlations and temporal advantages, it mines latent fraud features from multiple dimensions, accurately capturing the clustering, synergy, and temporal evolution of latent fraud features. These patterns cannot be obtained through single-dimensional analysis using existing technologies.

[0031] The extraction of core risk features is based on multi-dimensional reasoning results, specifically targeting three types of hidden fraud: deflection fraud, feigned cooperation fraud, and collusion fraud. The extraction process combines the essential characteristics of these three types of hidden fraud, selecting the most representative risk features that best reflect the core attributes of each type of hidden fraud from the fraud features obtained through multi-dimensional reasoning. This achieves the goal of extracting core features from a massive amount of fraud features, making the characterization of hidden fraud more accurate and focused.

[0032] The extracted core risk features are quantified and transformed into calculable and comparable quantitative indicators. A comprehensive risk feature matrix is ​​then constructed based on these indicators. This matrix is ​​not simply a feature quantification table, but rather a matrix with time-series, feature, and quantitative scoring dimensions, combining the full-cycle time-series characteristics of the collection target. This achieves a systematic, quantitative, and time-series representation of the core risk features of three types of hidden fraud, transforming previously abstract and unquantifiable hidden fraud features into concrete quantitative indicators. The matrix provides accurate and quantifiable feature inputs for subsequent model judgment.

[0033] Step 105 involves inputting the comprehensive risk feature matrix into the pre-trained risk assessment model to determine implicit fraud, and outputting a assessment result including the type of implicit fraud, fraud probability, and comprehensive risk level. This step is the final step in achieving accurate implicit fraud assessment in this application, overcoming the technical limitations of existing single assessment models that can only output simple assessment results.

[0034] The comprehensive risk feature matrix formed in step 104 is used as input. This input is not a single feature value or feature vector, but a matrix of data that includes time-series dimensions, core features of multiple types of fraud, and quantitative scores. It can provide the model with comprehensive, accurate, and quantitative feature information, enabling the model to judge hidden fraud from multiple dimensions. This solves the problem of low judgment accuracy caused by the single input feature dimension of the existing technology model.

[0035] A risk assessment model is employed for identification, specifically tailored to the identification of three types of covert fraud. Unlike existing binary fraud identification models, this model achieves accurate classification and identification of these three types of covert fraud, rather than a simple fraud or non-fraud determination. The model output is not a single conclusion, but a multi-dimensional result including the type of covert fraud, the probability of fraud, and the overall risk level. This accurately informs the debt collection agency of the specific type of covert fraud the target is facing, the probability of committing that type of fraud, and the overall fraud risk level. This provides the debt collection agency with precise and detailed risk information to develop differentiated and targeted collection strategies.

[0036] This application presents a complete technical solution for identifying hidden fraud in debt collection based on time-series graphs. It encompasses full-cycle, multi-dimensional data collection, graph-based representation, and then the fusion of specific fraud features and multi-dimensional reasoning quantification to ultimately achieve accurate judgment. Each step incorporates targeted technological innovations, overcoming several technical bottlenecks in existing intelligent debt collection fraud identification. From data collection, data representation, feature extraction, feature analysis, to result judgment, it solves the core problem that existing technologies cannot identify hidden fraud in debt collection, achieving accurate and efficient identification of three types of hidden fraud.

[0037] Full-cycle collection session association data forms the original data foundation for the entire method of identifying hidden fraud in collection. In some specific embodiments, full-cycle collection session association data includes semantic interaction data, non-semantic behavior data, time-series baseline data, entity association data, and performance record data. These five types of data complement each other and provide comprehensive coverage, together constituting the full-cycle collection session association data of the collection target, ensuring that the original dataset can fully reflect the interaction status, relationships, performance status, and time-series characteristics of the collection target.

[0038] Semantic interaction data refers to all semantic-related information of interactions between collection agents and target users throughout the entire collection cycle. It mainly includes the text-to-text transcripts of the voice conversations between the two parties, the original transcripts of the text conversations, the core themes of each round of conversations, semantic expression tendencies, and emotional characteristics.

[0039] Non-semantic behavioral data is behavioral information that does not contain conversational semantics but can reflect the user's debt collection interaction status. It mainly includes the user's answering, hanging up, and rejecting status in the debt collection conversation, the actual answering duration of each conversation, the number of times the user actively calls back the debt collector, and the switching of terminal devices used during the interaction.

[0040] Time-series baseline data is the basic information for the time dimension of all collection-related data. It mainly includes the initiation time and end time of each collection session, the time interval between different sessions for the same collection target, the start and end times of the entire collection cycle, and the collection stage time nodes corresponding to each session.

[0041] Entity-related data refers to information about various entities and their relationships related to the target user in debt collection. It mainly includes the target user's basic identity information, information about their relatives, colleagues, guarantors and other related persons, other credit accounts under the user's name, information about related companies, information about debt collection agencies and personnel, and the entity information to which the user's interactive terminal belongs.

[0042] Performance record data refers to the performance-related records of the target user during the credit lifecycle and collection cycle. It mainly includes the user's historical repayment records, repayment amount, overdue duration, number of overdue payments, the matching of the user's promised repayment time with the actual repayment behavior during the collection process, and supporting information related to the user's current performance ability.

[0043] In some specific embodiments, the acquisition of skeleton attributes includes: extracting semantic class attributes representing the core semantic logic of the session from semantic interaction data, extracting behavioral class attributes representing user collection interaction behavior from non-semantic behavior data, extracting time-series class attributes representing session time correlation features from time-series benchmark data, extracting entity class attributes representing subject correlation relationships from entity association data, and extracting performance class attributes representing user performance behavior status from performance record data; performing fraud identification correlation verification on the extracted attributes, eliminating redundant attributes, and obtaining skeleton attributes.

[0044] Obtaining skeleton attributes involves two steps: attribute extraction and correlation verification.

[0045] Semantic attributes extracted from semantic interaction data represent the core semantic logic of debt collection conversations. Specifically, they can extract the user's expressed willingness to repay, proposed repayment plans, core reasons for delaying repayment, and attitude towards collection requests during the interaction between the debt collector and the user. This information directly reflects the core intent of the user's conversation and is an important foundation for identifying deceptive and feigned cooperation fraud.

[0046] Behavioral attributes extracted from non-semantic behavioral data are used to characterize various behavioral features of users during debt collection interactions. They do not rely on the semantic content of the conversation, but only reflect the user's level of cooperation and abnormal tendencies at the behavioral level. Specifically, they can extract the duration of debt collection calls, the frequency of hanging up, the number of times the user proactively calls back, the time distribution of debt collection calls, and whether the user repeatedly rejects calls. These behavioral attributes can initially indicate the user's tendency to be uncooperative, providing a behavioral reference for subsequent fraud identification.

[0047] The time-series attributes extracted from the time-series baseline data characterize the temporal correlation features between each collection session, providing a temporal benchmark for subsequent construction of time-series graphs and the implementation of time-series fraud inference. Specifically, these attributes can extract the initiation time, end time, time interval between consecutive sessions, the total number and duration distribution of collection sessions for a single user throughout the entire cycle, and the distribution of sessions across different time periods. These time-series attributes can reflect temporal anomalies in user interactions, aiding in the discovery of temporal features of hidden fraud.

[0048] Entity-type attributes extracted from entity association data are used to characterize the relationships between various entities related to the target user in debt collection, providing a foundation for subsequent investigation of hidden fraud such as collusive fraud. Specifically, relationships can be extracted between the target user and related contacts, the user and the bank where the account is opened, the user and the guarantee institution, and the user and other overdue users. These entity-type attributes can help identify the potential risk of users colluding with others to commit fraud.

[0049] Performance attributes extracted from performance record data characterize a user's historical and current performance status, serving as a crucial basis for determining whether a user has fraudulent tendencies. Specifically, these attributes can include historical repayment amounts, whether repayments are overdue, the number of overdue days, the number of overdue payments, any history of malicious debt evasion, and whether any repayment commitments have been made but not fulfilled. These performance attributes directly reflect a user's creditworthiness, providing credit-level support for fraud identification.

[0050] All extracted attributes undergo fraud identification correlation verification, and redundant attributes that are irrelevant to or duplicate those for implicit fraud identification are removed. Common correlation analysis methods, such as Pearson correlation coefficient method and mutual information method, can be used to calculate the correlation degree between various attributes and historical implicit fraud cases. A reasonable correlation degree threshold is set, and attributes with a correlation degree lower than the threshold are judged as redundant attributes and removed. The various attributes that are ultimately retained together constitute the skeleton attributes.

[0051] The core of building the basic graph is to establish the collection target as the root node, and all nodes and relationships revolve around this root node to ensure that the graph can focus on the collection target's full-cycle collection interaction and avoid deviating from the core identification object.

[0052] The node set forms the basic skeleton of the graph, carrying various core entities and information related to the identification of hidden fraud in debt collection. Its composition strictly corresponds to the previously extracted skeleton attributes, achieving a precise mapping between attributes and graph nodes. The specific definitions, functions, and implementation methods of each node are as follows: User subject nodes are the core associated nodes of the graph, carrying the basic core information of the target user and serving as the hub for all relationships, used to connect other types of nodes; Associated entity nodes correspond to the entity class attributes in the skeleton attributes, carrying various associated subject information related to the debt collection target, realizing a graph-based representation of subject relationships; Interaction terminal nodes record terminal-related information used during debt collection interactions, helping to reflect the consistency and anomalies of user interaction behavior; Session content nodes are constructed based on the semantic class attributes in the skeleton attributes, transforming the extracted core semantic logic of the session into graph nodes, realizing a structured presentation of session semantic information, and providing support for subsequent mining of semantic-level fraud features; Performance behavior nodes are constructed based on the performance class attributes in the skeleton attributes, carrying information related to the user's performance behavior status, intuitively presenting the user's credit performance status, and providing credit-level node support for fraud identification.

[0053] Edge sets are the core of connecting nodes and constructing logical relationships between them. Based on the core logical relationships represented by various skeleton attributes, edge sets ensure that they can accurately reflect the inherent connections between different nodes and realize the linkage of node information. At the same time, in order to accurately represent the differences in the strength of the relationships between nodes, each edge needs to be assigned a quantitative correlation attribute. The setting of this quantitative correlation is based on the actual interaction rules of the debt collection scenario, combined with industry experience and historical data to set reasonable quantitative standards.

[0054] The temporal dimension layer is the core of the basic graph's temporal features. It uniformly binds all nodes and edges in the graph with a temporal stamp attribute extracted from the temporal class attribute in the skeleton attributes. This temporal stamp strictly corresponds to the temporal association features of the collection session, ensuring that each node and each edge has a clear time identifier. Through the temporal stamp attribute, it is possible to clearly trace the information generation time corresponding to each node, the occurrence time of the association relationship corresponding to each edge, and the evolutionary trajectory of different nodes and edges in the temporal dimension. This achieves temporal association and traceability of graph nodes and edges, allowing the basic graph to not only have structured logical associations but also clear temporal evolution features. Ultimately, it forms a basic semantic graph of session temporality that combines temporal features and logical association features, laying a solid structured foundation for subsequent fusion of specific fraud features and multi-dimensional temporal reasoning.

[0055] In some specific embodiments, the process of acquiring exclusive fraud features includes: preprocessing fine-grained attributes; based on the behavioral and semantic features of three types of implicit fraud in debt collection scenarios—evasive fraud, feigned cooperation fraud, and collusive fraud—feature clustering is performed on the preprocessed fine-grained attributes to extract the core fine-grained attributes corresponding to each type of implicit fraud as exclusive fraud features; all exclusive fraud features possess quantifiable, temporally sequential, and graph-mapped feature attributes. The entire process of acquiring exclusive fraud features achieves accurate classification and extraction of fraud features while ensuring that the features can adapt to the subsequent graph fusion and fraud identification process, solving the core problems of fraud feature generalization, inability to adapt to temporally sequential graphs, and difficulty in distinguishing implicit fraud types in existing technologies. Fine-grained attributes are more detailed and specific derived attributes mined from the skeleton attributes of the original dataset. These attributes may initially contain missing values, outliers, inconsistent units, and noise information unrelated to fraud detection. Directly using them for feature extraction would severely impact the accuracy and effectiveness of specific fraud features; therefore, preprocessing is essential. Based on the data characteristics of the debt collection scenario, targeted preprocessing methods are adopted.

[0056] By leveraging the inherent differences in behavioral and semantic features among three types of covert fraud—evasive fraud, feigned cooperation fraud, and collusive fraud—the preprocessed fine-grained attributes are grouped and clustered according to their correlation with each type of fraud. The most representative core fine-grained attributes for each type of fraud are then selected as exclusive fraud features for that type of fraud, enabling precise localization and classification of fraud features.

[0057] The behavioral and semantic characteristics of the three types of hidden fraud are clearly distinguishable and serve as the core basis for clustering: The core characteristics of evasive fraud are reflected in the semantic repeated evasion and avoidance of repayment responsibility, and in the behavior of not explicitly refusing but not cooperating with collection efforts. For example, in the conversational semantics, there are frequent expressions of evasion such as "I don't have money right now" and "wait a little longer", and in the behavior, they frequently hang up collection calls but do not block them or explicitly refuse to communicate.

[0058] The core characteristics of false cooperation fraud are that it semantically promises to repay the loan and shows a positive attitude of cooperation, but in practice it fails to fulfill its promises and engages in false cooperation. For example, in the conversational semantics, it clearly promises to "repay tomorrow" or "settle the debt within this week", but fails to fulfill the promise after the due date, and repeatedly makes promises without taking any actual action when it is urged to pay again.

[0059] The core characteristics of collusive fraud are that multiple debt collection targets have highly consistent semantic expressions, coordinated behaviors, and entity connections. For example, multiple users may have completely identical conversational semantics, highly consistent reasons for repayment, and simultaneously refuse to answer collection calls and make the same false repayment promises. They may also share contact numbers, have related bank accounts, and other entity connections.

[0060] Feature clustering can employ clustering algorithms suitable for the characteristics of debt collection scenarios, such as K-means clustering and hierarchical clustering. First, based on historical cases of hidden fraud in the debt collection industry, fine-grained attribute samples corresponding to various types of fraud are labeled, and the number of clusters is determined to be 3. Then, all preprocessed fine-grained attributes are input into the clustering model. The model will divide the fine-grained attributes into three clusters based on the similarity between the attributes and the features of various fraud types, with each cluster corresponding to one type of hidden fraud.

[0061] For each cluster, fine-grained attributes are analyzed for correlation and ranked for importance. Attributes with low correlation to the specific type of fraud are removed, and the most relevant and representative core fine-grained attributes are extracted as the exclusive fraud features for that type of fraud. These exclusive fraud features possess three key attributes: quantifiability, temporal sequence mapping, and graph mapping. These are crucial prerequisites for integrating exclusive fraud features into the basic graph and supporting subsequent multi-dimensional temporal reasoning and fraud detection.

[0062] In some specific embodiments, feature mapping associations are established with the node set and edge set of the basic graph based on the type and representation dimension of the specific fraud features; quantitative attribute fields of specific fraud features are added to the nodes, and fraud feature correlation and temporal coupling feature attributes are added to the edges between nodes; triggering temporal identifiers and temporal trajectories of specific fraud features are added at the temporal dimension layer to achieve deep fusion and graph representation of specific fraud features and the original skeleton attributes of the basic graph. The generation process of the temporal semantic graph is attached. Figure 3 As shown.

[0063] Based on the level of fraud risk represented by the exclusive fraud features, the graph carriers to which they belong are divided to ensure that the features are compatible with the graph structure. The exclusive fraud features that represent the fraud attributes of the node itself are directly mapped to the corresponding graph node set. For example, features that represent the fraud tendency of the user subject, such as "cumulative frequency of evasive statements" and "repayment commitment fulfillment rate", represent the fraud characteristics of the conversation content, such as "number of false commitments in this conversation", and represent the fraud characteristics of the performance behavior, such as "number of times the commitment was not fulfilled", are mapped to the user subject node, the conversation content node, and the performance behavior node, respectively.

[0064] After completing the feature mapping association, corresponding exclusive fraud feature-related attributes are added to the graph node set and edge set respectively, so as to realize the quantitative representation of fraud features in the graph structure.

[0065] For the specific fraud features mapped to the node set, quantitative attribute fields for the specific fraud features are added to the corresponding nodes. The quantitative values ​​of the specific fraud features are directly used as field values ​​and embedded into the original attribute system of the nodes. For example, quantitative attribute fields such as "cumulative frequency of evasive statements", "repayment commitment fulfillment rate" and "number of false cooperation" are added to user subject nodes. Quantitative attribute fields such as "percentage of evasive semantics in this session" and "number of false commitments in this session" are added to conversation content nodes. Quantitative attribute fields such as "frequency of false commitments after overdue" and "commitment fulfillment failure rate" are added to fulfillment behavior nodes, so that each node can intuitively reflect its own fraud risk characteristics.

[0066] For the exclusive fraud features mapped to the edge set, two types of attributes, fraud feature correlation and temporal coupling features, are added to the edges between corresponding nodes. The fraud feature correlation is used to quantify the strength of the association between nodes based on exclusive fraud features. For example, the edge between the user subject and the associated entity node is added with attributes such as "similarity of fraud features of associated entities" and "proportion of frequency of collaborative shirking". The higher the value, the tighter the association of fraud features between nodes. The temporal coupling feature attribute is used to characterize the synergy of fraud features between nodes in the time dimension. For example, attributes such as "fraud feature trigger time synchronization rate" and "time difference of fraud feature triggering sequence" are added to accurately reflect whether the fraud features of two nodes are triggered at the same time or at similar times. The addition of edge attributes makes the edges of the graph not only have the original logical correlation of collection interaction, but also reflect the fraud collaboration risk between nodes, realizing the quantitative representation of fraud features between nodes.

[0067] To achieve deep integration of exclusive fraud features with the temporal features of the basic graph, trigger timing identifiers and temporal trajectories of exclusive fraud features are added to the temporal dimension layer of the graph. This gives the fraud features clear temporal attributes and traceable evolutionary patterns, thus unifying them with the temporal system of the basic graph.

[0068] Triggering time sequence identifiers are added fraud feature attributes, including the node's exclusive fraud feature quantification field, the edge's fraud feature correlation degree, and time-series coupling feature attributes. They are bound to a unified timestamp set based on time-series attributes. This timestamp accurately corresponds to the specific collection session time node for the first trigger of the fraud feature and each subsequent trigger. The time-series trajectory, on the other hand, is based on each triggering time sequence identifier. It sorts out and records the numerical changes and evolution of the exclusive fraud feature over time, adding it as a new attribute in the time-series dimension layer in the form of a time-series sequence. For example, the numerical changes of the user's "frequency of making excuses" in each collection session are formed into a time-series trajectory with a numerical sequence corresponding to consecutive time nodes. The numerical changes of the "fraud feature triggering time synchronization rate" between nodes in different time periods are also formed into a time-series trajectory.

[0069] In some specific embodiments, multi-dimensional reasoning includes: performing anomaly determination and correlation mining on the exclusive fraud features added to each node, identifying fraud feature aggregation anomalies at the node level, and realizing node feature reasoning. The implementation involves first performing independent anomaly determination on each exclusive fraud feature of a node, combining business rules of the debt collection industry, historical implicit fraud case data, and the feature distribution of normal debt collection interactions to set a reasonable anomaly threshold range for each exclusive fraud feature. If the feature quantification value exceeds this range, it is determined to be a single feature anomaly. Then, correlation mining is performed on multiple exclusive fraud features within the same node to analyze the correlation between single feature anomalies. If two or more single feature anomalies appear in the same debt collection time sequence node, and the anomaly features belong to the same type of implicit fraud representation dimension, it is determined that the node has a fraud feature aggregation anomaly. For example, if a conversation content node simultaneously shows two feature anomalies: "the proportion of evasive semantics exceeds the threshold" and "the number of false promises exceeds the threshold," it is determined that the node has a feature aggregation anomaly of false cooperation or evasive fraud. By establishing an initial location of fraud risk at the node level, we can capture the concentrated manifestation patterns of fraud characteristics at a single node level, providing node-level anomaly clues for subsequent multi-dimensional reasoning.

[0070] Based on the correlation and temporal coupling characteristics of fraud features between nodes, this study mines the collaborative anomalies and transmission patterns of fraud features between nodes to achieve edge feature inference. The method involves first determining collaborative anomalies based on the correlation of fraud features, setting a collaborative anomaly threshold for the correlation of fraud features for different types of edges. If the quantified value of the correlation of fraud features between nodes exceeds this threshold, it indicates that the fraud features of the two nodes are highly similar, possessing the characteristic basis for collaborative fraud. Then, the effectiveness of collaborative anomalies is verified by combining temporal coupling characteristics. If not only does the correlation of fraud features between nodes exceed the threshold, but temporal coupling characteristics such as "fraud feature trigger time synchronization rate" and "fraud feature sequential trigger time difference" also exist, then the study will determine the validity of the collaborative anomalies. The combined features also conform to the temporal patterns of coordinated fraud. For example, if the core fraud features of both are triggered at the same collection session timeline and the time difference is within one hour, it is determined that there is a genuine coordinated fraud feature anomaly between the nodes. Simultaneously, based on the changes in edge features under consecutive timeline nodes, the transmission patterns of fraud features between nodes are mined, analyzing the transmission process and temporal patterns of fraud features from one node to its associated nodes. For example, if an associated entity node first exhibits a collusive fraud feature anomaly, and subsequently, in adjacent timeline nodes, the associated user entity node also exhibits the same collusive fraud feature anomaly, it is determined that the fraud feature has been transmitted from the associated entity node to the user entity node. The core function of this step is to overcome the limitations of single-node reasoning, capture the fraud risk of multi-node linkage, and especially provide core association-level reasoning basis for identifying collusive fraud.

[0071] This system traces the triggering time sequence of fraud features and their evolution time sequence to identify temporal abrupt changes and continuous evolution anomalies, thus enabling time-series feature processing. The implementation involves first comparing the changes in feature quantification values ​​at adjacent or consecutive time-series nodes based on the triggering time sequence of fraud features to identify temporal abrupt changes. If the quantification value of a fraud feature at a certain time-series node shows an irregular, significant jump or sudden change compared to the previous time-series node, and the magnitude of the change exceeds the feature fluctuation range of normal collection interactions, it is determined to be a temporal abrupt change anomaly. Next, based on the evolution time sequence of fraud features, the system analyzes the overall trend of feature changes at consecutive time-series nodes to identify continuous evolution anomalies. If a fraud feature exhibits regular abnormal changes across multiple consecutive time-series nodes and evolves towards a deeper level of fraud, it is determined to be a continuous evolution anomaly. For example, if a user's "number of false promises" continuously increases with the number of collection sessions, while the "commitment fulfillment rate" continuously decreases, forming an inverse continuous evolution, it is determined to be a continuous evolution anomaly. The core function of this step is to compensate for the shortcomings of static feature analysis, capture the dynamic evolution of fraud features from a time dimension, identify anomalies in the time dimension such as the escalation, persistence, and suddenness of fraudulent behavior, and provide temporal basis for distinguishing the behavioral stages of hidden fraud.

[0072] The inference results from the three dimensions are matched for correlation. If the fraud risk of the same collection target shows mutually corroborating characteristics in the three dimensions, that is, if there is an abnormal clustering of fraud features at the node level, a corresponding abnormal collaboration of fraud at the edge level, and a corresponding abnormal temporal evolution of fraud features at the time sequence level, then it is determined that the collection target has a real hidden fraud risk, and the abnormal patterns of the three dimensions are integrated to form a complete inference conclusion. If only a single dimension shows an abnormal fraud feature, and there are no corresponding abnormal features in the other dimensions to corroborate it, then it is determined that the abnormality is an isolated misjudgment and is excluded. If two dimensions show abnormalities and corroborate each other, and there is no obvious abnormality in the third dimension, then a comprehensive judgment is made in combination with the actual situation of the collection scenario, reasonable abnormal clues are retained and doubtful points are marked.

[0073] In some specific embodiments, based on multi-dimensional reasoning results, the core risk characteristics of three types of hidden fraud are extracted respectively: the core risk characteristics of evasive fraud focus on semantic ambiguity, behavioral evasion, and temporal inconsistency; the core risk characteristics of false cooperation fraud focus on contradictory statements, false behavior, and breach of promise; and the core risk characteristics of collusive fraud focus on entity relevance, behavioral coordination, and feature synchronization.

[0074] The core risk characteristics of deceptive fraud focus on semantic ambiguity, behavioral deception, and temporal inconsistency. These three sub-features correspond to anomalous results in the semantic features, behavioral features, and temporal evolution features of nodes in multi-dimensional reasoning, accurately reflecting the core essence of this type of fraud. Semantic ambiguity is extracted from semantically related anomalies such as conversation content nodes in node feature reasoning. It refers to users deliberately obscuring core information such as repayment time, repayment ability, and repayment plan in their debt collection interactions, without explicitly refusing to repay or providing specific and feasible repayment solutions. This is characterized by a conversation deception semantic proportion far exceeding the anomaly threshold and frequent use of vague expressions such as "temporarily out of money." Behavioral deception is extracted from behaviorally related anomalies such as user subject nodes and interaction terminal nodes in node feature reasoning. It refers to users exhibiting a clear tendency to deflect responsibility in debt collection interactions. While there is no explicit refusal or blocking, they deliberately avoid effective communication through various behaviors. This is characterized by frequent... The text discusses a cluster of abnormal behavioral characteristics specific to fraud, including frequent and brief hang-ups of collection calls, multiple unexplained interruptions in collection communication, and extremely short durations of effective collection communication. It also mentions temporal inconsistency, extracted from the abnormal evolution of fraud features in temporal feature processing. This refers to a significant contradiction between the user's semantic and behavioral characteristics over time, the absence of stable collection interaction patterns, and the unreliable and disordered semantic and behavioral characteristics that do not improve as the collection process progresses. This is characterized by a user vaguely stating "I'll be raising money to repay soon" at a certain time point, but showing no repayment preparation behavior at subsequent adjacent time points, with a sudden increase in the frequency of such evasive statements. The temporal evolution trajectory of semantics and behavior exhibits significant temporal abrupt changes or continuous evolutionary anomalies.

[0075] The core risk characteristics of fraudulent cooperation focus on contradictory statements, false behavior, and breach of promise. These three sub-features correspond to the abnormal results of node semantic features, node behavioral features, and performance behavior nodes and temporal evolution features in multi-dimensional reasoning, accurately reflecting the core essence of this type of fraud: feigned cooperation and repeated breach of promise. Contradictory statements are extracted from the semantic anomalies of conversation content nodes in the node feature reasoning. This refers to obvious self-contradictions in the semantic statements of users in debt collection interactions. Core information such as repayment plans, repayment ability, and sources of funds expressed in previous and subsequent statements cannot corroborate each other or are even completely contradictory. This is manifested in a previous conversation stating "I will repay next week when I receive my salary," while in a later conversation stating "I have no fixed income and currently have no ability to repay." The characteristics related to false statements in the conversation content nodes show clustering anomalies. False behavior is extracted from the abnormal behavior of user subject nodes in the node feature reasoning and the temporal evolution anomalies of behavior processed by temporal features. This refers to the user making a semantic statement of active cooperation but failing to take any actual behavior that matches the repayment promise, showing no behavioral representation of repayment preparation. This is manifested in the user promising "transferring repayment the next day," but then failing to do so the next day. The user failed to complete the transfer and did not proactively explain the situation. During subsequent collection efforts, the user continued to make positive promises but took no actual action. The temporal evolution of behavioral fraud characteristics showed a continuous anomaly. The promise default was extracted from the anomalies in the performance behavior nodes and the temporal evolution anomalies in the performance characteristics, derived from the node features. It is also the core characteristic of false cooperation fraud, indicating that the user repeatedly made clear repayment promises but never fulfilled them. The promise fulfillment rate was far below the anomaly threshold, and the default behavior showed a continuous trend as the collection process progressed. This was characterized by the user making more than ten repayment promises throughout the entire collection cycle but never actually fulfilling them. The specific fraud characteristics of the performance behavior nodes, such as the "promise fulfillment rate" and "number of times the promise was not fulfilled," far exceeded the anomaly threshold. The promise default-related characteristics in the temporal trajectory showed a continuously rising evolution anomaly.

[0076] The core risk characteristics of collusive fraud focus on entity correlation, behavioral coordination, and feature synchronization. These three sub-features correspond to anomalous results in the coupling of edge feature reasoning and temporal characteristics in multi-dimensional reasoning, aligning with the core essence of this type of fraud—"multi-entity linkage and collaborative fraud"—and distinguishing it from the previous two types of single-person fraud. Entity correlation, extracted from the anomalies in entity correlation between nodes in edge feature reasoning, is a fundamental characteristic of collusive fraud. It refers to anomalies in entity correlation between multiple debt collection target users related to debt collection interactions and repayment behaviors, rather than normal social relationships. This manifests as multiple users sharing the same interactive terminal, the same bank card number, or the same contact address, or having direct abnormal guarantees or fund transfers. The entity-like edge attributes between the user node and other related user nodes show obvious anomalous correlations. Behavioral coordination, extracted from the anomalies in behavioral coordination between nodes in edge feature reasoning, refers to multiple users with entity correlation exhibiting high coordination in debt collection interactions, with consistent anomalous patterns in their behavioral characteristics. This manifests as multiple related users simultaneously refusing debt collection calls, simultaneously interrupting conversations without cause, and engaging in similar behavior within the same specific time period. The phone calls were answered and the effective communication duration was consistent. The correlation of the fraudulent features between the user's main node and related user nodes far exceeded the threshold of collaborative anomaly. Feature synchronization is extracted from the temporal coupling feature anomaly of edge feature reasoning and the collaborative temporal evolution anomaly of temporal feature processing. It is the core identification feature of collusive fraud. It refers to the high synchronization of the exclusive fraud features of multiple users with entity correlation and behavioral collaboration in the time dimension. The triggering and evolution of fraud features have a consistent time pattern. It is characterized by multiple related users first appearing collusive fraud features at the same time node, the quantification value of fraud features jumping synchronously in adjacent time nodes, and the temporal evolution trajectory of fraud features being almost completely consistent. The "fraud feature triggering time synchronization rate" in the temporal coupling feature of the edge between nodes is close to 100%, which shows obvious temporal collaboration anomaly.

[0077] In some specific embodiments, the core risk characteristics of each of the three types of hidden fraud are quantified, and a quantitative score for each type of hidden fraud is calculated by combining the importance weights of each core risk characteristic. In the comprehensive risk characteristic matrix, each element value corresponds to the quantitative score of the collection target for the corresponding core risk characteristic at the corresponding time node. The last column of the matrix adds the summary score of a single type of hidden fraud risk and the cumulative risk score of the time series. The process of generating the feature matrix is ​​attached. Figure 4 As shown.

[0078] For each type of hidden fraud, three core risk characteristics are identified, and quantitative rules tailored to the collection scenario are developed, based on abnormal data from multi-dimensional reasoning results. Each core risk characteristic is converted into a standardized quantitative value of 0-100 points. While maintaining a unified dimension, the numerical value directly corresponds to the degree of abnormality of the risk characteristic; the higher the value, the more significant the fraud risk of that characteristic. For example, the "semantic ambiguity" of evasive fraud uses the proportion of conversational evasive semantics and the frequency of ambiguous expressions as core quantitative indicators, converting the proportion / frequency exceeding the abnormal threshold into a standardized score; the "promise default" of false cooperation fraud uses the promise fulfillment rate and the cumulative number of non-fulfillment instances as core indicators; the lower the fulfillment rate and the more non-fulfillment instances, the higher the standardized score; the "feature synchronization" of collusive fraud uses the synchronization rate of fraud feature triggering time and the similarity of feature quantitative values ​​as core indicators; the higher the synchronization rate and similarity, the higher the standardized score.

[0079] By combining the identification and core contribution of each core risk characteristic to the corresponding category of hidden fraud, differentiated importance weights are assigned to the three core risk characteristics of each type of fraud. The weight values ​​are determined through analysis of historical fraud cases in the debt collection industry and the analytic hierarchy process (AHP), with the sum of the weights of the three characteristics under a single type of fraud being 1, and core characteristics assigned higher weights. For example, in false cooperation fraud, "promise default" is the core identification characteristic, with a weight of 0.5, while "contradictory statements" and "false behavior" are auxiliary characteristics, each with a weight of 0.25; in evasive fraud, "evasive behavior" has a weight of 0.5, while "semantic ambiguity" and "temporal inconsistency" each have a weight of 0.25; in collusive fraud, "characteristic synchronicity" has a weight of 0.5, while "entity relevance" and "behavioral coordination" each have a weight of 0.25. Finally, the weighted calculation is completed by multiplying the standardized quantitative value of the core risk feature by the corresponding weight. The weighted results of the three features under a single type of fraud are summed to obtain the quantitative score of this type of hidden fraud. The score range is 0-100 points. The higher the score, the higher the risk of this type of hidden fraud for the collection target.

[0080] The vertical row dimension of the comprehensive risk feature matrix represents all collection session timeline nodes throughout the entire collection target cycle, arranged sequentially according to collection time to ensure that the evolution trajectory of risk features can be traced over time. The horizontal column dimension first sequentially displays all the core risk features of the three types of hidden fraud: semantic ambiguity, behavioral evasion, and temporal inconsistency in the form of evasiveness; contradictory statements, false behavior, and breach of promise in the form of false cooperation; and entity correlation, behavioral coordination, and feature synchronization in the form of collusion. Finally, two core scoring items are added to the last column of the matrix: the summary score of single-type hidden fraud risk and the cumulative risk score over time.

[0081] Each basic element value in the matrix corresponds to the standardized quantitative score of a core risk feature of the collection target at a specific time point. Based on the feature quantification results and time points mentioned above, the matrix is ​​filled in one by one to intuitively present the degree of abnormality of different core risk features at each time point.

[0082] The final column, the summary score for single-category hidden fraud risk, is a quantitative score calculated by weighting the three core risk characteristics of this type of fraud at the corresponding time point. Corresponding to the quantitative score for single-category fraud mentioned earlier, it can directly reflect the immediate risk level of a certain type of hidden fraud at this time point. The time-series cumulative risk score is the time-series cumulative value of the summary score for single-category hidden fraud from the first collection time point to the current point. To better reflect the risk evolution law of actual collection scenarios, a time decay coefficient can also be introduced, assigning lower weight to the long-term summary score and higher weight to the recent score before accumulating. This score can intuitively reflect the degree of accumulation and evolution trend of a certain type of hidden fraud risk as the collection process progresses.

[0083] The entire comprehensive risk feature matrix integrates scattered quantitative data in a structured way. It can clearly show the core risk characteristics and degree of individual risks of various hidden frauds of collection targets at a single time point, and can also trace the dynamic evolution of risks through the cumulative risk score over time.

[0084] In some specific embodiments, the comprehensive risk feature matrix is ​​input into the pre-trained risk judgment model. Through feature weight matching and multi-dimensional feature fusion calculation, the three types of hidden fraud are judged independently. At the same time, the time-series cumulative risk score is combined for cross-validation to complete the multi-dimensional judgment of hidden fraud of the collection target.

[0085] The model evaluation process mainly includes two key steps: feature weight matching and multi-dimensional feature fusion calculation. Feature weight matching involves accurately matching the model's preset feature weights with various features in the comprehensive risk feature matrix. Combining historical hidden fraud case data from the debt collection industry with model training experience, differentiated model-level weights are assigned to different core risk features and single-category summary scores in the matrix. This complements the importance weights of the core risk features mentioned earlier, emphasizing the weight proportion of various core fraud identification features and ensuring that the model prioritizes the most identifiable risk features.

[0086] Multi-dimensional feature fusion calculation integrates the scattered time-series features, single-node features, and aggregated scores in the matrix through model algorithms. This breaks down the barriers between time-series nodes and feature dimensions, fusing the risk feature quantification values ​​and single-category aggregated scores of each time-series node into a unified multi-dimensional judgment feature vector. This vector is then used to independently judge three types of hidden fraud: evasiveness, feigned cooperation, and collusion. The model sets independent judgment branches and thresholds for each type of hidden fraud. Based on the fused feature vector, the final judgment score for each type of fraud is calculated. If the judgment score is higher than the corresponding threshold, it is preliminarily determined that the debt collection target has this type of hidden fraud; if it is lower than the threshold, it is preliminarily determined that there is no such hidden fraud. The judgment processes for the three types of fraud are independent and do not interfere with each other, ensuring accurate identification of single or multiple types of hidden fraud.

[0087] The core of cross-validation relies on the time-series cumulative risk score for each type of hidden fraud in the comprehensive risk feature matrix. This score directly reflects the accumulation and evolution trend of a certain type of hidden fraud risk as the collection process progresses, and is a key basis for judging whether the model's independent judgment results are reliable. Compared to the single-category summary score of a single time-series node, the time-series cumulative risk score can avoid problems such as misjudgment due to single anomalies.

[0088] For each type of hidden fraud initially identified by the model, a two-way verification is performed using the corresponding category's time-series cumulative risk score. If the model initially identifies a type of hidden fraud in the collection target, and the time-series cumulative risk score for this type of fraud continues to rise and exceeds the preset cumulative risk threshold, it indicates that the fraud risk has formed a continuous evolution trend and is not an accidental anomaly. In this case, the identification result for this type of fraud is confirmed as valid. If the model initially identifies a type of hidden fraud, but the time-series cumulative risk score is extremely low, shows no obvious upward trend, or even shows a downward trend, it indicates that the identification result may originate from an accidental anomaly at a single time-series node. The identification result needs to be corrected to "not identified as existing" and marked as a suspicious risk for subsequent tracking and verification. If the model initially identifies no type of hidden fraud, but the time-series cumulative risk score for this type of fraud continues to rise and approaches the cumulative risk threshold, it indicates that the fraud risk is gradually accumulating. It needs to be marked as a potential risk and alerted for subsequent key tracking.

[0089] After completing cross-validation of all categories, the independent judgment results of the model are integrated with the cross-validation correction opinions to form the final multi-dimensional judgment result. This result clarifies whether the collection target is involved in hidden fraud, what type or types of hidden fraud exist, the risk level of each type of fraud, and whether there are any potential suspicious risks. Ultimately, this achieves an accurate, comprehensive, and reliable multi-dimensional judgment of hidden fraud by the collection target, providing a direct basis for the subsequent development of targeted collection strategies and fraud handling plans.

[0090] For example, the multi-class risk assessment model uses the XGBoost multi-class model. This model is adapted to the structured and quantitative feature attributes of the comprehensive risk feature matrix, has excellent nonlinear feature fusion capabilities and anti-overfitting characteristics, can accurately capture the complex correlation patterns of hidden fraud features in debt collection, and supports independent judgment of multi-class features. The loss function adopts multi-class cross-entropy, and the model evaluation index is based on the F1-score and recall rate of the three types of fraud.

[0091] The model training process consists of five steps. First, a training sample set is constructed based on historical full-cycle debt collection data from the debt collection industry. The comprehensive risk feature matrix corresponding to each sample is extracted and quantified as input features. Four-class labels are applied: "no fraud," "evasive fraud," "false cooperation fraud," and "collusion fraud." Sample cleaning, missing value imputation, and feature standardization are then completed. Second, the processed sample set is divided into training, validation, and test sets in a 7:2:1 ratio. Next, the XGBoost multi-class model is initialized and the initial parameters are loaded. Iterative training is performed using mini-batch gradient descent, with the multi-class F1-score of the validation set as the core monitoring indicator. Then, through 5... Cross-validation combined with grid search was used to optimize hyperparameters, adjusting parameters such as tree depth, learning rate, and sampling rate. An early stopping mechanism was used to stop training when the validation set metrics showed no improvement for 20 consecutive rounds, thus suppressing model overfitting. Finally, the model's generalization ability was validated on the test set, focusing on evaluating the recall and precision of three types of hidden fraud. After successful validation, the thresholds for the fraud judgment probabilities output by the model were calibrated according to the actual business rules of debt collection. Independent judgment probability thresholds were set for the three types of hidden fraud to meet the trade-off between missed and false positives in debt collection scenarios. At the same time, the feature weights of the time-series cumulative risk score were integrated into the model inference process, completing the final optimization of the model and its deployment.

[0092] A time-series graph-based system for identifying hidden fraud in debt collection is shown in the attached diagram. Figure 5 As shown, the system includes: Input unit 1 is used to collect and preprocess the collection session data related to the entire collection cycle of the collection target to form the original dataset; Basic graph unit 2 is used to extract skeleton attributes from the original dataset for building time series graphs, and to construct a basic graph with the target users of debt collection as the core based on the skeleton attributes; Temporal graph unit 3 is used to mine fine-grained attributes about skeleton attributes based on the original dataset and extract exclusive fraud features in the collection scenario. The exclusive fraud features are then fused into the basic graph to form a feature-enriched conversational temporal semantic graph. Feature matrix unit 4 is used to perform multi-dimensional reasoning based on the conversation temporal semantic graph, extract the core risk features of three types of hidden fraud: evasive fraud, false cooperation fraud, and collusion fraud, and form a comprehensive risk feature matrix based on the quantification of the core risk features; Output unit 5 is used to input the comprehensive risk feature matrix into the pre-trained risk judgment model for implicit fraud judgment, and output the judgment result including implicit fraud type, fraud probability, and comprehensive risk level.

[0093] Those skilled in the art will understand that the components of this application described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage system for execution by the computing system. Alternatively, they can be fabricated as separate integrated circuit components, or multiple components or steps can be fabricated as a single integrated circuit component. Thus, this application is not limited to any particular combination of hardware and software.

[0094] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

[0095] The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A method for identifying latent fraud in debt collection based on time-series graphs, characterized in that, include: Collect and preprocess data related to the entire collection cycle of the target debt collection session to form the raw dataset; Extract skeleton attributes for building a time series graph from the original dataset, and construct a basic graph centered on the target users of debt collection based on the skeleton attributes; Based on the original dataset, fine-grained attributes of the skeleton attributes are mined and specific fraud features for debt collection scenarios are extracted. These specific fraud features are then fused into the base graph to form a feature-enriched conversational temporal semantic graph. Based on the aforementioned conversational temporal semantic graph, multi-dimensional reasoning is performed to extract the core risk characteristics of three types of hidden fraud: deflection fraud, false cooperation fraud, and collusion fraud. A comprehensive risk characteristic matrix is ​​then formed based on these core risk characteristics. The comprehensive risk feature matrix is ​​input into the pre-trained risk judgment model to determine implicit fraud, and the judgment result includes implicit fraud type, fraud probability, and comprehensive risk level.

2. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, The full-cycle collection session associated data includes semantic interaction data, non-semantic behavior data, time-series benchmark data, entity association data, and performance record data.

3. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, The acquisition of skeleton attributes includes: extracting semantic attributes representing the core semantic logic of the session from semantic interaction data; extracting behavioral attributes representing user collection interaction behavior from non-semantic behavior data; extracting time-series attributes representing session time correlation features from time-series benchmark data; extracting entity attributes representing subject correlation relationships from entity correlation data; and extracting performance-related attributes representing user performance behavior status from performance record data. Fraud identification correlation verification is performed on the extracted attributes, redundant attributes are removed, and skeleton attributes are obtained.

4. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, The process of obtaining the specific fraud features includes: preprocessing the fine-grained attributes; based on the behavioral and semantic features of three types of implicit fraud in the debt collection scenario—evasive fraud, feigned cooperation fraud, and collusion fraud—feature clustering is performed on the preprocessed fine-grained attributes, and the core fine-grained attributes corresponding to each type of implicit fraud are extracted as specific fraud features; all specific fraud features have quantifiable, temporal, and graph-based mapping attributes.

5. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, Based on the type and representation dimension of the specific fraud features, feature mapping associations are established with the node set and edge set of the basic graph, respectively; Add quantitative attribute fields for exclusive fraud features to nodes, and add fraud feature correlation and temporal coupling feature attributes to the edges between nodes; add triggering time sequence identifiers and time sequence trajectories for exclusive fraud features to the temporal dimension layer, so as to achieve deep integration and graph representation of exclusive fraud features and the original skeleton attributes of the basic graph.

6. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, Multidimensional reasoning includes: Anomaly detection and correlation mining are performed on the exclusive fraud features added to each node to identify fraud feature aggregation anomalies at the node level and realize node feature reasoning. Based on the correlation and temporal coupling characteristics of fraud features between nodes, we can mine the collaborative anomalies and transmission patterns of fraud features between nodes to achieve edge feature reasoning. Tracing the temporal markers and evolutionary trajectories of fraud features, identifying temporal abrupt changes and continuous evolution anomalies of fraud features, and realizing temporal feature processing; By cross-validating the inference results from node feature inference, edge feature inference, and temporal feature processing, a multi-dimensional inference result is obtained.

7. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, Based on the results of multi-dimensional reasoning, the core risk characteristics of three types of hidden fraud are extracted respectively: the core risk characteristics of evasive fraud focus on semantic ambiguity, behavioral evasion and temporal inconsistency; The core risk characteristics of fraudulent cooperation focus on contradictory statements, false behavior, and breach of promises; The core risk characteristics of collusive fraud focus on the correlation between entities, the coordination of behaviors, and the synchronization of characteristics.

8. The method for identifying hidden fraud in debt collection according to claim 7, characterized in that, The core risk characteristics of the three types of hidden fraud are quantified, and the quantitative score of each type of hidden fraud is calculated by combining the importance weight of each core risk characteristic. In the comprehensive risk feature matrix, each element value corresponds to the quantitative score of the collection target and the corresponding core risk feature at the corresponding time node. The last column of the matrix adds the summary score of single-type hidden fraud risk and the time series cumulative risk score.

9. The method for identifying hidden fraud in debt collection according to claim 1, characterized in that, The comprehensive risk feature matrix is ​​input into the pre-trained risk judgment model. Through feature weight matching and multi-dimensional feature fusion calculation, the three types of hidden fraud are judged independently. At the same time, the time-series cumulative risk score is combined for cross-validation to complete the multi-dimensional judgment of hidden fraud of the collection target.

10. A debt collection covert fraud identification system based on time-series graphs, characterized in that, include: The input unit is used to collect and preprocess data related to the entire collection cycle of the target collection session to form the raw dataset. The basic graph unit is used to extract skeleton attributes from the original dataset for building a time series graph, and to construct a basic graph with the debt collection target user as the core based on the skeleton attributes; The temporal graph unit is used to mine fine-grained attributes about the skeleton attributes based on the original dataset and extract exclusive fraud features in the collection scenario. The exclusive fraud features are then fused into the base graph to form a feature-enriched conversational temporal semantic graph. The feature matrix unit is used to perform multi-dimensional reasoning based on the conversation temporal semantic graph, extract the core risk features of three types of hidden fraud: evasive fraud, false cooperation fraud, and collusion fraud, and quantify the core risk features to form a comprehensive risk feature matrix. The output unit is used to input the comprehensive risk feature matrix into the pre-trained risk judgment model for implicit fraud judgment and output the judgment result including implicit fraud type, fraud probability, and comprehensive risk level.