Method for intelligently generating exhaustion report of financial unfavorable assets

Through multimodal data fusion and context-aware modeling, combined with evolutionary computing and adaptive evaluation networks, the problem that traditional due diligence reports are difficult to track and predict financial non-performing assets in real time, and efficient and accurate risk identification and dynamic monitoring are achieved.

CN120198232AActive Publication Date: 2025-06-24SHANGHAI BAICHANG TECH GRP CO LTD

Patent Information

Application Number
CN202510682535.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Traditional due diligence reports are difficult to dynamically track real-time changes in debtor operating conditions, collateral value, legal environment and market environment, and lack effective prediction and early warning capabilities for risk evolution trends. At the same time, the iteration of existing AI systems depends on the perception of backend developers, and the iteration is weak and it is difficult to deal with unstructured and non-standardized data.

Method used

Through deep fusion of multimodal data and context-aware modeling, structured, semi-structured and unstructured data are collected to establish a context-aware model of financial non-performing assets. Evolutionary calculations are used to mine risk factors, build an adaptive risk assessment network, and use interpretability models to generate and report output.

Benefits of technology

Real-time dynamic monitoring and risk prediction of financial non-performing assets has been achieved, the sensitivity and accuracy of risk identification has been improved, the ability to continuously learn and evolve, and adapt to the dynamically changing market and risk environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198232A_ABST
    Figure CN120198232A_ABST
Patent Text Reader

Abstract

The invention provides a method for intelligently generating a complete dispatch report for financial unfavorable assets, relates to the field of management systems, and aims at deeply fusing multi-source heterogeneous data such as legal instruments and financial statements and forming a comprehensive context sensing model for target assets through multi-modal feature extraction, semantic alignment and knowledge graph construction technologies; a potential and non-dominant risk factor combination deeply coupled with asset characteristics is automatically mined by applying a genetic algorithm and other evolutionary calculation methods, and a self-adaptive risk assessment network is constructed to dynamically and quantitatively assess and predict the comprehensive risk level; carrying out attribution analysis on a risk assessment result by adopting an interpretable artificial intelligence model, and clearly revealing a key influence path and core data evidence; and according to a report logic framework and a narrative template which can be dynamically adjusted by a user, automatically outputting a financial non-performing asset full-duty survey report including deep analysis, risk early warning, diversified disposal suggestions and compliance review key points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of management systems, and specifically to a method for intelligently generating due diligence reports on non-performing financial assets. Background Art

[0002] Traditional due diligence reports are mostly static descriptions and analyses of the historical conditions and current points in time of assets, making it difficult to dynamically track the real-time changes of key factors such as the debtor's operating conditions, the value of collateral, the legal environment, and the market environment. Moreover, they lack the ability to effectively predict and early warn of the risk evolution trend. When the current publicly available practical methods combine AI or related intelligent systems for processing, the system iteration depends on the cognition and work of back-end developers, and the iteration timeliness is much weaker than real-time iteration and monitoring. At the same time, due to the fact that a single AI focuses on the processing of single or formatted files for data processing, it is difficult to process unstructured and non-standardized data. Summary of the Invention

[0003] Problems to be Solved

[0004] In view of the existing deficiencies, the present invention provides a method for intelligently generating due diligence reports on non-performing financial assets, which solves the problems of the prior art.

[0005] Technical Solution

[0006] To achieve the above objectives, the present invention is realized through the following solutions: A method for intelligently generating due diligence reports on non-performing financial assets, the method comprising:

[0007] Sp1. Deep fusion of multi-modal data and context-aware modeling:

[0008] Collect structured, semi-structured, and unstructured data related to the target non-performing financial assets to obtain heterogeneous data;

[0009] Map the heterogeneous data to a unified semantic space, identify and establish entities, attributes related to non-performing financial assets, and complex association relationships therebetween, and form a context-aware model of the non-performing financial assets;

[0010] Sp2. Risk factor mining and adaptive evaluation based on evolutionary computing:

[0011] Sp2.1. Initialize the risk factor candidate set: Based on the context-aware model, use the genetic algorithm evolutionary computing method to perform iterative optimization in a preset risk dimension space, and mine potential and non-explicit risk factor combinations that are coupled with the characteristics of the non-performing financial assets;

[0012] Sp2.2. Build a risk assessment network: The risk assessment network dynamically adjusts the weights and interaction relationships of various risk factors based on newly input data and historical assessment feedback, and combines an ensemble learning strategy to quantitatively evaluate the comprehensive risk level of non-performing financial assets, and outputs a risk profile and a confidence interval.

[0013] Sp3. Generate insights and report outputs based on an interpretable model:

[0014] Sp3.1. Use an interpretable intelligent model to perform attribution analysis on the output results of the risk assessment network, and identify the key influence paths and core data evidence that lead to specific risk assessment conclusions.

[0015] Sp3.2. According to a preset and dynamically adjustable report logic framework and narrative template by users, use natural language generation to output the key information in the context-aware model to generate a due diligence report on non-performing financial assets.

[0016] Preferably, in the multi-modal data deep fusion and context-aware modeling, it further includes: using a pre-trained language model enhanced by an attention mechanism to parse the long-distance dependencies and complex clause structures in legal texts, and encoding the key elements of the legal texts and their confidence scores into nodes and weighted edges in a knowledge graph.

[0017] Preferably, in the risk factor mining based on evolutionary computation, it further includes:

[0018] Quantify its interpretability by calculating the minimum description length of the decision rules generated by the risk factors;

[0019] Measure its prediction accuracy by evaluating the F1 score of the risk factor combination on historical default events in a backtest dataset;

[0020] Evaluate its correlation by calculating the cosine similarity of the vector space between the risk factor combination and the patterns in a preset typical non-performing asset risk pattern library verified by domain experts.

[0021] Preferably, the risk assessment network adopts a reinforcement learning mechanism, and the reinforcement learning mechanism further includes: constructing a deep deterministic policy gradient agent, where the state space of the agent represents the risk factors and assessment results of the current asset, and the action space corresponds to the adjustment strategy for the weights or activation functions of specific risk factors in the risk assessment network; quantifying the user's confirmation, correction, or rejection behavior of the risk assessment as a scalar reward signal to guide the policy learning of the agent.

[0022] Preferably, the insight generation of the interpretability model further includes: generating a visualization map of the risk conduction path, and the generating of the visualization map of the risk conduction path further includes: on the knowledge graph of the context-aware model, using a graph neural network inference algorithm based on attention weighting to identify and quantify the influence intensity and conduction probability between different risk entities and risk factors, and forming a conduction path.

[0023] Preferably, in the report output, the natural language generation further includes: using a conditional text generation model, taking the preset portrait of the target audience as the conditional input, combining the structured semantic representation extracted from the insights of the interpretability model, dynamically selecting narrative templates, adjusting the professional level of terms, and controlling the detailed level of the argument to generate a customized report text that meets the needs of specific audiences.

[0024] Preferably, the method further includes a continuous learning and model iteration module, and the continuous learning and model iteration module further includes:

[0025] A data and concept drift detection unit, which is used to monitor the statistical characteristics of the input data stream and the changes in the user feedback pattern in real time, and automatically trigger the model update process when significant drift is detected;

[0026] A differential knowledge graph update unit, which merges the newly added or changed entities, relationships and their confidence levels into the existing knowledge graph in an incremental manner, and applies the knowledge learned from historical data to the iterative upgrade of the new model using transfer learning.

[0027] Preferably, a system for intelligently generating a due diligence report on non-performing financial assets includes a processor and a memory coupled to the processor. Computer program instructions are stored in the memory, and the computer program instructions are executed by the processor. The system further includes:

[0028] A multi-modal data deep fusion and context-aware modeling engine:

[0029] Collect data related to the target non-performing financial assets from heterogeneous data sources;

[0030] Build a context-aware model for non-performing financial assets;

[0031] Sp2. An evolutionary computing-based risk factor mining and adaptive evaluation engine:

[0032] Perform iterative optimization in a preset risk dimension space;

[0033] Build and run a risk assessment network;

[0034] Sp3. An insight generation and report output engine based on an interpretability model:

[0035] Use interpretable intelligent models to conduct attribution analysis on risk assessment results;

[0036] Output due diligence report on financial non-performing assets.

[0037] Preferably, the pre-trained language model processing unit in the multimodal data deep fusion and context-aware modeling engine further includes:

[0038] Complex long sentence segmentation and dependency parsing module, used to accurately identify the master-slave structure and limiting conditions in contract terms;

[0039] The semantic role labeling module enhanced by domain knowledge is used to label specific roles of participants and core legal behaviors in financial transactions.

[0040] Preferably, the risk factor mining and adaptive assessment engine based on evolutionary computing further includes:

[0041] The feedback-driven model parameter adjustment module further comprises:

[0042] User feedback real-time capture and structured processing interface, used to receive and analyze user annotation information on risk factors and assessment results;

[0043] An incremental model training and version control unit, wherein the incremental model training and version control unit has an evolutionary algorithm embedded therein, and the incremental model training and version control unit supports online fine-tuning of a population initialization strategy of the evolutionary algorithm and local connection weights of the evaluation network;

[0044] Offline batch retraining scheduler, used to trigger global optimization of the entire model system after accumulating enough new data or feedback.

[0045] Beneficial Effects

[0046] The present invention provides a method for intelligently generating a due diligence report on financial non-performing assets. It has the following beneficial effects:

[0047] The present invention deeply integrates multimodal data, constructs a panoramic contextual knowledge graph, and uses evolutionary computing to mine hidden risk factors. It combines adaptive advanced evaluation networks with explainable AI insights, and can quickly extract accurate insights from heterogeneous data, reveal complex correlations and potential risks, and improve the efficiency of information collection and preliminary analysis to a new level. It can not only accurately quantify known risks, but also actively discover and evaluate non-explicit and combined risks, greatly improving the sensitivity and accuracy of risk identification.

[0048] The present invention has the capabilities of continuous learning and self - evolution, can adapt to the dynamically changing market and risk environment, realizes the digital precipitation and intelligent iteration of organizational knowledge and experience, and ensures the effectiveness and value of long - term application. Brief Description of the Drawings

[0049] Figure 1 It is a system composition diagram of the present invention;

[0050] Figure 2 It is a system flow chart of the present invention. Detailed Embodiment

[0051] Next, the solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Specific Embodiment:

[0053] As Figures 1 to 2 shown, a method for intelligently generating a due - diligence report on non - performing financial assets, the method includes obtaining data related to non - performing financial assets and generating a report structure; the method further includes the following steps:

[0054] Sp1. Deep fusion of multi - modal data and context - aware modeling lay the foundation for a comprehensive understanding of the target non - performing financial assets, involving the whole process from data collection, pre - processing to multi - modal feature extraction, alignment, and finally forming a context - aware model.

[0055] Data Collection and Pre - processing:

[0056] Data sources and types: including but not limited to: (1) Legal documents: including loan contracts, guarantee contracts, mortgage agreements, litigation / arbitration documents (complaints, judgments, mediation statements), bankruptcy and reorganization documents, etc., usually in PDF, Word or scanned image format; (2) Financial statements: balance sheets, income statements, cash flow statements, statements of changes in owners' equity and their notes of the debtor and related parties, usually in Excel, PDF or image format; (3) Market transaction data: secondary market transaction prices, trading volumes, relevant macroeconomic indicators, industry indices, interest rates, exchange rates, etc. of similar non-performing assets, usually structured time-series data; (4) Public opinion information: negative news, default records, litigation information, business anomalies, etc. about the debtor, guarantor, related parties, collateral on news portals, financial media, social platforms, industry forums, etc., usually unstructured text or semi-structured data; (5) Description of the physical state of assets: including appraisal reports, on-site inspection photos, videos, geographical location information, ownership certificates, etc. of collateral such as real estate and machinery and equipment, involving text, images, and geospatial data.

[0057] Preprocessing of data processing solutions: Legal documents and financial statements (text / image type):

[0058] OCR and layout analysis: For scanned or image-format documents, use a high-precision OCR engine (including Tesseract-OCR, optimized by combining with the deep learning model CRNN) for text recognition. Use layout analysis (including models based on Faster RCNN or Layout LM) to identify document structures, including titles, paragraphs, tables, seals, etc.

[0059] Table extraction: For tables in financial statements or contracts, use methods including OpenCV image processing combined with heuristic rules, or specialized table recognition models (including TableNet, Tab Transformer) to extract structured table data.

[0060] Text cleaning: Remove headers, footers, watermarks, irrelevant symbols, and garbled characters; perform traditional-simplified conversion and full-width and half-width unification; identify and process fill-in-the-blank items and handwritten supplementary parts in contracts.

[0061] Initial element extraction: Based on regular expressions and keyword libraries, initially extract core elements such as contract numbers, names of parties, amounts, dates, names of collateral, etc., as auxiliary information or verification basis for subsequent model processing.

[0062] Financial statements (structured data):

[0063] Subject standardization: Map financial subjects disclosed by different accounting standards and different enterprises to a unified standardized subject system (including based on XBRL classification standards or custom standard subject tables).

[0064] Data verification and cleaning: Check the reconciliation relationships in the reports (including Assets = Liabilities + Owner's Equity), identify and handle outliers (including using box plot method or Z-score method), and missing values (including filling with mean / median, regression imputation, or multiple imputation).

[0065] Calculation of financial indicators: Automatically calculate key financial indicators such as solvency ratios (including current ratio, quick ratio, debt-to-asset ratio), profitability ratios (including net profit margin, return on equity), and operating capacity ratios (including accounts receivable turnover, inventory turnover).

[0066] Market transaction data: API access and data parsing: Obtain data through API interfaces (including Bloomberg, Reuters, Wind) or web crawlers (subject to the robots.txt protocol). Parse data in formats such as JSON, XML, or CSV.

[0067] Data cleaning: Handle missing points (linear interpolation, spline interpolation) and abnormal fluctuations (moving average filtering, exponential smoothing) in time series data.

[0068] Time alignment and frequency conversion: Align data from different sources and with different frequencies to a unified time axis (including daily, weekly).

[0069] Public opinion information: Directed crawling and content extraction: Develop directed crawlers using frameworks such as Scrapy based on keywords such as debtors and related parties, and extract news text, release time, source, etc.

[0070] Data deduplication: Remove duplicate or highly similar public opinion information based on text similarity (including SimHash, MinHashLSH).

[0071] Initial judgment of sentiment tendency: Conduct a preliminary positive and negative sentiment scoring on public opinion text using a dictionary-based method (including HowNet sentiment dictionary) or a simple classification model (including Naive Bayes).

[0072] Description of the physical state of assets: Image data processing: Perform unified format conversion, size normalization, and image enhancement (including histogram equalization, denoising) on photos and videos.

[0073] Structuring of text descriptions: Extract key attributes from the text descriptions in the evaluation reports (including the area, location, and use of real estate; the model, purchase year, and depreciation status of equipment).

[0074] Multi-modal feature extraction and alignment: Map data from different sources and different modalities to a unified shared semantic space that can capture cross-modal semantic relationships.

[0075] Text feature extraction:

[0076] Model construction: Select a Transformer model pre-trained and fine-tuned for the financial or legal domain, including FinBERT (pre-trained on financial news and research reports), LawBERT / LegalBERT (pre-trained on legal documents and case precedents), or general models such as BERT, RoBERTa, ERNIE, etc. for domain-adaptive pre-training and downstream task fine-tuning on a specific non-performing asset-related corpus (including contracts, litigation documents, financial footnotes, market analysis reports, etc.). Downstream tasks may include: named entity recognition, relation extraction, text classification (including risk type judgment), and semantic similarity calculation. When fine-tuning, use the cross-entropy loss function to optimize the parameters.

[0077] Model application: Input the pre-processed text segments (including contract terms, financial footnote descriptions, news summaries), and the model outputs semantic vector representations (token-level embeddings, sentence / clause-level embeddings, document-level embeddings) of a fixed dimension (including 768 dimensions or 1024 dimensions). For long texts, segment processing followed by pooling (including mean / max pooling) or hierarchical Transformer (including HiBERT) can be used to obtain the overall representation.

[0078] Structured data feature extraction:

[0079] Processing solution: Numerical financial indicators, market data, etc., after cleaning and normalization / standardization (including Min-Max Scaling, Z-score Standardization), can be directly used as feature vectors. For categorical features (including industry classification, regional classification), one-hot encoding or mapping them to low-dimensional dense embedding vectors (Embedding-Layer, which can be obtained through end-to-end learning) can be used.

[0080] Image feature extraction:

[0081] Model Construction: Select a deep convolutional neural network (CNN) pre-trained on large image datasets (including ImageNet), such as ResNet-50 / 101, Efficient-Net (B0 - B7 series), or Vision Transformer (ViT). According to the characteristics of non-performing asset collateral images (including real estate appearance, equipment details, scanned copies of certificates), it can be fine-tuned on specific domain image data to better capture visual features related to asset evaluation.

[0082] Model Application: Input the preprocessed image, and the CNN / ViT model outputs a high-dimensional feature vector (the ResNet outputs 2048 dimensions, and the dimensions of Efficient-Net vary according to the series).

[0083] Multi-modal Alignment Model Construction:

[0084] Joint Embedding-based Model: Design a multi-input neural network architecture with the goal of mapping samples of different modalities but semantically related to neighboring positions in a shared semantic space. A certain clause description in a legal document and the corresponding collateral photo should have similar vector representations. During training, use a contrastive loss function or a triplet loss function. The contrastive loss aims to reduce the distance between positive sample pairs (including the vector of the "contract clause describing property A" and the vector of the "photo of property A"), and increase the distance between negative sample pairs. The triplet loss takes into account the anchor, positive sample, and negative sample simultaneously.

[0085] Co-training or Cross-modal Generation-based Method: Train a model to generate a pseudo-representation of image features from text descriptions and require it to be similar to the real image features; vice versa.

[0086] Training Data Construction: A large number of cross-modal alignment samples need to be constructed, including legal document fragments, associated financial data fragments, collateral description texts, collateral images, market news, and corresponding changes in company financial indicators.

[0087] Model Application: The original data from different modalities are finally converted into vector representations in the same high-dimensional semantic space through their respective feature extractors and alignment models. These vectors can be directly used for downstream tasks, including similarity calculation, clustering, risk assessment, etc.

[0088] Knowledge Graph Construction: Automatically identify entities, attributes, and their complex relationships related to non-performing financial assets from multi-modal data to form a structured knowledge network.

[0089] Entity Recognition Model Construction: A sequence labeling model is adopted, and the mainstream architecture is a pre-trained language model, using BERT + linear layer + conditional random field. The CRF layer can learn the constraint relationships between labels and improve the accuracy of entity boundary recognition.

[0090] Data Processing and Application: Define entity types in the field of non-performing assets, including debtor (DEBTOR), creditor (CREDITOR), guarantor (GUARANTOR), collateral (COLLATERAL), contract (CONTRACT), court (COURT), amount (AMOUNT), date (DATE), risk event (RISK_EVENT), etc.

[0091] Training Data: A large amount of labeled corpus (including BIO annotation format) is required. Active learning and semi-supervised learning can be used to reduce the annotation cost.

[0092] Model Application: Input text (including contracts, judgments), and output the recognized entities, their types, and their positions in the text. Identify Zhang San (DEBTOR), Li Si (CREDITOR), 1 million yuan (AMOUNT), and Wang Wu (GUARANTOR) from "Zhang San borrowed 1 million yuan from Li Si, and Wang Wu provided joint liability guarantee".

[0093] Relationship Extraction Model Construction: Based on the pipeline method: First perform NER, and then classify the relationships of the identified entities. The relationship classification model can be a BiLSTM based on the attention mechanism, a path-based graph convolutional network (Path-based GCN, using the dependency syntactic paths between entities), or a classifier based on a pre-trained language model (including adding a classification head to the [CLS] representation or entity pair representation).

[0094] Based on the joint extraction method (Joint-NER-and-RE): Design a model to complete entity recognition and relationship extraction tasks simultaneously, which can better utilize the dependency information between the two. A Transformer model based on multi-task learning.

[0095] Relationship Type Definition: Pre-define the relationship types between entities, including "borrows_from", "guarantees_for", "mortgages_to", "involved_in_lawsuit", etc.

[0096] Training Data: Triples that need to label entity pairs and their relationships are required.

[0097] Model Application: Input text containing identified entities and entity pairs, and output the type of relationship between them and its confidence level. (Zhang San, borrows_from, Li Si), (Wang Wu, guarantees_for, Zhang San).

[0098] Event Extraction (EE):

[0099] Model Construction: Identify specific events that occur in the text and their participants (arguments). The model usually includes two stages: event trigger word recognition and argument role annotation. Methods based on sequence labeling or machine reading comprehension can be adopted.

[0100] Data Processing and Application: Define event types (including "sign a contract", "initiate a lawsuit", "declare bankruptcy"), and the argument roles of each event type (including "signing party", "contract subject matter", "signing date" for the contract signing event).

[0101] Entity Alignment and Disambiguation Processing Solution: Solve the unification problem of the same entity with different data sources and different expressions (including "Company A" and "A Co., Ltd."), and the problem of different entities with the same name. Methods include those based on string similarity, entity attribute similarity, network structure similarity (including using link prediction in the knowledge graph), or entity linking methods based on pre-trained models.

[0102] Knowledge Storage and Reasoning: The graph database Neo4j is selected to store the constructed knowledge graph (entities as nodes, relationships as edges, and attributes as the characteristics of nodes / edges).

[0103] Inference Application: Use the graph query language CypherQL for complex queries. Conduct rule reasoning based on the graph using SWRL or link prediction and path discovery based on graph embedding.

[0104] Context-Aware Model Construction and Application Construction: With the constructed knowledge graph as the core framework, the multi-modal feature vectors extracted from each step in Sp1 (including text semantic vectors, image feature vectors, and structured data features) are used as the attributes of the corresponding entity nodes or relationship edges, or are associated with them. In addition to basic information, the debtor node is also associated with vectors of its financial statement analysis results, the latest public opinion sentiment scores, and key clause semantic vectors of relevant legal documents. This model provides a comprehensive, multi-dimensional, and structured view of non-performing assets and their related parties, related events, and potential risks.

[0105] Information Retrieval and Aggregation: Quickly retrieve all relevant information about a specific asset, regardless of its original modality and source. Reveal hidden association relationships, including common guarantee circles, indirect control relationships, and complex creditor-debtor chains.

[0106] Risk conduction analysis: Simulate the conduction path and potential impact scope of specific risk events (including the default of a core enterprise) in the knowledge graph network.

[0107] Feature engineering: Provide high-quality, context-rich input features (including graph embedding representations, path features, neighborhood aggregation features, etc.) for downstream risk assessment models.

[0108] Optimization options for extracting key elements from unstructured text data such as legal documents:

[0109] Application of pre-trained language models enhanced by attention mechanisms: Model selection and construction: Use the Reformer Transformer model for processing long-sequence text, and capture long-range dependencies in legal texts through sparse attention mechanisms (including a clause in a contract citing a definition dozens of pages earlier, or the comprehensive determination of multiple pieces of evidence in a judgment).

[0110] Analysis of complex clause structures: During the pre-training and fine-tuning processes, these models can better understand complex coordinate clauses, compound sentences, conditional clauses, restrictive clauses, etc. by learning a large amount of legal texts, so as to accurately identify the core rights and obligations, preconditions, excepted liabilities, etc.

[0111] Identification of implied warranty liabilities and contingent liabilities: Combine semantic understanding and knowledge graph reasoning. The model identifies typical sentence patterns such as "If Party B fails to repay the principal debt on schedule, then Party C agrees to assume joint and several liability for repayment", and combines the relationship between Party C and Party B in the knowledge graph (including parent company, actual controller) to infer the contingent liabilities of Party C. For more concealed clauses (including "Under specific market conditions, Party A has the right to require Party B to provide additional guarantees"), the model needs to have stronger context understanding and logical inference capabilities, and needs to combine specific sub-task models (including models based on reading comprehension to answer questions such as "Does Party C have warranty liability?") or rule-based post-processing.

[0112] Knowledge graph encoding of key elements and their confidence scores:

[0113] Confidence score calculation: When the model identifies entities, relationships, or extracts elements, it usually outputs a probability value or softmax score. This score can be used as the confidence after calibration (including Platt Scaling or Isotonic Regression).

[0114] Construction of weighted edges: The key elements of the extracted legal text (including "guarantee amount", "guarantee scope", and "lawsuit cause of action" in the contract) are used as node attributes or independent nodes in the knowledge graph. The weights of their associations with the core entities (including debtors and contracts) can be comprehensively determined based on the following factors:

[0115] Confidence score: The accuracy of extraction.

[0116] Calculate the importance of this element in the document through TF-IDF, Text Rank, or model attention weights. Preset the prior influence weights of different types of elements on risk assessment (including that "unlimited joint liability" has a higher weight than "general guarantee"). When the subsequent XAI model determines that this element contributes significantly to the risk assessment result, its weight can be increased in turn. These weighted edges can more accurately reflect the contribution degree of different information segments to the overall risk judgment in subsequent graph algorithms (including community discovery, centrality calculation, and risk conduction analysis) and risk assessment models.

[0117] Sp2. Mining of risk factors and adaptive assessment based on evolutionary computation: Automatically discover potential risk factor combinations related to non-performing asset risks and difficult to intuitively perceive from a high-dimensional and complex feature space, and construct a risk assessment model that can dynamically adapt to new situations.

[0118] Sp2.1. Initialization of the risk factor candidate set:

[0119] Detailed explanation of the evolutionary computation method - genetic algorithm: Iteratively optimize within a preset risk dimension space (including dimensions such as credit risk, market risk, operational risk, and legal risk, and each dimension contains numerous candidate atomic risk factors) to automatically mine potential and non-explicit risk factor combinations that are coupled with the characteristics of financial non-performing assets.

[0120] Sources of risk factors: Output of the context-aware model: Node attributes in the knowledge graph (including debtor financial ratios, public opinion scores), the presence or absence or weights of edges (including the existence of specific guarantee relationships, litigation relationships), and certain dimensions of the graph embedding vector. Multimodal features: Key dimensions of text semantic features, clustering results of image features. Expert-defined feature library: Known important risk points predefined by financial domain experts.

[0121] Details of GA construction: Chromosome encoding (individual representation): Each individual (chromosome) represents a candidate risk factor combination.

[0122] Randomly generate a group of individuals as the initial population. Domain knowledge can also be combined to include some known effective risk factors or combinations as part of the initial population (heuristic initialization) to accelerate convergence. The population size includes 50 - 200 individuals.

[0123] The selection operator selects excellent individuals according to their fitness values to enter the next generation.

[0124] The crossover operator simulates gene recombination in biological evolution and operates on the selected parental individuals with a certain crossover probability (0.6 - 0.9) to generate new offspring individuals.

[0125] The mutation operator randomly changes some gene positions of the offspring individuals with a certain mutation probability (0.01 - 0.1) to maintain population diversity and avoid falling into local optima.

[0126] Iteration stop conditions: reaching the preset maximum number of generations of evolution, the fitness function value not significantly improving for multiple consecutive generations, and finding a solution that meets specific conditions. The output of GA is a set (or an optimal) combination of risk factors, which will be used to construct the subsequent risk assessment network or directly form interpretable risk rules.

[0127] The fitness function of the evolutionary computing method further includes: the fitness function can balance the prediction accuracy, interpretability, simplicity, robustness, and domain relevance of the risk factor combination. The explanatory ability is quantified by calculating the minimum description length of the decision rules generated by the risk factors: the risk factor combination selected by GA is converted into a set of decision rules in the form of IF - THEN. For example, the causal risk factor combination is {Financial Indicator A < 0.5, Legal Risk B = True}, generating the rule "IF Financial Indicator A < 0.5 AND Legal Risk B = True THEN High Risk". The encoding length of the model itself (L(H)) and the encoding length of the data given the model D (L(D|H)). For decision rules, L(H) is quantified as the number of rules, the number of conditions in each rule, etc. L(D|H) is the encoding length required to misclassify samples under this rule. The goal is to minimize L(H)+L(D|H), where H represents the meaning of the encoding length.

[0128] The smaller the MDL value, the simpler the rules formed by the risk factor combination, the better it fits the data, and the stronger the interpretability. The fitness value can be set as the reciprocal of MDL or a certain constant minus MDL. The prediction accuracy of the risk factor combination for historical default events is measured by evaluating the F1 score on the backtest dataset:

[0129] Prepare a dataset containing historical non-performing asset cases. Each case includes the values of various risk factors for GA candidates (values at a certain point in time before the event) and the final true results (including whether there is a default, the degree of loss, the recovery rate level, etc.). The dataset needs to be divided into a training set, a validation set, and a test set. For each combination of risk factors generated by GA, use these factors as features to train a simple classification model (including logistic regression, support vector machine, decision tree) on the training set or directly build a scoring card model using these factors. Then predict historical default events on the validation set / test set.

[0130] The F1 score is the harmonic mean of precision and recall, which can balance the two well and is especially suitable for the imbalanced-class default prediction problem.

[0131] F1 score calculation formula:

[0132] ;

[0133] The F1 score is used to evaluate the accuracy of a classification model built based on a specific combination of risk factors in predicting non-performing events on the historical backtest dataset. It combines the precision and recall of the model.

[0134] F1: The F1 score value, ranging from 0 to 1, and a higher value indicates better prediction performance of the model.

[0135] Precision: Among the samples predicted as positive by the model, the proportion of samples that are actually positive. The calculation formula is:

[0136] ;

[0137] Recall: Among the samples that are actually positive, the proportion of samples successfully predicted as positive by the model. The calculation formula is:

[0138] ;

[0139] TP: True positive, the number of samples that are actually positive and predicted as positive by the model; FP: False positive (Type I error), the number of samples that are actually negative but predicted as positive by the model; FN: False negative (Type II error), the number of samples that are actually positive but predicted as negative by the model.

[0140] The higher the F1 score, the stronger the prediction ability of the risk factor combination. The correlation degree is evaluated by calculating the cosine similarity of the vector space between the risk factor combination and the patterns in a pre-set typical non-performing asset risk pattern library verified by domain experts:

[0141] Data processing is carried out by financial domain experts (including senior credit approvers and risk managers) to summarize typical risk patterns (including "over - expansion risk", "associated guarantee chain risk", "obsolete risk", etc.) based on experience or induction from historical cases. Each pattern can be described by a set of key risk factors and their typical manifestations. Both the risk patterns defined by each expert and the risk factor combinations generated by GA are represented as vectors. If the set of atomic risk factors is fixed, then each combination / pattern can be represented as a high - dimensional sparse vector, where the dimension corresponding to the existing factor is 1 or its weight, and 0 if it does not exist.

[0142] If a risk factor combination is highly similar to a typical risk pattern recognized by a certain (or certain) expert, it is considered to have good domain relevance and interpretability. The fitness function can consider the similarity to the most similar pattern, or the weighted average of the similarities to multiple patterns. The final fitness function is usually a weighted combination of the above indicators, and the weights are adjusted according to business requirements. The comprehensive fitness function formula of the genetic algorithm includes:

[0143] ;

[0144] It is used to evaluate the quality of each risk factor combination in the genetic algorithm. It has good interpretability, high prediction accuracy, and relevance to domain expert knowledge.

[0145] : The comprehensive fitness value of an individual (risk factor combination). The higher this value, the better the risk factor combination. , , : They are the weight coefficients of interpretability, prediction accuracy, and domain relevance respectively. The weights can be adjusted according to business requirements to focus on different optimization goals. If more importance is attached to prediction accuracy, then 's weight can be set higher. These weights usually take values from 0 to 1, and their sum is 1. : The value of the minimum description length. It quantifies the simplicity and interpretability of the decision rule generated by the current risk factor combination. The smaller the MDL value, the simpler the rule and the stronger the interpretability. Therefore, is used in the fitness function, so that the smaller the MDL value, the greater the contribution of this item. : The F1 - score. Here, it refers to an indicator of the prediction accuracy of the prediction model constructed based on the current risk factor combination, especially suitable for datasets with class imbalance (focusing on default prediction). It is the harmonic mean of precision and recall. The higher the F1 - score, the better the prediction accuracy. : Maximum cosine similarity. It measures the maximum vector space cosine similarity between the current risk factor combination and each pattern in the preset typical non-performing asset risk pattern library verified by domain experts. The higher this value, the more relevant the current risk factor combination is to the risk patterns recognized by experts, indicating better domain relevance.

[0146] Sp2.2. Build a risk assessment network: Based on the risk factor combinations mined by GA (or other more comprehensive feature sets), build a network model that can quantitatively evaluate the comprehensive risk level of financial non-performing assets and can be dynamically adjusted.

[0147] Model architecture selection and construction:

[0148] Multi-layer perceptron:

[0149] Input layer: Risk factors mined by GA (numerical, or vectors after encoding / embedding of categorical factors) and other important supplementary features (including macroeconomic indicators).

[0150] Hidden layer: One to multiple fully connected layers, and the number of neurons in each layer is determined according to the problem complexity and data volume (including 32, 64, 128, etc.). Activation functions such as Leaky ReLU and Tanh are selected. To prevent overfitting, a Dropout layer or L1 / L2 regularization can be added.

[0151] Output layer: Regression tasks (including predicting loss rate, recovery rate): A single neuron with a linear activation function, and the loss function is mean squared error or mean absolute error.

[0152] Classification tasks (including predicting risk level: low / medium / high, or default or not): The number of neurons is equal to the number of classes, the activation function is Softmax, and the loss function is cross-entropy loss.

[0153] Optimizer: Adam, RMSprop, SGD-with-momentum. Graph neural networks are suitable for risk assessment using knowledge graph information.

[0154] Context-aware models (knowledge graphs) in the form of the whole or subgraphs. Node features include entity attribute vectors and multi-modal feature vectors extracted in Sp1; edge features can include relationship types, weights, etc.

[0155] GNN layers: Adopt including GCN, GAT, Graph-SAGE. GCN updates node representations by aggregating neighbor node information. GAT introduces an attention mechanism to assign different learning weights to different neighbor nodes. Graph-SAGE designs multiple aggregation functions and supports inductive learning for unknown nodes.

[0156] The nodes of the GNN collect information from their neighborhoods through multiple rounds of iteration and update their own representations. Pooling layer: Perform graph-level pooling on the node representations to obtain the representation vectors of the entire graph or the subgraph of the target assets. The output layer is connected to a fully connected layer for the final risk scoring or classification. The GNN can automatically learn the complex interactions and dependencies between entities, capture the propagation patterns of risks in the network, and thus more accurately assess systemic risks and associated risks.

[0157] Dynamically adjust the weights and interaction relationships of each risk factor:

[0158] Online learning When there is new data (including asset performance updates, market changes) or user feedback, the model can perform incremental updates instead of completely retraining. For neural networks, mini-batch gradient descent can be used to continuously update the model parameters. For Bayesian networks, the Bayesian update method can be used to update the CPT. Adaptive learning rate algorithms (including variants of Adam) or learning rate decay strategies are adopted.

[0159] Adjustment based on feedback:

[0160] Direct feedback: Users (including risk analysts) can directly adjust the weights of certain risk factors or the evaluation results. These adjustments can be quantified and used to correct the model parameters (including by modifying the loss function, rewarding predictions consistent with user feedback and punishing those inconsistent).

[0161] Indirect feedback: The usage behavior of users on the report (including which parts are read intensively and which suggestions are adopted) can also be used as an indirect feedback signal.

[0162] Output the risk profile and confidence interval:

[0163] Risk profile: Present the comprehensive risk score and the sub-item scores in different risk dimensions (including credit, market, legal, operation) through visualization methods (including radar charts, dashboards, heat maps) to form an intuitive description of the risk status of non-performing assets.

[0164] Confidence interval calculation: Quantify the uncertainty of the evaluation results.

[0165] In Monte Carlo Dropout, the Dropout layer is also kept active during the prediction stage of the neural network. Multiple forward propagations are performed to obtain a set of prediction results. The mean and variance (or quantiles) are calculated based on the distribution of this set of results as the confidence interval.

[0166] The narrower the confidence interval, the more reliable the evaluation result. A wide confidence interval indicates a greater degree of uncertainty and requires more cautious decision-making or further investigation.

[0167] The risk assessment network adopts a reinforcement learning mechanism: enabling the risk assessment network to learn an optimal adjustment strategy through interaction with the environment (including user feedback), thereby continuously optimizing its risk sensitivity and generalization ability. Construction and application of deep deterministic policy gradient agents;

[0168] For reinforcement learning, soft update formula for target network parameters:

[0169] ;

[0170] Used for parameter updates of the target Actor network and target Critic network in the reinforcement learning algorithm. Soft updates are used to stabilize the learning process and avoid instability in training caused by overly rapid changes in the target network parameters.

[0171] : Parameters of the target network (target Actor network or target Critic network); : Parameters corresponding to the online network (online Actor network or online Critic network). : Soft update coefficient, a very small positive value ( , with a value range of 0.001 - 0.01), which controls the speed at which the online network parameters "transfer" to the target network parameters. The smaller it is, the slower the update of the target network and the more stable the learning process.

[0172] State space representation: The risk factor vector of the current asset (from the combination mined by GA, or a more comprehensive feature set), the current evaluation results of the risk assessment network (including risk scores, scores for each dimension), and even some parameters or confidence metrics of the model itself. It needs to be carefully designed to contain sufficient information to guide decision-making.

[0173] Data processing: Numerical and normalization processing.

[0174] The action space corresponds to adjustment strategies for specific risk factor weights or activation functions in the risk assessment network: The actions are continuous. Each dimension of the action vector can correspond to: The adjustment amount (increase / decrease percentage) of the weight of a certain key risk factor in the evaluation model. The adjustment of the slope or threshold of an activation function in a certain hidden layer. The adjustment of the weights of different sub-models in model integration.

[0175] Constraints: The action space needs to set reasonable boundaries to avoid instability of the model caused by excessive adjustments.

[0176] Reward signal design: Quantify the user's confirmation, correction, or rejection behavior of the risk assessment into a scalar reward signal:

[0177] Evaluation confirmation: If the user approves the evaluation result, give a positive reward (including +1).

[0178] Minor correction: The user made a small adjustment to the result, giving a small positive reward or zero reward (including +0.1, 0).

[0179] Significant correction / rejection: The user made a large adjustment to the result or completely negated it, giving a negative reward (including -1, -0.5).

[0180] Consistency of correction direction: If the direction of the user's correction is consistent with the direction that the RL agent attempts to adjust, even if it is a correction, a certain positive incentive can be given.

[0181] Delayed reward: Sometimes the true effect of the user's feedback becomes apparent after a period of time (including the comparison between the actual recovery situation after the disposal of non-performing assets and the prediction), and the distribution of delayed rewards needs to be considered.

[0182] Sparse reward problem: User feedback does not occur in every evaluation, and the sparse reward problem needs to be addressed, including using reward shaping or hierarchical reinforcement learning.

[0183] Loss function formula of the Critic network in reinforcement learning:

[0184] ;

[0185] Defines the loss function of the Critic network in the reinforcement learning algorithm. The purpose of the Critic network is to learn the state-action value function , that is, to evaluate the quality of performing the action in the state . Training the Critic network is to make its predicted value: , as close as possible to the target value . This loss function takes the form of mean squared error.

[0186] : Loss value of the Critic network. The goal of training is to minimize this loss value. : The number of experience samples in a mini batch. : Sum over all samples in the minibatch. : The target value of the th sample, also known as the TD target. Its calculation method is shown in the next formula. : The value predicted by the Critic network for the state and action of the th sample. : Parameters of the Critic network.

[0187] Objective of the Reinforcement Learning Critic Network Value calculation formula:

[0188]

[0189] Used to calculate the objective value during the learning of the Critic network in the DDPG algorithm Value . It is based on the Bellman equation and combines the immediate reward and the estimation of the future state value.

[0190] : The objective Q-value of the th sample; : In the th sample, the immediate reward obtained after executing the action in the state ; : Discount factor, with a value range between . It measures the importance of future rewards relative to the current reward. The closer it is to 1, the more the agent values long-term rewards; : Target Critic network 's evaluation of the next state and the action selected by the target Actor network in this state. This represents the estimation of the maximum expected return that may be obtained starting from the next state. : The next state in the th sample. : The stochastic action output by the target Actor network in the state . : Parameters of the target Actor network. : Parameters of the target Critic network.

[0191] Policy Gradient Formula of the Reinforcement Learning Actor Network:

[0192] ;

[0193] Describes the sampling policy gradient used for updating the parameters of the Actor network in the reinforcement learning algorithm. The purpose of the Actor network is to learn an optimal policy such that the action selected in any state Maximize the expected cumulative return. This gradient guides the update of the Actor network parameters in the direction of actions that can produce higher Q-values.

[0194] : The objective function of the Actor network (default is the expected cumulative return) with respect to its parameters The gradient. This is the direction and magnitude of the update of the Actor network parameters. : The number of experience samples in a mini-batch. : Sum over all samples in the minibatch. : The gradient of the output of the Critic network with respect to the action and evaluated at the current state and the action selected by the Actor network in that state. This gradient term indicates how much the value would change if the action selected by the Actor were changed slightly, indicating the direction of action improvement. : The gradient of the output of the Actor network with respect to its parameters and evaluated at the current state This gradient term shows how the parameters of the Actor network affect the actions it outputs; The action output by the Actor network in state according to the current parameters. The parameters of the Actor network; by multiplying these two gradient terms using the chain rule, the gradient of the objective function with respect to the parameters of the Actor network can be obtained, and then the parameters of the Actor network can be updated by gradient ascent to optimize the policy.

[0195] To allow the agent to explore different adjustment strategies, noise (including Ornstein-Uhlenbeck process noise or Gaussian noise) can be added to the actions output by the Actor network.

[0196] Application logic: The RL agent continuously observes the state of the risk assessment network, takes adjustment actions, receives user feedback as rewards, and continuously optimizes its adjustment strategy, so that the risk assessment network can adaptively improve its sensitivity and generalization ability to risks, and be closer to the judgment criteria of human experts and the actual business needs.

[0197] Sp3, Insight Generation and Report Output Based on Interpretability Models.

[0198] This step aims to transform the complex model evaluation results into human - understandable and credible insights, and generate high - quality due diligence reports based on normative logic and personalized requirements.

[0199] Sp3.1. Use an explainable artificial intelligence (XAI) model to conduct attribution analysis on the output results of the risk assessment network: identify the key influencing paths and core data evidence leading to specific risk assessment conclusions (including "high risk", "recommended not to pass"), and enhance model transparency and user trust.

[0200] XAI model selection and application: Generate perturbation samples within the local neighborhood of the sample to be explained, and use a decision tree to fit the local behavior of the risk assessment network.

[0201] Perturbation sample generation: For tabular data, randomly perturb the feature values; for text data, randomly mask or replace words; for graph data, perturb node features or edges.

[0202] Local interpretable model training: Use the perturbation samples and their corresponding complex model prediction results to train a weighted linear model, where the weights are based on the distance between the perturbation samples and the original samples.

[0203] Explanation output: The coefficients of the linear model can be regarded as the contribution degrees of each feature to this local prediction.

[0204] Application logic: LIME can explain a single prediction result and tell users "Why is this specific asset rated as high risk?".

[0205] Feature importance ranking: According to the coefficients of LIME, rank all the risk factors input into the risk assessment network, and identify the top - K factors that have the greatest impact on the current assessment conclusion.

[0206] Impact path tracing (combined with knowledge graph): If the risk assessment network is based on GNN, or its input features are associated with the knowledge graph, the key risk factors identified by XAI can be located to the nodes or edges in the knowledge graph. Then, use graph algorithms (including attention - weight - based path search, shortest key path algorithm) to trace in the knowledge graph how these key factors interact with each other through a series of association relationships (including guarantee chain, fund transfer, equity control) and finally lead to the conduction path of the risk event. Link the key risk factors and impact paths back to their original data sources. If a "contract clause risk" factor is identified as important, the system should be able to locate which clause of which specific contract and highlight it for the user. If an abnormal financial indicator is key, it should be able to link to the corresponding item and its context in the financial statements.

[0207] Provide users with a complete explanation chain from "what is the risk" to "why this risk" and then to "where is the evidence".

[0208] Sp3.2. Based on a preset and user-dynamically adjustable report logic framework and narrative template, embed the key information, risk portraits, risk factor combinations, key impact paths, and core data evidence in the context-aware model into the corresponding chapters of the report, and use natural language generation (NLG) to output a due diligence report on financial non-performing assets that includes in-depth analysis, risk warnings, disposal suggestions, and key points for compliance review:

[0209] Report logic framework design: Use XML Schema, JSON Schema, or domain-specific languages to define the hierarchical structure of the report (including cover, table of contents, abstract, each chapter of the main text, appendix), the titles of each chapter, the content modules to be included (including asset overview, debtor analysis, guarantee analysis, risk assessment summary, disposal suggestions, etc.), and the data types and sources of each module.

[0210] User dynamic adjustment: Provide a graphical configuration interface or parameterized interface that allows users to select or customize the chapters, modules, level of detail, and display style of the report according to different report purposes (including internal approval, external transfer, litigation support), audience types (including executives, salespersons, legal affairs), or asset characteristics.

[0211] Narrative template library construction:

[0212] Template types: Design diverse narrative templates for different analysis scenarios (including "short-term default caused by liquidity crisis", "systemic risks caused by over-guarantee", "substantial shrinkage of collateral value"), risk levels, asset types, and disposal strategies.

[0213] Template content: The template contains fixed text and dynamic placeholders. The placeholders will be filled by the NLG module according to the analysis results. A risk description template could be: "Debtor [debtor name], due to the influence of [key risk factor 1] and [key risk factor 2], the current assessed risk level is [risk level], mainly manifested in [description of specific data evidence]...".

[0214] Template management: Support the creation, editing, version control, and on-demand invocation of templates.

[0215] Encode the outputs of the upstream modules (including the list of key risk factors, sub-graphs of the knowledge graph, risk scores) into a format acceptable to the model, including structured input (key-value pairs), a linearized sequence of triples, or a text sequence combined with control codes. Generate the target report paragraphs or summaries.

[0216] A large amount of "input data - output text" parallel corpus is required. It can be extracted from existing written reports or constructed semi - automatically. Extract the "risk description" section from a large number of due diligence reports, and use the corresponding structured risk factors and financial data as input. Fine - tune the pre - trained model on the parallel corpus in a specific domain to adapt to the language style, professional terms, and logical structure of due diligence reports. The loss function is usually cross - entropy loss (word - by - word prediction). When generating text, a beam search decoding strategy is adopted to balance the fluency, diversity, and accuracy of the generated text. It can generate more natural, flexible, and context - coherent text, which is suitable for writing analytical comments, risk summaries, disposal suggestions, etc., which require complex logic and detailed expressions.

[0217] The NLG module fills each analysis result (in - depth analysis, risk warning, disposal suggestion, key points of compliance review) into the selected report logic framework and narrative template, and finally generates a preliminary draft of the due diligence report (including Word, PDF, HTML formats).

[0218] Generate a visual map of the risk conduction path:

[0219] Application of attention - weighted graph neural network inference algorithm:

[0220] If the knowledge graph itself includes historical data on risk conduction or expert - labeled conduction relationships, including GAT, the GNN can be trained to predict the conduction probability or influence intensity. On the constructed context - aware knowledge graph, use the pre - trained GNN for inference to identify potential influence paths from a certain risk source node to other nodes. The attention weights can be interpreted as proxies for influence intensity or conduction probability.

[0221] Combine the Dijkstra algorithm to search for high - weight (high - impact / high - probability) conduction paths starting from a specific risk event node on the graph with attention weights.

[0222] Input one or more initial risk nodes (including "Core debtor A has a liquidity crisis"), and the GNN infers and outputs a sub - graph, which contains risk conduction paths connected by high - attention - weight edges, as well as the quantified influence intensity / conduction probability of each node and edge on the path.

[0223] Visual map and interactive node drilling:

[0224] Front - end implementation: Use a graph visualization library to render the risk conduction path map on the Web interface. Visual elements such as node size, color, edge thickness, arrow direction, etc. can be used to represent risk levels, influence intensities, conduction directions, etc.

[0225] When the user clicks on any node in the graph (including debtors, collateral, risk events), the system should be able to dynamically request and display the detailed information, related attributes, associated original data fragments (including contract clause text, financial statement screenshots, news links) or textual evidence of that node. This requires close cooperation between the front-end and back-end APIs.

[0226] Preset portraits of the target audience: areas of concern (including whether more concerned about legal risks or market risks), professional levels (including junior analysts, senior experts, executives), and report purposes (including internal decision-making, external disclosure, regulatory reporting). Encode these discrete or continuous portrait dimensions into numerical vectors. Use an embedding layer to map each dimension to an embedding vector, and then concatenate these vectors or fuse them through a small network as additional conditional inputs to the Transformer model.

[0227] Dynamic adjustment mechanism: Conditional inputs guide the model to select or generate different high-level narrative structures. For example, if the audience is an executive, the model tends to generate a more general and conclusion-oriented text structure.

[0228] Adjustment of the professional level of terms:

[0229] Glossary control: Dynamically adjust the selection range of the glossary during decoding according to the professional level portrait (including restricting the use of overly professional terms or preferentially selecting easy-to-understand synonyms).

[0230] Style transfer: The model can be trained to convert between text styles of different professional levels or add style control codes during generation.

[0231] Model training: It is necessary to construct training data containing triples of (conditional portrait, structured semantic representation, target customized report text). This requires a large amount of writing and annotation, or the use of weak supervision learning, multi-task learning, etc.

[0232] The continuous learning and model iteration module ensures that the system can adapt to changing data distributions, risk patterns, and user needs, maintaining and improving its long-term performance.

[0233] Data and concept drift detection unit: Statistical characteristics of the input data stream: Monitor whether the distributions (including mean, variance, skewness, kurtosis, category frequencies) of various risk factors (numeric, categorical) have changed significantly. For high-dimensional data (including text / image embeddings), the distribution changes in the low-dimensional projection space can be monitored.

[0234] Model prediction performance: Monitor whether key performance indicators (including accuracy, recall, F1 score, AUC, stability of risk scores) decline over time.

[0235] Changes in the user feedback mode: The magnitude of the user's correction of the evaluation results, the correction frequency, the newly proposed risk points, etc.

[0236] Model relationship drift: The true relationship between features and target variables changes. It is usually indirectly detected by monitoring the continuous decline of model performance. It can also be judged by comparing the differences between models trained on new and old data.

[0237] Trigger the model update process: When significant drift is detected (including the statistical test value being less than the threshold, or the performance decline exceeding the preset magnitude), the system automatically performs the following actions:

[0238] 1. Issue an alarm: Notify the operation and maintenance personnel and the model maintainer.

[0239] 2. Collect new data: Start collecting new data after the drift occurs for model retraining or adjustment.

[0240] 3. Schedule the retraining task: Automatically or semi - automatically start the process of retraining, fine - tuning, or structural adjustment of the model.

[0241] Differential knowledge graph update unit: Data source monitoring: Continuously monitor the changes in the original data sources (including legal document libraries, financial databases, public opinion APIs), and capture newly added, modified, or deleted data.

[0242] Change detection and extraction: Re - run the information extraction process in Sp1 on the changed data to obtain new entities, relationships, attributes, or their changes. Incorporate these changes into the existing knowledge graph incrementally. Support efficient addition / deletion / modification of nodes, edges, and their attributes. Define conflict resolution strategies, including based on timestamp (latest valid), based on source credibility, or arbitration. Ensure the atomicity, consistency, isolation, and durability (ACID properties) of the update process, especially in the scenario of concurrent updates.

[0243] Transfer learning effectively utilizes the knowledge learned from historical data during model iteration and upgrade, accelerates the convergence of the new model, reduces the dependence on newly labeled data, and improves the performance of the model on new tasks or new data distributions.

[0244] Parameter transfer: When training a new model, use the parameters of the old model trained on historical data as the initial weights (or the weights of some layers). It is applicable to the situation where the new and old tasks are similar or the new data is less. Based on the FinBERT model, for example, when it is necessary to identify risk factors for a new type of non - performing asset, the weights of the general FinBERT can be used as a starting point and fine - tuned on new data instead of starting from random initialization.

[0245] Feature representation transfer directly uses the feature representations learned by the old model (including text embeddings, image embeddings, and graph embeddings) as the input features of the new model or as part of the features of the new model. The vector representations of non-performing assets learned by the old risk assessment model can be used as node features by the new and more complex assessment models (including GNNs).

[0246] Domain adaptation: When there are differences in the data distributions of the source domain (historical data) and the target domain (new data) but the tasks are the same, the model can be made to adapt to the target domain by adjusting the weights based on instances. The historical data mainly comes from non-performing assets in the real estate industry, while the new data mainly comes from the manufacturing industry. Domain adaptation can enable the risk assessment model to better generalize to the manufacturing industry.

[0247] Its overall operation process is as Figure 2 shown:

[0248] [Start] Task reception and initialization, data collection and preprocessing phase, context-aware modeling phase, risk factor mining phase, risk assessment phase, insight generation phase, automatic report generation phase, user interaction and report review phase, report finalization and output phase, continuous learning and model iteration phase, and finally end, task completed, resources released.

[0249] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a reference structure" does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0250] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for intelligently generating a due diligence report on non-performing financial assets, characterized in that, The method includes: Sp1, Deep multi-modal data fusion and context-aware modeling: Collect structured, semi-structured, and unstructured data related to target non-performing financial assets to obtain heterogeneous data; Map the heterogeneous data to a unified semantic space, identify and establish entities, attributes, and complex association relationships related to non-performing financial assets, and form a context-aware model for the non-performing financial assets; Sp2, Risk factor mining and adaptive assessment based on evolutionary computation: Sp2.1, Initialize the risk factor candidate set: Based on the context-aware model, use the genetic algorithm evolutionary computation method to perform iterative optimization in a preset risk dimension space to mine potential and non-explicit risk factor combinations that are coupled with the characteristics of the non-performing financial assets; Sp2.2, Construct a risk assessment network: The risk assessment network dynamically adjusts the weights and interaction relationships of each risk factor according to newly input data and historical assessment feedback, and combines an ensemble learning strategy to quantitatively evaluate the comprehensive risk level of non-performing financial assets, and outputs a risk portrait and a confidence interval; Sp3, Insight generation and report output based on an interpretable model: Sp3.1, Use an interpretable intelligent model to perform attribution analysis on the output results of the risk assessment network, and identify the key impact paths and core data evidence that lead to specific risk assessment conclusions; Sp3.2, According to a preset and user-dynamically adjustable report logic framework and narrative template, use natural language generation for the key information in the context-aware model to output a due diligence report on non-performing financial assets.

2. The method for intelligently generating a due diligence report on non-performing financial assets according to claim 1, characterized in that, In the deep multi-modal data fusion and context-aware modeling, it further includes: using a pre-trained language model enhanced by an attention mechanism to parse long-distance dependencies and complex clause structures in legal texts, and encoding the key elements of the legal texts and their confidence scores together as nodes and weighted edges in a knowledge graph.

3. A method for intelligently generating a due diligence report on non-performing financial assets according to claim 1, characterized in that, In the risk factor mining based on evolutionary computation, it further includes: Quantify its interpretability by calculating the minimum description length of the decision rules generated by the risk factors; Measure its prediction accuracy by evaluating the F1 score of the risk factor combination for historical default events on a backtest dataset; Evaluate its correlation by calculating the cosine similarity of the vector space between the risk factor combination and the patterns in a preset typical non-performing asset risk pattern library verified by domain experts.

4. A method for intelligently generating a due diligence report on non-performing financial assets according to claim 1, characterized in that, The risk assessment network adopts a reinforcement learning mechanism, and the reinforcement learning mechanism further includes: constructing a deep deterministic policy gradient agent, where the state space of the agent represents the risk factors and assessment results of the current asset, and the action space corresponds to the adjustment strategy for the weights or activation functions of specific risk factors in the risk assessment network; quantifying the user's confirmation, correction, or rejection behavior of the risk assessment as a scalar reward signal to guide the policy learning of the agent.

5. A method for intelligently generating a due diligence report on non-performing financial assets according to claim 1, characterized in that, The insight generation of the interpretability model further includes: generating a visual map of the risk conduction path, and the generating of the visual map of the risk conduction path further includes: on the knowledge graph of the context-aware model, using a graph neural network inference algorithm based on attention weighting to identify and quantify the influence intensity and conduction probability between different risk entities and risk factors, and forming a conduction path.

6. The method for intelligently generating a due diligence report on non-performing financial assets according to claim 1, wherein, In the report output, the natural language generation further includes: using a conditional text generation model, taking the preset portrait of the target audience as the conditional input, and combining the structured semantic representation extracted from the insights of the interpretability model to dynamically select narrative templates, adjust the professional level of terms, and control the detailed level of the argumentation to generate customized report texts that meet the needs of specific audiences.

7. A method for intelligently generating a due diligence report on non-performing financial assets according to claim 1, characterized in that, The method further includes a continuous learning and model iteration module, and the continuous learning and model iteration module further includes: A data and concept drift detection unit, which is used to monitor the statistical characteristics of the input data stream and the changes in the user feedback mode in real time, and automatically triggers the model update process when significant drift is detected; A differential knowledge graph update unit, which merges the newly added or changed entities, relationships and their confidence levels into the existing knowledge graph in an incremental manner, and applies the knowledge learned from historical data to the iterative upgrade of the new model using transfer learning.

8. A system for intelligently generating a due diligence report on non-performing financial assets, which includes a processor and a memory coupled to the processor, and computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, it is characterized in that, The system further includes: A multi-modal data deep fusion and context-aware modeling engine: Collecting data related to target non-performing financial assets from heterogeneous data sources; Constructing a context-aware model for non-performing financial assets; Sp2. A risk factor mining and adaptive evaluation engine based on evolutionary computation: Performing iterative optimization in a preset risk dimension space; Constructing and running a risk assessment network; Sp3. An insight generation and report output engine based on an interpretability model: Performing attribution analysis on the risk assessment results using an interpretable intelligent model; Outputting a due diligence report on non-performing financial assets.

9. The system for intelligently generating a due diligence report on non-performing financial assets according to claim 8, characterized in that, The pre-trained language model processing unit in the multi-modal data deep fusion and context-aware modeling engine further includes: A complex long sentence segmentation and dependency parsing module, which is used to accurately identify the main-subordinate structure and restrictive conditions in contract terms; A semantic role annotation module enhanced by domain knowledge, which annotates specific participant roles and core legal behaviors in financial transactions.

10. A system for intelligently generating due diligence reports on non-performing financial assets according to claim 8, characterized in that, The risk factor mining and adaptive evaluation engine based on evolutionary computation further includes: A feedback-driven model parameter adjustment module, and the feedback-driven model parameter adjustment module further includes: A user feedback real-time capture and structured processing interface, which is used to receive and parse the annotation information of users on risk factors and evaluation results; An incremental model training and version control unit, which is embedded with an evolutionary algorithm, and the incremental model training and version control unit supports online fine-tuning of the population initialization strategy of the evolutionary algorithm and the local connection weights of the evaluation network; An offline batch re-training scheduler, which is used to trigger the global optimization of the entire model system after accumulating enough new data or feedback.

Citation Information

Patent Citations

  • Unhealthy asset valuation algorithm based on multi-modal high-dimensional features

    CN117217807A

  • Risk drug identification method and system based on drug feedback mining

    CN118230983A

  • Legal document information extraction method

    CN118627619A

  • Electronic contract management method and system based on deep learning model

    CN118761735A

  • Complex ecological smart brain-driven data knowledge graph construction method and system

    CN119719388A

Cited By

  • Data resource migration risk prediction method and system, terminal equipment and storage medium

    CN120448161A

  • Enterprise supply chain risk prediction method based on artificial intelligence

    CN120706915A

  • Data structure adaptive visualization method based on large language model

    CN120744145A

  • Large model driven enterprise exhaustion report generation method and system

    CN120765387A

  • Multi-source heterogeneous data fusion type financial analysis report generation method and system and medium

    CN120781987A