Credit evaluation method and system based on multi-modal data

CN121213229BActive Publication Date: 2026-08-21HANGZHOU MONEY BAG DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511430847.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2026-08-21
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

传统评估系统通常仅考虑财务报表等单一维度的数据,而忽视了企业运营过程中产生的文本报告、市场舆情、供应链关系以及图像视频等多模态信息

Benefits of technology

[0013] As can be seen from the above, the credit assessment method and system based on multimodal data provided in this application integrate multimodal data through cross-modal feature alignment and fusion algorithms, extract related features and temporal evolution patterns by combining dynamic knowledge graphs, and generate a credit assessment report that includes risk transmission path analysis. To a certain extent, this solves the problem that traditional methods cannot effectively utilize multimodal data and quantify related risks. It has the advantages of improving the accuracy, comprehensiveness and foresight of the assessment, effectively capturing corporate related risks, and dynamically reflecting changes in credit status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121213229B_ABST
    Figure CN121213229B_ABST
Patent Text Reader

Abstract

The application discloses a credit evaluation method and system based on multi-modal data, and belongs to the technical field of computers.A credit evaluation method and system based on multi-modal data are provided, multi-modal data is integrated through a cross-modal feature alignment and fusion algorithm, associated features and time sequence evolution patterns are extracted in combination with a dynamic knowledge graph, a credit evaluation report containing risk transmission path analysis is generated, the problem that traditional methods cannot effectively utilize multi-modal data and quantify associated risks is solved to some extent, and the credit evaluation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a credit assessment method and system based on multimodal data. Background Technology

[0002] Existing corporate credit assessment methods primarily rely on structured financial data, exhibiting significant limitations in utilizing unstructured multimodal data. Traditional assessment systems typically consider only single-dimensional data such as financial statements, neglecting multimodal information generated during corporate operations, including text reports, market sentiment, supply chain relationships, and images and videos. The heterogeneity of data structures and feature spaces across different modalities leads to inconsistent feature distributions, making it difficult for traditional feature engineering methods to effectively align and fuse features across modalities. This limitation makes it difficult for assessment models to capture the full picture of a company's creditworthiness, resulting in the omission of key risk signals and biased assessment results. Particularly when dealing with corporate association risks and credit transmission effects, existing methods lack effective knowledge graph modeling capabilities, failing to accurately quantify the risk transmission paths and concentrations between related companies. Furthermore, static assessment frameworks struggle to reflect the temporal evolution of corporate creditworthiness, causing risk assessments to lag behind actual operational changes. These technical deficiencies severely restrict the accuracy and foresight of credit assessments. Summary of the Invention

[0003] This application provides a credit assessment method and system based on multimodal data, which can improve the accuracy of credit assessment. The technical solution is as follows: On the one hand, a credit assessment method based on multimodal data is provided, the method comprising: The multimodal data of the target enterprise is extracted and fused using a cross-modal feature alignment and fusion algorithm to obtain multimodal features. The multimodal data includes financial data, text data, market data, supply chain data, and image and video data. Based on the multimodal features and a pre-set initial knowledge graph, the enterprise credit association features of the target enterprise are determined. The enterprise credit association features are high-order latent features used for credit assessment, and the pre-set initial knowledge graph is the initial data structure used in the field of credit assessment. Based on the multimodal features and enterprise credit association features, the credit score of the target enterprise is determined. Based on the multimodal features, enterprise credit association features, and credit score, a credit assessment report of the target enterprise is generated.

[0004] Furthermore, this application proposes to extract and fuse features from the multimodal data of a target enterprise using a cross-modal feature alignment and fusion algorithm to obtain multimodal features of the multimodal data. This includes: extracting modal features of each modal data and semantic association weights between each modal data through a cross-modal attention mechanism, and generating an inter-modal attention map of the multimodal data based on the semantic association weights between each modal data; using the inter-modal attention map to guide the alignment process of each modal feature, and using a gradient inversion layer and a domain discriminator in an adversarial training framework to process each modal feature to eliminate distribution differences between modal features; and fusing the processed modal features to obtain the multimodal features of the multimodal data.

[0005] Furthermore, this application proposes to extract modal features of each modality in multimodal data and semantic association weights between each modality through a cross-modal attention mechanism, and to generate an inter-modal attention map of multimodal data based on the semantic association weights between each modality. This includes: mapping each modality to a shared semantic space based on an attention mechanism to obtain the modal features of each modality; determining the correlation score between any two modal features to obtain the semantic association weights between different modalities; dynamically weighting and fusing the modal features of each modality based on the semantic association weights to obtain the inter-modal attention map; using the inter-modal attention map to guide the alignment process of each modality feature, and employing an adversarial training framework. The gradient inversion layer and domain discriminator process the modal features to eliminate distribution differences between them. This includes: extracting semantic association weights from the inter-modal attention map and using these weights as the attention mask for the domain discriminator, enabling it to focus on semantically related feature regions for domain classification; adding a gradient inversion layer between the feature extractor and the domain discriminator, inputting each modal feature into the feature extractor, and inverting the backpropagation gradient of the domain discriminant loss to drive the feature extractor to process each modal feature, resulting in multiple reference modal features; and inputting these multiple reference modal features into the domain discriminator, which then aligns the distribution of semantically related feature regions to improve the distribution consistency among the multiple reference modal features.

[0006] Furthermore, this application proposes determining the corporate credit association characteristics of a target enterprise based on multimodal features and a pre-defined initial knowledge graph, including: generating a dynamic multimodal knowledge graph based on the multimodal features and the initial knowledge graph; extracting the association features and temporal evolution patterns of the target enterprise based on the dynamic multimodal knowledge graph, whereby the association features describe the strength of the relationship between enterprises and the temporal evolution patterns describe the trend of changes in the enterprise's credit status; and determining the corporate credit association characteristics of the target enterprise based on the association features and temporal evolution patterns.

[0007] Furthermore, this application proposes generating a dynamic multimodal knowledge graph based on multimodal features and an initial knowledge graph, including: updating the node attributes and edge weights of the initial knowledge graph based on entity and relationship information in the multimodal features; adding timestamp attributes to the nodes and edges in the updated initial knowledge graph to obtain a reference knowledge graph with temporal information; dynamically updating the node attributes of the reference knowledge graph through the message passing mechanism of a graph neural network to obtain the dynamic multimodal knowledge graph; and extracting the association features and temporal evolution patterns of target enterprises based on the dynamic multimodal knowledge graph, including: extracting inter-enterprise relationships from the dynamic multimodal knowledge graph. The study identifies the following characteristics: Inter-enterprise association strength features, including degree centrality and betweenness centrality; temporal variation features of corporate credit status are extracted from a dynamic multimodal knowledge graph, including trend and fluctuation features; inter-enterprise association strength features are defined as association features, and temporal variation features are defined as temporal evolution patterns; based on the association features and temporal evolution patterns of the target enterprise, the study determines the enterprise credit association features of the target enterprise, including: performing principal component analysis on the association features of the target enterprise to obtain the first key feature related to credit assessment; and fusing the first key feature and the temporal evolution pattern to obtain the enterprise credit association features.

[0008] Furthermore, this application proposes a method for determining the credit score of a target enterprise based on multimodal features and enterprise credit association features, including: fusing multimodal features with enterprise credit association features to obtain the fused features of the target enterprise; determining the static qualification score and dynamic association risk score of the target enterprise based on the fused features; and determining the credit score of the target enterprise based on the static qualification score and dynamic association risk score, wherein the credit score is used to represent the static qualification and dynamic association risk of the target enterprise.

[0009] Furthermore, this application proposes generating a credit assessment report for a target enterprise based on multimodal features, enterprise credit association features, and credit scores. This includes: extracting a second key feature influencing the credit score and its corresponding contribution based on multimodal features; determining the risk transmission path and risk concentration in the target enterprise's enterprise association network based on enterprise credit association features; decomposing the credit score into a static qualification score and a dynamic association risk score; and generating a structured credit assessment report based on the second key feature, its corresponding contribution, risk transmission path, risk concentration, static qualification score, and dynamic association risk score. The credit assessment report includes risk source analysis, association impact assessment, and improvement suggestions.

[0010] Furthermore, this application proposes extracting a second key feature affecting credit scores and its corresponding contribution based on multimodal features, including: decomposing multimodal features into multiple sub-features; analyzing the causal effect of each sub-feature on credit scores through a causal inference model, where the causal effect includes direct and indirect effects. The direct effect represents the direct impact of a sub-feature on credit scores without the influence of other sub-features, while the indirect effect represents the indirect impact of a sub-feature on credit scores by influencing one or more other sub-features; extracting a second key feature affecting credit scores and its corresponding contribution based on the causal effect of each sub-feature on credit scores, where the second key feature belongs to multiple sub-features; and based on enterprise credit association... The study identifies risk transmission paths and risk concentration within the enterprise network of a target enterprise. This includes: constructing an enterprise network graph based on enterprise credit association characteristics, where nodes represent enterprises, edges represent inter-enterprise relationships, and edge weights represent association strength; employing a path analysis algorithm based on a risk propagation model to calculate all potential risk transmission paths from high-risk nodes to the target enterprise; extracting the top K risk transmission paths with the highest risk using the K-shortest path algorithm and calculating the risk transmission probability for each path; and determining the risk concentration of the target enterprise within the enterprise network graph based on target indicators, including betweenness centrality and proximity centrality, which quantify the degree to which the target enterprise is influenced by other enterprises in the enterprise network graph.

[0011] Furthermore, this application proposes generating a structured credit assessment report based on the second key feature, its corresponding contribution, risk transmission path, risk concentration, static creditworthiness score, and dynamic associated risk score. This includes: constructing a multi-dimensional assessment matrix, mapping the static creditworthiness score and dynamic associated risk score to three assessment dimensions: financial health, operational stability, and associated risk exposure, forming three-dimensional scores; generating a risk topology map based on the risk transmission path and risk concentration, whereby the risk topology map identifies key risk sources and major risk transmission paths, with the key risk source being the risky enterprise; converting the second key feature and its corresponding contribution into a structured text description; generating an initial credit assessment report based on the three-dimensional scores, the risk topology map, and the structured text description; and adding credit trend predictions and improvement suggestions based on a time-series evolution pattern to the initial credit assessment report, thus generating a structured credit assessment report.

[0012] Furthermore, this application proposes mapping static qualification scores and dynamic associated risk scores to three assessment dimensions: financial health, operational stability, and associated risk exposure, forming a three-dimensional score. This includes: mapping static qualification scores to the dimensions of financial health and operational stability, and determining the dimensional scores for these dimensions based on financial data; mapping dynamic associated risk scores to the dimension of associated risk exposure, and calculating the dimensional score for this dimension based on risk transmission probability and risk concentration; and generating a risk topology graph based on risk transmission paths and risk concentration, including: extracting key nodes and edge relationships from the top K risk transmission paths with the highest risk; calculating the risk contribution of each key node and the risk transmission strength of the edges; and using the graph... Visualization technology generates a risk topology map and marks high-risk areas in the map as a heatmap. The second key feature and its corresponding contribution are converted into structured text descriptions, including: generating a standardized risk description template based on the feature type and influence direction of the second key feature; filling the risk description template with contribution values ​​to obtain the structured text description; and generating an initial credit assessment report based on the three-dimensional scores, the risk topology map, and the structured text description, including: inputting the three-dimensional scores, the risk topology map, and the structured text description into a report generation model; processing the three-dimensional scores, the risk topology map, and the structured text description through the report generation model to generate the initial credit assessment report. The report generation model is a large language model.

[0013] As can be seen from the above, the credit assessment method and system based on multimodal data provided in this application integrate multimodal data through cross-modal feature alignment and fusion algorithms, extract related features and temporal evolution patterns by combining dynamic knowledge graphs, and generate a credit assessment report that includes risk transmission path analysis. To a certain extent, this solves the problem that traditional methods cannot effectively utilize multimodal data and quantify related risks. It has the advantages of improving the accuracy, comprehensiveness and foresight of the assessment, effectively capturing corporate related risks, and dynamically reflecting changes in credit status.

[0014] On the one hand, a credit assessment system based on multimodal data is provided, the system comprising: The feature extraction and fusion module is used to extract and fuse features from the multimodal data of the target enterprise through cross-modal feature alignment and fusion algorithms to obtain the multimodal features of the multimodal data, which includes financial data, text data, market data, supply chain data and image and video data. The feature determination module is used to determine the enterprise credit association features of the target enterprise based on the multimodal features and the preset initial knowledge graph. The enterprise credit association features are high-order latent features used for credit assessment, and the preset initial knowledge graph is an initial data structure used in the field of credit assessment. The credit scoring determination module is used to determine the credit score of the target enterprise based on the multimodal features and the enterprise credit association features. The report generation module is used to generate a credit assessment report for the target enterprise based on the multimodal features, the enterprise credit association features, and the credit score.

[0015] On one hand, a computer device is provided, the computer device including one or more processors and one or more memories, the one or more memories storing at least one computer program, the computer program being loaded and executed by the one or more processors to implement the credit assessment method based on multimodal data.

[0016] On the one hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the credit assessment method based on multimodal data.

[0017] On the one hand, a computer program product or computer program is provided, which includes program code stored in a computer-readable storage medium. The processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the aforementioned credit assessment method based on multimodal data. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the implementation environment of a credit assessment method based on multimodal data provided in an embodiment of this application; Figure 2 This is a flowchart of a credit assessment method based on multimodal data provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the determination of multimodal features provided in an embodiment of this application; Figure 4 This is a flowchart illustrating how to determine corporate credit association characteristics, as provided in an embodiment of this application. Figure 5 This is a flowchart of a method for determining a credit score provided in an embodiment of this application; Figure 6This is a flowchart of generating a credit assessment report provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a credit assessment system based on multimodal data provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0021] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0022] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain better results.

[0023] Normalization: Mapping sequences of values ​​with different ranges to the interval (0, 1) to facilitate data processing. In some cases, normalized values ​​can be directly expressed as probabilities.

[0024] Embedded coding, mathematically speaking, represents a correspondence, that is, mapping data in space X to space Y using a function F. This function F is injective, and the mapping result preserves the structure. An injective function means that the mapped data uniquely corresponds to the original data, and preserving the structure means that the size relationship between the original and mapped data is the same. For example, if there are data X1 and X2 before mapping, after mapping we get Y1 corresponding to X1 and Y2 corresponding to X2. If the original data X1 > X2, then correspondingly, the mapped data Y1 > Y2. For words, this means mapping words to another space to facilitate subsequent machine learning and processing.

[0025] Attention weights represent the importance of a piece of data during training or prediction. Importance indicates the magnitude of the influence of input data on output data. Data with high importance corresponds to higher attention weights, while data with low importance corresponds to lower attention weights. The importance of data varies in different scenarios, and training the model to assign attention weights is essentially the process of determining data importance.

[0026] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0027] Figure 1 This is a schematic diagram illustrating the implementation environment of a credit assessment method based on multimodal data provided in this application embodiment. See also... Figure 1 The implementation environment may include terminal 110 and server 140.

[0028] Terminal 110 is connected to server 140 via a wireless or wired network. Optionally, terminal 110 may be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. Terminal 110 has an application installed and running that supports credit assessment based on multimodal data.

[0029] Server 140 is a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Server 140 can provide background services for applications running on terminal 110.

[0030] In related technologies, corporate credit assessment has long relied on structured financial data, with limited capabilities for processing unstructured data. Traditional methods struggle to integrate heterogeneous information such as text, images, and supply chain relationships, resulting in a single assessment dimension. Different modalities of data cannot be effectively aligned due to differences in feature spaces, easily overlooking key risk signals. For example, when a company's financial data is strong but negative public opinion appears on social media, related technologies fail to capture this contradictory information, leading to biased assessment results.

[0031] To address the aforementioned issues, the inventors discovered that the key to multimodal data fusion lies in eliminating differences in feature distribution. Analysis revealed that cross-modal attention mechanisms can capture semantic relationships, but simple weighted fusion cannot solve the problem of deep distribution inconsistencies. Further research proposed introducing adversarial training into the feature alignment process, using gradient inversion layers to force the extraction of modality-invariant features. Simultaneously, the inventors recognized that corporate credit depends not only on its own qualifications but also on the influence of relational networks. By introducing dynamic knowledge graphs, multimodal features are mapped to the semantic space of credit assessment, thus mitigating the limitations of isolated feature analysis to some extent.

[0032] Therefore, this application proposes a credit assessment method based on multimodal data, see [link to relevant documentation]. Figure 2 Taking the server as the executing entity as an example, the method includes the following steps: 201. The server extracts and fuses features from the multimodal data of the target enterprise through a cross-modal feature alignment and fusion algorithm to obtain multimodal features of the multimodal data, which includes financial data, text data, market data, supply chain data, and image and video data. 202. Based on multimodal features and a preset initial knowledge graph, the server determines the enterprise credit association features of the target enterprise. The enterprise credit association features are high-order latent features used for credit assessment, and the preset initial knowledge graph is the initial data structure used in the field of credit assessment. 203. The server determines the credit score of the target enterprise based on multimodal features and enterprise credit association features; 204. The server generates a credit assessment report for the target company based on multimodal features, corporate credit association features, and credit scores.

[0033] Cross-modal feature alignment and fusion algorithms refer to techniques that eliminate modal differences through attention mechanisms and adversarial training. Specifically, this can be achieved by combining cross-modal attention map generation with gradient inversion layers to address the inconsistency in feature spaces of heterogeneous data. Multimodal data includes financial indicators, text records, market dynamics, supply chain relationships, and image and video information. Data cleaning and standardization preprocessing methods can be used for preprocessing to construct a feature system covering multidimensional information about the enterprise.

[0034] In some embodiments, financial data includes current ratio, quick ratio, cash ratio, debt-to-equity ratio, equity ratio, accounts receivable turnover, inventory turnover, total asset turnover, net profit margin on total assets, return on equity, operating profit margin, return on investment, fixed asset growth rate, total asset growth rate, return on equity growth rate, net profit growth rate, operating revenue growth rate, financial leverage, and operating leverage ratio. Text data includes text records of the company's credit history, repayment history, and historical defaults. Market data includes market dynamics, market feedback, and social media comments; supply chain data includes supply chain relationships and the creditworthiness of partner companies; image and video data includes images of the company's business premises, product images, and promotional videos. The initial knowledge graph refers to a knowledge base in the credit assessment domain that includes industry risk rules and entity relationships. This can be achieved through expert knowledge construction and industry data mining, and is used to provide a semantic framework for credit assessment. Enterprise credit association features refer to a set of high-order features optimized from the knowledge graph. This can be achieved through graph neural networks and principal component analysis methods, and is used to extract network and temporal features relevant to the credit assessment objective.

[0035] Specifically, after standardization, financial data is cross-modal aligned with text sentiment features, and semantic association weights between different modalities are calculated using an attention mechanism. Market dynamics data and supply chain relationship data are distributed and aligned within an adversarial training framework to eliminate feature bias caused by different data sources. Image and video data, after extracting visual features through convolutional neural networks, are dynamically fused with text semantic features. The fused multimodal features are input into a dynamic knowledge graph for node updates, capturing trends in corporate credit changes through temporal evolution patterns. The combined assessment of static qualification scores and dynamic associated risk scores reflects both the company's current operating status and reveals its risk exposure within the industry chain. During the generation of the structured assessment report, key feature contribution analysis is combined with risk transmission path visualization to provide traceable evidence for decision-making.

[0036] Compared to related technologies, traditional methods rely on single-modal data, resulting in a lack of risk assessment dimensions. This solution, however, achieves deep integration of multi-source heterogeneous data through cross-modal adversarial training. Related technologies neglect the influence of enterprise relationship networks; this solution models complex relationships between enterprises using dynamic knowledge graphs. Traditional credit scoring models lack interpretability; this solution enhances the credibility of the results through feature contribution decomposition and risk transmission path analysis.

[0037] Through the above technical solutions, this application effectively addresses, to some extent, the assessment bias caused by inconsistent feature distributions in multimodal data, and enables collaborative analysis of financial data and unstructured data. The introduction of dynamic knowledge graphs can identify hidden risk factors, such as supply chain risk transmission, that are difficult to capture using traditional methods. The scoring model considers both static creditworthiness and dynamic associated risks, improving the comprehensiveness and foresight of credit assessment results. The structured report generation mechanism transforms complex features into interpretable decision-making criteria, addressing, to some extent, the black-box nature of traditional assessment methods.

[0038] This application further proposes the following technical solutions for determining multimodal features, see [link to relevant documentation]. Figure 3 Taking the server as the executing entity as an example, the method includes the following steps.

[0039] 301. The server extracts the modal features of each modal data in the multimodal data and the semantic association weights between each modal data through a cross-modal attention mechanism, and generates an intermodal attention map of the multimodal data based on the semantic association weights between each modal data. 302. The server uses intermodal attention graphs to guide the alignment process of each modality feature, and uses gradient inversion layer and domain discriminator in adversarial training framework to process each modality feature to eliminate the distribution differences between modality features; 303. The server fuses the processed modal features to obtain the multimodal features of the multimodal data.

[0040] Among them, the cross-modal attention mechanism refers to an algorithm that dynamically captures cross-modal semantic relationships by calculating the correlation scores between data from different modalities. Specifically, it can be implemented using a multi-head attention mechanism combined with a shared semantic space mapping, to address the difficulty in quantifying semantic relationships between modalities. The inter-modal attention graph refers to graph-structured data constructed through semantic association weights, specifically implemented using a graph attention network to generate an adjacency matrix, used to characterize the interaction strength between features from different modalities. The gradient reversal layer is a network layer that keeps the data unchanged during forward propagation and reverses the gradient direction during backpropagation. Specifically, it can be implemented using a sign reversal operation combined with a backpropagation mechanism, used to drive the feature extractor to generate domain-invariant features. The domain discriminator is a classifier used to determine the modality from which features originate, specifically implemented using a multilayer perceptron combined with an attention mask mechanism, used to constrain the alignment of the distribution of features from different modalities.

[0041] Specifically, in the feature extraction stage, different modal data are first mapped to a shared semantic space. A semantic association weight matrix is ​​generated by calculating cross-modal attention scores, which dynamically reflects the semantic interaction strength between different modalities. In the feature alignment stage, the semantic association weights are used as attention masks for the domain discriminator, guiding adversarial training to focus on highly associated regions. Simultaneously, a gradient inversion layer forces the feature extractor to generate domain-invariant features independent of modality. In the feature fusion stage, a weighted summation method is used to fuse the distributed aligned modal features, where the weights are determined by the dynamic association strength of the inter-modal attention maps.

[0042] Compared to related technologies, traditional multimodal fusion methods typically employ simple concatenation or static weighting strategies, failing to consider the dynamic semantic associations and distribution differences between modalities. While existing adversarial training frameworks can eliminate domain differences, they lack focus on semantically related regions, leading to overgeneralization of key features. This approach introduces a cross-modal attention mechanism in conjunction with the attention masking mechanism of the domain discriminator. During adversarial training, it prioritizes aligning the feature distribution of semantically related regions while preserving the modality-specific features of non-related regions.

[0043] Through the above technical solutions, this application effectively addresses the issue of inconsistent feature distribution caused by the heterogeneity of multimodal data to a certain extent, achieving semantic alignment and distribution consistency optimization of cross-modal features. Specifically, the cross-modal attention mechanism dynamically captures the potential semantic relationships between modalities, the attention masking mechanism in the adversarial training framework ensures the alignment accuracy of key feature regions, and the synergistic effect of the gradient inversion layer and the domain discriminator significantly reduces the distribution differences between modalities. The resulting multimodal features retain the uniqueness of each modality while strengthening cross-modal semantic consistency, providing highly discriminative feature representations for subsequent credit assessment.

[0044] This application further proposes to extract modal features of each modality in multimodal data and semantic association weights between them through a cross-modal attention mechanism. Based on these semantic association weights, an inter-modal attention map is generated. Each modality is mapped to a shared semantic space using the attention mechanism to obtain its modal features. The correlation score between any two modal features is determined to obtain the semantic association weights between different modalities. Based on these semantic association weights, the modal features of each modality are dynamically weighted and fused to obtain the inter-modal attention map. The inter-modal attention map guides the alignment process of each modal feature, and a laddered approach is used within an adversarial training framework. The gradient inversion layer and domain discriminator process each modality feature to eliminate distribution differences between modality features. Semantic association weights are extracted from the inter-modality attention map and used as the attention mask of the domain discriminator, enabling the domain discriminator to focus on feature regions with high semantic association for domain classification. A gradient inversion layer is added between the feature extractor and the domain discriminator, and each modality feature is input into the feature extractor. The backpropagation gradient of the domain discriminant loss is inverted to drive the feature extractor to process each modality feature to obtain multiple reference modality features. The multiple reference modality features are input into the domain discriminator, and the domain discriminator performs alignment constraints on the distribution of semantically related feature regions to improve the distribution consistency among multiple reference modality features.

[0045] Cross-modal attention mechanisms refer to techniques for semantic alignment by calculating the correlation weights between features of different modalities. Specifically, this can be implemented using a multi-head attention module, which calculates cross-attention scores after mapping different modal data to a shared semantic space. Inter-modal attention maps are matrix structures representing the strength of semantic associations between features of different modalities. These can be generated by dynamically weighting and fusing features from different modalities and are used to guide subsequent feature alignment processes. Gradient reversal layers are network layers that maintain data invariance during forward propagation but reverse the gradient direction during backward propagation. These layers can be inserted between the feature extractor and the domain discriminator to drive the feature extractor to generate domain-invariant features. Domain discriminators are classifiers used to distinguish the sources of features from different modalities. These can be implemented using convolutional neural networks, focusing on highly correlated regions for classification through an attention masking mechanism.

[0046] Specifically, firstly, different modalities such as financial data, text data, and market data are input into a shared semantic space mapping module, which transforms them into feature vectors of a unified dimension through linear transformation. Then, a cross-attention mechanism is used to calculate the relevance score between financial and text data; for example, similarity between feature vectors is calculated through dot product operations, and the similarity score is normalized into semantic association weights. Based on the calculated semantic association weights, financial and text features are dynamically weighted and summed to generate an inter-modal attention map. During adversarial training, the weight matrix in the inter-modal attention map is used as the attention mask for the domain discriminator, causing the domain discriminator to classify only features in highly correlated regions. The gradient reversal layer reverses the domain discriminant loss gradient during backpropagation, forcing the feature extractor to generate reference modal features that are difficult for the domain discriminator to distinguish. The domain discriminator constrains the distribution consistency of highly correlated regions, ensuring that financial and text features have similar feature distributions in semantically related regions.

[0047] Compared to related techniques, traditional methods typically employ global feature distribution alignment strategies, leading to the destruction of feature information in semantically related regions. For example, related techniques directly use the maximum mean difference algorithm to align the global distribution of different modalities, but ignore the differences in local semantic associations between different modalities. This scheme introduces an attention masking mechanism, causing the domain discriminator to focus only on the feature distribution of highly correlated regions, thus preserving the feature specificity of non-correlated regions while achieving alignment of key semantic regions. Furthermore, the adversarial training process in related techniques lacks guidance on feature correlation, easily leading to over-alignment problems. This scheme, however, effectively avoids the loss of semantic information by dynamically adjusting the focus of adversarial training through inter-modal attention maps.

[0048] Through the above technical solutions, this application addresses to some extent the problem of inconsistent feature distributions caused by the heterogeneity of multimodal data, achieving accurate capture of cross-modal semantic associations. By sharing semantic space mapping and dynamic weighted fusion, the structural differences between different modalities are eliminated, enabling effective association between numerical features in financial data and semantic features in text data. Utilizing an attention mask-guided adversarial training mechanism, while preserving the unique information of each modality, the feature distribution of key semantic regions is aligned, avoiding feature distortion caused by traditional global alignment methods. The final generated reference modality features maintain semantic consistency between modalities while retaining the unique information of each modality, providing high-quality feature input for subsequent credit assessment.

[0049] This application further proposes the following technical solutions for determining corporate credit association characteristics, see [link to relevant documentation]. Figure 4 Taking the server as the executing entity as an example, the method includes the following steps.

[0050] 401. Generate a dynamic multimodal knowledge graph based on multimodal features and an initial knowledge graph; 402. The server extracts the association features and temporal evolution patterns of target enterprises based on dynamic multimodal knowledge graphs. The association features are used to describe the strength of the relationship between enterprises, and the temporal evolution patterns are used to describe the trend of changes in the credit status of enterprises. 403. The server determines the enterprise credit association characteristics of the target enterprise based on the association characteristics and time-series evolution patterns of the target enterprise.

[0051] Among them, the association feature refers to the features extracted from the dynamic multimodal knowledge graph used to quantify the strength of the relationship between the target enterprise and other entities. Specifically, it can be implemented using graph structure calculations, including degree centrality, proximity centrality, and betweenness centrality, and neighbor node features are aggregated through a graph neural network. This feature can explicitly quantify the risks implicit in the relationship network, such as supply chain dependence and risk spillover effects. The temporal evolution pattern refers to the pattern of credit status changes identified from the temporal changes of the dynamic multimodal knowledge graph. Specifically, it can be achieved by analyzing multiple time slice snapshots, including calculating the slope and volatility of key indicators, and encoding the node embedding sequence using a temporal graph neural network. This pattern can capture the dynamic trends of credit risk, such as early signals of a company's marginalization.

[0052] Specifically, the dynamic multimodal knowledge graph integrates multimodal features with the node attributes and edge weights of the initial knowledge graph, and adds timestamp attributes, enabling the graph to reflect the real-time status of enterprises. The message passing mechanism of the graph neural network further updates node attributes, ensuring the dynamic nature of the graph. The extraction of associated features relies on the computation of the graph's topology and the aggregation of neighbor features, such as calculating the connection strength between the target enterprise and its upstream and downstream enterprises. The temporal evolution pattern identifies trends in credit scores and event sequence patterns by analyzing snapshots of the graph at different time windows. Finally, principal component analysis reduces the dimensionality of associated features and merges them with the temporal pattern to form higher-order latent features reflecting enterprise credit risk.

[0053] Compared to related technologies, traditional methods rely on static knowledge graphs, failing to capture the dynamic changes in enterprise relationships, and their multimodal data is separated from the graph information. This solution achieves real-time evolutionary analysis of the knowledge graph by dynamically updating node attributes and edge weights, combined with timestamp attributes and graph neural networks. Furthermore, the joint modeling of association features and temporal patterns compensates for the shortcomings of static data in risk transmission and trend prediction.

[0054] Through the above technical solutions, this application addresses to some extent the problem of insufficient integration of multimodal data and knowledge graph information, enabling the complete extraction of risk dependencies in enterprise relationship networks and accurate identification of temporal changes in credit status. The construction of a dynamic multimodal knowledge graph allows for more comprehensive extraction of enterprise relationship features, and the analysis of temporal evolution patterns enhances the foresight of credit assessment, thereby reducing the risk of misjudgment due to missing information.

[0055] This application further proposes a method for generating a dynamic multimodal knowledge graph based on multimodal features and an initial knowledge graph. Specifically, this includes: updating the node attributes and edge weights of the initial knowledge graph based on entity and relationship information from the multimodal features; adding timestamp attributes to the nodes and edges in the updated initial knowledge graph to obtain a reference knowledge graph with temporal information; dynamically updating the node attributes of the reference knowledge graph through a graph neural network message passing mechanism to obtain a dynamic multimodal knowledge graph; extracting inter-enterprise association strength features and temporal change features from the dynamic multimodal knowledge graph; performing principal component analysis on the association features to obtain the first key feature; and fusing the first key feature with the temporal evolution pattern to obtain enterprise credit association features.

[0056] Among them, dynamic multimodal knowledge graphs refer to knowledge networks that integrate multimodal data and update in real time. Specifically, they can be implemented using graph databases combined with streaming data processing frameworks to address the problem of lagging data updates in traditional static knowledge graphs. Timestamp attributes refer to metadata recording the time points of data changes. Specifically, they can be implemented using distributed time-series databases to build the basic data structure for time-series evolution analysis. The message passing mechanism of graph neural networks refers to algorithms that update features through information interaction between nodes. Specifically, they can be implemented using GraphSAGE or GAT models to capture dynamic changes in enterprise relationships. Degree centrality is a measure of the number of direct connections between nodes. Specifically, it can be calculated using adjacency matrices to quantify the direct influence of enterprises in the supply chain network. Betweenness centrality refers to the importance of a node as an intermediary bridge. Specifically, it can be calculated using shortest path statistics to identify key nodes in risk transmission. Trend features are long-term directional indicators of credit status changes over time. Specifically, they can be extracted using linear regression or moving average algorithms to predict the evolution trend of credit ratings. Volatility features are short-term indicators of the magnitude of changes in credit status. Specifically, they can be calculated using standard deviation or ARCH models to assess the likelihood of sudden credit risk.

[0057] Specifically, in the knowledge graph construction phase, after the entity relationship information in the multimodal data is extracted, the node attribute values ​​and edge weight values ​​of the knowledge graph are rewritten through incremental updates. Timestamp attributes are embedded in the metadata of each node and edge, forming a time-series knowledge graph with traceable historical states. Graph neural networks iteratively update the attribute representation of the target enterprise by aggregating the features of adjacent nodes and edge weights, enabling the knowledge graph to reflect real-time changes in business relationships. In the feature extraction phase, degree centrality calculates the number of directly associated partners of the target enterprise, and betweenness centrality counts the frequency with which the enterprise acts as an information bridge in the network; both together constitute the association strength feature. Trend features capture long-term change directions by analyzing the first derivative of historical credit scores, while volatility features measure short-term uncertainty by calculating the variance of the credit score sequence. Principal component analysis orthogonally transforms the high-dimensional association features, retaining the low-dimensional projection with the largest variance as the key feature, eliminating redundant information interference. Finally, the dimensionality-reduced spatial association features and temporal evolution features are tensor-concatenated to form a high-order credit representation that simultaneously contains network topology characteristics and temporal patterns.

[0058] Compared to related technologies, traditional methods rely on static knowledge graphs, which fail to reflect changes in enterprise cooperation relationships in a timely manner. This solution, however, achieves real-time data fusion through temporal attribute annotation and dynamic updates via graph neural networks. Related technologies handle spatial relationships or time-series features individually, resulting in missing evaluation dimensions. This solution, however, constructs a multi-dimensional evaluation system by extracting both association strength and trend fluctuations. Traditional principal component analysis only reduces the dimensionality of static features, while this solution fuses spatiotemporal features before dimensionality compression, effectively preserving the interactive information of cross-modal data.

[0059] Through the above technical solutions, this application achieves deep integration of multimodal data and enterprise relationship networks, which to some extent solves the problem of single data dimension in traditional credit assessment. By using the temporal attribute annotation of dynamic knowledge graphs and the update mechanism of graph neural networks, it effectively captures the real-time changing characteristics of enterprise cooperative relationships. Combining the dual feature extraction of spatial correlation strength and temporal evolution patterns overcomes the shortcomings of static analysis methods in identifying periodic risk fluctuations. Finally, through the synergistic optimization of principal component analysis and feature fusion, it reduces data dimensionality while retaining key credit influencing factors, improving the interpretability and accuracy of the credit assessment model.

[0060] This application further proposes a method for determining credit scores, see [link to relevant documentation]. Figure 5 Taking the server as the executing entity as an example, the following steps are included.

[0061] 501. The server fuses multimodal features with enterprise credit association features to obtain the fused features of the target enterprise; 502. Based on the fusion characteristics, the server determines the static qualification score and dynamic associated risk score of the target enterprise; 503. The server determines the target company's credit score based on the target company's static qualification score and dynamic associated risk score. The credit score is used to represent the target company's static qualification and dynamic associated risk.

[0062] Multimodal features refer to a set of heterogeneous data features after eliminating distributional differences through cross-modal alignment techniques. Specifically, this can be achieved using cross-modal attention mechanisms and adversarial training frameworks, used to integrate static enterprise information from different modalities, such as financial and market data. Enterprise credit association features refer to the association network features extracted from dynamic knowledge graphs. Specifically, this can be achieved by extracting the strength of inter-enterprise associations and risk transmission paths through graph neural networks, used to characterize dynamic external influencing factors such as supply chain risks. Fusion features refer to a combined representation of multimodal features and credit association features. Specifically, this can be achieved using feature concatenation and attention-weighted fusion methods, used to construct a unified feature space containing internal and external enterprise information. Static qualification scores refer to quantitative indicators reflecting the financial health of an enterprise. Specifically, this can be achieved by calculating inherent attribute scores through linear weighting or machine learning models. Dynamic association risk scores refer to indicators that quantify the impact of external risk transmission. Specifically, this can be achieved by combining risk propagation models with topological structure analysis.

[0063] Specifically, multimodal features are aligned using cross-modal attention graphs to eliminate distributional differences between textual and financial data. Enterprise credit association features capture risk transmission paths between supply chain enterprises through dynamic knowledge graphs. During the fusion process, an attention mechanism is used to weight and combine static and dynamic features, generating fused features that incorporate both the enterprise's own strengths and the influence of the external environment. When calculating static creditworthiness scores, key indicators such as the debt-to-equity ratio and cash flow stability in the enterprise's financial statements are analyzed. When calculating dynamic association risk scores, the degree of default impact of associated enterprises is quantified based on risk transmission probability and path length. Finally, the credit score linearly combines the static and dynamic scores using preset weighting coefficients, which are dynamically adjusted according to industry risk characteristics.

[0064] Compared to related technologies, traditional credit scoring models only use financial statement data to calculate a single-dimensional score, failing to consider dynamic factors such as the transmission of supply chain default risks. This solution adds a quantitative assessment dimension of risk transmission between upstream and downstream enterprises to the scoring model by integrating dynamically related features extracted from knowledge graphs. Simple feature concatenation methods in related technologies cannot solve the problem of semantic inconsistency in multimodal data. This solution achieves feature space alignment through a cross-modal attention mechanism, enabling the fusion of textual descriptions and financial indicators within a unified semantic space.

[0065] Through the above technical solutions, this application effectively addresses, to a certain extent, the problem of biased scoring caused by the single data modality in traditional credit assessment. By integrating a company's own multimodal data with related network characteristics, it can simultaneously quantify internal operational stability and external environmental risks. The dynamic risk scoring module can identify the transmission paths of high-risk nodes in the supply chain and provide early warnings of the chain reactions caused by defaults of related companies. The feature alignment mechanism eliminates the semantic gap between different modalities of data, enabling unstructured text data and structured financial data to collaboratively participate in credit assessment decisions.

[0066] This application further proposes a method for generating credit assessment reports, see [link to relevant documentation]. Figure 6 Taking the server as the executing entity as an example, the method includes the following steps.

[0067] 601. The server extracts the second key feature that affects the credit score and its corresponding contribution based on multimodal features; 602. Based on the characteristics of enterprise credit association, the server determines the risk transmission path and risk concentration in the enterprise association network of the target enterprise; 603. The server breaks down the credit score into a static qualification score and a dynamically associated risk score; 604. Based on the second key feature, the contribution of the second key feature, the risk transmission path, the risk concentration, the static qualification score, and the dynamic correlation risk score, the server generates a structured credit assessment report. The credit assessment report includes risk source analysis, correlation impact assessment, and improvement suggestions.

[0068] The second key feature refers to the multimodal sub-features with significant causal effects on credit scoring, identified through causal inference models. This can be achieved by analyzing the direct and indirect effects of each sub-feature using a counterfactual reasoning framework, used to accurately pinpoint core influencing factors. The risk transmission path refers to the potential risk propagation links from high-risk nodes to the target enterprise in the enterprise network graph. This can be achieved using the K-shortest path algorithm combined with risk propagation probability calculations, used to identify sources of systemic risk. Risk concentration refers to the degree of risk exposure of the target enterprise in the network of connections. This can be achieved through betweenness centrality and proximity centrality indices, used to quantify the impact intensity of external associated risks. The structured credit assessment report is an analytical document containing multi-dimensional scores, risk topology graphs, and natural language descriptions. This can be generated by integrating multi-source information through a large language model, used to provide interpretable decision support.

[0069] Specifically, when generating a credit assessment report, the multimodal features are first decomposed into multiple sub-features. A causal inference model is used to quantify the direct and indirect impact of each sub-feature on the credit score, and the second key feature with high contribution is selected. Simultaneously, an enterprise association graph is constructed based on the enterprise's credit association features. A risk propagation model is used to calculate potential transmission paths, extracting the top K high-risk paths and their probabilities. The credit score is decomposed into a static qualification score reflecting the enterprise's own strength and a dynamic association risk score reflecting external risks. Subsequently, the static score is mapped to the dimensions of financial health and operational stability, and the dynamic score is mapped to the dimension of risk exposure, forming a multi-dimensional assessment matrix. Risk transmission paths and concentration information are used to generate a risk topology map with heatmap annotations using graph visualization technology. Finally, combining the contribution of key features, multi-dimensional scores, and the risk map, a large language model is used to automatically generate a structured report containing risk tracing, trend prediction, and improvement suggestions.

[0070] Compared to related technologies, traditional methods for generating credit assessment reports typically rely solely on financial data, lacking deep integration of unstructured multimodal data and failing to reveal risk transmission mechanisms within interconnected networks. Related technologies struggle to distinguish between direct and indirect factors influencing scores, resulting in reports with limited analytical dimensions. Furthermore, traditional report generation processes depend on human experience, making it difficult to visualize dynamic risk paths. This solution combines causal reasoning with graph computation, enabling not only refined attribution analysis of multimodal features but also the construction of a dynamic risk transmission model. This allows the report to encompass both internal and external risk factors within the enterprise, while automated generation technology enhances analytical efficiency.

[0071] Through the aforementioned technical solutions, this application addresses to some extent the problem of insufficient integration of multi-source information in the credit assessment report generation process, achieving interpretable traceability of key risk factors. By separating static qualification from dynamic associated risk scoring, it can more comprehensively reflect the credit status of enterprises. Quantitative analysis of risk transmission paths provides a basis for early warning of systemic risks in the associated network, while the structured report generation mechanism significantly enhances the operability and decision support value of credit assessment conclusions.

[0072] This application further proposes a technical solution based on multimodal feature extraction to identify the second key feature affecting credit scoring and its contribution, and using an enterprise association graph to identify risk transmission paths and risk concentration. Specifically, this includes: decomposing multimodal features into multiple sub-features; analyzing the direct and indirect effects of each sub-feature on credit scoring using a causal inference model; extracting key features and their contribution based on causal effects; constructing an enterprise association graph; calculating potential risk transmission paths using a risk propagation model; selecting the top K high-risk paths using a K-shortest path algorithm; and quantifying risk concentration by combining betweenness centrality and proximity centrality.

[0073] Among them, the causal inference model refers to a statistical method that analyzes the causal relationship between variables through counterfactual inference. Specifically, it can be implemented using dual machine learning or structural causal models to eliminate spurious correlations between features. The risk propagation model is a mathematical model that simulates the diffusion of risk among network nodes. Specifically, it can be implemented using contagion dynamics methods based on probabilistic graphical models to quantify the probability of risk transmission. The K-shortest path algorithm is a search algorithm that finds the first K shortest paths between two points in a graph. Specifically, it can be implemented using the Yen algorithm or the Eppstein algorithm to balance computational efficiency and critical path identification accuracy. Betweenness centrality is an indicator of the frequency of a node's appearance in all shortest paths. Specifically, it can be calculated using the Brandes algorithm to measure the pivotal role of a target enterprise in the network. Proximity centrality is an indicator of the average shortest path length from a node to other nodes. Specifically, it can be calculated using the Floyd-Warshall algorithm to assess susceptibility to risk propagation.

[0074] Specifically, after multimodal features are decomposed into sub-features such as financial, textual, and supply chain features, the causal inference model uses counterfactual inference to separate the independent impact of each sub-feature on credit scoring. For example, order fulfillment rate in supply chain data may directly affect credit scoring, while indirectly changing the score by affecting cash flow stability. By quantifying direct and indirect effects, key features with contributions exceeding a threshold are selected. In the enterprise association graph, the risk propagation model simulates the dynamic process of default events spreading along association edges, calculating the propagation probability of each path. The K-shortest path algorithm extracts the top K paths with the highest propagation probability; for example, K can be set to 5 or 10 to avoid computational complexity explosion while retaining key risk links. Betweenness centrality and proximity centrality quantify the risk exposure of target enterprises from a network topology perspective; for example, enterprises with high betweenness centrality are more likely to become transit nodes for risk propagation.

[0075] Compared to related technologies, traditional methods typically use Pearson correlation coefficients or logistic regression to analyze the correlation between features and scores, failing to distinguish between direct causality and indirect association, leading to misselection of key features. For example, company size may show a statistical correlation with credit scores, but the actual causal effect may be determined by the operational stability behind the size. Existing risk transmission analyses are mostly limited to directly related companies, ignoring network cascading effects. For instance, they only analyze the risk of first-tier suppliers, failing to trace the impact paths of second-tier suppliers or customers. Furthermore, traditional methods rely on expert experience to manually screen risk paths, lacking quantitative model support.

[0076] Through the above technical solutions, this application addresses to some extent the problem of incomplete extraction of key risk features due to the heterogeneity of multimodal data, and accurately identifies the core factors that truly influence the scoring through causal effect analysis. Simultaneously, it overcomes the deficiency of inaccurate identification of risk transmission paths in enterprise networks, achieving quantitative tracking of risk links based on dynamic network models and path search algorithms. This improves the interpretability of credit assessment results and provides a clear basis for risk tracing and prevention.

[0077] This application further proposes a method for generating structured credit assessment reports based on the second key feature, its corresponding contribution, risk transmission path, risk concentration, static creditworthiness score, and dynamic correlation risk score. Specifically, this includes: constructing a multi-dimensional assessment matrix, mapping the static creditworthiness score and dynamic correlation risk score to three assessment dimensions: financial health, operational stability, and correlation risk exposure, forming a three-dimensional score; generating a risk topology map based on the risk transmission path and risk concentration; converting the second key feature and its contribution into a structured text description; generating an initial credit assessment report based on the three-dimensional score, the risk topology map, and the structured text description; and adding credit trend predictions and improvement suggestions based on a time-series evolution pattern to the initial report.

[0078] Among these, the multi-dimensional assessment matrix refers to a quantitative tool that decomposes credit scores into multiple assessment dimensions. Specifically, it can use a weighted allocation matrix to map static creditworthiness scores to dimensions of financial health and operational stability, and dynamically correlated risk scores to dimensions of correlated risk exposure, thus addressing the problem of single assessment dimensions in traditional methods. The risk topology map refers to a visualized network structure that identifies key risk sources and transmission paths. Specifically, it can use graph neural networks to extract high-risk nodes and edge relationships, and use heatmaps to annotate high-risk areas, providing a clear visual representation of risk transmission paths. Structured text description refers to a standardized template that transforms technical features into natural language expressions. Specifically, it can use predefined descriptive frameworks combined with contribution values ​​to fill in the blanks, bridging the semantic gap between technical indicators and business interpretation. The time-series evolution pattern refers to the trend characteristics of a company's credit status over time. Specifically, it can use time series analysis models to extract trend and fluctuation characteristics, providing data support for trend prediction.

[0079] Specifically, the credit score is first decomposed into three dimensions—financial health, operational stability, and associated risk exposure—by constructing a multi-dimensional assessment matrix. The financial health score is calculated based on financial data such as debt-to-equity ratio and cash flow; the operational stability score is calculated based on indicators such as market share volatility; and the associated risk exposure score combines risk transmission probability and risk concentration indicators. Then, the top K high-risk transmission paths are extracted from the enterprise association graph, and a risk topology graph containing key nodes and edge weights is generated using graph visualization technology, with different color gradients used to indicate the risk transmission intensity. Next, the second key feature and its contribution value are input into a standardized description template to generate a structured text paragraph containing feature type and direction of influence. The scoring data from the three dimensions, the risk topology graph, and the structured text are then input into a large language model for report integration, generating an initial assessment report containing data charts and textual analysis. Finally, a predictive model is trained based on historical time-series data, and a credit trend prediction curve for the next 12 months is added to the initial report. Combined with risk transmission path analysis, specific improvement suggestions for supply chain optimization and financial structure adjustment are generated.

[0080] Compared to related technologies, existing credit assessment report generation methods typically present data in single-dimensional tables, lacking dynamic visualization of risk transmission, and separating technical indicators from business descriptions. This solution quantifies different risk types using a multi-dimensional matrix, visualizes risk transmission paths using a graph structure, and achieves natural language translation of technical features using standardized templates, thus addressing to some extent the technical shortcomings of traditional methods, such as fragmented information presentation and unintuitive risk analysis. Furthermore, by predicting credit trends through a time-series evolution model, the assessment report possesses dynamic early warning capabilities, overcoming the limitations of static assessment.

[0081] Through the aforementioned technical solutions, this application can integrate scattered, multi-dimensional assessment data into a unified structured report. It clearly visualizes risk transmission paths through visual graphs and enhances the interpretability of technical indicators using standardized text templates. Specifically, the quantitative scoring of the financial health dimension helps identify corporate financial vulnerabilities, the heatmap annotations in the risk topology graph can quickly locate high-risk supplier nodes, the contribution values ​​in the structured text descriptions clearly indicate the degree of impact of key risk factors, and the time-series forecast curves provide a forward-looking decision-making basis for risk prevention and control.

[0082] This application further proposes mapping static qualification scores and dynamic associated risk scores to three assessment dimensions: financial health, operational stability, and associated risk exposure, to form dimensional scores. A risk topology map is generated based on risk transmission paths and risk concentration. The second key feature and its contribution are transformed into structured text descriptions. An initial credit assessment report is generated based on the dimensional scores, risk topology map, and structured text. Credit trend predictions and improvement suggestions are added to the report to form a structured credit assessment report.

[0083] Financial health refers to indicators calculated from financial data, specifically using metrics such as current ratio and debt-to-equity ratio, used to assess a company's financial stability. Operational stability reflects a company's ability to continue operating, calculated using non-financial data such as market share and supply chain stability, used to measure operational risk. Related risk exposure quantifies a company's exposure to risks from related networks, calculated using a weighted average of risk transmission probability and risk concentration, used to assess the potential impact of external risk transmission. A risk topology map visualizes the risk propagation paths within a company's related networks, using a graph database to store node and edge relationships and a graph layout algorithm to generate a visual map, used to locate high-risk areas. Structured text description refers to explanatory text generated following a pre-defined template, achieved by filling feature contribution values ​​into a standardized template using a natural language generation model, used to explain the impact mechanism of key features on credit scoring.

[0084] Specifically, by decomposing static creditworthiness scores into financial health and operational stability dimensions, and combining financial and non-financial data to calculate dimensional scores separately, credit assessment is expanded from a single numerical value to a multi-dimensional quantitative system. Dynamic correlation risk scores are mapped to the correlation risk exposure dimension. By calculating the probability of risk transmission paths and the centrality index of the enterprise in the correlation network, a quantitative assessment of risk exposure is formed. When generating the risk topology map, key nodes and edge relationships of the top K high-risk transmission paths are extracted. Combined with graph visualization technology, a map with heatmap annotations is generated to visually display key hub nodes and high-risk areas in the risk propagation path. The second key feature and its contribution are input into a standardized risk description template. Interpretable structured text is generated through numerical filling; for example, the supply chain stability feature is described as "supply chain concentration is 30% higher than the industry average, resulting in a 15% decrease in operational stability score." The initial credit assessment report is automatically generated by integrating dimensional scores, risk maps, and structured text using a large language model. Finally, credit trends predicted based on time-series evolution patterns and improvement suggestions are added, forming a complete report containing data support, risk positioning, and decision-making recommendations.

[0085] Compared to related technologies, existing credit assessment methods typically compress multimodal data into a single score, failing to distinguish the impact of static creditworthiness and dynamic risk, and relying on manually written text reports leads to insufficient interpretability. This solution breaks through the limitations of traditional single-dimensional methods by decomposing complex credit indicators into a quantifiable assessment system through multidimensional scoring mapping. The risk topology map is generated using graph databases and visualization technology, providing a more intuitive display of networked risk transmission paths compared to traditional tabular data. Structured text descriptions, through standardized templates and numerical input, address the issue of ambiguous feature interpretation in traditional reports, while the introduction of a large language model enables fully automated generation from data to report, significantly improving assessment efficiency.

[0086] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0087] Through the aforementioned technical solutions, this application addresses to some extent the problem of insufficient multimodal data fusion leading to a single evaluation dimension. It achieves the separation and quantification of static and dynamic indicators through three-dimensional scoring mapping. The generation of the risk topology map enables the visualization and localization of complex enterprise-related network risks, overcoming the limitation of traditional text descriptions in failing to demonstrate network topology relationships. Structured text descriptions combined with contribution value filling clearly reveal the impact path and extent of key features on credit scoring, improving the interpretability of the evaluation results. The final credit assessment report integrates quantitative scoring, risk maps, and causal explanations, providing enterprises with a decision-making basis that includes data support, risk tracing, and improvement suggestions.

[0088] Figure 7 This is a schematic diagram of the structure of a credit assessment system based on multimodal data provided in an embodiment of this application. See also... Figure 7 The system includes: The feature extraction and fusion module 701 is used to extract and fuse features from the multimodal data of the target enterprise through cross-modal feature alignment and fusion algorithms to obtain multimodal features of the multimodal data, which includes financial data, text data, market data, supply chain data and image and video data. The feature determination module 702 is used to determine the enterprise credit association features of the target enterprise based on multimodal features and a preset initial knowledge graph. The enterprise credit association features are high-order latent features used for credit assessment, and the preset initial knowledge graph is an initial data structure used in the field of credit assessment. The credit scoring determination module 703 is used to determine the credit score of a target enterprise based on multimodal features and enterprise credit association features. The report generation module 704 is used to generate a credit assessment report for the target company based on multimodal features, corporate credit association features, and credit scores.

[0089] It should be noted that the credit assessment device based on multimodal data provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the credit assessment device based on multimodal data provided in the above embodiments and the credit assessment method embodiments based on multimodal data belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0090] Through the aforementioned technical solutions, this application addresses to some extent the problem of insufficient multimodal data fusion leading to a single evaluation dimension. It achieves the separation and quantification of static and dynamic indicators through three-dimensional scoring mapping. The generation of the risk topology map enables the visualization and localization of complex enterprise-related network risks, overcoming the limitation of traditional text descriptions in failing to demonstrate network topology relationships. Structured text descriptions combined with contribution value filling clearly reveal the impact path and extent of key features on credit scoring, improving the interpretability of the evaluation results. The final credit assessment report integrates quantitative scoring, risk maps, and causal explanations, providing enterprises with a decision-making basis that includes data support, risk tracing, and improvement suggestions.

[0091] Figure 8 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 800 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 801 and one or more memories 802. The one or more memories 802 store at least one computer program, which is loaded and executed by the one or more processors 801 to implement the methods provided in the various method embodiments described above. Of course, the server 800 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 800 may also include other components for implementing device functions, which will not be elaborated upon here.

[0092] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program that can be executed by a processor to perform the credit assessment method based on multimodal data in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0093] In an exemplary embodiment, a computer program product or computer program is also provided, which includes program code stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium and executes the program code, causing the computer device to perform the aforementioned credit assessment method based on multimodal data.

[0094] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0095] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0096] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A credit assessment method based on multimodal data, characterized in that, The method includes: Modal features and semantic association weights between various modal data in the target enterprise's multimodal data are extracted using a cross-modal attention mechanism. Based on these semantic association weights, an inter-modal attention map is generated to characterize the interaction strength between different modal features. The inter-modal attention map guides the alignment process of each modal feature, and a gradient inversion layer and domain discriminator in an adversarial training framework are used to process each modal feature to eliminate distribution differences between modal features. The processed modal features are then fused to obtain the multimodal features of the multimodal data, which includes financial data, text data, market data, supply chain data, and image / video data. Based on the multimodal features and the preset initial knowledge graph, the enterprise credit association features of the target enterprise are determined. The enterprise credit association features are high-order latent features used for credit assessment, and the preset initial knowledge graph is an initial data structure used in the field of credit assessment. The multimodal features are fused with the enterprise credit association features to obtain the fused features of the target enterprise; based on the fused features, the static qualification score and dynamic association risk score of the target enterprise are determined; based on the static qualification score and dynamic association risk score of the target enterprise, the credit score of the target enterprise is determined, wherein the credit score is used to represent the static qualification and dynamic association risk of the target enterprise, wherein the static qualification score is a quantitative indicator reflecting the financial health of the enterprise, and the dynamic association risk score is an indicator quantifying the impact of external risk transmission; Based on the multimodal features, a second key feature influencing the credit score and its corresponding contribution are extracted; based on the enterprise credit association features, the risk transmission path and risk concentration in the enterprise association network of the target enterprise are determined; the credit score is decomposed into a static qualification score and a dynamic association risk score; a multi-dimensional evaluation matrix is ​​constructed, mapping the static qualification score and the dynamic association risk score to three evaluation dimensions: financial health, operational stability, and association risk exposure, forming a three-dimensional score; based on the risk transmission path and the risk concentration, a risk topology map is generated, which is used to identify key risk sources and main risk transmission paths, where the key risk sources are risky enterprises; the second key feature and its corresponding contribution are converted into a structured text description; based on the three-dimensional score, the risk topology map, and the structured text description, an initial credit assessment report is generated; credit trend prediction and improvement suggestions based on time-series evolution patterns are added to the initial credit assessment report to generate a structured credit assessment report.

2. The method according to claim 1, characterized in that, The step of extracting modal features of each modal data in the multimodal data and semantic association weights between each modal data through a cross-modal attention mechanism, and generating an intermodal attention map of the multimodal data based on the semantic association weights between each modal data, includes: Based on the attention mechanism, each modal data is mapped to a shared semantic space to obtain the modal features of each modal data; the correlation score between any two modal features is determined to obtain the semantic association weight between different modal data; the modal features of each modal data are dynamically weighted and fused based on the semantic association weight to obtain the intermodal attention map; The process of aligning modal features using the inter-modal attention map and processing each modal feature using a gradient inversion layer and a domain discriminator in an adversarial training framework to eliminate distribution differences between modal features includes: Semantic association weights are extracted from the inter-modal attention map and used as the attention mask for the domain discriminator, enabling the domain discriminator to focus on feature regions with high semantic association for domain classification. A gradient inversion layer is added between the feature extractor and the domain discriminator, and each modality feature is input into the feature extractor. By inverting the backpropagation gradient of the domain discriminant loss, the feature extractor is driven to process each modality feature, resulting in multiple reference modality features. The multiple reference modality features are input into the domain discriminator, which performs alignment constraints on the distribution of semantically related feature regions to improve the distribution consistency among the multiple reference modality features.

3. The method according to claim 1, characterized in that, The step of determining the enterprise credit association features of the target enterprise based on the multimodal features and the preset initial knowledge graph includes: Based on the multimodal features and the initial knowledge graph, a dynamic multimodal knowledge graph is generated; Based on the dynamic multimodal knowledge graph, the association features and temporal evolution patterns of the target enterprise are extracted. The association features are used to describe the strength of the relationship between enterprises, and the temporal evolution patterns are used to describe the trend of changes in the enterprise's credit status. Based on the association characteristics and time-series evolution patterns of the target enterprise, the enterprise credit association characteristics of the target enterprise are determined.

4. The method according to claim 3, characterized in that, The process of generating a dynamic multimodal knowledge graph based on the multimodal features and the initial knowledge graph includes: Based on the entity and relation information in the multimodal features, the node attributes and edge weights of the initial knowledge graph are updated; timestamp attributes are added to the nodes and edges in the updated initial knowledge graph to obtain a reference knowledge graph with temporal information; the node attributes of the reference knowledge graph are dynamically updated through the message passing mechanism of the graph neural network to obtain the dynamic multimodal knowledge graph. The step of extracting the association features and temporal evolution patterns of the target enterprise based on the dynamic multimodal knowledge graph includes: The dynamic multimodal knowledge graph extracts inter-enterprise association strength features, including degree centrality and betweenness centrality; it also extracts temporal variation features of enterprise credit status, including trend features and fluctuation features; the inter-enterprise association strength features are identified as the association features, and the temporal variation features are identified as the temporal evolution modes. The determination of the enterprise credit association characteristics of the target enterprise based on its association characteristics and time-series evolution patterns includes: Principal component analysis was performed on the correlation characteristics of the target enterprise to obtain the first key feature related to credit assessment. The enterprise credit association feature is obtained by fusing the first key feature and the time-series evolution pattern.

5. The method according to claim 1, characterized in that, The step of extracting a second key feature affecting the credit score and its corresponding contribution based on the multimodal features includes: The multimodal features are decomposed into multiple sub-features; the causal effect of each sub-feature on the credit score is analyzed using a causal inference model. The causal effect includes direct and indirect effects. The direct effect represents the direct impact of a certain sub-feature on the credit score without the influence of other sub-features, and the indirect effect represents the indirect impact of a certain sub-feature on the credit score by influencing one or more other sub-features. Based on the causal effect of each sub-feature on the credit score, a second key feature affecting the credit score and its corresponding contribution are extracted. The second key feature belongs to the multiple sub-features. The step of determining the risk transmission path and risk concentration in the enterprise association network of the target enterprise based on the enterprise credit association characteristics includes: Based on the aforementioned corporate credit association characteristics, a corporate association graph is constructed. Nodes in the graph represent companies, edges represent inter-company relationships, and edge weights represent association strength. A path analysis algorithm based on a risk propagation model is used to calculate all potential risk transmission paths from high-risk nodes to the target company. The K shortest path algorithm is used to extract the top K risk transmission paths with the highest risk, and the risk transmission probability of each path is calculated. Based on target indicators, the risk concentration of the target company in the corporate association graph is determined. These target indicators include betweenness centrality and proximity centrality, and are used to quantify the degree to which the target company is influenced by other companies in the corporate association graph.

6. A credit assessment system based on multimodal data, characterized in that, include: The feature extraction and fusion module is used to extract modal features of each modality in the multimodal data of the target enterprise and the semantic association weights between each modality through a cross-modal attention mechanism. Based on the semantic association weights between each modality, an inter-modal attention map is generated for the multimodal data. The inter-modal attention map is used to characterize the interaction strength between different modal features. The inter-modal attention map guides the alignment process of each modality feature, and the gradient inversion layer and domain discriminator in the adversarial training framework are used to process each modality feature to eliminate the distribution differences between modality features. The processed modality features are then fused to obtain the multimodal features of the multimodal data, which includes financial data, text data, market data, supply chain data, and image and video data. The feature determination module is used to determine the enterprise credit association features of the target enterprise based on the multimodal features and the preset initial knowledge graph. The enterprise credit association features are high-order latent features used for credit assessment, and the preset initial knowledge graph is an initial data structure used in the field of credit assessment. The credit scoring determination module is used to fuse the multimodal features with the enterprise credit association features to obtain the fused features of the target enterprise; Based on the fusion features, the static qualification score and dynamic associated risk score of the target enterprise are determined; Based on the static qualification score and dynamic associated risk score of the target enterprise, the credit score of the target enterprise is determined. The credit score is used to represent the static qualification and dynamic associated risk of the target enterprise. The static qualification score is a quantitative indicator that reflects the financial health of the enterprise, and the dynamic associated risk score is an indicator that quantifies the impact of external risk transmission. The report generation module is used to extract the second key feature affecting the credit score and its corresponding contribution based on the multimodal features; determine the risk transmission path and risk concentration in the enterprise association network of the target enterprise based on the enterprise credit association features; and decompose the credit score into a static qualification score and a dynamic association risk score. A multi-dimensional assessment matrix is ​​constructed, mapping the static qualification score and the dynamic associated risk score to three assessment dimensions: financial health, operational stability, and associated risk exposure, forming a three-dimensional score. Based on the risk transmission path and the risk concentration, a risk topology map is generated, which is used to identify key risk sources and main risk transmission paths, with the key risk sources being risky enterprises. The second key feature and its corresponding contribution are converted into a structured text description. Based on the three-dimensional score, the risk topology map, and the structured text description, an initial credit assessment report is generated. Add credit trend predictions and improvement suggestions based on time-series evolution patterns to the initial credit assessment report to generate a structured credit assessment report.

Citation Information

Patent Citations

  • Commercial credit evaluation and supervision method based on multi-modal coevolution algorithm

    CN119250963A

  • Multi-modal enterprise credit risk assessment method and device based on knowledge graph

    CN120509958A