Enterprise credit risk assessment method based on big data acquisition
By building a spatiotemporal fusion engine and a cross-modal attention mechanism, combining conditional independent verification algorithms and risk conduction dynamic models, the shortcomings of multimodal data fusion and risk conduction model are solved, efficient enterprise credit risk assessment is achieved, and the accuracy and timeliness of the assessment are improved.
Patent Information
- Application Number
- CN202510214149.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has shortcomings in the fusion of multimodal data and the modeling of risk conduction mechanisms, which limits the accuracy and practicality of the evaluation method.
By building a space-time fusion engine, combining cross-modal attention mechanism for data fusion, and building a causal graph skeleton through a conditional independent verification algorithm, quantifying the intensity of causal effects, building a risk transmission dynamic model, dynamically computing the weight of risk indicators, and finally generating an enterprise credit risk score.
It realizes efficient fusion of multimodal data, improves the accuracy and robustness of feature expression, dynamically captures the conduction law of risks in the time and space dimensions, and enhances the timeliness and predictiveness of risk assessment.
Smart Images

Figure CN120146992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and in particular to an enterprise credit risk assessment method based on big data collection. Background Art
[0002] With the rapid development of big data technology, enterprise credit risk assessment methods have gradually shifted from traditional single - data - source analysis to multi - modal data fusion analysis. Multi - modal data includes financial indicators, logistics trajectories, satellite images, and industry macro - data, etc. These data can reflect the operating conditions and risk levels of enterprises from different dimensions. In recent years, credit risk assessment methods based on big data have been widely used in fields such as finance and supply chain management. For example, by analyzing financial data through machine learning models, the default risk of enterprises can be predicted; by logistics trajectory data, the supply chain stability of enterprises can be evaluated; by satellite image data, the actual operation of enterprises can be monitored; by industry macro - data, the impact of the external environment on enterprises can be analyzed. However, there are still many deficiencies in the existing technology in terms of the fusion of multi - modal data and the modeling of risk conduction mechanisms, which limit the accuracy and practicality of the assessment methods.
[0003] The main deficiencies of the existing technology are as follows: First, the fusion methods of multi - modal data usually adopt simple weighted average or linear combination, failing to fully consider the interaction relationships and spatio - temporal characteristics between different modal data, resulting in the loss or redundancy of fused feature information. Second, the existing risk conduction models are mostly based on static causal relationships, failing to effectively capture the dynamic conduction laws of risks in the time and space dimensions, resulting in the lack of timeliness and predictability of risk assessment results. For example, although traditional causal graph models can describe the causal relationships between nodes, they cannot quantify the risk conduction intensity and ignore the indirect conduction effects of intermediate nodes. In addition, the existing methods also have limitations in eliminating data distribution differences and optimizing the weights of risk indicators, and it is difficult to adapt to the complex and changeable business environment. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an enterprise credit risk assessment method based on big data collection to solve the problems of insufficient multi - modal data fusion, insufficient dynamic modeling of risk conduction, and limitations in eliminating data distribution differences and optimizing the weights of risk indicators.
[0006] To solve the above - mentioned technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides an enterprise credit risk assessment method based on big data collection, which includes collecting multimodal data and performing preprocessing, and constructing a spatio-temporal feature matrix through Kafka; the multimodal data includes financial indicators, logistics trajectories, satellite images, and industry macro data; constructing a spatio-temporal fusion engine, combining a cross-modal attention mechanism for fusion, generating spatio-temporal fusion features, and performing adversarial training through a gradient reversal layer and a domain classifier to eliminate the distribution differences of spatio-temporal fusion features; based on the spatio-temporal fusion features, using a conditional independence test algorithm to construct a causal graph skeleton, quantifying the causal effect strength between nodes through a machine learning model, constructing a risk conduction dynamic model, and predicting the risk conduction strength based on the causal effect strength; based on the risk conduction strength, calculating the risk index weights through a dynamic game network, combining the spatio-temporal fusion features to generate an enterprise risk score, and generating a dynamic risk score through a time series neural network.
[0008] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: the preprocessing includes data cleaning, data conversion, timestamp alignment, and spatial alignment.
[0009] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: the steps of constructing the spatio-temporal feature matrix through Kafka are as follows,
[0010] Adopt the Kafka sub-Topic mechanism to create independent KafkaTopics for each data type in the multimodal data;
[0011] Convert the preprocessed multimodal data into JSON format through Jackson and publish it to the corresponding KafkaTopic for each data;
[0012] In the KafkaTopic, align the multimodal data according to the time window according to the timestamp, and perform spatial alignment on the multimodal data according to the geographical location;
[0013] Based on the aligned multimodal data, extract financial indicator features, logistics trajectory features, satellite image features, and industry macro data features through feature engineering methods;
[0014] Organize all the extracted features according to the time and space dimensions to generate a spatio-temporal feature matrix.
[0015] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: the steps of constructing the spatio-temporal fusion engine are as follows,
[0016] The input layer receives the spatio-temporal feature matrix through the API;
[0017] The feature embedding layer maps all features in the spatio-temporal feature matrix to a unified vector space through FCN;
[0018] The spatio-temporal modeling layer models all the mapped features in the time dimension and the space dimension through LSTM and CNN;
[0019] A spatio-temporal fusion engine is constructed based on the input layer, the feature embedding layer and the spatio-temporal modeling layer.
[0020] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: the cross-modal attention mechanism is combined for fusion to generate spatio-temporal fusion features, and adversarial training is carried out through the gradient reversal layer and the domain classifier to eliminate the distribution difference. The specific steps are as follows.
[0021] The cross-modal interaction weights are identified through the cross-attention mechanism;
[0022] Based on the cross-modal interaction weights, the local attention features within each modality are identified through the self-attention mechanism, and the global attention features between global modalities are identified through the multi-head attention mechanism;
[0023] The weighted average is performed on the local attention features and the global attention features, and non-linear enhancement is carried out through ReLU, and finally the cross-modal attention mechanism is formed;
[0024] The spatio-temporal feature matrix is input into the spatio-temporal fusion engine, and combined with the cross-modal attention mechanism for fusion to generate the spatio-temporal fusion feature F;
[0025] The spatio-temporal fusion feature F is passed to the domain classifier through GRL. The gradient reversal layer flips the gradient during backpropagation. At the same time, the domain classifier receives the spatio-temporal fusion feature F and updates through the gradient to perform adversarial training to eliminate the distribution difference of the spatio-temporal fusion feature F.
[0026] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: based on the spatio-temporal fusion features, a causal graph skeleton is constructed by using the conditional independence test algorithm, and the causal effect strength between nodes is quantified through a machine learning model. The specific steps are as follows.
[0027] Based on the spatio-temporal fusion feature F, an initial graph is constructed assuming that all features are fully connected;
[0028] Through the chi-square test method, it is checked whether there is independence between any two features in the initial graph. If independent, the independent features are gradually eliminated, and finally a preliminary causal graph skeleton is generated;
[0029] Based on the acquisition time sequence of the multi-modal data, the undirected edges in the preliminary causal graph skeleton are oriented as directed edges to generate the final causal graph skeleton;
[0030] Each node in the causal graph skeleton represents a feature, and the edges between nodes represent the causal relationships between features;
[0031] Based on the causal graph skeleton, a structural equation model combined with a Bayesian network is used to quantify the strength of the causal effect between nodes.
[0032] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: the construction of the risk conduction dynamic model, predicting the risk conduction intensity based on the strength of the causal effect, the specific steps are as follows,
[0033] Based on the causal graph skeleton, define the conduction mechanism of risk between nodes;
[0034] Based on the strength of the causal effect between nodes, define the weight of the risk conduction coefficient;
[0035] Based on historical enterprise credit risk data, define the risk conduction coefficient;
[0036] Through dynamic weighting, combine the conduction mechanism, the weight of the risk conduction coefficient and the risk conduction coefficient to construct a risk conduction dynamic model;
[0037] Input the strength of the causal effect between nodes into the risk conduction dynamic model to predict the risk conduction intensity R ij .
[0038] As a preferred solution of the enterprise credit risk assessment method based on big data collection according to the present invention, wherein: based on the risk conduction intensity, calculate the risk index weight through a dynamic game network, generate an enterprise risk score by combining spatio-temporal fusion features, and generate a dynamic risk score through a time series neural network, the specific steps are as follows,
[0039] Take the nodes in the causal graph skeleton as game participants, and the risk conduction intensity R ij as the payoff matrix of the game, and define the strategy space of each node;
[0040] Based on the Nash equilibrium theory, identify the optimal strategy combination of the strategy space of each node;
[0041] According to the optimal strategy combination, extract the strategy strength of each node as the risk index weight;
[0042] Perform weighted fusion on the spatio-temporal fusion feature F and the risk index weight to generate a weighted spatio-temporal feature F';
[0043] Based on the weighted spatio-temporal feature F', generate an enterprise credit risk score S through a non-linear activation function;
[0044] Based on the enterprise credit risk score S, construct a time series data {S 1 , S 2 , …, S t}, and use LSTM as a time series neural network to capture the time dynamics of the risk score and generate a dynamic risk score S' t ;
[0045] The risk indicators include financial indicators, logistics track features, satellite image features, and industry macro data features.
[0046] In a second aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the enterprise credit risk assessment method based on big data collection as described in the first aspect of the present invention is implemented.
[0047] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the enterprise credit risk assessment method based on big data collection as described in the first aspect of the present invention is implemented.
[0048] The beneficial effects of the present invention are as follows: By constructing a spatio-temporal fusion engine and combining a cross-modal attention mechanism, the efficient fusion of multi-modal data is realized, and the accuracy and robustness of feature expression are improved; at the same time, by constructing a causal graph skeleton through a conditional independence test algorithm and quantifying the causal effect strength, a risk conduction dynamic model is constructed, and the conduction law of risk in the time and space dimensions is dynamically captured, enhancing the timeliness and predictability of risk assessment. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a flowchart of the enterprise credit risk assessment method based on big data collection in Embodiment 1.
[0051] Figure 2 It is a flowchart of constructing a spatio-temporal fusion engine in Embodiment 1. Detailed Embodiments
[0052] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification.
[0053] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0054] Secondly, as used herein, "an embodiment" or "embodiments" refer to specific features, structures, or characteristics that may be included in at least one implementation of the present invention. The appearances of "in an embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other.
[0055] Example 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an enterprise credit risk assessment method based on big data collection, including the following steps:
[0056] S1. Collect multimodal data and perform preprocessing, and construct a spatio-temporal feature matrix through Kafka.
[0057] The multimodal data includes financial indicators, logistics trajectories, satellite images, and industry macro data.
[0058] It should be noted that the financial indicator data is mainly extracted from the enterprise's financial statements, bank transaction records, and tax information, and is automatically collected using data interfaces or web crawler technology; the logistics trajectory data is obtained in real time through the GPS positioning system, RFID tags, and logistics management platform, recording the transportation path and status of the goods; the satellite image data is taken by remote sensing satellites and processed in combination with geographic information system (GIS) technology to obtain the geographical information and environmental changes of the enterprise's operation site; the industry macro data is collected from government statistical departments, industry association reports, and public market data, covering economic indicators, policies and regulations, and industry trends.
[0059] The preprocessing includes data cleaning, data transformation, timestamp alignment, and spatial alignment.
[0060] It should be noted that data cleaning ensures the integrity and accuracy of the data by identifying and removing outliers, missing values, and duplicate values; data transformation standardizes or normalizes the original data, unifying the data format and dimension for subsequent analysis; timestamp alignment adjusts the time series of multimodal data based on a unified time reference to ensure consistency in the time dimension; spatial alignment maps the multimodal data to a unified spatial coordinate system based on geographical location information to ensure consistency in the spatial dimension.
[0061] Adopt the partitioned Topic mechanism of Kafka to create independent KafkaTopics for each data type in the multimodal data;
[0062] The data types include numerical, text, time series, and spatial geographical types in multimodal data.
[0063] It should be noted that by using the Kafka topic partitioning mechanism, independent Kafka topics are created for financial indicators, logistics tracks, satellite images, and industry macro data in multimodal data. For example, a topic named "financial_data" is created for financial indicators, a topic named "logistics_track" is created for logistics tracks, a topic named "satellite_image" is created for satellite images, and a topic named "industry_macro" is created for industry macro data. Each data type is published to the corresponding topic by a producer, and the consumer pulls data from the message queue according to the subscribed topic, ensuring that each data type is independently stored and processed in the message queue.
[0064] The preprocessed multimodal data is converted into JSON format by Jackson and published to the corresponding Kafka topic for each type of data.
[0065] It should be noted that the preprocessed multimodal data is converted into JSON format by Jackson. For example, financial indicator data is converted into a JSON object containing fields such as timestamp, revenue, and profit; logistics track data is converted into a JSON object containing fields such as timestamp, latitude, and longitude; satellite image data is converted into a JSON object containing fields such as timestamp, image link, and coordinates; industry macro data is converted into a JSON object containing fields such as timestamp, GDP, and industry growth rate. Then, this JSON-formatted data is published to the corresponding Kafka topic. For example, financial indicator data is published to the financial data topic, logistics track data is published to the logistics track topic, satellite image data is published to the satellite image topic, and industry macro data is published to the industry macro topic.
[0066] In the Kafka topic, the multimodal data is aligned by time window according to the timestamp and spatially aligned according to the geographical location.
[0067] It should be noted that in the Kafka topic, the multimodal data is aligned by time window according to the timestamp. For example, the timestamps of financial indicators, logistics tracks, satellite images, and industry macro data are unified to a specific time window to ensure that the data within the same time window can be processed synchronously. The multimodal data is spatially aligned according to the geographical location. For example, the longitude and latitude of the logistics track are matched with the coordinates of the satellite image to ensure that the data at the same geographical location can be analyzed in association.
[0068] Based on the aligned multimodal data, feature engineering is used to extract financial indicator features, logistics trajectory features, satellite image features, and industry macro data features;
[0069] It should be noted that the financial indicator characteristics are extracted by calculating key financial ratios such as revenue growth rate and profit margin; the logistics trajectory characteristics are extracted by analyzing behavioral patterns such as transportation routes, residence time and speed changes; the satellite image characteristics are extracted by identifying geographic information such as building density, vegetation coverage and land use type; the industry macro data characteristics are extracted by counting macro indicators such as GDP growth rate, industry prosperity index and policy impact.
[0070] All extracted features are organized according to time and space dimensions to generate a spatiotemporal feature matrix.
[0071] It should be noted that the extracted financial indicator characteristics, logistics trajectory characteristics, satellite image characteristics and industry macro data characteristics are integrated according to timestamps and geographic locations. For example, the financial indicator characteristics (such as revenue growth rate), logistics trajectory characteristics (such as transportation routes), satellite image characteristics (such as building density) and industry macro data characteristics (such as GDP growth rate) located in a certain longitude and latitude area within the time window of 12:00:00 on October 1, 2023 are combined into a multidimensional vector and arranged according to time series and spatial distribution to generate a spatiotemporal feature matrix.
[0072] S2. Build a spatiotemporal fusion engine, combine it with the cross-modal attention mechanism for fusion, generate spatiotemporal fusion features, and perform adversarial training through the gradient reversal layer and domain classifier to eliminate the distribution differences of spatiotemporal fusion features.
[0073] The input layer receives the spatiotemporal feature matrix through the API;
[0074] The feature embedding layer maps all features in the spatiotemporal feature matrix to a unified vector space through FCN;
[0075] It should be noted that the feature embedding layer maps the financial indicator features, logistics trajectory features, satellite image features and industry macro data features in the spatiotemporal feature matrix to a unified vector space through a fully connected network. For example, features such as income growth rate, transportation path, building density and GDP growth rate are respectively input into the fully connected network, and after linear transformation and nonlinear activation function processing, they are converted into vector representations of the same dimension to ensure that all features are comparable and computable in a unified vector space, which is convenient for subsequent spatiotemporal modeling and feature fusion.
[0076] The spatiotemporal modeling layer uses LSTM and CNN to model all mapped features in the temporal and spatial dimensions.
[0077] It should be noted that in the spatio-temporal modeling layer, the mapped features are first input into the LSTM to capture the changing trends and patterns of the data over time, ensuring the ability to understand and remember the long-term dynamic changes of the data. At the same time, the CNN is used to process the spatial dimensions of these features, scanning the input features through convolutional kernels to identify local patterns and structures in the spatial distribution. For example, when analyzing logistics trajectory data, the LSTM can track the change of the moving path of goods from the shipping point to the receiving point over time, while the CNN focuses on the correlations and patterns between different geographical locations, such as more frequent transportation activities between certain regions. The combination of the two enables the spatio-temporal fusion engine to not only understand the spatial distribution characteristics at a single time point, but also grasp how these characteristics evolve over time, thus comprehensively capturing the time dynamic changes and spatial distribution laws of the data.
[0078] A spatio-temporal fusion engine is constructed based on the input layer, the feature embedding layer, and the spatio-temporal modeling layer.
[0079] It should be noted that the spatio-temporal fusion engine constructed based on the input layer, the feature embedding layer, and the spatio-temporal modeling layer first maps the features from different data sources into a common space for subsequent processing. Within this framework, the cross-attention mechanism is used to identify the interaction weights between modalities, which means that for each pair of different data types (such as financial indicators and logistics trajectories), their mutual importance is evaluated to determine which modality has a greater impact on the other modality.
[0080] Identify the interaction weights between modalities through the cross-attention mechanism;
[0081] It should be noted that based on the interaction weights between modalities, the local attention features within each modality are calculated through the self-attention mechanism. This step focuses on each individual data type and identifies the importance of the key parts or elements within it. For example, when examining satellite image data, the changes in certain specific regions may be more indicative of the changes in the enterprise's operating conditions than those in other regions, so these regions will be assigned higher attention values. At the same time, the multi-head attention mechanism is used to identify the global attention features between global modalities.
[0082] Based on the interaction weights between modalities, identify the local attention features within each modality through the self-attention mechanism, and identify the global attention features between global modalities through the multi-head attention mechanism;
[0083] It should be noted that the weighted average is performed based on local attention features and global attention features, and non-linear enhancement is performed through the ReLU activation function, ultimately constituting a cross-modal attention mechanism. The process here involves comprehensively considering the relationships within and between all modalities to generate a feature representation that integrates various key information. For example, when evaluating enterprise risks, such comprehensive features can reflect key factors and their interrelationships in multiple aspects such as financial indicators, logistics trajectories, satellite images, and industry macro data, providing a comprehensive and in-depth risk perspective.
[0084] Perform a weighted average on the local attention features and global attention features, and perform non-linear enhancement through ReLU, ultimately constituting a cross-modal attention mechanism;
[0085] It should be noted that first, a weighted average is performed on the local attention features and global attention features. This process aims to balance the importance of internal details within different modalities and the interactions between modalities. For example, when evaluating enterprise credit risks, the subtle changes in logistics trajectories (local features) and the correlations between financial indicators and other data types (global features) are comprehensively considered according to their respective importance. Subsequently, non-linear enhancement is performed on the weighted average features through the ReLU activation function. The result of such processing is a strengthened feature representation that not only contains the key information of each modality but also reflects the complex interaction patterns between these modalities, thus providing a solid foundation for subsequent enterprise risk assessment.
[0086] Input the spatio-temporal feature matrix into the spatio-temporal fusion engine, and perform fusion in combination with the cross-modal attention mechanism to generate the spatio-temporal fusion feature F, and the expression is:
[0087]
[0088] Among them, H 1 is the local attention feature identified by the self-attention mechanism, H 2 is the global attention feature identified by the multi-head attention mechanism, W 1 is the local attention feature weight matrix, W 2 is the weight matrix of the global attention feature, b 1 is the bias term of the local attention feature, b 2 is the bias term of the global attention feature, and F is the spatio-temporal fusion feature;
[0089] Pass the spatio-temporal fusion feature F to the domain classifier through GRL. The gradient reversal layer flips the gradient during backpropagation. At the same time, the domain classifier receives the spatio-temporal fusion feature F and updates through the gradient to perform adversarial training to eliminate the distribution differences of the spatio-temporal fusion feature F.
[0090] It should be noted that the spatio-temporal fusion feature F is passed to the domain classifier through the Gradient Reversal Layer (GRL). During the backpropagation process, the Gradient Reversal Layer flips the received gradients, enabling the domain classifier to distinguish different domain features while forcing the spatio-temporal fusion feature F to learn a more domain-invariant representation. Specifically, when dealing with enterprise credit risk assessment, if the spatio-temporal fusion feature F contains data from different industries or regions, the goal of the domain classifier is to identify these differences. However, due to the effect of the Gradient Reversal Layer, the learning process actually minimizes the feature differences between these domains, promoting the generated spatio-temporal fusion feature F to be more generally applicable to different scenarios without being limited by the distribution characteristics of specific domains. In this way, after adversarial training, the distribution differences in the spatio-temporal fusion feature F can be effectively eliminated.
[0091] S3. Based on the spatio-temporal fusion feature, construct the causal graph skeleton using the conditional independence test algorithm, quantify the causal effect strength between nodes through a machine learning model, construct a risk conduction dynamic model, and predict the risk conduction strength based on the causal effect strength.
[0092] Based on the spatio-temporal fusion feature F, assume that all features are fully connected to construct an initial graph;
[0093] It should be noted that based on the spatio-temporal fusion feature F, an initial graph is first constructed by assuming that all features are fully connected. This means that each feature is considered to potentially be associated with all other features, forming a fully connected network. For example, when analyzing urban traffic conditions, this includes all features from different data sources, such as bus GPS trajectories, taxi operation data, the popularity of discussions on traffic congestion on social media, and meteorological conditions, etc., with each item directly connected to all others to form the initial graph.
[0094] Through the chi-square test method, check whether there is independence between any two features in the initial graph. If independent, gradually remove the independent features to finally generate a preliminary causal graph skeleton;
[0095] It should be noted that for any two features in the initial graph, calculate the degree of dependence between them. If it is found that a certain pair of features is statistically independent, that is, there is no significant correlation or causal relationship between them, then gradually remove the connection between these two features. For example, when analyzing enterprise credit risk, assume that after the chi-square test, the logistics trajectory and a specific industry macro data indicator are shown to be independent, then the direct connection between the two will be removed. In this way, continuously streamline the connections in the initial graph until all the remaining connections represent statistically significant relationships, finally forming a more concise and effective preliminary causal graph skeleton.
[0096] Orient the undirected edges in the preliminary causal graph skeleton as directed edges based on the chronological order of multi-modal data collection to generate the final causal graph skeleton;
[0097] It should be noted that by analyzing the relationship between data changes over time among different features, it is determined which features change first and which change later. For example, when examining the relationship between consumer behavior and sales volume, if it is found that consumer behavior changes first after a specific advertisement campaign and then the sales volume changes accordingly, it can be inferred that there is a causal flow from the advertisement campaign to the sales volume. Thus, the undirected edge representing the association between the two is oriented as a directed edge pointing from the advertisement campaign to the sales volume. And so on, gradually assign directions to each undirected edge in the causal graph skeleton to ensure that the entire graph can accurately reflect the causal relationships and influence paths among features, ultimately forming a causal graph skeleton with clear directionality.
[0098] Each node in the causal graph skeleton represents a feature, and the edges between nodes represent the causal relationships between features;
[0099] Based on the causal graph skeleton, use structural equation models combined with Bayesian networks to quantify the intensity of causal effects between nodes. The expression is:
[0100]
[0101] where, C(X i →X j ) represents the total causal effect intensity of node X i on node X j , β ij represents the regression coefficient of node X i on node X j , β kj represents the regression coefficient of node X k on node X j , is the partial derivative symbol, represents the change in node X k , represents the change in node X i .
[0102] Based on the causal graph skeleton, define the risk conduction mechanism between nodes;
[0103] It should be noted that, first of all, based on the directed edges in the causal graph skeleton, identify which nodes are the direct or indirect causes of other nodes, thereby establishing a risk transmission path. For example, in enterprise credit risk assessment, if it is found that changes in financial health can affect the stability of the enterprise's supply chain, and the supply chain stability further affects the enterprise's market performance, then a risk transmission link from financial health to supply chain stability and then to market performance can be defined. Then, according to these paths, analyze how risks spread from one node to another along the chain, and consider the possible amplification or attenuation effects of intermediate nodes on risks, so as to construct a comprehensive risk transmission mechanism.
[0104] Define the weights of the risk transmission coefficients based on the causal effect intensity between nodes;
[0105] It should be noted that first, it is necessary to evaluate the strength of the causal relationship represented by each directed edge. For example, in enterprise credit risk analysis, if changes in the logistics track have a significant impact on financial health, then this edge will be assigned a higher weight, indicating that it plays an important role in the risk transmission process. Specifically, by quantifying the causal effect intensity between each pair of directly connected nodes, determine the impact strength when risks are transmitted from one node to another. This process takes into account the interaction between features and their potential influence differences, ensuring that those paths with stronger causal effects occupy a more important position in the risk transmission network. Therefore, according to the different causal effect intensities between nodes, assign corresponding weights to each edge.
[0106] Define the risk transmission coefficients based on historical enterprise credit risk data;
[0107] It should be noted that first, it is necessary to analyze past enterprise credit events and their impact paths to identify the specific patterns of risk transmission between different features. For example, when examining how past financial distress affects the enterprise's market performance, by analyzing the changes in various indicators before and after the enterprise encounters a financial crisis, determine the impact mode and degree of the deterioration of financial health on supply chain stability, customer trust, and ultimately market share. In this way, use the data in a large number of historical cases to measure the risk transmission effect from one node to another, so as to define specific risk transmission coefficients for each directed edge.
[0108] Through dynamic weighting, combine the transmission mechanism, the weights of the risk transmission coefficients, and the risk transmission coefficients to construct a risk transmission dynamic model;
[0109] First, determine the basic conduction framework according to the defined risk conduction paths and the risk conduction coefficients on each path. Then, adjust the weight of each path based on the causal effect intensity between different nodes to ensure that those paths with stronger causal associations play a greater role in risk propagation. For example, when analyzing the impact of the internal financial health of an enterprise on its overall market performance, not only consider the impact brought by direct changes in financial data, but also comprehensively consider the changes in multiple factors such as supply chain stability and customer satisfaction and their interactions on the final result. Through this dynamic weighting method, the model can flexibly reflect the differences in the importance of different factors over time, thus more accurately simulating the actual propagation process of risks in complex networks, and forming a risk conduction dynamic model that considers both static structures and adapts to dynamic changes.
[0110] Input the causal effect intensity between nodes into the risk conduction dynamic model to predict the risk conduction intensity. The expression is:
[0111] R ij =σ(w ij ·α ij ·C(X i →X j )+∑ k≠i γ k ·R kj );
[0112] Where, R ij represents the risk conduction intensity from node X i to node X j , w ij is the dynamic weight of the risk conduction coefficient, α ij is the risk conduction coefficient from node X i to node X j , γ k represents the indirect risk conduction intensity transmitted through the intermediate node X k , R kj represents the risk conduction intensity from node X k to node X j , and σ is a non-linear activation function.
[0113] S4. Based on the risk conduction intensity, calculate the risk index weights through a dynamic game network, generate an enterprise risk score by combining spatio-temporal fusion features, and generate a dynamic risk score through a time series neural network.
[0114] Take the nodes in the causal graph skeleton as game participants, and the risk conduction intensity R ij between nodes as the payoff matrix of the game, and define the strategy space of each node;
[0115] It should be noted that when the nodes in the causal diagram skeleton are regarded as game participants, each node represents a specific feature or variable, such as the financial health of an enterprise, the stability of the supply chain, etc. The risk conduction intensity R ij serves as the payoff matrix in the game, which quantifies the degree of mutual influence between different nodes. For example, when evaluating the credit risk of an enterprise, if the change in the logistics track significantly affects the financial health, the risk conduction intensity between these two nodes reflects the magnitude of this influence and serves as the payoff value for the game between them. Based on this payoff matrix, a strategy space is defined for each node, that is, each node can choose different actions or states to respond to the risk conduction from other nodes. These strategies include adjusting resource allocation, changing management strategies, or other measures that can mitigate the impact of risks. In this way, the dynamic interaction between various features and its impact on the overall risk are simulated and analyzed using game theory methods.
[0116] Based on the Nash equilibrium theory, identify the optimal strategy combination in the strategy space of each node;
[0117] It should be noted that in the context of enterprise credit risk assessment, assuming that each node represents different aspects of an enterprise, such as financial health, supply chain stability, and market performance, etc., each node has its specific strategy space, including various possible actions to respond to risks. By analyzing the interaction between nodes and the risk conduction intensity, it can be determined that when all nodes have chosen a certain specific strategy, there is no node that can benefit by unilaterally changing its own strategy, and this is the Nash equilibrium point. For example, if the interaction between the financial health node and the supply chain stability node of an enterprise reaches a balance, that is, the efforts of both parties in risk management make it impossible for either party to further reduce risks or increase benefits through additional unilateral adjustments, then this strategy combination is considered the optimal strategy combination reaching the Nash equilibrium.
[0118] According to the optimal strategy combination, extract the strategy intensity of each node as the risk index weight;
[0119] It should be noted that in enterprise credit risk analysis, if a certain node (such as financial health) significantly reduces the impact of risk conduction to other nodes (such as supply chain stability or market performance) under its optimal strategy, then the strategy intensity of this node reflects its importance in overall risk management. In this way, based on the strategy effects of each node in the equilibrium state, its contribution to overall risk control is quantified, and then this contribution or strategy intensity is converted into the corresponding risk index weight. For example, if the financial health of an enterprise can greatly alleviate the risks brought by market fluctuations under the optimal strategy, then the risk index weight of this feature will be relatively high, indicating that it occupies a more critical position in enterprise risk assessment.
[0120] The spatio-temporal fusion feature F and the risk index weights are weighted and fused to generate a weighted spatio-temporal feature F', and the expression is:
[0121]
[0122] where w v is the weight coefficient of the v-th risk index, and n is the number of features participating in the weighted fusion;
[0123] Based on the weighted spatio-temporal feature F', a corporate credit risk score is generated through a non-linear activation function, and the expression is:
[0124] S = Sigmoid(W 3 ·F'+b);
[0125] where S is the corporate credit risk score, W 3 is the weight matrix of the weighted spatio-temporal feature, and b 3 is the bias term of the weighted spatio-temporal feature;
[0126] Based on the corporate credit risk score S, a time series data {S 1 , S 2 , …, S t} is constructed. Using LSTM as a time series neural network to capture the time dynamic changes of the risk score and generate a dynamic risk score, and the expression is:
[0127] S' t = LSTM(S t-1 , S t-2 , …, S 1 );
[0128] where S' t represents the risk score generated at time step t, S t-1 is the risk score sequence from time step 1 to t - 1, and t is the current time step, indicating the time point at which the dynamic risk score needs to be generated;
[0129] The risk indicators include financial indicators, logistics track features, satellite image features, and industry macro data features.
[0130] This embodiment also provides a computer device applicable to the situation of the corporate credit risk assessment method based on big data collection, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the corporate credit risk assessment method based on big data collection as proposed in the above embodiment.
[0131] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, carrier networks, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0132] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the enterprise credit risk assessment method based on big data collection as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, magnetic disks, or optical discs.
[0133] In summary, through the following steps: constructing a spatio-temporal fusion engine and combining it with a cross-modal attention mechanism, the present invention realizes the efficient fusion of multi-modal data, improves the accuracy and robustness of feature expression; at the same time, by constructing a causal graph skeleton through a conditional independence test algorithm and quantifying the causal effect strength, a risk conduction dynamic model is constructed, which dynamically captures the conduction law of risks in the time and space dimensions, and enhances the timeliness and predictability of risk assessment.
[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for assessing corporate credit risk based on big data collection, characterized by: include, Collect and preprocess multimodal data, and build a spatiotemporal feature matrix through Kafka; the multimodal data includes financial indicators, logistics trajectories, satellite images, and industry macro data; Build a spatiotemporal fusion engine, combine it with the cross-modal attention mechanism to generate spatiotemporal fusion features, and perform adversarial training through the gradient reversal layer and domain classifier to eliminate the distribution differences of spatiotemporal fusion features; Based on the spatiotemporal fusion features, the conditional independence test algorithm is used to construct the causal graph skeleton, the causal effect strength between nodes is quantified through the machine learning model, and the risk conduction dynamic model is constructed to predict the risk conduction strength based on the causal effect strength; Based on the intensity of risk transmission, the risk indicator weights are calculated through a dynamic game network, and the enterprise risk score is generated by combining the time-space fusion characteristics, and the dynamic risk score is generated through a time series neural network.
2. The enterprise credit risk assessment method based on big data collection as claimed in claim 1, characterized in that: The preprocessing includes data cleaning, data conversion, timestamp alignment and space alignment.
3. The enterprise credit risk assessment method based on big data collection as claimed in claim 2, characterized in that: The specific steps of constructing the spatiotemporal feature matrix through Kafka are as follows: Adopt Kafka's topic-based mechanism to create an independent Kafka topic for each data type in multimodal data; The preprocessed multimodal data is converted into JSON format through Jackson and published to the KafkaTopic corresponding to each type of data; In KafkaTopic, multimodal data is aligned in time windows based on timestamps, and multimodal data is spatially aligned based on geographic locations; Based on the aligned multimodal data, feature engineering is used to extract financial indicator features, logistics trajectory features, satellite image features, and industry macro data features; All extracted features are organized according to time and space dimensions to generate a spatiotemporal feature matrix.
4. The enterprise credit risk assessment method based on big data collection as claimed in claim 3, characterized in that: The specific steps of constructing the space-time fusion engine are as follows: The input layer receives the spatiotemporal feature matrix through the API; The feature embedding layer maps all features in the spatiotemporal feature matrix to a unified vector space through FCN; The spatiotemporal modeling layer uses LSTM and CNN to model all mapped features in the temporal and spatial dimensions. A spatiotemporal fusion engine is built based on the input layer, feature embedding layer, and spatiotemporal modeling layer.
5. The enterprise credit risk assessment method based on big data collection as claimed in claim 4, characterized in that: The cross-modal attention mechanism is combined for fusion to generate spatiotemporal fusion features, and adversarial training is performed through the gradient reversal layer and domain classifier to eliminate distribution differences. The specific steps are as follows: Identify the interaction weights between modalities through the cross-attention mechanism; Based on the interaction weights between modalities, the local attention features within each modality are identified through the self-attention mechanism, and the global attention features between global modalities are identified through the multi-head attention mechanism; The local attention features and global attention features are weighted averaged and nonlinearly enhanced through ReLU to finally form a cross-modal attention mechanism; Input the spatiotemporal feature matrix into the spatiotemporal fusion engine, combine it with the cross-modal attention mechanism to generate the spatiotemporal fusion feature F; The spatiotemporal fusion feature F is passed to the domain classifier through GRL, and the gradient reversal layer flips the gradient during back propagation. At the same time, the domain classifier receives the spatiotemporal fusion feature F and performs adversarial training through gradient update to eliminate the distribution difference of the spatiotemporal fusion feature F.
6. The enterprise credit risk assessment method based on big data collection as claimed in claim 5, characterized in that: Based on the spatiotemporal fusion features, the conditional independence test algorithm is used to construct the causal graph skeleton, and the causal effect strength between nodes is quantified through the machine learning model. The specific steps are as follows: Based on the spatiotemporal fusion feature F, assuming that all features are fully connected, an initial graph is constructed; The chi-square test method is used to test whether there is independence between any two features in the initial graph. If so, the independent features are gradually eliminated to finally generate a preliminary causal graph skeleton. Based on the time sequence of multimodal data collection, the undirected edges in the preliminary causal graph skeleton are oriented to directed edges to generate the final causal graph skeleton; Each node in the causal graph skeleton represents a feature, and the edges between nodes represent the causal relationship between features; Based on the causal graph skeleton, the structural equation model combined with the Bayesian network is used to quantify the strength of the causal effects between nodes.
7. The enterprise credit risk assessment method based on big data collection as claimed in claim 6, characterized in that: The risk transmission dynamic model is constructed to predict the risk transmission intensity based on the causal effect intensity. The specific steps are as follows: Based on the causal graph skeleton, define the risk transmission mechanism between nodes; Based on the strength of the causal effect between nodes, the weight of the risk transmission coefficient is defined; Based on historical corporate credit risk data, define the risk transmission coefficient; Through dynamic weighting, the transmission mechanism, the weight of the risk transmission coefficient and the risk transmission coefficient are combined to construct a risk transmission dynamic model; The causal effect strength between nodes is input into the risk transmission dynamic model to predict the risk transmission strength R ij .
8. The enterprise credit risk assessment method based on big data collection as claimed in claim 7, characterized in that: Based on the risk transmission intensity, the risk indicator weight is calculated through the dynamic game network, the enterprise risk score is generated by combining the spatiotemporal fusion characteristics, and the dynamic risk score is generated through the time series neural network. The specific steps are as follows: The nodes in the causal graph skeleton are regarded as game participants, and the risk transmission intensity R between nodes ij As the payoff matrix of the game, define the strategy space of each node; Based on Nash equilibrium theory, identify the optimal strategy combination in the strategy space of each node; According to the optimal strategy combination, the strategy strength of each node is extracted as the risk indicator weight; Perform weighted fusion of the spatiotemporal fusion feature F and the risk indicator weight to generate the weighted spatiotemporal feature F'; Based on the weighted spatiotemporal features F', the enterprise credit risk score S is generated through a nonlinear activation function; Based on the enterprise credit risk score S, construct the time series data {S1,S2,…,S t }, LSTM is used as a temporal neural network to capture the temporal dynamic changes of risk scores and generate dynamic risk scores S' t ; The risk indicators include financial indicators, logistics trajectory characteristics, satellite image characteristics and industry macro data characteristics.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the enterprise credit risk assessment method based on big data collection according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the enterprise credit risk assessment method based on big data collection according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Enterprise technology achievement evaluation method, device, equipment and medium
CN120764862A
Insurance online payment transaction security real-time detection method and system
CN120765247A
Credit assessment method and system based on multi-modal data
CN121213229A
Intelligent financing debt management system and method based on multi-modal data
CN121707701A