Multi-source heterogeneous small and micro enterprise operation data fusion credit evaluation method
Patent Information
- Application Number
- CN202611215520.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-12
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]本发明的目的在于提供一种多源异构小微企业经营数据融合信用评估方法,用于解决现有技术难以从异构数据流交互中量化经营失序信号,以及无法依据失序演化趋势动态校正信用评分置信区间的问题
通过在每个滑动窗口内计算不同数据源之间的双向耦合扰动势,构建出扰动传递矩阵。该矩阵刻画了企业经营中不同维度活动之间的波动传导关系,其中任意两个数据源间的互协方差与自协方差比值被量化为单向扰动强度,交换方向后构成完整的双向交互度量。这一过程将传统的孤立特征分析转变为对数据流间动态耦合结构的建模。当企业经营状态发生变化时,各数据源间的扰动传递关系随之改变,这种改变直接反映在非对角元素的离散程度上。从扰动传递矩阵的非对角元素集合提取信息熵,生成核心经营熵序列。核心经营熵量化了多源数据交互结构的无序程度,其数值越高,表明企业内部不同经营活动间的协同性越弱、紊乱度越高。通过持续追踪每个滑动窗口位置上的核心经营熵,将企业信用基础的微观侵蚀过程转化为一个可观测的单调序列,使评估不再依赖静态的财务比率,而是直接捕获经营结构从有序向无序演变的动态信号。这种从跨源扰动关系中直接萃取结构失序度量的方式,实现了对企业隐性信用风险的深层表征。构建以核心经营熵为因变量的二阶差分方程,拟合信用衰减动力学模型,得到内在信用势能。从核心经营熵序列的前置窗口子序列拟合线性回归直线,提取初始衰减速率,从后置窗口子序列计算相邻元素差值的均值,作为衰减加速度。衰减加速度反映了经营失序过程是趋于加速恶化还是在收敛改善,它约束了二阶差分方程的二阶导数项,使得衰减曲线函数并非简单的惯性外推,而是受内部结构演化快慢调节的轨迹。信用衰减动力学模型将风险判断从静态阈值比较提升为对信用状态动力学行为的推演,内在信用势能包含了衰减的历史速率和加速度信息。进一步根据内在信用势能在末端窗口的漂移方向与幅度,计算置信区间漂移校正因子。末端窗口是由多源数据最新时刻截取的子矩阵,该子矩阵的整体能量与衰减曲线函数在最新位置的预测值之间存在偏差。计算这一偏差的方向和幅度,将其作为校正因子直接叠加到由衰减曲线函数累积得到的内在信用势能基准值上。这种末端能量相对模型预测的漂移校正,使得信用评分在长期趋势预测基础上,能够对最新的多源联合状态突变做出即时响应,避免了单纯依赖模型外推带来的滞后误判。
Smart Images

Figure CN122736348A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing and credit assessment technology, specifically a method for credit assessment by fusing multi-source heterogeneous micro and small enterprise operating data. Background Technology
[0002] Credit assessment for micro and small enterprises has long faced the problems of data fragmentation and information silos. Existing multi-source data fusion assessment techniques typically involve simply piecing together operational data from different channels and inputting it into a static scoring model. This approach fails to capture the transmission effect of dynamic, non-linear interactive disturbances between different data streams on the enterprise's credit status. When a fluctuation occurs in one operational dimension, its chain reaction on another dimension cannot be quantitatively characterized, resulting in assessment results that are insensitive to local risk signals. Furthermore, existing methods often rely on fitting historical data to a fixed credit trend curve, failing to fully consider the directional changes in the degree of disorder in the enterprise's internal operational structure over time. This structural disorder reflects the degradation of the coordination of the enterprise's multi-source operational activities, essentially a decline in the enterprise's endogenous credit driving force. However, conventional techniques struggle to extract quantitative indicators characterizing this disorder from multimodal time-series data, nor can they transform it into a dynamic mechanism describing the speed and acceleration of credit decay. Therefore, how to decipher the signal of dissipation in the overall operational order of the enterprise from subtle disturbances in cross-source data streams, and dynamically adjust the confidence range of credit prediction based on the accelerating decay trend of the signal, has become a current challenge. The problem this method aims to solve is to quantify the fluctuation transmission effect between heterogeneous data sources to clarify the degree of disorder in the business structure, and to inject the disorder trend into the credit decay process to obtain a credit score with drift awareness. Summary of the Invention
[0003] The purpose of this invention is to provide a credit assessment method that integrates multi-source heterogeneous business data of micro and small enterprises, in order to solve the problems of existing technologies that make it difficult to quantify business disorder signals from heterogeneous data stream interactions and cannot dynamically correct the confidence interval of credit scores based on the trend of disorder evolution.
[0004] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a multi-source heterogeneous micro and small enterprise operating data fusion credit assessment method. This method collects digital asset records of the target enterprise from multiple heterogeneous data sources, including transaction records, tax invoices, electricity consumption, and social security payment records. All digital asset records are aligned along a time axis to construct a multimodal data flow matrix with time as the horizontal axis and data source type as the vertical axis. A sliding window is performed on the multimodal data flow matrix, and the bidirectional coupling perturbation potential between different data sources is calculated within each sliding window to obtain a perturbation transfer matrix. Based on the distribution dispersion of the off-diagonal elements in the perturbation transfer matrix, the core operating entropy sequence of the target enterprise is extracted. Based on the decay gradient of the core operating entropy sequence within continuous sliding windows, a credit decay dynamics model of the target enterprise is fitted to obtain the intrinsic credit potential. According to the drift direction and amplitude of the intrinsic credit potential in the final window of the multimodal data flow matrix, a confidence interval drift correction factor is calculated. The confidence interval drift correction factor is used to correct the intrinsic credit potential, generating a credit score.
[0005] As a preferred technical solution of the present invention, the process of constructing a multimodal data flow matrix includes: extracting the timestamp field from each digital asset record, arranging all digital asset records in ascending order according to the timestamp field to generate a unified timeline; determining the total time span based on the minimum and maximum timestamps in the unified timeline, and dividing the total time span into multiple equally spaced time slices; within each time slice, extracting numerical features from the digital asset records of each heterogeneous data source, and filling the corresponding cells of the initial matrix with the numerical features; performing linear interpolation to fill the cells in the initial matrix that are missing numerical features, thereby obtaining a multimodal data flow matrix indexed by time on the horizontal axis and by data source type on the vertical axis. This method ensures that business records from different sources and with different collection frequencies are strictly aligned in the time dimension, eliminating deviations caused by inconsistencies in the timeline, and providing a well-organized data foundation for subsequent multimodal coupling analysis.
[0006] Preferably, the process of sliding window cutting and perturbation transfer matrix generation includes: setting a sliding window of fixed length, moving the sliding window along the horizontal axis of the multimodal data flow matrix with a fixed step size, and extracting a sub-matrix at each sliding window position; in each sub-matrix, selecting row vectors corresponding to any two different data source types, calculating the conditional fluctuation transfer rate of the first row vector relative to the second row vector, wherein the conditional fluctuation transfer rate is equal to the cross-covariance of the two row vectors in the sub-matrix divided by the autocovariance of the second row vector, using the conditional fluctuation transfer rate as the unidirectional perturbation intensity from the second data source to the first data source, and calculating the reverse unidirectional perturbation intensity by swapping the roles of the two row vectors; combining the bidirectional unidirectional perturbation intensities between all different data source pairs in the sub-matrix into a square matrix to obtain the perturbation transfer matrix. The bidirectional coupling perturbation potential is further calculated based on mutual information or transfer entropy to more comprehensively measure the degree of nonlinear information transfer and coupling between data sources and improve the sensitivity to capturing abnormal operational fluctuations.
[0007] Preferably, the step of extracting the core operating entropy sequence specifically involves: extracting all off-diagonal elements in the perturbation transfer matrix whose row indices are not equal to their column indices, forming an off-diagonal element set; calculating the information entropy of all elements in the off-diagonal element set, where the information entropy is equal to the negative of the sum of the products of the probability mass of each element and the logarithm of that probability mass, and the probability mass is determined by the ratio of the element value to the sum of the off-diagonal element set; using the information entropy as the core operating entropy at the current sliding window position, and arranging the core operating entropies at each sliding window position according to the sliding window's movement order to generate a core operating entropy sequence. This core operating entropy sequence effectively characterizes the overall uncertainty of coupled perturbations between multi-source operating data of an enterprise and can sensitively reflect systematic changes in operating status.
[0008] As a further improvement of the present invention, the process of fitting the credit decay dynamics model includes: extracting the preceding core operating entropy subsequences corresponding to the preceding multiple consecutive sliding windows from the core operating entropy sequence; fitting a linear regression line from the preceding core operating entropy subsequences; calculating the slope of the linear regression line as the initial decay rate; extracting the subsequent core operating entropy subsequences corresponding to the following multiple consecutive sliding windows from the core operating entropy sequence; calculating the difference between adjacent elements in the subsequent core operating entropy subsequences; using the average of all differences as the decay acceleration; constructing a second-order difference equation with the window number as the independent variable and the core operating entropy as the dependent variable; using the initial decay rate as the initial condition for the first derivative and the decay acceleration as the constraint condition for the second derivative; solving the second-order difference equation to obtain the decay curve function, which serves as the credit decay dynamics model. The credit decay dynamics model includes exponential decay terms or power-law decay terms, which can truly reflect the nonlinear decay law of credit potential energy over time windows and improve the prediction accuracy of credit deterioration trajectories.
[0009] In a preferred embodiment of the present invention, the steps of calculating the confidence interval drift correction factor and generating a credit score include: taking the submatrix corresponding to the last sliding window in the multimodal data stream matrix, calculating the sum of all elements in the submatrix as the terminal energy value; extracting the function value of the decay curve function corresponding to the intrinsic credit potential at the position of the last sliding window, and calculating the ratio of the terminal energy value to the function value as the drift amplitude; comparing the magnitude of the terminal energy value and the function value, if the terminal energy value is greater than the function value, the drift direction is recorded as the positive direction, if the terminal energy value is less than the function value, the drift direction is recorded as the negative direction, and if the two are equal, the drift direction is recorded as the zero direction; and using the product of the drift amplitude and the sign value of the drift direction as the confidence interval drift correction factor, wherein the sign value of the positive direction is positive one, the sign value of the negative direction is negative one, and the sign value of the zero direction is zero. Then, the function values of the decay curve function at all sliding window positions are accumulated to obtain the cumulative potential energy benchmark value; the confidence interval drift correction factor is multiplied by a preset correction strength coefficient to obtain the weighted drift amount; the cumulative potential energy benchmark value and the weighted drift amount are added to obtain the corrected total potential energy; the benchmark credit threshold of the target company's industry is obtained; the difference between the corrected total potential energy and the benchmark credit threshold is divided by the benchmark credit threshold to obtain the normalized ratio; the normalized ratio is mapped to a preset credit scoring interval to generate a credit score. The confidence interval drift correction factor is dynamically updated using an adaptive sliding window or exponentially weighted moving average method, enabling the credit score to adapt to the latest business dynamics of the company and changes in industry benchmarks.
[0010] In a more preferred embodiment of the present invention, after obtaining the intrinsic credit potential, the operational response delay index of the target enterprise is calculated based on the phase synchronization between the row vectors of all data source types in the multimodal data flow matrix, and the decay acceleration of the credit decay dynamics model is corrected according to the operational response delay index. Specifically, the process of calculating the operational response delay index includes: performing a Hilbert transform on the row vectors of each data source type in the multimodal data flow matrix to obtain the instantaneous phase sequence of each row vector; selecting any two instantaneous phase sequences of different data source types, calculating the phase difference between the two instantaneous phase sequences on the same time slice, and dividing the absolute value of the phase difference by pi to obtain the normalized phase difference; using the mean of the normalized phase differences between all different data source pairs as the phase synchronization index, calculating and subtracting the phase synchronization index to obtain the initial terminal delay index; and multiplying the initial terminal delay index by the total number of time slices in the horizontal direction of the multimodal data flow matrix to obtain the operational response delay index. When correcting the attenuation acceleration, the operational response delay index is compared with a preset response threshold. If the operational response delay index is greater than the response threshold, the attenuation acceleration in the credit attenuation dynamics model is extracted and multiplied by an amplification factor greater than one to obtain the corrected attenuation acceleration. If the operational response delay index is less than the response threshold, the attenuation acceleration is multiplied by a reduction factor less than one to obtain the corrected attenuation acceleration. If the operational response delay index is equal to the response threshold, the attenuation acceleration remains unchanged. The attenuation acceleration in the credit attenuation dynamics model is replaced with the corrected attenuation acceleration to obtain the corrected credit attenuation dynamics model. This scheme introduces a multi-source data phase synchronization measure, enabling the attenuation model to consider the response lags between different operational dimensions of the enterprise, significantly improving the ability to identify internal coordination imbalances and implicit credit risks.
[0011] The technical effects and advantages provided by the present invention in the above technical solution are as follows: A perturbation transfer matrix is constructed by calculating the bidirectional coupling perturbation potential between different data sources within each sliding window. This matrix characterizes the fluctuation transmission relationship between different dimensions of business operations. The ratio of cross-covariance to auto-covariance between any two data sources is quantified as the unidirectional perturbation intensity, and after exchanging directions, they form a complete bidirectional interaction metric. This process transforms traditional isolated feature analysis into modeling the dynamic coupling structure between data flows. When the business status changes, the perturbation transfer relationship between various data sources changes accordingly, and this change is directly reflected in the dispersion of off-diagonal elements. Information entropy is extracted from the off-diagonal element set of the perturbation transfer matrix to generate a core operating entropy sequence. Core operating entropy quantifies the degree of disorder in the multi-source data interaction structure; the higher the value, the weaker the synergy and the higher the disorder among different business activities within the enterprise. By continuously tracking the core operating entropy at each sliding window position, the micro-erosion process of the enterprise's credit foundation is transformed into an observable monotonic sequence, allowing the assessment to no longer rely on static financial ratios but directly capture the dynamic signal of the evolution of the business structure from order to disorder. This method of directly extracting structural disorder measures from cross-source disturbance relationships achieves a deep characterization of implicit credit risk in enterprises. A second-order difference equation with core operating entropy as the dependent variable is constructed and fitted to a credit decay dynamics model to obtain the intrinsic credit potential. A linear regression line is fitted from the preceding window subsequence of the core operating entropy sequence to extract the initial decay rate. The mean of the differences between adjacent elements is calculated from the following window subsequence as the decay acceleration. The decay acceleration reflects whether the operational disorder process tends to accelerate deterioration or converge and improve. It constrains the second derivative term of the second-order difference equation, making the decay curve function not a simple inertial extrapolation, but a trajectory adjusted by the speed of internal structural evolution. The credit decay dynamics model elevates risk assessment from static threshold comparison to the deduction of the dynamic behavior of credit status. The intrinsic credit potential includes information on the historical rate and acceleration of decay. Furthermore, based on the drift direction and amplitude of the intrinsic credit potential in the final window, a confidence interval drift correction factor is calculated. The terminal window is a sub-matrix extracted from the latest multi-source data. The overall energy of this sub-matrix deviates from the predicted value of the decay curve function at the latest position. The direction and magnitude of this deviation are calculated and used as a correction factor, directly superimposed onto the intrinsic credit potential benchmark value accumulated from the decay curve function. This drift correction of terminal energy relative to model predictions enables credit scoring to respond instantly to the latest multi-source joint state changes based on long-term trend predictions, avoiding the lag and misjudgment caused by simply relying on model extrapolation. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0013] Figure 1 This is a flowchart of a multi-source heterogeneous micro and small enterprise operating data fusion credit assessment method; Figure 2 This is a flowchart of the multi-source digital asset record processing and disturbance transmission matrix construction process; Figure 3 This is a flowchart of the credit score generation process based on confidence interval drift correction. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] See Figure 1 This invention provides a method for credit assessment of micro and small enterprises by fusing multi-source heterogeneous operating data, comprising: collecting digital asset records of the target enterprise from multiple heterogeneous data sources, the digital asset records including transaction records, tax invoices, electricity consumption, and social security payment records; aligning all digital asset records according to the time axis to construct a multimodal data flow matrix with time as the horizontal axis and data source type as the vertical axis; performing sliding window segmentation on the multimodal data flow matrix, calculating the bidirectional coupling perturbation potential between different data sources within each sliding window to obtain a perturbation transfer matrix; extracting the core operating entropy sequence of the target enterprise based on the distribution dispersion of off-diagonal elements in the perturbation transfer matrix; fitting a credit decay dynamic model of the target enterprise based on the decay gradient of the core operating entropy sequence within continuous sliding windows to obtain the intrinsic credit potential; calculating a confidence interval drift correction factor based on the drift direction and amplitude of the intrinsic credit potential in the final window of the multimodal data flow matrix, using the confidence interval drift correction factor to correct the intrinsic credit potential, and generating a credit score.
[0016] Example 1: In the specific implementation process, refer to Figure 2The system collects digital asset records from multiple heterogeneous data sources for the target enterprise. These records include transaction logs, tax invoices, electricity consumption, and social security payment records. The timestamp field is extracted from each digital asset record, and all records are sorted in ascending order by timestamp field to generate a unified timeline. The timestamp field is a marker indicating the moment the data was generated in the digital asset record; it is uniformly converted to Unix timestamp or ISO8601 standard format before sorting.
[0017] The total time span is determined based on the minimum and maximum timestamps in a unified timeline, and then divided into multiple equally spaced time slices. The length of each time slice is set to a fixed value, determined by the minimum sampling frequency of the data source: if the longest interval between records from any data source is one month, the time slice length is set to one month; or it can be set to one week to accommodate higher-frequency data updates. During the division process, the time slices are numbered sequentially from left to right. The first time slice covers the time interval from the minimum timestamp to the minimum timestamp plus one time slice length, and subsequent time slices are shifted forward accordingly.
[0018] Within each time slice, numerical features are extracted from the digitized asset records of each heterogeneous data source. For transaction log data sources, the extracted numerical features are the total transaction amount and the number of transactions for all transaction log records within the time slice. The total transaction amount is then taken as the numerical feature for that data source in that time slice. For tax invoice data sources, the extracted numerical feature is the total invoice amount for all tax invoice records within the time slice. For electricity consumption data sources, the extracted numerical feature is the total electricity consumption for all electricity consumption records within the time slice. For social security payment record data sources, the extracted numerical feature is the total actual payment amount for all social security payment records within the time slice. The extracted numerical features are then filled into the corresponding cells of the initial matrix. The rows of the initial matrix are indexed by the data source type on the vertical axis, with the data source types being transaction logs, tax invoices, electricity consumption, and social security payment records, totaling four rows. The columns of the initial matrix are indexed by the time slice number on the horizontal axis, with the number of columns equal to the total number of time slices.
[0019] Linear interpolation is performed to fill cells in the initial matrix that lack numerical features. When a cell is empty due to a lack of records for a certain data source type within a time slice, the two nearest non-empty cells are searched forward and backward along the row containing that data source type. A straight line is fitted using the values from these two non-empty cells and the time slice number. The calculation formula is as follows: ,in, Indicates the time slice number Numerical feature values to be filled This represents the numeric feature value of the nearest non-empty cell found during the forward search. This represents the numeric feature value of the nearest non-empty cell found during the backward search. This indicates the time slice number where the non-empty cell found during the forward search is located. This indicates the time slice number where the non-empty cell found during the backward search is located. satisfy If a missing cell is located at the beginning or end of a row, nearest neighbor interpolation is used, meaning the value of the nearest non-empty cell is directly assigned. The filled initial matrix serves as a multimodal data flow matrix, where each cell represents a numerical feature.
[0020] A fixed-length sliding window is set, with its length determined by a preset number of time slices. This preset number of time slices is determined based on the total number of columns in the multimodal data flow matrix, rounded up to one-third of the total number of columns, and not less than 3. The fixed step size of the sliding window is set to the length of one time slice. The sliding window starts from the leftmost position of the multimodal data flow matrix and moves one fixed step to the right along the horizontal axis each time until the right boundary of the sliding window aligns with the rightmost column of the multimodal data flow matrix. A submatrix is extracted at each sliding window position. The submatrix retains all four rows of the multimodal data flow matrix, and the columns are taken from the time slice range covered by the sliding window, forming a 4-row, L-column submatrix, where L is the length of the sliding window.
[0021] In each submatrix, select any two row vectors corresponding to different data source types. The first selected row vector is denoted as the first source row vector, and the second as the second source row vector. Both the first and second source row vectors are sequences containing L values. Calculate the conditional fluctuation transitivity of the first source row vector relative to the second source row vector. The conditional fluctuation transitivity is obtained by dividing the cross-covariance of the first and second source row vectors within the submatrix by the autocovariance of the second source row vector. The cross-covariance is calculated as follows: calculate the deviations of the L values in the first source row vector from their respective means, calculate the deviations of the L values in the second source row vector from their respective means, sum the products of the deviations at corresponding positions, and divide by L. The autocovariance is calculated as follows: calculate the sum of the squares of the deviations of the L values in the second source row vector from their own means, and divide by L. The conditional fluctuation transitivity is used as the unidirectional perturbation strength from the data source type corresponding to the second source row vector to the data source type corresponding to the first source row vector.
[0022] The roles of the first and second source row vectors are swapped, i.e., the conditional fluctuation transfer rate of the original second source row vector relative to the original first source row vector is calculated, which is used as the reverse unidirectional perturbation strength. This calculation is performed on all data source pairs composed of different data source types in the submatrix, resulting in twelve unidirectional perturbation strengths. These twelve unidirectional perturbation strengths are arranged in a fixed order according to the data source types, forming a 4x4 matrix. The element in the p-th row and q-th column of the matrix represents the unidirectional perturbation strength from the q-th data source type to the p-th data source type. When p equals q, the diagonal elements of the matrix are set to 0. This matrix is the perturbation transfer matrix for the current sliding window position. The above operation is performed on all sliding window positions to obtain the perturbation transfer matrix corresponding to each sliding window position.
[0023] Example 2: In the specific implementation process, a multimodal data stream matrix is constructed and a sliding window is performed on the multimodal data stream matrix. After extracting a sub-matrix at each sliding window position, the bidirectional coupling perturbation potential between row vectors corresponding to any two different data source types in the sub-matrix is calculated. In some implementations, the bidirectional coupling perturbation potential is calculated based on the transfer entropy. For the first source row vector and the second source row vector selected in the sub-matrix, the first source row vector and the second source row vector correspond to two different data source types, respectively denoted as the first data source type and the second data source type. The numerical sequence of the first source row vector in L time slices in the sub-matrix is taken as the source signal. The numerical sequence of the second source row vector across L time slices in the submatrix is taken as the target signal. For the target signal Construct its past state set, the past state set is composed of The values of the previous time slice are used to construct the historical values, which are delayed by one time slice. This applies to the source signal. Similarly, its past state set is constructed. Within the time window of the submatrix, the calculation is performed on the known target signal. Given its past state, the source signal is then added. After the past state, the target signal The reduction in the uncertainty of the current state is represented by the difference in conditional entropy under a discrete probability distribution. Specifically, the target signal... Current value, target signal Past values and source signals The past values are symbolized, and each numerical sequence is discretized using an equal interval binning method. The number of bins is set to a fixed integer value, which is determined based on the length L of the submatrix and is no greater than 1. The largest integer is 3 and at least 3. The joint probability distribution and conditional probability distribution after statistical discretization are calculated, and the transfer entropy value is used as the unidirectional perturbation strength from the data source type corresponding to the first source row vector to the data source type corresponding to the second source row vector. The roles of the source signal and the target signal are swapped, and the transfer entropy from the second source row vector to the first source row vector is calculated as the reverse unidirectional perturbation strength.
[0024] In other implementations, the bidirectional coupled perturbation potential is calculated based on mutual information. The numerical sequences of the first and second source row vectors are also discretized into equal-interval bins. The joint probability distribution and individual marginal probability distributions of the two discretized sequences are statistically analyzed, and the mutual information value between the first and second source row vectors is calculated. The mutual information value is normalized by dividing it by the larger of the information entropy of the first and second source row vectors. This normalized mutual information value is used as the unidirectional perturbation strength from the second source row vector to the first source row vector, and the same normalized mutual information value is used as the unidirectional perturbation strength from the first source row vector to the second source row vector. The above calculation based on transfer entropy or mutual information is performed on all data source pairs in the submatrix composed of different data source types to obtain all unidirectional perturbation strength values. All unidirectional perturbation intensity values are arranged in a fixed order according to the data source type and combined into a square matrix of four rows and four columns. The element in the p-th row and q-th column of the square matrix represents the unidirectional perturbation intensity from the q-th data source type to the p-th data source type. When p equals q, the diagonal elements of the square matrix are set to zero. This square matrix serves as the perturbation transfer matrix for the current sliding window position.
[0025] Extract all off-diagonal elements from the perturbation transfer matrix whose row indices are not equal to their column indices, forming an off-diagonal element set. This off-diagonal element set contains twelve elements. Calculate the information entropy of all elements in the off-diagonal element set. The formula for calculating information entropy is: ,in, Represents information entropy. This represents the total number of elements in the off-diagonal element set. . Represents the set of off-diagonal elements. The probability mass of each element, The value of is an integer from 1 to 12. (Probability mass) By the The value of each off-diagonal element is obtained by dividing the sum of the values of all elements in the off-diagonal element set. All probability masses are obtained when the sum of the values of all elements in the off-diagonal element set is zero. Set to zero, at which point the information entropy... It is zero. Taking the base of the logarithm to 2, the information entropy is... The unit is bits.
[0026] The calculated information entropy The core operating entropy is used as the value for the current sliding window position. Following the order of the sliding window positions generated by moving the sliding window along the horizontal axis of the multimodal data flow matrix, the core operating entropy corresponding to each sliding window position is arranged sequentially to generate a core operating entropy sequence. The core operating entropy sequence is a one-dimensional vector, the length of which is equal to the total number of sliding window positions, and each element corresponds to the core operating entropy value for that sliding window position.
[0027] Example 3: In the specific implementation process, after extracting the core operating entropy sequence from the disturbance transfer matrix, a credit decay dynamics model for the target enterprise is fitted based on the decay gradient of the core operating entropy sequence within a continuous sliding window. The core operating entropy sequence is a one-dimensional vector of length M, where M equals the total number of sliding window positions. The m-th element in the core operating entropy sequence is denoted as... The value of m is an integer from 1 to M, and each element... The core operational entropy value corresponding to a sliding window position.
[0028] Extract the preceding core operating entropy subsequences corresponding to the first multiple consecutive sliding windows from the core operating entropy sequence. The preceding core operating entropy subsequence contains elements from index 1 to K in the core operating entropy sequence, where K is one-third of the total number of sliding window positions M, rounded up, and K is not less than 3. The elements of the preceding core operating entropy subsequence are denoted as follows: Using the index of the element in the aforementioned core operational entropy subsequence as the independent variable, and the element value as the independent variable... As the dependent variable, a linear regression line is fitted. The expression for the linear regression line is: ,in This represents the intercept of the linear regression line. Let represent the slope of the linear regression line, and m represent the index of the element in the preceding core operating entropy subsequence. The intercept is solved using the least squares method. and slope , making Take the minimum value. Use the slope obtained from the solution. As the initial decay rate. For real numbers, when the initial decay rate A positive value indicates that the core operating entropy is increasing in the initial stage, and the initial decay rate... A negative value indicates that the core operating entropy is decreasing in the initial stage.
[0029] Extract the subsequent core operating entropy subsequences corresponding to multiple consecutive sliding windows from the core operating entropy sequence. The subsequent core operating entropy subsequences contain elements from index M-K+1 to M in the core operating entropy sequence, and the elements of the subsequent core operating entropy subsequences are denoted as follows: The difference between adjacent elements in the post-core operating entropy subsequence is calculated using the following formula: Where w takes values from 2 to K. All differences... The arithmetic mean of the decay acceleration is denoted as . Decaying acceleration Calculated using the following formula: ,in, represents the decay acceleration, K represents the number of elements in the subsequent core operating entropy subsequence, which is the same as the number of elements in the preceding core operating entropy subsequence. w represents the index of the difference, which is an integer ranging from 2 to K. This represents the difference between the w-th element and the previous element in the subsequent core operating entropy subsequence.
[0030] Construct a second-order difference equation with window number as the independent variable and core operating entropy as the dependent variable. The window number is represented by a continuous variable t, ranging from 1 to M. The function representing the change of core operating entropy with window number is denoted as... The second-order difference equation is in the form of: That is, the second derivative of the core operating entropy with respect to the window number is everywhere equal to the decay acceleration. The initial decay rate As the initial condition for the first derivative, i.e. Decaying acceleration As a constraint condition for the second derivative, the first element value of the preceding core operating entropy subsequence is... As the initial value of the function at t=1, that is Solve the second-order difference equation by performing two integrations. The first integration yields the first-order derivative function. Substitute the initial conditions Find the integral constant The second integral yields the decay curve function. Substitute the initial value Find the integral constant The obtained decay curve function is used as a credit decay dynamic model.
[0031] In some implementations, the credit decay dynamics model includes an exponential decay term. The resulting decay curve function is then described as follows: ,in Indicates the initial amplitude parameter. This represents the decay rate parameter. Let represent the steady-state offset parameter, and e represent the natural constant. Initial amplitude parameters are obtained by fitting data from the preceding and following core operating entropy subsequences using a nonlinear least squares method. Attenuation rate parameter and steady-state offset parameters The numerical value. In other implementations, the credit decay dynamics model includes a power-law decay term. The decay curve function constructed in this case is of the form... ,in This represents the amplitude scaling parameter. This represents the power-law exponent parameter. This represents the baseline offset parameter. Similarly, the values of each parameter are determined through nonlinear fitting using data from the aforementioned two subsequences. Regardless of the function form used, the fitted function is the credit decay dynamics model, which outputs a function showing the continuous change of core operating entropy with window number t.
[0032] Example 4: In the specific implementation process, refer to Figure 3 After fitting the credit decay dynamics model and obtaining the decay curve function, the submatrix corresponding to the last sliding window in the multimodal data stream matrix is taken. The total number of sliding window positions in the multimodal data stream matrix is M, and the index of the last sliding window is M. This submatrix retains all four rows of the multimodal data stream matrix, and the columns are taken from the time slice range covered by the Mth sliding window. The submatrix is a 4-row, L-column numerical matrix. The sum of all elements in this submatrix is calculated, that is, all 4×L numerical features in the submatrix are added together, and the resulting scalar is used as the terminal energy value.
[0033] Extract the function value of the decay curve function corresponding to the intrinsic credit potential at the last sliding window position. The decay curve function takes the window index t as its independent variable, and the value of t is a continuous real number from 1 to M. Substitute t=M into the decay curve function. The function value is calculated. Calculate the terminal energy value and function value. The ratio of the observed energy to the predicted energy is used as the drift magnitude. The drift magnitude represents the factor by which the actual observed energy at the end deviates from the model's predicted potential energy.
[0034] Compare the terminal energy value with the function value The magnitude of the energy at the end. If the terminal energy value is greater than the function value. Then the drift direction is recorded as the positive direction. If the terminal energy value is less than the function value... Then the drift direction is recorded as the negative direction. If the terminal energy value equals the function value... If the drift direction is zero, then the drift direction is designated as the zero direction. A sign value is assigned to the drift direction: positive for a positive direction, negative for a negative direction, and zero for the zero direction. The drift amplitude is multiplied by the sign value of the drift direction; the product is used as the confidence interval drift correction factor. The sign of the confidence interval drift correction factor reflects the direction in which the actual energy deviates from the model trend, and its absolute value reflects the degree of deviation.
[0035] The cumulative potential energy baseline value is obtained by summing the function values of the decay curve function corresponding to the intrinsic credit potential energy at all sliding window positions. The function values of the decay curve function at all sliding window positions t=1,2,...,M are as follows: Add the above M function values together, and the sum is the cumulative potential energy baseline value.
[0036] The confidence interval drift correction factor is multiplied by a preset correction strength coefficient to obtain the weighted drift. The correction strength coefficient is a positive real number used to adjust the strength of the drift correction and avoid overcorrection. The value of the correction strength coefficient is set between 0.1 and 2.0, and its specific value is determined based on the historical credit performance data of the target company's industry. The principle for determining the value is to maximize the default discrimination ability of the corrected credit score on the test sample set. Without industry customization, the correction strength coefficient uses a default value of 0.5. This default value ensures that the adjustment of the confidence interval drift correction factor to the cumulative potential energy benchmark value does not exceed half of the fluctuation range of the terminal energy value, thereby maintaining score stability while responding to data freshness. The cumulative potential energy benchmark value is added to the weighted drift to obtain the total corrected potential energy.
[0037] Obtain the benchmark credit threshold for the target company's industry. The benchmark credit threshold is derived from the quantiles obtained through statistical analysis of the modified total potential distribution of a large number of companies within the same industry. The median of the modified total potential of the sample of companies in that industry is taken as the benchmark credit threshold. The difference between the modified total potential and the benchmark credit threshold is calculated, and this difference is divided by the benchmark credit threshold to obtain the normalized ratio. This normalized ratio reflects the degree of deviation of the target company's credit potential from the typical level of the industry.
[0038] The normalized ratio is mapped to a preset credit score range. This preset range is a closed interval [300, 850], representing the standard credit score range. A linear transformation is used to linearly extend the distribution range of the normalized ratio to the preset credit score range. A normalized ratio of -1 maps to 300 points, +1 maps to 850 points, and 0 maps to 575 points. The specific calculation formula is: Credit Score = 575 + Normalized Ratio × 275. For normalized ratios exceeding the range [-1, 1], a boundary truncation is applied: values less than -1 are taken as 300 points, and values greater than +1 are taken as 850 points. The mapping result serves as the target company's credit score.
[0039] The confidence interval drift correction factor is dynamically updated using an adaptive sliding window method. After each credit score is generated, when new business data is generated and added to the multimodal data stream matrix, the sliding window moves to the right, generating a new last sliding window submatrix. Based on the new submatrix, the terminal energy value is recalculated, and the comparison between the terminal energy value and the terminal value of the decay curve function is re-performed to obtain a new confidence interval drift correction factor. In other implementations, an exponentially weighted moving average method is used to dynamically update the confidence interval drift correction factor. Specifically, the historical confidence interval drift correction factor sequence is denoted as... The newly calculated correction factor currently in use is denoted as Updated confidence interval drift correction factor according to Calculate, where the smoothing coefficient is... The smoothing coefficient is set to 0.3 to give higher weight to recent data while maintaining historical information. The value of is chosen to balance response speed and smoothness when there are many sliding windows; 0.3 ensures that each new observation contributes 30% of the weight. The weighted drift is recalculated using the updated confidence interval drift correction factor, and the cumulative potential energy benchmark value is corrected to generate a dynamically adjusted credit score.
[0040] Example 5: In the specific implementation process, based on the phase synchronization between the row vectors of all data source types in the multimodal data flow matrix, the operational response delay index of the target enterprise is calculated, and the decay acceleration of the credit decay dynamics model is corrected according to the operational response delay index. The row vectors of the multimodal data flow matrix correspond to the data source types, with a total of four data source types: transaction flow data source, tax invoice data source, electricity consumption data source, and social security payment record data source. Each row vector is a time series of length T, where T is equal to the number of columns in the multimodal data flow matrix, i.e., the total number of time slices.
[0041] A Hilbert transform is performed on the row vectors of each data source type in the multimodal data stream matrix. The Hilbert transform is implemented using a Fourier transform. First, a fast Fourier transform is performed on the row vectors of length T to obtain the spectrum. The amplitude of the positive frequency components in the spectrum is doubled, the negative frequency components are set to zero, and the DC component remains unchanged. Then, an inverse fast Fourier transform is performed to obtain a Hilbert transform result orthogonal to the original sequence. An analytic signal is constructed using the original row vectors as the real part and the Hilbert transform result as the imaginary part. The argument of the analytic signal at each time slice position is calculated. The argument is obtained using the arctangent function of the real and imaginary parts, and the argument ranges from (-π, +π) radians. The argument sequence is phase-unwrapped to eliminate discontinuities where the argument jump between adjacent time slices exceeds π radians, resulting in an expanded instantaneous phase sequence. The above operation is performed on the row vectors of the four data source types respectively, resulting in four instantaneous phase sequences.
[0042] Select any two instantaneous phase sequences from different data source types, and denote the first instantaneous phase sequence as... The second instantaneous phase sequence is denoted as Where t is the time slice number, and t takes the value of an integer from 1 to T, and p and q represent the indices of two different data source types. The phase difference between two instantaneous phase sequences on the same time slice t is calculated; the phase difference is equal to... minus Divide the absolute value of the phase difference by pi (π) to obtain the normalized phase difference for that time slice t. Calculate the normalized phase difference for all time slices t=1 to T, obtaining T values. Then calculate the average of these T values as the mean phase difference between data source type p and data source type q. The mean phase difference ranges from 0 to 1; a smaller value indicates that the instantaneous phases of the two data source types are more consistent.
[0043] The phase difference mean calculation was performed on all data source pairs composed of different data source types, resulting in six phase difference mean values. The arithmetic mean of these six phase difference mean values was used as the phase synchronization index. The phase synchronization index ranges from 0 to 1; a smaller value indicates higher phase synchronization among all data source types. The initial terminal latency index was obtained by subtracting the phase synchronization index from the calculated value. The initial terminal latency index ranges from 0 to 1; a larger value indicates more severe operational response delays.
[0044] The operational response latency index is obtained by multiplying the initial terminal latency index by the total number of time slices T along the horizontal axis of the multimodal data stream matrix. The operational response latency index is a real number greater than or equal to zero, and its dimensions are the same as the number of time slices.
[0045] The operational response delay index is compared with a preset response threshold. The response threshold is set based on the historical quantile of the operational response delay index of the target company's industry. This historical quantile is taken from the median of the distribution of operational response delay indices of multiple normally operating companies in the same industry. When industry data is unavailable, the response threshold uses a default value, which is one-third of the total number of time slices T and rounded up, to ensure that the operational response delay index of most normally operating companies is below this threshold.
[0046] If the operational response latency index is greater than the response threshold, it indicates that the response latency between various data sources within the target enterprise is higher than normal, reflecting a blockage in operational coordination. In this case, the decay acceleration in the credit decay dynamics model should be extracted. Decaying acceleration Multiply by the amplification factor to obtain the corrected decay acceleration. The amplification factor is set to a fixed value of 1.5. This fixed value is based on the fact that when the response delay exceeds the threshold, the actual credit decay is often masked by the delay effect. The decay acceleration in the credit decay dynamics model tends to underestimate the true decay rate numerically. Amplifying the decay acceleration by 1.5 times can make the corrected dynamics model closer to the risk exposure trajectory in high-latency scenarios.
[0047] If the operational response latency index is less than the response threshold, it indicates that the response latency between various data sources within the target enterprise is below normal levels, reflecting smooth operational flow. In this case, the decay acceleration in the credit decay dynamics model can be extracted. Decaying acceleration Multiplying by a reduction factor yields the corrected attenuation acceleration. The reduction factor is set to a fixed value of 0.6. This fixed value is based on the fact that when the response delay is small, the credit attenuation dynamics model may overestimate the attenuation acceleration due to short-term fluctuations. Multiplying by a reduction factor of 0.6 can reduce the impact of noise and make the credit assessment results match the actual lower default risk of low-latency enterprises.
[0048] If the operational response delay index equals the response threshold, then the decay acceleration in the credit decay dynamics model is maintained. The corrected decay acceleration remains unchanged and equals the original decay acceleration. .
[0049] The decay acceleration in the credit decay dynamics model The modified decay acceleration is then replaced to obtain the corrected credit decay dynamics model. After the modification process is complete, the second derivative constraint in the decay curve function is updated to the value of the modified decay acceleration, while the first derivative initial condition remains the initial decay rate. The initial value is still the value of the first element of the previous core operating entropy subsequence. The second-order difference equation is then resolved to obtain the corrected decay curve function form, and the intrinsic credit potential is recalculated. The corrected intrinsic credit potential is then used to continue the subsequent steps of calculating the confidence interval drift correction factor and generating the credit score.
[0050] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A method for credit assessment by fusing multi-source heterogeneous operating data of micro and small enterprises, characterized in that, Includes the following steps: Collect the target enterprise's digital asset records from multiple heterogeneous data sources. These digital asset records include transaction records, tax invoices, electricity consumption, and social security payment records. All digital asset records are aligned according to the timeline to construct a multimodal data flow matrix with time as the horizontal axis and data source type as the vertical axis; A sliding window segmentation is performed on the multimodal data stream matrix, and the bidirectional coupling perturbation potential between different data sources is calculated within each sliding window to obtain the perturbation transfer matrix; Based on the distribution dispersion of off-diagonal elements in the perturbation transfer matrix, the core operating entropy sequence of the target enterprise is extracted; Based on the decay gradient of the core operating entropy sequence within a continuous sliding window, a credit decay dynamics model of the target enterprise is fitted to obtain the intrinsic credit potential. Based on the drift direction and magnitude of the intrinsic credit potential in the end window of the multimodal data stream matrix, the confidence interval drift correction factor is calculated. The intrinsic credit potential is then corrected using the confidence interval drift correction factor to generate a credit score.
2. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 1, characterized in that, The steps of aligning all digitized asset records according to the time axis and constructing a multimodal data flow matrix with time as the horizontal axis and data source type as the vertical axis include: Extract the timestamp field from each digital asset record, sort all digital asset records in ascending order according to the timestamp field, and generate a unified timeline; The total time span is determined based on the minimum and maximum timestamps in the unified timeline, and then divided into multiple time slices at equal intervals. Within each time slice, numerical features are extracted from the digital asset records of each heterogeneous data source, and these numerical features are filled into the corresponding cells of an initial matrix indexed by time on the horizontal axis and by data source type on the vertical axis. Linear interpolation is performed to fill cells in the initial matrix that are missing numerical features, and the filled initial matrix is used as a multimodal data stream matrix.
3. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 2, characterized in that, The steps of performing sliding window segmentation on the multimodal data stream matrix, calculating the bidirectional coupling perturbation potential between different data sources within each sliding window, and obtaining the perturbation transfer matrix include: Set a sliding window of fixed length, move the sliding window along the horizontal axis of the multimodal data stream matrix with a fixed step size, and extract a sub-matrix at each sliding window position; In each submatrix, select any two row vectors corresponding to different data source types, and calculate the conditional volatility transfer rate of the first row vector relative to the second row vector. The conditional volatility transfer rate is equal to the cross covariance of the two row vectors in the submatrix divided by the autocovariance of the second row vector. The conditional fluctuation transitivity is used as the unidirectional perturbation strength from the second data source to the first data source. The reverse unidirectional perturbation strength is calculated by swapping the roles of the two row vectors. The bidirectional unidirectional perturbation intensities between all different data source pairs in the submatrix are combined into a square matrix, which is then used as the perturbation transfer matrix.
4. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 1, characterized in that, The bidirectional coupled perturbation potential is calculated based on mutual information or transfer entropy.
5. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 3, characterized in that, The steps for extracting the core operating entropy sequence of the target enterprise based on the distribution dispersion of off-diagonal elements in the perturbation transfer matrix include: Extract all off-diagonal elements in the perturbation transfer matrix whose row indices are not equal to their column indices, and form a set of off-diagonal elements. Calculate the information entropy of all elements in the off-diagonal set. The information entropy is equal to the negative of the sum of the products of the probability mass of each element and the logarithm of the probability mass. The probability mass is determined by the ratio of the element value to the sum of the off-diagonal set. Using information entropy as the core operating entropy at the current sliding window position, the core operating entropy of each sliding window position is arranged according to the moving order of the sliding window to generate a core operating entropy sequence.
6. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 5, characterized in that, The steps for fitting a credit decay dynamics model of the target enterprise based on the decay gradient of the core operating entropy sequence within a continuous sliding window include: Extract the preceding core operating entropy subsequences corresponding to the first multiple consecutive sliding windows from the core operating entropy sequence, fit a linear regression line from the preceding core operating entropy subsequences, and calculate the slope of the linear regression line as the initial decay rate. Extract the subsequent core operating entropy subsequences corresponding to multiple consecutive sliding windows from the core operating entropy sequence, calculate the difference between adjacent elements in the subsequent core operating entropy subsequences, and use the average of all differences as the decay acceleration. A second-order difference equation is constructed with window number as independent variable and core operating entropy as dependent variable. The initial decay rate is used as the initial condition for the first derivative, and the decay acceleration is used as the constraint condition for the second derivative. The decay curve function is obtained by solving the second-order difference equation. The decay curve function is used as a dynamic model of credit decay.
7. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 1, characterized in that, The credit decay dynamics model includes either an exponential decay term or a power-law decay term.
8. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 6, characterized in that, The step of calculating the confidence interval drift correction factor based on the drift direction and magnitude of the intrinsic credit potential in the end window of the multimodal data stream matrix includes: Take the submatrix corresponding to the last sliding window in the multimodal data stream matrix, and calculate the sum of all elements in the submatrix as the terminal energy value; Extract the function value of the decay curve function corresponding to the intrinsic credit potential at the last sliding window position, and calculate the ratio of the terminal energy value to the function value as the drift amplitude; Compare the terminal energy value with the function value. If the terminal energy value is greater than the function value, the drift direction is recorded as the positive direction. If the terminal energy value is less than the function value, the drift direction is recorded as the negative direction. If the two are equal, the drift direction is recorded as the zero direction. The product of the drift amplitude and the sign value of the drift direction is used as the confidence interval drift correction factor, where the sign value in the positive direction is positive one, the sign value in the negative direction is negative one, and the sign value in the zero direction is zero.
9. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 8, characterized in that, The step of using a confidence interval drift correction factor to correct the intrinsic credit potential and generate a credit score includes: The cumulative potential energy baseline value is obtained by summing the function values of the decay curve function corresponding to the intrinsic credit potential energy at all sliding window positions. Multiply the confidence interval drift correction factor by the preset correction intensity coefficient to obtain the weighted drift amount, and add the cumulative potential energy benchmark value to the weighted drift amount to obtain the corrected total potential energy. Obtain the benchmark credit threshold of the target company's industry, and divide the difference between the corrected total potential energy and the benchmark credit threshold by the benchmark credit threshold to obtain the normalized ratio. The normalized ratio is mapped to a preset credit score range, and the mapping result is taken as the credit score.
10. The method for credit assessment of multi-source heterogeneous micro and small enterprise operating data fusion according to claim 1, characterized in that, The confidence interval drift correction factor is dynamically updated using an adaptive sliding window or exponentially weighted moving average method.