Order perception-based credit data dynamic arrangement and strategy optimization system

CN122656748APending Publication Date: 2026-08-28SHANG ANXIN (SHANGHAI) ENTERPRISE DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610797800.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

对于外部数据采购,常规方案多依赖人工审核或者按照固定的供应商目录逐一调用接口获取数据,未对内部存量数据的状态进行量化评估,也未将数据获取成本、接口响应时间与当前分析需求进行动态关联与联合优化

Benefits of technology

1.通过对信用报告订单解析提取信用分析精度阈值,结合内部存量数据的更新频率、行业波动率和风险等级加权计算时效性分值,将时效性分值与精度阈值比对以识别需补充的外部数据维度,并利用多目标优化算法对供应商评价矩阵求解获取最优供应商组合,实现了按需动态调配数据源,避免了盲目采购,使数据获取成本与数据质量在满足分析精度的前提下达到量化平衡,解决了数据获取成本与分析需求失配的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122656748A_ABST
    Figure CN122656748A_ABST
Patent Text Reader

Abstract

The present application relates to the field of strategy optimization, in particular to a credit data dynamic arrangement and strategy optimization system based on order perception. The credit report order is parsed to extract enterprise identification, analysis dimension and precision threshold; the internal historical database is searched, and the timeliness score is calculated based on the update frequency, industry volatility and risk level weighting; the timeliness score is compared with the precision threshold, and the gap analysis model is triggered to identify the external data dimension that needs to be supplemented in response to the condition not being met; the supplier selection engine is started based on the external data dimension, the multi-objective optimization algorithm is used to solve the supplier evaluation matrix to obtain the optimal supplier combination; the optimal supplier combination is called to obtain external data, which is fused with historical records after verification to generate a report. The present application dynamically allocates data sources on demand, avoids blind purchasing, and makes the data acquisition cost and data quality reach a quantitative balance under the condition of meeting the analysis accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of strategy optimization, specifically to a system for dynamic orchestration and strategy optimization of credit data based on order awareness. Background Technology

[0002] Existing enterprise credit analysis systems typically use a direct call to fixed data sources when processing credit report requests. This means that upon receiving a request, they directly request the full amount of enterprise data from a pre-set external database or perform a simple keyword matching query within the internal database. For external data procurement, conventional solutions often rely on manual review or calling interfaces one by one according to a fixed supplier directory to obtain data. There is no quantitative assessment of the status of existing internal data, nor is there any dynamic correlation or joint optimization between data acquisition costs, interface response time, and current analytical needs.

[0003] The aforementioned model of directly calling fixed data sources and making procurement decisions has a core technical problem: it fails to quantify and compare the accuracy of order demand analysis with the timeliness of internal stock data, and fails to optimize external data acquisition costs and data quality through multi-objective joint optimization. This results in the inability to dynamically allocate data sources as needed when processing credit analysis orders, leading to a mismatch between data acquisition costs and analysis requirements. Summary of the Invention

[0004] The purpose of this invention is to provide a dynamic orchestration and strategy optimization system for credit data based on order awareness, which can solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A credit data dynamic orchestration and strategy optimization system based on order awareness includes: an order parsing component configured to receive credit report orders, parse the credit report orders, and extract target enterprise identifiers, analysis dimensions, and credit analysis accuracy thresholds; an internal verification component configured to retrieve matching historical records from an internal historical database based on the target enterprise identifier, and perform weighted calculations based on the data update frequency of the historical records, the industry volatility of the target enterprise's industry, and the risk level of the target enterprise to obtain a timeliness score for the historical records; and a gap analysis component configured to compare the timeliness score with the credit analysis accuracy threshold, and respond to situations where the timeliness score does not meet the credit score threshold. If the analysis accuracy threshold or the historical records do not cover the analysis dimension, a data gap analysis model is triggered to identify the external data dimensions that need to be supplemented. A supplier selection component is configured to launch a supplier selection engine based on the external data dimensions, using a multi-objective optimization algorithm to solve a supplier evaluation matrix that includes single-call cost, data field coverage, historical data accuracy, average interface response time, and data compliance score, to obtain the optimal supplier combination. A data fusion component is configured to call the interface of the optimal supplier combination to obtain external data, perform field-level balance checks on the external data and the historical records, and fuse the external data that passes the checks with the historical records to generate a credit analysis report.

[0006] Preferably, the order parsing component is further configured to: semantically decompose the analysis dimension, extract the implicit temporal dependencies and entity associations in the analysis dimension; based on the temporal dependencies and entity associations, map the target enterprise identifier and the analysis dimension into a multi-hop graph structure to construct a credit demand graph; the nodes in the credit demand graph represent data entities, the edges represent the temporal dependencies and entity associations, and the node attributes include the credit analysis accuracy threshold and data timeliness weight; generate a structured data request instruction based on the credit demand graph, the structured data request instruction carrying the topological constraints of the credit demand graph.

[0007] Preferably, the internal verification component is further configured to: extract the timestamp sequence of the historical records and construct a time decay function; obtain a first adjustment coefficient corresponding to the industry volatility and a second adjustment coefficient corresponding to the risk level; input the time decay function, the first adjustment coefficient, and the second adjustment coefficient into a multidimensional spatiotemporal decay model to calculate the local timeliness score of each data field in the historical records; perform nonlinear fusion on the local timeliness scores to obtain the global timeliness score of the historical records; and use the global timeliness score as the timeliness score, wherein the multidimensional spatiotemporal decay model is constructed based on the product of the exponential decay function and the first adjustment coefficient and the second adjustment coefficient.

[0008] Preferably, the gap analysis component is further configured to: convert the analysis dimension into a demand feature vector and the historical record into a stock feature vector; calculate the cosine similarity between the demand feature vector and the stock feature vector to obtain the dimension coverage rate; trigger the data gap analysis model in response to the dimension coverage rate being lower than a preset coverage rate threshold or the timeliness score being lower than the credit analysis accuracy threshold; the data gap analysis model extracts uncovered feature vectors based on the difference operation between the demand feature vector and the stock feature vector; and map the uncovered feature vectors to a preset external data label space to generate the external data dimension to be supplemented.

[0009] Preferably, the supplier selection component is further configured to: construct a set of constraints based on the external data dimensions to be supplemented, the set of constraints including a lower limit constraint on data field coverage, an upper limit constraint on average interface response time, and a lower limit constraint on data compliance score; construct an objective function by combining the single-call cost and the historical data accuracy; perform Pareto optimization on the objective function and the set of constraints using the multi-objective optimization algorithm to generate a Pareto front solution set; and select the optimal combination of the weighted sum of the single-call cost and the historical data accuracy from the Pareto front solution set as the optimal supplier combination.

[0010] Preferably, the data fusion component is further configured to: extract numeric and text fields from the external data; perform cross-table logical topology verification based on reconciliation for the numeric fields to identify calculation deviations between the numeric fields; perform conflict detection based on entity alignment for the text fields to identify homonymous and heteronymous entities; in response to the calculation deviation exceeding a preset deviation threshold or the existence of conflicting entities identified by the conflict detection, trigger a data rollback mechanism to reacquire the external data; in response to the calculation deviation being within the preset deviation threshold and the absence of conflicting entities, perform primary key mapping and overwrite merging of the external data and the historical records.

[0011] Preferably, the process of constructing a credit demand graph by the order parsing component is further configured as follows: identifying the in-degree and out-degree of nodes in the credit demand graph, defining nodes with an in-degree greater than a preset in-degree threshold as core entity nodes; extracting the credit analysis accuracy threshold corresponding to the core entity nodes, and assigning dynamic weights to the core entity nodes; monitoring the time span of the temporal dependencies, and injecting temporal decay weights into the edges connecting the core entity nodes and edge entity nodes in response to the time span being greater than a preset time window; performing a pruning operation on the credit demand graph based on the dynamic weights and the temporal decay weights, removing edges with temporal decay weights lower than a preset decay threshold, and generating a simplified credit demand graph.

[0012] Preferably, the internal verification component is further configured to: acquire a macroeconomic volatility index and industry policy change events; encode the macroeconomic volatility index and industry policy change events into an external shock vector; input the external shock vector into the multidimensional spatiotemporal decay model to dynamically correct the first adjustment coefficient and the second adjustment coefficient; update the local timeliness score based on the corrected first adjustment coefficient and the second adjustment coefficient; extract the time interval change rate of adjacent update cycles in the historical record, use the time interval change rate as a decay acceleration factor, and superimpose it into the time decay function to adjust the convergence speed of the global timeliness score.

[0013] Preferably, the supplier selection component is further configured to: obtain historical order data acquisition logs, the data acquisition logs including actual call cost, actual response time, and data quality backtracking score; use the actual call cost, actual response time, and data quality backtracking score as feedback reward signals; update the data compliance score and historical data accuracy in the supplier evaluation matrix based on the feedback reward signals; and introduce a dynamic penalty factor in the Pareto optimization process, applying the dynamic penalty factor to the single call cost in the objective function in response to the data compliance score falling below the compliance red line threshold, so as to exclude supplier combinations that do not meet compliance requirements.

[0014] Preferably, the process of the data fusion component performing cross-table logical topology verification based on reconciliation relationships is further configured as follows: constructing a directed graph of accounting equations between the balance sheet and the income statement, wherein the nodes of the directed graph of accounting equations are financial accounts and the edges are operational relationships; substituting the numerical fields in the external data into the directed graph of accounting equations for propagation calculation to obtain the calculated values ​​of the terminal nodes; comparing the calculated values ​​with the actual values ​​of the corresponding fields in the external data to obtain the difference ratio; in response to the difference ratio exceeding the logical self-consistency threshold, tracing back along the edges of the directed graph of accounting equations to locate the source financial account that introduced the calculation deviation; marking the source financial account with an abnormal label, and correcting and replacing it based on the values ​​of the corresponding historical financial accounts in the historical records.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By parsing credit report orders to extract credit analysis accuracy thresholds, and combining the update frequency of internal stock data, industry volatility, and risk level to calculate timeliness scores, the timeliness scores are compared with accuracy thresholds to identify external data dimensions that need to be supplemented. Furthermore, a multi-objective optimization algorithm is used to solve the supplier evaluation matrix to obtain the optimal supplier combination. This enables dynamic allocation of data sources on demand, avoids blind procurement, and achieves a quantitative balance between data acquisition costs and data quality while meeting analytical accuracy requirements. This solves the technical problem of mismatch between data acquisition costs and analytical needs.

[0016] 2. By constructing a credit demand graph and performing pruning operations based on dynamic weights and time-series decay weights, a structured expression of order demand and the removal of redundant information were achieved. By performing cross-table logical topology verification based on reconciliation relationships and conflict detection based on entity alignment on the acquired external data, and by using a directed graph of accounting equations for propagation calculation and reverse tracing, the source financial accounts were located and corrected and replaced, ensuring the logical consistency and numerical accuracy of multi-source data fusion. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall workflow of the order-aware credit data dynamic orchestration and strategy optimization system of the present invention. Figure 2 This is a flowchart illustrating the credit demand graph construction and structured instruction generation process of the order parsing component of the present invention. Figure 3 This is a flowchart illustrating the calculation of the historical data timeliness score for the internal verification component of this invention. Figure 4 This is a flowchart of the data gap identification and external dimension generation process of the gap analysis component of the present invention; Figure 5 The flowchart for selecting the optimal supplier combination for the components of this invention is shown below; Figure 6 This is a flowchart of the multi-source data verification and fusion generation process of the data fusion component of the present invention. Detailed Implementation

[0018] In one embodiment, a credit data dynamic orchestration and strategy optimization system based on order awareness includes an order parsing component, an internal verification component, a gap analysis component, a supplier selection component, and a data fusion component. (Reference) Figure 1 During system operation, the order parsing component receives credit report orders from clients. These orders are transmitted in structured JSON format, containing request headers, request bodies, and signature information. The order parsing component first verifies the order's signature to ensure its integrity and legitimacy. After successful verification, it extracts the target enterprise identifier, analysis dimensions, and credit analysis precision threshold from the order. The target enterprise identifier includes at least one of the following: Unified Social Credit Code, enterprise name, and business registration number. The analysis dimensions include at least one of the following: basic enterprise information, financial status, operational risks, legal proceedings, intellectual property, and market competitiveness. The credit analysis precision threshold is a value between 0 and 1, with higher values ​​indicating higher requirements for data timeliness and completeness.

[0019] refer to Figure 2 The order parsing component semantically decomposes the analysis dimensions, extracting the implicit temporal dependencies and entity relationships within them. Semantic decomposition employs an entity relationship extraction method based on a pre-trained language model. The analysis dimension text is input into the pre-trained language model, which outputs a set of entities and a set of relationships. The entity set includes target companies, subsidiaries, parent companies, suppliers, customers, legal representatives, shareholders, etc., while the relationship set includes holding relationships, transaction relationships, employment relationships, guarantee relationships, etc. Temporal dependencies are extracted by recognizing time adverbs and tense information in the text. For example, the financial status of the past three years implies a temporal dependency on financial data for 2023, 2024, and 2025, and the latest legal proceedings imply a temporal dependency on legal proceedings data within the six months prior to the current time.

[0020] Based on the extracted temporal dependencies and entity associations, the order parsing component maps the target enterprise identifier and analysis dimensions into a multi-hop graph structure, constructing a credit demand graph. Nodes in the credit demand graph represent data entities, edges represent temporal dependencies and entity associations, and node attributes include a credit analysis accuracy threshold and a data timeliness weight. The data timeliness weight is determined based on the time span of the temporal dependencies; the shorter the time span, the higher the data timeliness weight. For example, for operating data from the past month, the data timeliness weight is set to 0.9; for financial data from the past three years, the data timeliness weight is set to 0.6; and for basic information since the company's inception, the data timeliness weight is set to 0.3.

[0021] The order parsing component generates structured data request instructions based on the credit demand graph. These instructions carry the topological constraints of the credit demand graph. The topological constraints include node type constraints, edge type constraints, and path length constraints, used to limit the scope of subsequent data acquisition and processing. For example, for the financial status of a target company and its first-tier subsidiaries, the topological constraints would be: node type is "company," edge type is "controlling relationship," and path length is 1.

[0022] The internal verification component receives structured data request instructions from the order parsing component and retrieves matching historical records from its internal historical database based on the target company identifier. The internal historical database employs a distributed storage architecture and stores relevant data for all processed credit report orders over the past five years, including basic company information, financial data, legal proceedings data, and intellectual property data. The internal verification component first performs a precise match based on the target company identifier. If multiple historical records are matched, they are sorted in descending order of timestamp, with the most recent record selected as the primary historical record and the rest as secondary historical records.

[0023] refer to Figure 3 The internal validation component extracts the timestamp sequence of historical records and constructs a time decay function. The timestamp sequence includes the creation timestamp of the historical record and the last update timestamp of each data field. The time decay function uses an exponential decay form to quantify the value loss of data over time. The expression for the time decay function is: in, The value of the time decay function. The attenuation coefficient is... This is the current timestamp. This is the last update timestamp of the data field. Decay coefficient. The determination depends on the data type. For rapidly changing data types, such as operational data, Set to 0.01; for data types that change slowly, such as basic enterprise information, Set to 0.001.

[0024] The internal verification component obtains the first adjustment coefficient corresponding to the industry volatility of the target company's industry and the second adjustment coefficient corresponding to the target company's risk level. Industry volatility is obtained by calculating the standard deviation of the data update frequency of all companies in that industry over the past 12 months; the higher the industry volatility, the larger the first adjustment coefficient. The risk level is determined based on historical credit analysis results and is divided into three levels: low risk, medium risk, and high risk, with corresponding second adjustment coefficients of 0.8, 1.0, and 1.2, respectively.

[0025] The internal verification component inputs the time decay function, the first adjustment coefficient, and the second adjustment coefficient into the multidimensional spatiotemporal decay model to calculate the local timeliness score of each data field in the historical record. The multidimensional spatiotemporal decay model is constructed based on the product of the exponential decay function and the first and second adjustment coefficients, and its expression is: in, For the first The local timeliness score of each data field. The first adjustment coefficient, This is the second adjustment coefficient. For the first The time decay function value of each data field.

[0026] The internal validation component performs non-linear fusion on local timeliness scores to obtain a global timeliness score based on historical records. The non-linear fusion uses a weighted geometric average method, with weights assigned to each data field based on its importance score in the credit analysis. Importance scores are determined through expert scoring; for example, financial data has an importance score of 0.4, legal litigation data 0.3, basic enterprise information 0.1, intellectual property data 0.1, and market competitiveness data 0.1. The expression for the global timeliness score is: Among them, represents the global timeliness score. For the first Importance scores for each data field This represents the total number of data fields. The internal validation component uses the global timeliness score as the timeliness score for historical records.

[0027] refer to Figure 4 The gap analysis component receives timeliness scores from the internal verification component and structured data request instructions from the order parsing component, comparing the timeliness scores with the credit analysis accuracy threshold. The gap analysis component transforms the analysis dimensions into demand feature vectors and historical records into stock feature vectors. Both demand and stock feature vectors are binary vectors, with each dimension corresponding to an analysis dimension. If the dimension is included in the analysis dimensions or historical records, the corresponding position is 1; otherwise, it is 0.

[0028] The gap analysis component calculates the cosine similarity between the demand feature vector and the stock feature vector to obtain the dimensional coverage. The expression for cosine similarity is: in, For the demand feature vector, For existing feature vectors, The dot product of two vectors. and These are the magnitudes of the two vectors. Dimensional coverage is equal to the cosine similarity value, ranging from 0 to 1.

[0029] In response to a dimensional coverage rate falling below a preset coverage threshold, or a timeliness score falling below a credit analysis accuracy threshold, the gap analysis component triggers the data gap analysis model. The preset coverage threshold is determined based on the credit analysis accuracy threshold; the higher the credit analysis accuracy threshold, the higher the preset coverage threshold. For example, when the credit analysis accuracy threshold is 0.8, the preset coverage threshold is set to 0.9; when the credit analysis accuracy threshold is 0.5, the preset coverage threshold is set to 0.7.

[0030] The data gap analysis model extracts uncovered feature vectors based on the difference operation between the demand feature vector and the stock feature vector. The expression for the difference operation is: in, If the feature vector is not covered, then One of its dimensions is 1 and If the corresponding dimension is 0, then The corresponding dimension is 1, otherwise it is 0.

[0031] The gap analysis component maps the uncovered feature vectors to a predefined external data label space, generating the external data dimensions that need to be supplemented. The predefined external data label space contains labels for all data dimensions that can be obtained from external vendors, with each label corresponding to one or more external data fields. For example, if the legal litigation dimension in the uncovered feature vectors is 1, it will be mapped to labels such as court judgments, information on judgment debtors, information on dishonest judgment debtors, and information on restrictions on high-level consumption in the external data label space, generating external data dimensions that need to be supplemented as legal litigation-related data.

[0032] The supplier selection component receives the required external data dimensions from the gap analysis component and launches the supplier selection engine based on these dimensions. The supplier selection engine maintains a supplier database storing basic information, interface information, data field coverage, and historical call records of all cooperating external data suppliers. The supplier selection component first filters suppliers that can provide at least one required data field based on the required external data dimensions, forming a candidate supplier set.

[0033] refer to Figure 5The supplier selection component constructs a set of constraints based on the external data dimensions that need to be supplemented. This set of constraints includes a lower limit constraint on data field coverage, an upper limit constraint on average interface response time, and a lower limit constraint on data compliance score. The lower limit constraint on data field coverage is the minimum proportion of the required data fields provided by the candidate supplier combination to the total required data fields. The upper limit constraint on average interface response time is the upper limit of the weighted average response time of the candidate supplier combination. The lower limit constraint on data compliance score is the lowest data compliance score of the candidate supplier.

[0034] The supplier selection component uses the cost per call and historical data accuracy as its objective function. The expression for the objective function is: in, The objective function value, The total cost per call for the candidate supplier combination. This represents the weighted average historical data accuracy of the candidate supplier portfolio. and These are the weighting coefficients, and The weighting coefficients are determined based on the credit analysis accuracy threshold; the higher the credit analysis accuracy threshold, the higher the weighting coefficients. The larger, The smaller the value. For example, when the credit analysis accuracy threshold is 0.8, Set to 0.3, Set to 0.7; when the credit analysis accuracy threshold is 0.5, Set to 0.7, Set it to 0.3.

[0035] The supplier selection component utilizes a multi-objective optimization algorithm to perform Pareto optimization on the objective function and constraint set, generating a Pareto front solution set. The multi-objective optimization algorithm employs Non-Dominated Sorting Genetic Algorithm II (NSGA-II), which efficiently generates a uniformly distributed Pareto front solution set through fast non-dominated sorting, crowding calculation, and an elite retention strategy. Each solution in the Pareto front solution set corresponds to a supplier combination, and no other solution can reduce the cost per call without decreasing historical data accuracy, or improve historical data accuracy without increasing the cost per call.

[0036] The supplier selection component chooses the optimal supplier combination from the Pareto front solution set, based on the weighted sum of single-call cost and historical data accuracy. The expression for the weighted sum is: in, For weighted sums, select The supplier combination corresponding to the smallest solution is the optimal supplier combination.

[0037] In this embodiment, Table 1 is an example of a supplier evaluation matrix, showing the scores of the five candidate suppliers on various evaluation indicators.

[0038] Table 1 Example of Supplier Evaluation Matrix

[0039] In this embodiment, it is assumed that the external data dimension to be supplemented includes 10 data fields, the lower limit constraint for data field coverage is 90%, the upper limit constraint for average interface response time is 350ms, and the lower limit constraint for data compliance score is 90 points. Based on these constraints, candidate supplier S004 is excluded because its data compliance score is below 90 points. Among the remaining candidate suppliers S001, S002, S003, and S005, none of the individual suppliers have a data field coverage rate of 90%, therefore, a supplier combination needs to be selected. Pareto optimization was performed using the NSGA-II algorithm. The resulting Pareto front solution set includes the following three combinations: Combination 1 (S001+S003), with a single call cost of 270 yuan, data field coverage of 98%, historical data accuracy of 93.5%, and an average interface response time of 350ms; Combination 2 (S002+S003), with a single call cost of 240 yuan, data field coverage of 92%, historical data accuracy of 91.5%, and an average interface response time of 325ms; and Combination 3 (S001+S005), with a single call cost of 230 yuan, data field coverage of 90%, historical data accuracy of 91%, and an average interface response time of 290ms. Assume weight coefficients... , The weighted sums of the three combinations are as follows: Combination 1: 0.5×270+0.5×(1-0.935)×100=135+3.25=138.25; Combination 2: 0.5×240+0.5×(1-0.915)×100=120+4.25=124.25; Combination 3: 0.5×230+0.5×(1-0.91)×100=115+4.5=119.5. Therefore, the combination with the smallest weighted sum, combination 3 (S001+S005), is selected as the optimal supplier combination.

[0040] The data fusion component receives optimal supplier combination information from the supplier selection component and calls the interface of the optimal supplier combination to obtain external data. The data fusion component first generates an interface call request based on the interface information of the optimal supplier combination, carrying the target enterprise identifier, the external data dimensions to be supplemented, and authentication information. The data fusion component uses an asynchronous call method, sending interface call requests to multiple suppliers simultaneously and setting a timeout period. If an interface call from a supplier times out or returns an error, a backup supplier call mechanism is triggered, selecting a backup supplier from the candidate supplier set that can provide the same data fields for the call.

[0041] refer to Figure 6 After acquiring external data, the data fusion component performs field-level balance checks on the external data and historical records. Field-level balance checks include data type checks, data format checks, and data range checks. Data type checks verify that the data type of the external data fields matches the preset data type; for example, financial data should be numeric, and company names should be strings. Data format checks verify that the format of the external data fields conforms to preset format requirements; for example, the unified social credit code should be 18 characters, and the date format should be YYYY-MM-DD. Data range checks verify that the values ​​of the external data fields are within a reasonable range; for example, total assets should be non-negative, and the debt-to-asset ratio should be between 0 and 1.

[0042] For external data that passes validation, the data fusion component extracts numeric and text fields. For numeric fields, a cross-table logical topology validation based on reciprocal relationships is performed to identify calculation discrepancies between numeric fields. Reciprocal relationships refer to the inherent logical relationships between different financial statements or different accounts within the same financial statement, such as Assets = Liabilities + Owner's Equity, Net Profit = Total Profit - Income Tax Expense, etc. The cross-table logical topology validation based on reciprocal relationships first constructs a reciprocal relationship rule base, then substitutes the numeric fields from the external data into the rule base for validation. If any discrepancies in reciprocal relationships are found, the calculation discrepancy is calculated.

[0043] For text-based fields, conflict detection based on entity alignment is performed to identify synonyms and heteronyms. Entity alignment employs a hybrid method based on string similarity and semantic similarity. First, the string similarity between two entity names is calculated. If the string similarity is higher than a preset threshold, the semantic similarity between the two entities is further calculated. Semantic similarity is obtained by calculating the cosine similarity between the vector representations of the entity names using a pre-trained language model. If both the string similarity and semantic similarity of two entities are higher than the preset threshold, they are identified as heteronyms; if two entities have the same name but their semantic similarity is lower than the preset threshold, they are identified as synonyms.

[0044] In response to a calculation deviation exceeding a preset deviation threshold or the presence of conflicting entities identified by conflict detection, the data fusion component triggers a data rollback mechanism to reacquire external data. The rollback mechanism first records any anomalies in the currently acquired external data, then selects the next optimal supplier combination from the candidate supplier set for invocation. If multiple invocations fail to acquire the required external data, a data acquisition failure message is returned to the client.

[0045] In response to calculation deviations within a preset deviation threshold and the absence of conflicting entities, the data fusion component performs primary key mapping and overwriting merging between external data and historical records. Primary key mapping uses the target enterprise's Unified Social Credit Code as the primary key, matching each data field in the external data with its corresponding data field in the historical records. Overwriting merging employs a new data priority principle: if a data field in the external data is identical to a historical record, the value in the external data overwrites the value in the historical record; if a data field in the external data is not present in the historical record, that field is added to the historical record.

[0046] After data fusion is completed, the data fusion component generates a credit analysis report. The report includes basic enterprise information, financial status analysis, operational risk analysis, legal proceedings analysis, intellectual property analysis, market competitiveness analysis, and a comprehensive credit rating. The report is output in both PDF and JSON formats. The PDF format is used for user presentation, while the JSON format is used for subsequent data processing and storage.

[0047] This embodiment extracts credit analysis accuracy thresholds by parsing credit report orders. It then calculates a timeliness score by weighting the update frequency of internal stock data, industry volatility, and risk level. The timeliness score is compared with the accuracy threshold to identify external data dimensions that need supplementation. A multi-objective optimization algorithm is used to solve the supplier evaluation matrix to obtain the optimal supplier combination. This enables dynamic allocation of data sources on demand, avoiding blind procurement and achieving a quantitative balance between data acquisition costs and data quality while meeting analytical accuracy requirements. A credit demand graph is constructed to achieve a structured expression of order demand. Field-level balance checks, cross-table logical topology checks based on reciprocal relationships, and conflict detection based on entity alignment are performed on the acquired external data to ensure accuracy and consistency during multi-source data fusion.

[0048] In a preferred embodiment, after constructing a credit demand graph, the order parsing component further identifies the in-degree and out-degree of nodes in the credit demand graph, defining nodes with an in-degree greater than a preset in-degree threshold as core entity nodes. The in-degree of a node refers to the number of edges pointing to that node, and the out-degree refers to the number of edges originating from that node. The preset in-degree threshold is determined based on the size of the credit demand graph; for example, for a credit demand graph containing 100 nodes, the preset in-degree threshold is set to 5. Core entity nodes are typically entities most relevant to credit analysis, such as the target company itself, the target company's parent company, the target company's major suppliers, and customers.

[0049] The order parsing component extracts the credit analysis accuracy threshold corresponding to the core entity nodes and assigns dynamic weights to these nodes. The dynamic weights are positively correlated with the credit analysis accuracy thresholds; the higher the accuracy threshold, the greater the dynamic weight. The expression for the dynamic weights is: in, The dynamic weights of the core entity nodes. This is the proportionality coefficient. This is the threshold for credit analysis accuracy. (Scale factor) The ratio is determined based on the type of core entity node. For example, the ratio of the target company itself is set to 1.2, the ratio of the target company's parent company is set to 1.0, and the ratio of the target company's main suppliers and customers is set to 0.8.

[0050] The order parsing component monitors the time span of time-series dependencies. In response to a time span exceeding a preset time window, it injects time-series decay weights into edges connecting core entity nodes and edge entity nodes. Edge entity nodes are those with both in-degree and out-degree less than a preset threshold, typically entities with low relevance to credit analysis. The preset time window is determined based on the analysis dimension; for example, a 3-year window is used for financial status analysis, while a 1-year window is used for legal proceedings analysis. The time-series decay weight is negatively correlated with the time span; the larger the time span, the smaller the decay weight. The expression for the time-series decay weight is: in, For time-series decay weights, The attenuation coefficient is... The time span of the temporal dependency. This is a preset time window. hour, Set to 1.0.

[0051] The order parsing component prunes the credit demand graph based on dynamic weights and time-series decay weights, removing edges with time-series decay weights below a preset decay threshold to generate a simplified credit demand graph. The preset decay threshold is set to 0.5, meaning edges with time-series decay weights below 0.5 will be removed. After pruning, the size of the credit demand graph is significantly reduced while retaining the entities and relationships most relevant to credit analysis, improving the efficiency of subsequent data processing.

[0052] The internal validation component further acquires macroeconomic volatility indices and industry policy change events. Macroeconomic volatility indices include GDP growth rate, Consumer Price Index (CPI), Producer Price Index (PPI), and Purchasing Managers' Index (PMI), which are obtained from the National Bureau of Statistics and relevant economic research institutions. Industry policy change events include the introduction of industry regulatory policies, adjustments to tax policies, and the release of industry support policies, which are obtained through web crawling from government websites, news media, and industry association websites.

[0053] The internal validation component encodes macroeconomic volatility indices and industry policy change events into external shock vectors. The encoding process combines one-hot encoding and numerical encoding. For numerical macroeconomic volatility indices, their values ​​are directly used as the corresponding dimension of the vector. For event-based industry policy change events, one-hot encoding is used; if a certain type of policy change event occurs, the corresponding dimension is 1, otherwise it is 0. The dimension of the external shock vector is determined based on the number of macroeconomic volatility indices and the number of types of industry policy change events.

[0054] The internal verification component inputs the external impact vector into the multidimensional spatiotemporal decay model, dynamically correcting the first and second adjustment coefficients. The expressions for the corrected first and second adjustment coefficients are as follows: in, and These are the corrected first and second adjustment coefficients, respectively. For the external impact vector, and These are the impact weight vectors for the first and second adjustment coefficients, respectively. The impact weight vectors are trained using historical data and are used to quantify the impact of different external shocks on industry volatility and risk levels. For example, when a strict regulatory policy is introduced for an industry, its volatility will increase significantly, resulting in a larger dimension for the corresponding impact weight vector.

[0055] The internal verification component updates the local timeliness score based on the revised first and second adjustment coefficients. The expression for the updated local timeliness score is: in, For the updated number The local timeliness score of each data field.

[0056] The internal validation component extracts the rate of change of time intervals between adjacent update cycles from the historical records. This rate of change is used as a decay acceleration factor and added to the time decay function to adjust the convergence speed of the global timeliness score. The expression for the rate of change of time intervals is: in, The rate of change of the time interval. For the first The update cycle and the first The time interval between update cycles For the first The update cycle and the first The time interval between update cycles. If This indicates that the update cycle is longer, the data update frequency is reduced, the decay acceleration factor is greater than 1, and the convergence speed of the time decay function is accelerated; if This indicates that the update cycle is shortened, the data update frequency is increased, the decay acceleration factor is less than 1, and the convergence speed of the time decay function is slowed down.

[0057] The expression for the time decay function after adding the decay acceleration factor is: in, This is the value of the time decay function after adding the decay acceleration factor.

[0058] In this embodiment, Table 2 is a table showing the correlation between volatility and adjustment coefficients for different industries, displaying the industry volatility, first adjustment coefficient, and adjusted first adjustment coefficient under different macroeconomic shocks for 10 major industries.

[0059] Table 2. Correspondence between volatility and adjustment coefficient for different industries

[0060] In this embodiment, it is assumed that the target company belongs to the real estate industry, with an industry volatility of 0.20 and a first adjustment coefficient of 1.20. When the macroeconomy experiences a 1% drop in GDP growth, the adjusted first adjustment coefficient becomes 1.32; when real estate industry regulatory policies are simultaneously introduced, the adjusted first adjustment coefficient becomes 1.45. The target company's risk level is medium risk, and the second adjustment coefficient is 1.0. If the last update time of a certain data field is 6 months ago, the decay coefficient... The value is 0.01, representing the current timestamp. With the last update timestamp of the data field If the difference is 180 days, then the time decay function value is... Uncorrected local timeliness score Corrected local timeliness score It can be seen that the timeliness score of the data has decreased due to the impact of macroeconomic shocks and changes in industry policies, reflecting the impact of changes in the external environment on the value of the data.

[0061] The internal verification component recalculates the global timeliness score based on the updated local timeliness score and sends the recalculated global timeliness score to the gap analysis component. The gap analysis component re-performs timeliness comparison and gap analysis based on the updated global timeliness score. If a new data gap is found, it triggers the supplier selection component to reselect the optimal supplier combination.

[0062] This embodiment improves data processing efficiency and accuracy by pruning the credit demand graph, eliminating redundant information. By introducing macroeconomic fluctuation indices and industry policy change events to dynamically adjust the adjustment coefficients, the timeliness score calculation becomes more realistic and accurately reflects the true value of the data. Furthermore, by introducing a decay acceleration factor, the convergence speed of the time decay function can be adjusted according to changes in data update frequency, further improving the accuracy of the timeliness score calculation.

[0063] In a preferred embodiment, the supplier selection component further acquires historical order data acquisition logs, which include actual call costs, actual response times, and data quality backtracking scores. These logs are stored in a distributed log system, recording detailed information on all past external data calls, including call time, supplier ID, interface name, request parameters, response result, actual call cost, actual response time, and data quality score. The data quality backtracking score is determined based on subsequent credit analysis and manual review. If the external data matches the actual situation, the data quality backtracking score is 1; if the external data contains errors, a score between 0 and 1 is assigned based on the severity of the error.

[0064] The supplier selection component uses actual call cost, actual response time, and data quality backtesting score as feedback reward signals. Based on these signals, the data compliance score and historical data accuracy in the supplier evaluation matrix are updated. The update process employs an exponential moving average method, which smooths out historical data fluctuations while assigning higher weight to recent data. The expression for the updated historical data accuracy is: in, To update the accuracy of historical data, The data quality backtesting score is given as follows: The accuracy of historical data before the update. This is the smoothing coefficient, with a value ranging from 0 to 1. The frequency of data acquisition determines the likelihood of higher data acquisition frequencies. The larger the number of calls, the better. For example, for a vendor with more than 100 calls per month, Set to 0.3; for suppliers that make fewer than 10 calls per month, Set it to 0.1.

[0065] The updated expression for the data compliance score is: in, For the updated data compliance score, For this data compliance score, This is the data compliance score before the update. This data compliance score is determined based on the compliance during the data acquisition process. If the supplier strictly adheres to data privacy protection regulations and contractual agreements, then… The score is 100 points; if there is a violation, the score will be between 0 and 100 points depending on the severity of the violation.

[0066] In the Pareto optimization process, the supplier selection component introduces a dynamic penalty factor. In response to a data compliance score falling below the compliance red line threshold, this dynamic penalty factor is applied to the cost per call in the objective function to exclude supplier combinations that do not meet compliance requirements. The compliance red line threshold is set at 80 points; suppliers with a data compliance score below 80 points are considered high-risk suppliers. The expression for the dynamic penalty factor is: in, As a dynamic penalty factor, To meet the compliance red line threshold, Assign a score to the supplier's data compliance, with a penalty index set to 2. When hour, Set to 1.0, meaning no penalty is applied; when hour, ,and The smaller, The larger.

[0067] The objective function after applying the dynamic penalty factor is expressed as follows: in, Let be the objective function value after applying the dynamic penalty factor, and be the supplier combination. For suppliers The dynamic penalty factor, For suppliers The cost per call.

[0068] The data fusion component's process of performing cross-table logical topology verification based on reconciliation relationships is further configured to construct a directed graph of accounting equations between the balance sheet and the income statement. The nodes of this directed graph represent financial accounts, and the edges represent operational relationships. The directed graph is constructed based on corporate accounting standards and includes all major financial accounts in the balance sheet and income statement, along with the operational relationships between them. For example, accounts such as cash and cash equivalents, trading financial assets, notes receivable, and accounts receivable are connected to the total current assets account through addition operations; total current assets and total non-current assets are connected to the total assets account through addition operations; and operating profit is obtained by subtracting operating costs, taxes and surcharges, selling expenses, administrative expenses, and financial expenses from operating revenue.

[0069] In this embodiment, Table 3 is a definition table of nodes and edges of the directed graph of accounting equations, showing some financial subject nodes and the operational relationship edges between them.

[0070] Table 3. Definition of Nodes and Edges in the Directed Graph of the Accounting Equation

[0071] The data fusion component substitutes numerical fields from external data into the directed graph of the accounting equation for propagation calculations, obtaining the calculated values ​​of the terminal nodes. The propagation calculation starts from the leaf nodes and proceeds along the edges, calculating the values ​​of each intermediate and terminal node sequentially. Leaf nodes are nodes without incoming edges, such as cash, trading financial assets, and operating revenue. Terminal nodes are nodes without outgoing edges, such as total assets and net profit.

[0072] The data fusion component compares the calculated value with the actual value of the corresponding field in the external data to obtain the difference ratio. The expression for the difference ratio is: in, For the proportion of difference, For calculated values, This is the actual value.

[0073] In response to a discrepancy exceeding the logical consistency threshold, the data fusion component performs reverse tracing along the edges of the directed graph of the accounting equation to locate the source financial account introducing the calculation error. The logical consistency threshold is set to 5%, meaning that a discrepancy exceeding 5% is considered a logical error. Reverse tracing starts from the terminal node with the discrepancy and sequentially checks the difference between the calculated and actual values ​​of each intermediate and leaf node in the reverse direction of the edges until the first node with a discrepancy is found; this node is the source financial account introducing the calculation error.

[0074] The data fusion component marks source financial items with anomaly tags and corrects and replaces them based on the corresponding historical financial item values ​​in the historical records. The correction and replacement uses a weighted average method, with the weights being the timeliness score of the historical data. The expression for the corrected value is:

[0075] in, These are the corrected values. For the first The timeliness score of the corresponding financial item in the historical record. For the first The value of the corresponding financial item in the historical record. This represents the number of historical records.

[0076] In this embodiment, assume the following financial data of a company in the external data: cash and cash equivalents of 10 million yuan, trading financial assets of 2 million yuan, notes receivable of 3 million yuan, accounts receivable of 5 million yuan, prepayments of 1 million yuan, inventory of 4 million yuan, total current assets of 24 million yuan, total non-current assets of 16 million yuan, total assets of 40 million yuan, total liabilities of 20 million yuan, and total equity of 18 million yuan. Substituting these values ​​into the directed graph of the accounting equation for propagation calculation, the calculated value of total current assets is 1000+200+300+500+100+400=25 million yuan. The difference between this and the actual value of 24 million yuan is (2500-2400) / 2400×100%≈4.17%, which does not exceed the logical self-consistency threshold. The calculated total assets were 25 million + 16 million = 41 million yuan, differing from the actual value of 40 million yuan by (41 million - 40 million) / 40 million × 100% = 2.5%, which did not exceed the logical consistency threshold. The calculated value of the accounting equation was 20 million + 18 million = 38 million yuan, differing from the actual value of 40 million yuan by (40 million - 38 million) / 40 million × 100% = 5%, which was exactly equal to the logical consistency threshold. Tracing back along the directed graph of the accounting equation, it was found that the actual value of total equity of 18 million yuan differed from the calculated value. Further examination of the components of total equity revealed an error in the actual value of retained earnings. Based on the historical value of retained earnings, a correction was made, resulting in a total equity of 20 million yuan. The accounting equation became 20 million + 20 million = 40 million yuan, consistent with the total assets, thus logically consistent.

[0077] After the data fusion component completes the data correction, it merges the corrected external data with historical records to generate a credit analysis report. The credit analysis report includes a data correction explanation, detailing the location, cause, and correction methods of the abnormal data, so that users can understand the data processing procedure.

[0078] This embodiment updates the supplier evaluation matrix by introducing a historical feedback mechanism, making supplier selection more accurate and reliable. By introducing a dynamic penalty factor, it can effectively exclude supplier combinations that do not meet compliance requirements, reducing compliance risks in data acquisition. By constructing a directed graph of the accounting equation for propagation calculation and reverse tracing, it can accurately locate the source of logical deviations in financial data and correct and replace them, ensuring the logical consistency and accuracy of the financial data.

Claims

1. A credit data dynamic orchestration and strategy optimization system based on order awareness, characterized in that, include: The order parsing component is configured to receive credit report orders, parse the credit report orders, and extract the target enterprise identifier, analysis dimension, and credit analysis accuracy threshold. An internal verification component is configured to retrieve matching historical records from an internal historical database based on the target company's identifier, and to perform a weighted calculation based on the data update frequency of the historical records, the industry volatility of the target company's industry, and the risk level of the target company to obtain a timeliness score for the historical records. The gap analysis component is configured to compare the timeliness score with the credit analysis accuracy threshold, and in response to the timeliness score not meeting the credit analysis accuracy threshold or the historical records not covering the analysis dimension, trigger the data gap analysis model to identify the external data dimensions that need to be supplemented. The supplier selection component is configured to launch a supplier selection engine based on the external data dimension, and use a multi-objective optimization algorithm to solve the supplier evaluation matrix, which includes single call cost, data field coverage, historical data accuracy, average interface response time and data compliance score, to obtain the optimal supplier combination. The data fusion component is configured to call the interface of the optimal supplier combination to obtain external data, perform field-level balance verification on the external data and the historical records, and fuse the external data that passes the verification with the historical records to generate a credit analysis report.

2. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 1, characterized in that, The order parsing component is further configured to: semantically decompose the analysis dimension and extract the implicit temporal dependencies and entity associations in the analysis dimension; Based on the temporal dependency relationship and the entity association relationship, the target enterprise identifier and the analysis dimension are mapped into a multi-hop graph structure to construct a credit demand graph; In the credit demand graph, nodes represent data entities, edges represent the temporal dependencies and the associations between the entities, and node attributes include the credit analysis accuracy threshold and the data timeliness weight. A structured data request instruction is generated based on the credit demand graph, and the structured data request instruction carries the topological constraints of the credit demand graph.

3. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 1, characterized in that, The internal verification component is further configured to: extract the timestamp sequence of the historical records and construct a time decay function; Obtain the first adjustment coefficient corresponding to the industry volatility and the second adjustment coefficient corresponding to the risk level; The time decay function, the first adjustment coefficient, and the second adjustment coefficient are input into the multidimensional spatiotemporal decay model to calculate the local timeliness score of each data field in the historical record. The local timeliness scores are nonlinearly fused to obtain the global timeliness scores of the historical records; The global timeliness score is used as the timeliness score, wherein the multidimensional spatiotemporal decay model is constructed based on the product of the exponential decay function and the first adjustment coefficient and the second adjustment coefficient.

4. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 1, characterized in that, The gap analysis component is further configured to: convert the analysis dimension into a demand feature vector and convert the historical record into a stock feature vector; Calculate the cosine similarity between the demand feature vector and the stock feature vector to obtain the dimension coverage. The data gap analysis model is triggered in response to the dimensional coverage rate being lower than a preset coverage threshold, or the timeliness score being lower than the credit analysis accuracy threshold. The data gap analysis model extracts the uncovered feature vector based on the difference operation between the demand feature vector and the stock feature vector; The uncovered feature vectors are mapped to a preset external data label space to generate the external data dimensions that need to be supplemented.

5. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 4, characterized in that, The supplier selection component is further configured to: construct a set of constraints based on the external data dimensions to be supplemented, the set of constraints including a lower limit constraint on data field coverage, an upper limit constraint on average interface response time, and a lower limit constraint on data compliance score; The cost of a single call and the accuracy of historical data are used to construct an objective function; The multi-objective optimization algorithm is used to perform Pareto optimization on the objective function and the set of constraints to generate a Pareto front solution set; The optimal supplier combination is selected from the Pareto front solution set by weighting the single call cost and the historical data accuracy.

6. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 1, characterized in that, The data fusion component is further configured to extract numeric fields and text fields from the external data; For the numeric fields, perform cross-table logical topology validation based on reconciliation relationships to identify calculation discrepancies between the numeric fields; For the text field, perform conflict detection based on entity alignment to identify synonyms and heteronyms; In response to the calculation deviation exceeding a preset deviation threshold or the existence of a conflict entity identified by the conflict detection, a data rollback mechanism is triggered to reacquire the external data. In response to the calculation deviation being within the preset deviation threshold and the absence of conflicting entities, the external data and the historical records are mapped and overwritten and merged using primary keys.

7. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 2, characterized in that, The process of constructing a credit demand graph by the order parsing component is further configured as follows: identifying the in-degree and out-degree of nodes in the credit demand graph, and defining nodes with an in-degree greater than a preset in-degree threshold as core entity nodes; Extract the credit analysis accuracy threshold corresponding to the core entity node, and assign dynamic weights to the core entity node; Monitor the time span of the time-series dependency, and in response to the time span being greater than a preset time window, inject time-series decay weights into the edges connecting the core entity nodes and the edge entity nodes; The credit demand graph is pruned based on the dynamic weight and the time-series decay weight, removing edges whose time-series decay weight is lower than a preset decay threshold, thereby generating a simplified credit demand graph.

8. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 3, characterized in that, The internal verification component is further configured to: acquire macroeconomic volatility index and industry policy change events; The macroeconomic volatility index and the industry policy change event are encoded into an external shock vector; The external impact vector is input into the multidimensional spatiotemporal decay model to dynamically correct the first adjustment coefficient and the second adjustment coefficient. The local timeliness score is updated based on the corrected first adjustment coefficient and the second adjustment coefficient; Extract the time interval change rate of adjacent update cycles in the historical records, and use the time interval change rate as a decay acceleration factor to be superimposed on the time decay function in order to adjust the convergence speed of the global timeliness score.

9. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 5, characterized in that, The supplier selection component is further configured to: obtain historical order data acquisition logs, the data acquisition logs including actual call costs, actual response times, and data quality backtracking scores; The actual call cost, the actual response time, and the data quality backtracking score are used as feedback reward signals. The data compliance score and historical data accuracy in the supplier evaluation matrix are updated based on the feedback reward signal. In the Pareto optimization process, a dynamic penalty factor is introduced. In response to the data compliance score falling below the compliance red line threshold, the dynamic penalty factor is applied to the single call cost in the objective function to exclude supplier combinations that do not meet compliance requirements.

10. The order-aware credit data dynamic orchestration and strategy optimization system according to claim 6, characterized in that, The process of the data fusion component performing cross-table logical topology verification based on reconciliation relationships is further configured as follows: constructing a directed graph of accounting equations between the balance sheet and the income statement, wherein the nodes of the directed graph of accounting equations are financial accounts and the edges are operational relationships; The numerical fields in the external data are substituted into the directed graph of the accounting equation for propagation calculation to obtain the calculated values ​​of the end nodes; Compare the calculated value with the actual value of the corresponding field in the external data to obtain the difference ratio; In response to the discrepancy ratio exceeding the logical self-consistency threshold, reverse tracing is performed along the edges of the directed graph of the accounting equation to locate the source financial account that introduced the calculation deviation. The source financial account is marked with an anomaly label, and the corresponding historical financial account value is corrected and replaced based on the historical record.