Intelligent market supervision data management system and method based on multi-stage data sharing and exchange
By introducing multi-level data sharing and exchange modules and other management modules into the market supervision data governance system, the complex authority management and insufficient privacy protection of market supervision data in cross-departmental, cross-regional, and cross-level sharing are solved, and the security, efficient sharing and high-quality analysis of data are achieved, and data governance and analysis capabilities are improved.
Patent Information
- Application Number
- CN202510079085.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In the sharing and collaboration between cross-departmental, cross-regional and cross-level market supervision data, there are problems such as complex permission management, insufficient privacy protection, low data access efficiency, inconsistent data standards, uneven data quality and lack of efficient analytical models.
An intelligent market supervision data governance system based on multi-level data sharing and exchange is proposed, including data sharing and exchange module, data standard management module, data quality management module, data model management module and data asset management module. Through technical means such as strategy-driven data access control, dynamic resource directory management, standardized management of the entire life cycle, closed-loop quality governance process, optimization and analysis of multi-star data models, and dynamic storage adjustment based on access popularity, data security, efficient sharing and high-quality analysis are achieved.
It realizes efficient sharing and exchange of market supervision data across departments, regions and levels, ensures data security and controllability, improves the reliability and analysis efficiency of data quality, fully taps the value of data, and supports accurate and reliable decision-making.
Smart Images

Figure CN120013333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance and market supervision information technology, and specifically to an intelligent market supervision data governance system and method based on multi-level data sharing and exchange. Background Art
[0002] With the rapid development of informatization, the amount and complexity of market supervision data are increasing, involving various departments, regions and levels. The decentralized storage and isolated management of data make cross-departmental, cross-regional and cross-level data sharing and collaboration difficult. At the same time, due to the lack of unified data standards, uneven quality and lack of efficient analysis models, market supervision data faces the following major problems in use:
[0003] The data sharing process has problems such as complex permission management, insufficient privacy protection, and low data access efficiency;
[0004] The diversity and heterogeneity of data lead to a lack of unified standards for shared data, which cannot meet the actual needs of market supervision;
[0005] Due to the diversity of data sources and the lack of an effective governance mechanism, the identification, governance, and tracking of problematic data are inefficient, affecting the accuracy of subsequent analysis;
[0006] Cross-system and cross-business data correlation analysis relies on complex model construction, and existing technologies are difficult to support multi-dimensional and high-frequency data integration and analysis;
[0007] Storage strategies and access efficiency are not combined with dynamic adjustment mechanisms, and there is a lack of effective means for tracking and change management of data links, making it difficult to fully tap the value of assets. Summary of the invention
[0008] In response to the above problems, the present invention proposes an intelligent market supervision data governance system and method based on multi-level data sharing and exchange, providing a full-process solution from data sharing, standardization, quality governance to model analysis and asset management.
[0009] The present invention achieves the above-mentioned purpose through the following technical solutions:
[0010] An intelligent market supervision data governance system based on multi-level data sharing and exchange, the system comprising:
[0011] The data sharing and exchange module is used to share and exchange market supervision data across departments, regions and levels, and implements data access control and dynamic resource directory management in combination with policy-driven. The data resource directory and shared data generated by the data sharing and exchange module are used by other modules;
[0012] The data standard management module is used to standardize the entire life cycle of market supervision data and connect with the data sharing and exchange module to ensure that shared data meets standardization requirements. It optimizes the standardization process by dynamically updating data standards and version difference analysis, and provides standardized data for processing by the data quality management module;
[0013] The data quality management module is used to manage the quality of market supervision data throughout its life cycle. It receives standardized data and forms a closed-loop data quality management process through the identification, allocation, management and re-testing of problematic data, ensuring high-quality data for use by the data model management module.
[0014] The data model management module is used to optimize and analyze multidimensional data based on the multi-satellite data model, receive high-quality data to build a correlation model across business systems, realize data integration analysis, and provide the analysis results to the data asset management module to improve data processing efficiency through scheduling optimization;
[0015] The data asset management module is used to manage data assets in an all-round way, store, optimize and display the analysis results output by the data model management module, including dynamically adjusting the data storage strategy based on access popularity, improving data usage efficiency in combination with data link analysis, and providing data change impact analysis function.
[0016] As a preferred solution of the present invention, the data sharing and exchange module includes:
[0017] The policy-driven unit is used to define data sharing policies, including:
[0018] Data access strategy: Maintain an independent private blockchain ledger through the Byzantine fault-tolerant consensus protocol for cross-departmental, cross-regional, and cross-level data operation records. Maintain a cross-domain public blockchain ledger through a lightweight proof-of-work mechanism, only record the checkpoint information of cross-domain operations, and use smart contracts in the blockchain network to automatically define data access rights, privacy protection constraints, and data integrity verification rules to achieve automated permission control for cross-domain data sharing;
[0019] Data processing strategy: Privacy-preserving data modeling is achieved through a federated learning module, and collaborative modeling is performed among multiple institutions. The federated learning module adjusts the model update weight according to the data volume and data quality score of the participants, groups similarly distributed data through a clustering algorithm, improves the model's adaptability to non-independent and identically distributed data, combines differential privacy technology and a compressed parameter transmission mechanism to reduce the communication cost of updating the federated learning model, and adopts a compute-to-data mode that allows computing tasks to be executed at the data source and only outputs aggregated statistical results to ensure data security.
[0020] Data retention policy: used to limit the retention period of market regulatory data and automatically delete or archive the data after the period has expired;
[0021] Approval and verification strategy: used to define the approval process and rules before data sharing. After verifying the sharing request through the rule engine, the approval process is generated;
[0022] The optimization unit is used to build the data resource directory into a graph structure based on the graph database and graph algorithm, optimize the data retrieval path using the shortest path algorithm, dynamically optimize the data exchange rules using the DQN algorithm based on the dynamic rule adjustment function of reinforcement learning, and realize real-time load balancing.
[0023] As a preferred solution of the present invention, the lightweight proof-of-work mechanism specifically includes:
[0024] Each node in the network organizes the current transaction records into a Merkle tree and calculates the root hash value as the basis for the checkpoint record;
[0025] The difficulty of each proof-of-work task is dynamically adjusted according to the following formula:
[0026]
[0027] Where D is the difficulty of the current task; min and D max are the lower and upper limits of difficulty respectively; D prev is the difficulty of the last task; T target is the target time, the default value is 1 minute; T actual is the actual time to complete the previous task;
[0028] The node attempts to calculate a hash value that meets the difficulty requirement. The first node to complete the calculation broadcasts the checkpoint record and the proof result. Other nodes check whether the hash value meets the difficulty requirement through batch verification and use multi-threaded asynchronous processing. If the verification passes, the checkpoint record is stored on the public blockchain, and only the operation summary information is recorded;
[0029] In cross-domain operations, nodes in the local domain quickly reach consensus through Byzantine fault-tolerant consensus, and then submit cross-domain checkpoints to the public blockchain to form a hierarchical consensus process;
[0030] Based on smart contracts, only checkpoint records with high access frequency are retained. The high access frequency is judged by setting a time window to count the number of accesses and setting a threshold. The statistical results are updated every 30 minutes, sorted by the PageRank algorithm, and the storage strategy is dynamically adjusted to reduce redundant data.
[0031] As a preferred solution of the present invention, the calculation to data mode is specifically:
[0032] The requester sends the task to the data source through a secure channel, including the calculation logic, input parameters, and privacy protection requirements;
[0033] The data source uses a dynamic scheduling algorithm to allocate tasks according to the current resource load, supports multi-task parallel execution, prioritizes high-priority tasks, and runs tasks in an isolated environment;
[0034] Use DBSCAN or adaptive density clustering algorithm to automatically select clustering parameters according to local data distribution, including the maximum distance ε between sample points and the minimum number of sample points MinPts. The minimum number of sample points MinPts is set to twice the sample dimension in the data set. The maximum distance ε between sample points is estimated by the following formula:
[0035] ε=mean(d)+k×std(d);
[0036] In the formula, d is the distance set between sample points in the data set, mean(d) represents the average value of the distance between all sample points in the data set, std(d) represents the standard deviation of the distance between sample points in the data set; k is the adjustment coefficient;
[0037] The intermediate output data of the task is perturbed by differential privacy technology, adding noise to prevent the leakage of sensitive information. The calculation formula of the noise intensity b is: Where △f is the function sensitivity and δ is the privacy budget;
[0038] Prune or quantize the task results, and use homomorphic encryption technology that supports addition and multiplication to encrypt the results for transmission;
[0039] After receiving the result, the requester decrypts and verifies the correctness of the data. If the verification passes, the result is stored in the data asset library for subsequent analysis;
[0040] The data source cleans or archives temporary data according to the data retention policy to optimize resource utilization efficiency and prevent sensitive data from being retained for a long time.
[0041] As a preferred solution of the present invention, the data standard management module includes:
[0042] Dynamic standard maintenance unit, used to achieve real-time optimization and dynamic maintenance of data standards. It records the version information of each update through process management based on the smallest data unit, generates standard optimization reports based on version difference analysis, and automatically identifies and feedbacks deviations between data standards and market regulatory requirements.
[0043] The non-standard data detection unit is used to perform automatic detection of metadata and data content according to preset standard specifications, generate a conformity analysis report based on the detection results, and provide improvement suggestions for problematic data through difference analysis with current standards;
[0044] The master data management unit is used to identify and track the master data of multi-department collaboration through centralized coding and dynamic management strategies, automatically evaluate the impact of master data changes on related data and business processes, and output optimization plans to ensure consistency and efficient sharing of cross-department data;
[0045] The standard fit analysis unit is used to generate a scoring report by comparing the degree of fit between the data table instance and the data standard, dynamically adjust the standard content based on the scoring results and actual usage feedback, and generate actionable improvement suggestions for the weak links in the implementation of the standard.
[0046] As a preferred solution of the present invention, the data quality management module includes:
[0047] The problem data identification unit is used to perform automatic detection on the data in the system according to the preset data quality rules, dynamically adjust the detection rule weights in combination with the improved multi-level rule engine algorithm, and build a hierarchical weighted model with data characteristics as factors, which is expressed as:
[0048]
[0049] Where W i is the detection rule weight; α i is the priority value of the data feature, n is the total number of data features; β i is the rule fitness;
[0050] The data quality closed-loop governance unit is used to manage problem data throughout its life cycle. It introduces a dynamic progress prediction algorithm P=f(Q,T,R), where P is the estimated governance completion time, Q is the data quality score, T is the data volume, R is the resource allocation efficiency, and f represents the function mapping relationship. It optimizes the governance task allocation process to ensure that problem data is managed in the shortest time possible, and adjusts the governance priority in real time during the re-detection phase to form a closed-loop governance process.
[0051] The data quality scoring unit is used to quantitatively score the data quality based on the defined data detection rules and governance results, and adopts a dynamic fuzzy comprehensive evaluation algorithm to combine the data characteristics with the governance results and generate a quality detection report. The dynamic fuzzy comprehensive evaluation algorithm is expressed as:
[0052]
[0053] In the formula, S is the total score of data quality; μ iis the weight of data feature i; τ ij is the scoring result of data feature i under rule j; w ij is the weight of rule j on data feature i; m is the total number of evaluation rules;
[0054] The manual data governance support unit is used to assign problematic data to specific personnel for verification and processing based on the data's region of origin, business type, or custom label, combined with a workload balancing algorithm Where L is the workload per person, W is the total task volume, and N is the number of people. Tasks are dynamically allocated to avoid waste of resources, and batch export of problem data and monitoring of governance progress are supported.
[0055] As a preferred solution of the present invention, the data model management module includes:
[0056] Multi-star data model design unit, used to build a unified model across business systems based on multiple fact tables and shared dimension tables, supporting multi-dimensional data analysis and integration, expanding the semantic association capabilities of the model through knowledge graph technology, and combining relationship inference algorithms to achieve efficient integration of heterogeneous data;
[0057] Model conversion and optimization unit, which is used to dynamically switch between the star model and the snowflake model according to specific business needs, optimize data query performance through adaptive scheduling mechanism, and improve data processing efficiency in combination with dynamic partitioning strategy;
[0058] The scenario-driven analysis unit is used to perform risk prediction, trend analysis and hotspot detection based on the lazy data model. The lazy data model is implemented through a lazy loading mechanism. Specifically, the required data fragments are dynamically loaded only when a specific analysis task is received. The unrequested data is kept in the storage and does not occupy memory resources. After loading, the cache strategy is combined to avoid repeated loading. In addition, the generative adversarial network technology is combined to simulate business boundary scenarios, evaluate the adaptability of the model under extreme conditions, and provide decision support for market supervision through analysis results.
[0059] Data lineage analysis and early warning unit, used to trace data sources and change links in real time during data processing, display data processing paths through visualization, monitor model output quality in combination with dynamic rules, and trigger early warning mechanisms when abnormalities occur, providing feedback through system announcements, SMS reminders, and email notifications;
[0060] The data processing configuration and scheduling unit is used to provide a browser-based visual interface, combined with an automatic scheduling mechanism driven by reinforcement learning, to dynamically adjust task order and resource allocation according to real-time task load and priority.
[0061] As a preferred solution of the present invention, the multi-star data model design unit introduces a dynamic semantic expansion technology based on deep learning, optimizes the semantic association strength of the knowledge graph in real time through the attention mechanism, and mines implicit relationships in combination with multi-hop path analysis to achieve deep association of cross-domain data;
[0062] Using knowledge graph embedding technology, nodes and relationships are mapped to low-dimensional vector space, and efficient relationship inference is achieved through a distributed computing framework, supporting semantic integration of heterogeneous data sources;
[0063] The semantic conflicts generated during the integration process are dynamically adjusted using a reinforcement learning algorithm. The semantic conflict adjustment is implemented based on the DQN algorithm. The conflict is identified by building a conflict discrimination function. The reinforcement learning model updates the reward function in each adjustment to maximize the consistency of the association relationship, ensuring the rationality and consistency of the semantic association results and generating optimized high-order semantic relationships.
[0064] The dynamic adjustment mechanism of semantic priority and real-time expansion function can improve the adaptability and integration efficiency of the model in complex business scenarios.
[0065] As a preferred solution of the present invention, the data asset management module includes:
[0066] The data resource directory unit is used to form a hierarchical directory through the metadata description of the data resource to support the retrieval and acquisition of the data resource;
[0067] Data asset heat analysis unit, which is used to dynamically calculate the heat of data based on the access frequency, citation count, and browsing count of data, and adjust the storage and access strategies to optimize system performance;
[0068] The data link analysis unit is used to combine SQL logs and stored procedure analysis to automatically generate a link relationship map of metadata. The update cycle of the link relationship map is dynamically adjusted according to the system log traffic. The optimized analysis process is achieved through distributed log analysis technology, and version change notification and difference analysis functions are provided.
[0069] A data governance method for a smart market supervision data governance system based on multi-level data sharing and exchange, the method comprising:
[0070] Share and exchange market supervision data across departments, regions and levels, and implement data access control and dynamic resource catalog management in a policy-driven manner;
[0071] Conduct standardized management of market supervision data throughout its life cycle to ensure that shared data meets standardization requirements, and optimize the standardization process through dynamic updates of data standards and version difference analysis;
[0072] Conduct quality management for the entire life cycle of market supervision data, receive standardized data, and form a closed-loop data quality management process through the identification, allocation, management and re-testing of problematic data;
[0073] Optimize and analyze multidimensional data based on multi-satellite data models, receive high-quality data to build cross-business system correlation models, and realize data integration analysis;
[0074] Comprehensively manage data assets, store, optimize and display analysis results of data models, including dynamically adjusting data storage strategies based on access popularity, and improving data utilization efficiency in combination with data link analysis.
[0075] The beneficial effects of the present invention are as follows: through the data sharing and exchange module, efficient sharing and exchange of market supervision data across departments, regions and levels is realized, and the security and controllability of the data sharing process are ensured by combining policy-driven authority management and dynamic resource directory functions, providing a unified data call basis for other modules; the data standard management module ensures that the shared data meets the unified standard requirements by dynamically updating the data standard and version difference analysis, providing high-quality and standardized data input for the data quality management module, and building a closed-loop quality governance process for standardized data, including the identification, governance and re-detection of problem data, thereby improving the reliability of data quality; the data model management module realizes the correlation analysis across business systems based on the multi-star data model, efficiently integrates multidimensional data, and provides optimized analysis results for the data asset management module; the data asset management module dynamically adjusts the storage strategy using the access heat, and combines the data link analysis function to optimize the data storage and management efficiency, and fully tap the data value. Through the collaborative work between modules, the system realizes the full process coverage from data sharing to governance, analysis and management, improves the governance capacity and intelligent level of market supervision data, and provides accurate and reliable decision support for supervision. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0077] in:
[0078] Figure 1 It is a system modular structure diagram of the present invention;
[0079] Figure 2 A schematic diagram of a process of data quality management in an embodiment of the present invention;
[0080] Figure 3 This is a schematic diagram of the interface of data asset heat analysis in an embodiment of the present invention;
[0081] Figure 4 4 is a flow chart of a method in an embodiment of the present invention. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution of the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all of the embodiments. Based on the described embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.
[0083] like Figure 1-Figure 3 As shown in FIG. 1 , an embodiment of the present invention is provided, which provides an intelligent market supervision data governance system based on multi-level data sharing and exchange, including:
[0084] (1) Data sharing and exchange module
[0085] It is used to share and exchange market supervision data across departments, regions and levels, and implement data access control and dynamic resource directory management in combination with policy-driven. The data resource directory and shared data generated by the data sharing and exchange module are used by other modules to achieve safe and efficient data circulation.
[0086] In one embodiment, the data sharing exchange module includes:
[0087] Policy-driven unit, used to define data sharing policies, including data access policies, data processing policies, data retention policies, approval and verification policies, etc.
[0088] Data access policies are used to dynamically define the access scope, access purpose and access object of data sharing, and control data access according to permission rules. Specifically, independent private blockchain ledgers are maintained through the Byzantine Fault Tolerant consensus protocol for cross-departmental, cross-regional and cross-level data operation records, and cross-domain public blockchain ledgers are maintained through lightweight proof of work, only the checkpoint information of cross-domain operations is recorded, and data access rights, privacy protection constraints and data integrity verification rules are defined through smart contracts in the blockchain network to achieve automated permission control for cross-domain data sharing.
[0089] The data processing strategy is used to perform anonymization, desensitization or aggregation processing on the data during the data sharing process to protect data privacy. Specifically: privacy-preserving data modeling is achieved through the federated learning module, and collaborative modeling is performed among multiple institutions;
[0090] The federated learning module adjusts the model update weight according to the data volume and data quality score of the participants, groups data with similar distribution through clustering algorithm, improves the model's adaptability to non-independent and identically distributed data, combines differential privacy technology and compressed parameter transmission mechanism to reduce the communication cost of updating the federated learning model, and adopts the compute-to-data mode, allowing computing tasks to be executed at the data source and only outputting aggregated statistical results to ensure data security.
[0091] Data retention and deletion policies are used to limit the retention period of data, and automatically delete or archive the data after the period has expired. For example, transaction records are only allowed to be retained for 6 months and must be deleted after the expiration date.
[0092] Approval and verification policies are used to define the approval process and rules before data sharing. After verifying the sharing request through the rule engine, the approval process is generated. For example, the detailed financial data of a company can only be accessed after joint approval by the local market supervision department and the State Administration for Market Regulation.
[0093] Data sharing policies are dynamically configured by system administrators or data owners and stored in the system's policy management library. When data is shared or processed, the system automatically calls the data sharing policy to ensure that the operation complies with predefined rules.
[0094] The optimization unit is used to build the data resource directory into a graph structure based on the graph database and graph algorithm, optimize the data retrieval path using the shortest path algorithm, dynamically optimize the data exchange rules based on the dynamic rule adjustment function of reinforcement learning, and realize real-time load balancing using the DQN (Deep Q-Network) algorithm.
[0095] Furthermore, the lightweight proof-of-work mechanism specifically includes:
[0096] Each node in the network organizes the current transaction records into a Merkle tree and calculates the root hash value as the basis for the checkpoint record;
[0097] The difficulty of each proof-of-work task is dynamically adjusted according to the following formula:
[0098]
[0099] Where D is the difficulty of the current task; min and D max are the lower and upper limits of difficulty respectively; D prev is the difficulty of the last task; T target is the target time, the default value is 1 minute; T actual is the actual time to complete the previous task;
[0100] The node attempts to calculate a hash value that meets the difficulty requirement. The first node to complete the calculation broadcasts the checkpoint record and the proof result. Other nodes check whether the hash value meets the difficulty requirement through batch verification and use multi-threaded asynchronous processing. If the verification passes, the checkpoint record is stored on the public blockchain, and only the operation summary information is recorded;
[0101] In cross-domain operations, nodes in the local domain quickly reach consensus through Byzantine fault-tolerant consensus, and then submit cross-domain checkpoints to the public blockchain to form a hierarchical consensus process;
[0102] Based on smart contracts, only checkpoint records with high access frequency are retained. The high access frequency is judged by setting a time window to count the number of accesses and setting a threshold. The statistical results are updated every 30 minutes, sorted by the PageRank algorithm, and the storage strategy is dynamically adjusted to reduce redundant data.
[0103] The specific data model is calculated as follows:
[0104] The requester sends the task to the data source through a secure channel, including the calculation logic, input parameters, and privacy protection requirements;
[0105] The data source uses a dynamic scheduling algorithm to allocate tasks according to the current resource load, supports multi-task parallel execution, prioritizes high-priority tasks, and runs tasks in an isolated environment;
[0106] Use DBSCAN or adaptive density clustering algorithm to automatically select clustering parameters according to local data distribution, including the maximum distance ε between sample points and the minimum number of sample points MinPts. The minimum number of sample points MinPts is set to twice the sample dimension in the data set. The maximum distance ε between sample points is estimated by the following formula:
[0107] ε=mean(d)+k×std(d);
[0108] In the formula, d is the distance set between sample points in the data set, mean(d) represents the average value of the distance between all sample points in the data set, reflecting the average size of the overall distance between samples; std(d) represents the standard deviation of the distance between sample points in the data set, measuring the degree of dispersion of the distance distribution between samples; k is the adjustment coefficient, which is used to adjust the size of ε. By increasing or decreasing k, the adaptability of the clustering radius to the sample points can be adjusted;
[0109] By calculating the mean and standard deviation of the distance between sample points and combining it with the adjustment coefficient k, the value of ε is dynamically estimated, so that the clustering algorithm can better adapt to the distribution characteristics of the data;
[0110] The intermediate output data of the task is perturbed by differential privacy technology, adding noise to prevent the leakage of sensitive information. The calculation formula of the noise intensity b is: Where △f is the function sensitivity, which refers to the maximum change in the output of the objective function. When the input data changes by one sample, the maximum change in the output result may occur; δ is the privacy budget, which is a key parameter in differential privacy protection and reflects the trade-off between privacy protection and data accuracy. A smaller δ provides stronger privacy protection;
[0111] Prune or quantize the task results, and use homomorphic encryption technology that supports addition and multiplication to encrypt the results for transmission;
[0112] After receiving the result, the requester decrypts and verifies the correctness of the data. If the verification passes, the result is stored in the data asset library for subsequent analysis;
[0113] The data source cleans or archives temporary data according to the data retention policy to optimize resource utilization efficiency and prevent sensitive data from being retained for a long time.
[0114] (2) Data Standard Management Module
[0115] It is used to conduct standardized management of market supervision data throughout its life cycle, and to connect with the data sharing and exchange module to ensure that shared data meets standardization requirements. It optimizes the standardization process through dynamic updating of data standards and version difference analysis, and provides standardized data for processing by the data quality management module.
[0116] Specifically, the data standard management module includes:
[0117] Dynamic standard maintenance unit, used to achieve real-time optimization and dynamic maintenance of data standards. It records the version information of each update through process management based on the smallest data unit, generates standard optimization reports based on version difference analysis, automatically identifies and feedbacks deviations between data standards and market supervision requirements, and ensures the dynamic adaptability and pertinence of standards.
[0118] The non-standard data detection unit is used to perform automatic detection of metadata and data content according to preset standard specifications, generate a conformity analysis report based on the detection results, and provide improvement suggestions for problematic data through difference analysis with current standards, supporting subsequent standardization governance and quality improvement;
[0119] The master data management unit is used to identify and track the master data of multi-department collaboration through centralized coding and dynamic management strategies, automatically evaluate the impact of master data changes on related data and business processes, and output optimization plans to ensure consistency and efficient sharing of cross-department data;
[0120] The standard conformity analysis unit is used to generate a scoring report by comparing the degree of conformity between data table instances and data standards, dynamically adjust the standard content based on the scoring results and actual usage feedback, and generate actionable improvement suggestions for the weak links in the implementation of the standard, so as to improve the accuracy and effectiveness of standardization implementation.
[0121] (3) Data quality management module
[0122] It is used to perform quality management over the entire life cycle of market supervision data, receive standardized data, and form a closed-loop data quality management process through the identification, allocation, management and re-testing of problematic data, ensuring high-quality data for use by the data model management module, thereby improving data integrity and accuracy.
[0123] like Figure 2 As shown in the figure, it is a data quality management process based on business data, which mainly includes the following steps:
[0124] Data quality requirements analysis: determine the quality requirements of business data, formulate data quality rules, and provide a basis for subsequent testing and governance;
[0125] Data verification and problem data identification: Based on data quality rules, business data is verified to identify problematic data;
[0126] Data quality report generation: Organize the test results into a data quality report, including the distribution, scoring and governance recommendations of the problematic data;
[0127] Governance progress monitoring: Allocate and claim problematic data, and combine progress monitoring and feedback mechanisms to ensure efficient completion of data governance tasks;
[0128] Data governance and repair: Store problematic data in the governance repository, correct the data automatically or manually, and return the corrected data to the business system;
[0129] Data reuse: The corrected data flows back into the business system to support subsequent business operations.
[0130] In one embodiment, the data quality management module includes:
[0131] The problem data identification unit is used to perform automatic detection on the data in the system according to the preset data quality rules, dynamically adjust the detection rule weights in combination with the improved multi-level rule engine algorithm, and build a hierarchical weighted model with data characteristics as factors, which is expressed as:
[0132]
[0133] Where W iis the detection rule weight, which is used to indicate the importance of the rule in data quality detection; α i is the priority value of the data feature, reflecting the influence of different data features on the importance of the rule; n is the total number of data features; β i The rule adaptability indicates the matching degree between the current rule and the target data characteristics. The weight of the detection rule is adjusted dynamically so that it can change flexibly according to the data characteristics and rule matching situation, thus optimizing the detection process.
[0134] The data quality closed-loop governance unit is used to manage problem data throughout its life cycle. It introduces a dynamic progress prediction algorithm P=f(Q,T,R), where P is the expected governance completion time (used to estimate the total time of the task), Q is the data quality score (indicating the current quality status of the data), T is the data volume (indicating the scale of the data to be processed), R is the resource allocation efficiency (indicating the efficiency of the computing resources allocated in the governance task), and f represents the function mapping relationship. It optimizes the governance task allocation process to ensure that the problem data is governed in the shortest time possible, and adjusts the governance priority in real time during the re-detection stage to form a closed-loop governance process.
[0135] The data quality scoring unit is used to quantitatively score the data quality based on the defined data detection rules and governance results, and adopts a dynamic fuzzy comprehensive evaluation algorithm to combine the data characteristics with the governance results and generate a quality detection report. The dynamic fuzzy comprehensive evaluation algorithm is expressed as:
[0136]
[0137] In the formula, S is the total score of data quality, which is used to comprehensively evaluate the data quality status; μ i is the weight of data feature i, indicating the importance of different data features in the overall score; τ ij is the scoring result of data feature i under rule j, indicating the quality of data features under a specific rule; w ij is the weight of rule j on data feature i, indicating the matching degree of the rule to the feature; m is the total number of evaluation rules;
[0138] The manual data governance support unit is used to assign problematic data to specific personnel for verification and processing based on the data's region of origin, business type, or custom label, combined with a workload balancing algorithm Where L is the workload per person (indicating the amount of tasks after amortization), W is the total amount of tasks, and N is the number of people. Tasks are dynamically allocated to avoid waste of resources, and support batch export of problem data and monitoring of governance progress.
[0139] (4) Data model management module
[0140] It is used to optimize and analyze multidimensional data based on the multi-satellite data model, receive high-quality data to build correlation models across business systems, realize data integration analysis, and provide the analysis results to the data asset management module to improve data processing efficiency through scheduling optimization.
[0141] In one embodiment, the data model management module includes:
[0142] Multi-star data model design unit, used to build a unified model across business systems based on multiple fact tables and shared dimension tables, supporting multi-dimensional data analysis and integration, expanding the semantic association capabilities of the model through knowledge graph technology, and combining relationship inference algorithms to achieve efficient integration of heterogeneous data;
[0143] Specifically, the multi-star data model design unit introduces dynamic semantic expansion technology based on deep learning, optimizes the semantic association strength of the knowledge graph in real time through the attention mechanism, and combines multi-hop path analysis to mine implicit relationships, thus achieving deep association of cross-domain data;
[0144] Using knowledge graph embedding technology, nodes and relationships are mapped to low-dimensional vector space, and efficient relationship inference is achieved through a distributed computing framework, supporting semantic integration of heterogeneous data sources;
[0145] The semantic conflicts generated during the integration process are dynamically adjusted using a reinforcement learning algorithm. The semantic conflict adjustment is implemented based on the DQN algorithm. The conflict is identified by building a conflict discrimination function. The reinforcement learning model updates the reward function in each adjustment to maximize the consistency of the association relationship, ensuring the rationality and consistency of the semantic association results and generating optimized high-order semantic relationships.
[0146] Improve the adaptability and integration efficiency of the model in complex business scenarios through the dynamic adjustment mechanism of semantic priority and real-time expansion function;
[0147] Model conversion and optimization unit, which is used to dynamically switch between the star model and the snowflake model according to specific business needs, optimize data query performance through adaptive scheduling mechanism, and improve data processing efficiency in combination with dynamic partitioning strategy;
[0148] The scenario-driven analysis unit is used to perform risk prediction, trend analysis and hotspot detection based on the lazy data model. The lazy data model is implemented through a lazy loading mechanism. Specifically, the required data fragments are dynamically loaded only when a specific analysis task is received. The unrequested data is kept in the storage and does not occupy memory resources. After loading, the cache strategy is combined to avoid repeated loading. In addition, the generative adversarial network technology is combined to simulate business boundary scenarios, evaluate the adaptability of the model under extreme conditions, and provide decision support for market supervision through analysis results.
[0149] Data lineage analysis and early warning unit, used to trace data sources and change links in real time during data processing, display data processing paths through visualization, monitor model output quality in combination with dynamic rules, and trigger early warning mechanisms when abnormalities occur, providing feedback through system announcements, SMS reminders, and email notifications;
[0150] The data processing configuration and scheduling unit is used to provide a browser-based visual interface, combined with an automatic scheduling mechanism driven by reinforcement learning, to dynamically adjust task order and resource allocation according to real-time task load and priority.
[0151] (5) Data asset management module
[0152] It is used to manage data assets in an all-round way, store, optimize and display the analysis results output by the data model management module, including dynamically adjusting data storage strategies based on access popularity, improving data usage efficiency in combination with data link analysis, and providing data change impact analysis functions.
[0153] The data asset management module includes:
[0154] The data resource directory unit is used to form a hierarchical directory through the metadata description of the data resource to support the retrieval and acquisition of the data resource;
[0155] The data asset heat analysis unit is used to dynamically calculate the heat of data based on the access frequency, citation count, and browsing count of the data, and adjust the storage and access strategies to optimize system performance, such as Figure 3 As shown;
[0156] The data link analysis unit is used to combine SQL logs and stored procedure analysis to automatically generate a link relationship map of metadata. The update cycle of the link relationship map is dynamically adjusted according to the system log traffic. The optimized analysis process is achieved through distributed log analysis technology, and version change notification and difference analysis functions are provided.
[0157] like Figure 4 As shown in FIG. 1 , another embodiment of the present invention provides a data governance method of an intelligent market supervision data governance system based on multi-level data sharing and exchange, comprising the following steps:
[0158] Share and exchange market supervision data across departments, regions and levels, and implement data access control and dynamic resource catalog management in a policy-driven manner;
[0159] Conduct standardized management of market supervision data throughout its life cycle to ensure that shared data meets standardization requirements, and optimize the standardization process through dynamic updates of data standards and version difference analysis;
[0160] Conduct quality management for the entire life cycle of market supervision data, receive standardized data, and form a closed-loop data quality management process through the identification, allocation, management and re-testing of problematic data;
[0161] Optimize and analyze multidimensional data based on multi-satellite data models, receive high-quality data to build cross-business system correlation models, and realize data integration analysis;
[0162] Comprehensively manage data assets, store, optimize and display analysis results of data models, including dynamically adjusting data storage strategies based on access popularity, and improving data utilization efficiency in combination with data link analysis.
[0163] In summary, the present invention realizes efficient sharing and exchange of market supervision data across departments, regions and levels through the introduction of data sharing and exchange modules. Combined with the policy-driven implementation method, the system can flexibly define data access rights and dynamically manage resource directories to ensure the security and controllability of data during the sharing process. The generated data resource directory is further called by other modules, building a data sharing foundation between modules and providing support for subsequent standardization, quality governance and model analysis.
[0164] The data standard management module conducts standardized management of the entire life cycle of market supervision data and seamlessly connects with the data sharing and exchange module to ensure that shared data meets unified standard requirements. By dynamically updating data standards and optimizing the standardization process through version difference analysis, data can quickly adapt to changes in market supervision needs and provide high-quality input for the data quality management module based on standardized data.
[0165] The data quality management module conducts quality management for the entire life cycle of standardized data, forming a closed-loop management process, including the identification, allocation, management and re-detection of problem data. The construction of a closed-loop process greatly improves the efficiency and accuracy of problem data processing, and ensures reliable and high-quality data input for the data model management module, thus laying a solid foundation for subsequent analysis.
[0166] The data model management module builds a cross-business system association model based on the multi-star data model, which can efficiently integrate multi-dimensional data and improve data processing efficiency through optimization analysis. The module uses the association model to realize the integrated analysis of multi-system data and efficiently passes the analysis results to the data asset management module for use, greatly improving the depth and efficiency of data analysis.
[0167] The data asset management module provides a full range of data asset management services based on analysis results. Through dynamic storage adjustment strategies based on access popularity, the system can optimize data storage and access efficiency; combined with data link analysis functions, the system provides impact analysis and dynamic management functions for data changes, so that the value of data assets can be fully explored and utilized.
[0168] Through the collaborative work of multiple modules, a complete closed-loop process has been formed from data sharing, standardization, quality governance to model analysis and asset management, which significantly improved the governance, analysis and intelligence capabilities of market supervision data and provided accurate, fast and reliable decision-making support for market supervision.
[0169] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An intelligent market supervision data governance system based on multi-level data sharing and exchange, characterized by: The system includes: a data sharing and exchange module, which is used to share and exchange market supervision data across departments, regions and levels, and implement data access control and dynamic resource directory management in combination with policy-driven, and the data resource directory and shared data generated by the data sharing and exchange module are used by other modules; The data standard management module is used to standardize the entire life cycle of market supervision data and connect with the data sharing and exchange module to ensure that shared data meets standardization requirements. It optimizes the standardization process by dynamically updating data standards and version difference analysis, and provides standardized data for processing by the data quality management module; The data quality management module is used to manage the quality of market supervision data throughout its life cycle. It receives standardized data and forms a closed-loop data quality management process through the identification, allocation, management and re-testing of problematic data, ensuring high-quality data for use by the data model management module. The data model management module is used to optimize and analyze multidimensional data based on the multi-satellite data model, receive high-quality data to build a correlation model across business systems, realize data integration analysis, and provide the analysis results to the data asset management module to improve data processing efficiency through scheduling optimization; The data asset management module is used to manage data assets in an all-round way, store, optimize and display the analysis results output by the data model management module, including dynamically adjusting the data storage strategy based on access popularity, improving data usage efficiency in combination with data link analysis, and providing data change impact analysis function.
2. According to claim 1, the intelligent market supervision data management system based on multi-level data sharing and exchange is characterized in that: The data sharing and exchange module includes: The policy-driven unit is used to define data sharing policies, including: Data access strategy: Maintain an independent private blockchain ledger through the Byzantine fault-tolerant consensus protocol for cross-departmental, cross-regional, and cross-level data operation records. Maintain a cross-domain public blockchain ledger through a lightweight proof-of-work mechanism, only record the checkpoint information of cross-domain operations, and use smart contracts in the blockchain network to automatically define data access rights, privacy protection constraints, and data integrity verification rules to achieve automated permission control for cross-domain data sharing; Data processing strategy: Privacy-preserving data modeling is achieved through a federated learning module, and collaborative modeling is performed among multiple institutions. The federated learning module adjusts the model update weight according to the data volume and data quality score of the participants, groups similarly distributed data through a clustering algorithm, improves the model's adaptability to non-independent and identically distributed data, combines differential privacy technology and a compressed parameter transmission mechanism to reduce the communication cost of updating the federated learning model, and adopts a compute-to-data mode that allows computing tasks to be executed at the data source and only outputs aggregated statistical results to ensure data security. Data retention policy: used to limit the retention period of market regulatory data and automatically delete or archive the data after the period has expired; Approval and verification strategy: used to define the approval process and rules before data sharing. After verifying the sharing request through the rule engine, the approval process is generated; The optimization unit is used to build the data resource directory into a graph structure based on the graph database and graph algorithm, optimize the data retrieval path using the shortest path algorithm, dynamically optimize the data exchange rules using the DQN algorithm based on the dynamic rule adjustment function of reinforcement learning, and realize real-time load balancing.
3. According to claim 2, the intelligent market supervision data management system based on multi-level data sharing and exchange is characterized in that: The lightweight proof-of-work mechanism specifically includes: Each node in the network organizes the current transaction records into a Merkle tree and calculates the root hash value as the basis for the checkpoint record; The difficulty of each proof-of-work task is dynamically adjusted according to the following formula: Where D is the difficulty of the current task; min and D max are the lower and upper limits of difficulty respectively; D prev is the difficulty of the last task; T target is the target time, the default value is 1 minute; T actual is the actual time to complete the previous task; The node attempts to calculate a hash value that meets the difficulty requirement. The first node to complete the calculation broadcasts the checkpoint record and the proof result. Other nodes check whether the hash value meets the difficulty requirement through batch verification and use multi-threaded asynchronous processing. If the verification passes, the checkpoint record is stored on the public blockchain, and only the operation summary information is recorded; In cross-domain operations, nodes in the local domain quickly reach consensus through Byzantine fault-tolerant consensus, and then submit cross-domain checkpoints to the public blockchain to form a hierarchical consensus process; Based on smart contracts, only checkpoint records with high access frequency are retained. The high access frequency is judged by setting a time window to count the number of accesses and setting a threshold. The statistical results are updated every 30 minutes, sorted by the PageRank algorithm, and the storage strategy is dynamically adjusted to reduce redundant data.
4. The intelligent market supervision data management system based on multi-level data sharing and exchange according to claim 2 is characterized in that: The calculation to data mode is specifically: The requester sends the task to the data source through a secure channel, including the calculation logic, input parameters, and privacy protection requirements; The data source uses a dynamic scheduling algorithm to allocate tasks according to the current resource load, supports multi-task parallel execution, prioritizes high-priority tasks, and runs tasks in an isolated environment; Use DBSCAN or adaptive density clustering algorithm to automatically select clustering parameters according to local data distribution, including the maximum distance ε between sample points and the minimum number of sample points MinPts. The minimum number of sample points MinPts is set to twice the sample dimension in the data set. The maximum distance ε between sample points is estimated by the following formula: ε=mean(d)+k×std(d); In the formula, d is the distance set between sample points in the data set, mean(d) represents the average value of the distance between all sample points in the data set, std(d) represents the standard deviation of the distance between sample points in the data set; k is the adjustment coefficient; The intermediate output data of the task is perturbed by differential privacy technology, adding noise to prevent the leakage of sensitive information. The calculation formula of the noise intensity b is: Where △f is the function sensitivity and δ is the privacy budget; The task results are pruned or quantized, and the results are encrypted and transmitted using homomorphic encryption technology that supports addition and multiplication. After receiving the results, the requester decrypts and verifies the correctness of the data. If the verification passes, the results are stored in the data asset library for subsequent analysis. The data source cleans or archives temporary data according to the data retention policy to optimize resource utilization efficiency and prevent sensitive data from being retained for a long time.
5. The intelligent market supervision data management system based on multi-level data sharing and exchange according to claim 1 is characterized in that: The data standard management module includes: Dynamic standard maintenance unit, used to achieve real-time optimization and dynamic maintenance of data standards. It records the version information of each update through process management based on the smallest data unit, generates standard optimization reports based on version difference analysis, and automatically identifies and feedbacks deviations between data standards and market regulatory requirements. The non-standard data detection unit is used to perform automatic detection of metadata and data content according to preset standard specifications, generate a conformity analysis report based on the detection results, and provide improvement suggestions for problematic data through difference analysis with current standards; the master data management unit is used to identify and track multi-departmental collaborative master data through centralized coding and dynamic management strategies, automatically evaluate the impact of master data changes on related data and business processes, and output optimization plans to ensure consistency and efficient sharing of cross-departmental data; The standard fit analysis unit is used to generate a scoring report by comparing the degree of fit between the data table instance and the data standard, dynamically adjust the standard content based on the scoring results and actual usage feedback, and generate actionable improvement suggestions for the weak links in the implementation of the standard.
6. The intelligent market supervision data management system based on multi-level data sharing and exchange according to claim 1 is characterized in that: The data quality management module includes: The problem data identification unit is used to perform automatic detection on the data in the system according to the preset data quality rules, dynamically adjust the detection rule weights in combination with the improved multi-level rule engine algorithm, and build a hierarchical weighted model with data characteristics as factors, which is expressed as: Where W i is the detection rule weight; α i is the priority value of the data feature, n is the total number of data features; β i is the rule adaptability; the data quality closed-loop governance unit is used to manage problem data throughout its life cycle, and introduces a dynamic progress prediction algorithm P=f(Q,T,R), where P is the expected completion time of governance, Q is the data quality score, T is the data volume, R is the resource allocation efficiency, and f represents the function mapping relationship. It optimizes the governance task allocation process to ensure that problem data is governed in the shortest time possible, and adjusts the governance priority in real time during the re-detection stage to form a closed-loop governance process; The data quality scoring unit is used to quantitatively score the data quality based on the defined data detection rules and governance results, and adopts a dynamic fuzzy comprehensive evaluation algorithm to combine the data characteristics with the governance results and generate a quality detection report. The dynamic fuzzy comprehensive evaluation algorithm is expressed as: In the formula, S is the total score of data quality; μ i is the weight of data feature i; τ ij is the scoring result of data feature i under rule j; w ij is the weight of rule j on data feature i; m is the total number of evaluation rules; The manual data governance support unit is used to assign problematic data to specific personnel for verification and processing based on the data's region of origin, business type, or custom label, combined with a workload balancing algorithm Where L is the workload per person, W is the total task volume, and N is the number of people. Tasks are dynamically allocated to avoid waste of resources, and batch export of problem data and monitoring of governance progress are supported.
7. The intelligent market supervision data management system based on multi-level data sharing and exchange according to claim 1 is characterized in that: The data model management module includes: Multi-star data model design unit, used to build a unified model across business systems based on multiple fact tables and shared dimension tables, supporting multi-dimensional data analysis and integration, expanding the semantic association capabilities of the model through knowledge graph technology, and combining relationship inference algorithms to achieve efficient integration of heterogeneous data; The model conversion and optimization unit is used to dynamically switch between the star model and the snowflake model according to specific business needs, optimize data query performance through an adaptive scheduling mechanism, and improve data processing efficiency in combination with a dynamic partitioning strategy; the scenario-driven analysis unit is used to perform risk prediction, trend analysis, and hotspot detection based on an inert data model, where the inert data model is implemented through a delayed loading mechanism, specifically: only when a specific analysis task is received, the required data fragments are dynamically loaded, and the unrequested data is kept in the storage and does not occupy memory resources. After loading is completed, the cache strategy is combined to avoid repeated loading; and the generative adversarial network technology is combined to simulate business boundary scenarios, evaluate the adaptability of the model under extreme conditions, and provide decision support for market supervision through analysis results; Data lineage analysis and early warning unit, used to trace data sources and change links in real time during data processing, display data processing paths through visualization, monitor model output quality in combination with dynamic rules, and trigger early warning mechanisms when abnormalities occur, providing feedback through system announcements, SMS reminders, and email notifications; The data processing configuration and scheduling unit is used to provide a browser-based visual interface, combined with an automatic scheduling mechanism driven by reinforcement learning, to dynamically adjust task order and resource allocation according to real-time task load and priority.
8. The intelligent market supervision data management system based on multi-level data sharing and exchange according to claim 7 is characterized in that: In the multi-star data model design unit, a dynamic semantic expansion technology based on deep learning is introduced to optimize the semantic association strength of the knowledge graph in real time through the attention mechanism, and multi-hop path analysis is combined to mine implicit relationships to achieve deep association of cross-domain data; Using knowledge graph embedding technology, nodes and relationships are mapped to low-dimensional vector space, and efficient relationship inference is achieved through a distributed computing framework, supporting semantic integration of heterogeneous data sources; The semantic conflicts generated during the integration process are dynamically adjusted using a reinforcement learning algorithm. The semantic conflict adjustment is implemented based on the DQN algorithm. The conflict is identified by building a conflict discrimination function. The reinforcement learning model updates the reward function in each adjustment to maximize the consistency of the association relationship, ensuring the rationality and consistency of the semantic association results and generating optimized high-order semantic relationships. The dynamic adjustment mechanism of semantic priority and real-time expansion function can improve the adaptability and integration efficiency of the model in complex business scenarios.
9. The intelligent market supervision data management system based on multi-level data sharing and exchange according to claim 1 is characterized in that: The data asset management module includes: The data resource directory unit is used to form a hierarchical directory through the metadata description of the data resource to support the retrieval and acquisition of the data resource; Data asset heat analysis unit, which is used to dynamically calculate the heat of data based on the access frequency, citation count, and browsing count of data, and adjust the storage and access strategies to optimize system performance; The data link analysis unit is used to combine SQL logs and stored procedure analysis to automatically generate a link relationship map of metadata. The update cycle of the link relationship map is dynamically adjusted according to the system log traffic. The optimized analysis process is achieved through distributed log analysis technology, and version change notification and difference analysis functions are provided.
10. The data governance method of the intelligent market supervision data governance system based on multi-level data sharing and exchange according to any one of claims 1 to 9 is characterized in that: The method comprises: Share and exchange market supervision data across departments, regions and levels, and implement data access control and dynamic resource catalog management in a policy-driven manner; Conduct standardized management of market supervision data throughout its life cycle to ensure that shared data meets standardization requirements, and optimize the standardization process through dynamic updates of data standards and version difference analysis; Conduct quality management for the entire life cycle of market supervision data, receive standardized data, and form a closed-loop data quality management process through the identification, allocation, management and re-testing of problematic data; Optimize and analyze multidimensional data based on multi-satellite data models, receive high-quality data to build cross-business system correlation models, and realize data integration analysis; Comprehensively manage data assets, store, optimize and display analysis results of data models, including dynamically adjusting data storage strategies based on access popularity, and improving data utilization efficiency in combination with data link analysis.
Citation Information
Patent Citations
Data exchange method, apparatus and system
CN101404609A
A data sharing exchange system and method based on a block chain
CN109729168A
Treatment method for cement production and operation data
CN114298550A
Intelligent fusion sharing system of market supervision data exchange platform
CN118735560A
Multi-source computing power data integration and intelligent scheduling system and method
CN118916147A
Cited By
Intelligent contract driven workflow engine automatic execution method and system based on block chain
CN120179419A
Quality treatment and real-time data verification method based on data sandbox
CN120596369A
Quality governance and real-time data verification method based on data sandbox
CN120596369B
Intelligent operation and maintenance monitoring method and system for data center
CN120602308A
Tea industry chain operation monitoring system based on big data sharing and exchange
CN120688099A