A distributed data governance system based on a new pipeline-based scheduling algorithm

By designing a distributed data governance system based on a novel pipeline-based scheduling algorithm, the problem of insufficient management of the data exchange process in traditional data governance systems is solved. This enables efficient and secure data transmission and management, supports high-concurrency access by large users, and promotes the development of enterprise information systems.

CN113553381BActive Publication Date: 2026-04-28CHINA NAT BUILDING MATERIALS TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA NAT BUILDING MATERIALS TECH CO LTD
Filing Date
2021-07-28
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional enterprise data governance systems lack management of the data exchange process, resulting in poor data quality, ineffective collection and scheduling, and affecting the quality of enterprise information system construction.

Method used

Design a distributed data governance system based on a novel pipeline scheduling algorithm, including an infrastructure unit, a pipeline scheduling unit, a unified service unit, and a data management unit. The system is connected via network communication to achieve inter-process communication, data quality control, and unified standardized management.

Benefits of technology

It improves the transmission efficiency of the data exchange process, reduces data loss and leakage, enhances data integrity and security, supports high-concurrency access by large users, and promotes the construction of enterprise information systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113553381B_ABST
    Figure CN113553381B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data governance, in particular to a distributed data governance system based on a new type of scheduling algorithm of a pipeline. The system comprises a basic framework unit, a pipeline scheduling unit, a unified service unit and a data management unit; the basic framework unit is used for managing a network framework supporting system operation; the pipeline scheduling unit is used for managing scheduling of an inter-process communication pipeline; the unified service unit is used for unified standardized management of data; and the data management unit is used for comprehensive governance of data. The application can realize distributed data governance, improve the fusion degree of the system and enhance the extensibility of the system; can improve the transmission effect of a data exchange process, improve the integrity and security of data; through centralized scheduling management of data, the framework is convenient for node expansion of the system, can support high-concurrency access of a large user, and can realize visual management of a data full life cycle, shorten a data management cycle and reduce cost waste.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data governance technology, and more specifically, to a distributed data governance system based on a novel pipeline-based scheduling algorithm. Background Technology

[0002] For businesses, the most important asset is data; its core value can be understood as core business value. Realizing data value requires business data analysis and value mining, and only clean and complete data can effectively demonstrate this value. Most traditional enterprises lack awareness of data resource quality management and don't understand how to conduct data governance, resulting in poor data quality, such as missing data, data silos, and data distortion, preventing them from leveraging data resources to obtain greater business value. A major reason for poor data quality is the lack of effective data transmission channels and management during data exchange. This leads to problems like missing data and information leaks, resulting in time-consuming, labor-intensive, and inefficient subsequent data governance operations. Currently, there is a lack of data governance system platforms that focus on managing the data exchange process, preventing data from being effectively collected and managed, thus affecting the quality of enterprise information system construction and hindering enterprise development. Summary of the Invention

[0003] The purpose of this invention is to provide a distributed data governance system based on a novel pipeline scheduling algorithm to solve the problems mentioned in the background art.

[0004] To address the aforementioned technical problems, one objective of this invention is to provide a distributed data governance system based on a novel pipeline-based scheduling algorithm, comprising:

[0005] The system comprises an infrastructure unit, a pipeline scheduling unit, a unified service unit, and a data management unit. These units are sequentially connected via network communication. The infrastructure unit provides and manages the basic network topology supporting system operation. The pipeline scheduling unit provides and manages different inter-process communication pipelines and performs allocation and scheduling. The unified service unit provides multiple unified service functions to achieve standardized management and analysis of data for the enterprise. The data management unit performs quality control and comprehensive governance of the data.

[0006] The infrastructure unit includes an application platform module, a data source module, a technical support module, and an algorithm management module;

[0007] The pipeline scheduling unit includes an inter-process communication module, an anonymous pipeline module, a named pipeline module, and a round-robin scheduling module;

[0008] The unified service unit includes a metadata management module, a service gateway module, a data storage module, and an identity authentication module;

[0009] The data management unit includes a data integration module, a data exchange module, a data governance module, a master data management module, and a data application module.

[0010] As a further improvement to this technical solution, the application platform module, the data source module, the technical support module, and the algorithm management module are sequentially connected via network communication and operate in parallel. The application platform module is used to build a multi-functional data governance application platform based on big data and blockchain technology to achieve human-computer interaction. The data source module is used to establish signal connection and data transmission channels between the system and various data sources and to manage and allocate the data source platform. The technical support module is used to load various intelligent technologies to support and improve the functionality of the system. The algorithm management module is used to encapsulate various intelligent algorithms that support system functions.

[0011] Among these, intelligent technologies include, but are not limited to, big data technology, big data analytics technology, blockchain technology, and front-end / back-end separation technology.

[0012] Among them, intelligent algorithms include, but are not limited to, intelligent scheduling algorithms, encryption algorithms, consensus algorithms, and matching algorithms.

[0013] As a further improvement to this technical solution, the signal output terminal of the inter-process communication module is connected to the signal input terminals of the anonymous pipe module and the named pipe module. The anonymous pipe module and the named pipe module run in parallel. The signal output terminals of the anonymous pipe module and the named pipe module are connected to the signal input terminal of the polling scheduling module. The inter-process communication module is used to allocate a buffer in the kernel using memory as a medium to realize inter-process data exchange communication. The anonymous pipe module is used to create anonymous pipes between related processes by calling the pipe function to provide one-way communication. The named pipe module is used to create system-visible named pipes between any two processes by calling the mknod or mkfifo command to realize communication between the two processes. The polling scheduling module is used to schedule pipes in a polling manner to achieve load balancing of pipe scheduling.

[0014] As a further improvement to this technical solution, the round-robin scheduling module includes a round-robin scheduling algorithm and an improved weighted round-robin scheduling algorithm. The round-robin scheduling algorithm is suitable for situations where all servers in a server group have the same hardware and software configuration and the average service requests are relatively balanced. The weighted round-robin scheduling algorithm is suitable for situations where the configurations, installed business applications, and processing capabilities of the servers in the server group are different. Therefore, the calculation expression for weight allocation in the weighted round-robin scheduling algorithm is:

[0015]

[0016] In the formula, degree k Use load on the server.

[0017] As a further improvement to this technical solution, the metadata management module, the service gateway module, the data storage module, and the identity authentication module communicate sequentially via the network and operate in parallel. The metadata management module provides unified management of information description metadata and information authorization metadata, as well as metadata description, metadata interfaces, and a high-performance metadata access mechanism. The service gateway module provides a unified data service gateway to achieve secure and controllable connection and service between internal and external systems, and supports full lifecycle management of APIs. The data storage module provides a unified distributed data storage service, enabling unified storage and access to various data types such as structured data, unstructured data, and files, and can achieve self-description and information authorization through metadata management information. The identity authentication module provides unified user management and identity authentication services, supports multiple access methods, and supports the use of third-party ID cards to meet the identity authentication needs of different application systems.

[0018] The unified metadata management system includes a rich set of built-in acquisition adapters, enabling end-to-end automated acquisition and one-click metadata analysis. It can quickly clean up data resources, understand the origin and development of data, and build data maps to provide basic support for data standardization and data quality. The metadata management functions include, but are not limited to, metadata acquisition, metadata retrieval, data mapping, lineage analysis, and impact analysis.

[0019] During the unified identity authentication process, login methods include, but are not limited to, mobile phone, email, QQ, WeChat, Weibo, etc.

[0020] As a further improvement to this technical solution, the data integration module, the data exchange module, the data governance module, the master data management module, and the data application module are sequentially connected and interleaved via network communication. The data integration module is used to load, clean, transform, and integrate cross-block data, and supports custom scheduling and graphical monitoring, thereby achieving unified scheduling and unified monitoring to meet the needs of operation and maintenance visualization. The data exchange module is used to realize the transmission and sharing of data or files between several business subsystems, and can integrate data collection, processing, distribution, and exchange transmission. The data governance module is used to integrate and govern data from multiple aspects. The master data management module is used to establish a unified view and centralized management of data that needs to be shared, thereby providing optimal data for business system data calls. The data application module is used to provide users with high-quality, well-governed data for retrieval and application through deep mining, statistical reports, and other formats.

[0021] The data integration functions include, but are not limited to, real-time data collection, cleaning and transformation, encryption and desensitization.

[0022] The master data management functions include, but are not limited to, application identification, master data retrieval, master data storage, sharing and publishing, and master data monitoring.

[0023] As a further improvement to this technical solution, the data exchange module includes a data transmission module, a node management module, a pipeline matching module, and an exchange approval module. The data transmission module, node management module, pipeline matching module, and exchange approval module are sequentially connected via communication. The data transmission module is used to implement secure data transmission in various ways through multiple data exchange components. The node management module is used for visual configuration of data transmission and exchange nodes, controlling the data transmission status of nodes, supporting the transmission of various data formats, and shielding data type differences between systems. The pipeline matching module is used to select the inter-process pipeline to use based on the correlation analysis results between business data and processes using metadata, through a matching algorithm. The exchange approval module is used to support the design of data exchange task methods and to realize the detection and approval of data exchange tasks through flexible data extraction and exchange.

[0024] During the data exchange process, data exchange components include, but are not limited to, table exchange, file transfer, SFTP upload and download, and HTTP components; data transmission methods include, but are not limited to, encrypted data transmission, power outage resume transmission, de-identification algorithms, dual control of data permissions and function permissions, data partitioning, and parallel loading technology.

[0025] During node management, data formats include, but are not limited to, mainstream databases, text files, Excel files, API interfaces, and WebService services.

[0026] As a further improvement to this technical solution, in the pipeline matching module, the correlation-based matching algorithm is used to determine the relationship between process relationships and pipeline types, and the NCC normalized cross-correlation coefficient is used to represent the degree of correlation between the two. Its calculation expression is as follows:

[0027]

[0028] In the formula, X and Y are two random variables, μ X μ Y Let σ be the mean of two random variables. X σ Y Let be the standard deviation of two random variables;

[0029] In the above formula, the denominator is the standard deviation of the two random variables, which serves as a normalization function. The mean of the two random variables is also subtracted from the numerator, which is called the centering.

[0030] Normalization and centering are not the essence of the correlation coefficient. After separating these two, what remains is the expectation of the product of two random variables, that is, the inner product of two vectors, which is the essence of the correlation coefficient.

[0031] Specifically, it can be understood as follows: if random variables X and Y are sampled multiple times and the sampled samples are placed into two vectors, then the inner product of these two vectors is the expectation of the product of X and Y; conversely, it can be understood as follows: if the two vectors are regarded as the joint distribution of two random variables, then the inner product of the two vectors is the expectation of the product of the two random variables.

[0032] As a further improvement to this technical solution, the data governance module includes a standard management module, a quality management module, a security management module, and a lifecycle module; the standard management module, the quality management module, the security management module, and the lifecycle module are sequentially connected via network communication; the standard management module is used to provide a comprehensive and complete data standard management process and method to determine and establish a single, accurate, and authoritative source of fact, to achieve complete, effective, consistent, standardized, open, and shared management of data, and to provide standard basis for data quality inspection and data security management; the quality management module is used for guided, visual, and other operations. The approach integrates the work environment of quality assessment, quality inspection, quality rectification, and quality reporting into a complete data quality management closed loop, using data standards as the basis for data inspection and metadata as the object of data inspection. The security management module provides various data security management measures throughout the data governance process, such as encryption, desensitization, obfuscation, and database authorization monitoring of privacy data, to achieve comprehensive protection of data security operations. The lifecycle module records the entire flow of data from its creation and initial storage to its deletion when it becomes obsolete, and performs near-line archiving, offline archiving, destruction, and full lifecycle monitoring of data during its existence.

[0033] The data standards management functions include, but are not limited to, standard publishing, standard mapping, and standard querying.

[0034] Data quality management includes, but is not limited to, rule management, quality reports, and problem rectification.

[0035] The data security management functions include, but are not limited to, change tracking, anomaly monitoring, data anonymization, secure transmission, and de-privacy protection.

[0036] The data lifecycle management functions include, but are not limited to, data policies, level agreements, classification and assessment, data archiving, and data destruction.

[0037] The second objective of this invention is to provide a method for operating a distributed data governance system based on a novel pipeline scheduling algorithm, comprising the following steps:

[0038] S1. Based on blockchain and big data technologies, build a distributed data governance system and connect the application platform with multiple data sources;

[0039] S2, Load the corresponding inter-process communication pipe creation method and round-robin scheduling algorithm;

[0040] S3. Implement unified management of the data governance system, pre-manage metadata, and monitor and adjust the standard quality of metadata;

[0041] S4. When exchanging data between the data source and the data integration module, based on the metadata's provisions for different processes and the results of lineage analysis, and according to the correlation-based matching algorithm, select and create the corresponding inter-process pipelines to achieve secure data transmission.

[0042] S5. Implement security management and lifecycle management for data, and manage data collection, transformation, cleaning, and archiving in a unified manner to provide high-quality data.

[0043] S6. Apply the well-managed high-quality data to front-end business operations;

[0044] S7. Users access the system's application platform through various methods and log in after unified identity authentication, obtain high-quality data from the system, and should call relevant data for application.

[0045] The third objective of this invention is to provide an operating device for a distributed data governance system based on a novel pipeline-based scheduling algorithm, comprising a processor, a memory, and a computer program stored in the memory and running on the processor. The processor executes the computer program to implement any of the aforementioned novel pipeline-based scheduling algorithms for distributed data governance.

[0046] The fourth objective of this invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the aforementioned novel pipeline-based scheduling algorithms in a distributed data governance system.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] 1. This distributed data governance system based on a novel pipeline-based scheduling algorithm is grounded in big data and blockchain. It can build a wide-ranging and stable parallel system architecture, realize distributed data governance, improve system integration, and enhance system scalability.

[0049] 2. This distributed data governance system based on a novel pipeline scheduling algorithm manages the creation, selection, and scheduling of inter-process communication, thereby improving the transmission efficiency of data exchange, reducing data loss and leakage, and enhancing data integrity and security.

[0050] 3. This distributed data governance system based on a novel pipeline scheduling algorithm centrally collects and schedules data. Its distributed architecture facilitates node expansion, supports high-concurrency access by large users, and enables visualized management of the entire data lifecycle. This shortens the data management cycle, reduces cost waste, and promotes the construction and development of enterprise information systems. Attached Figure Description

[0051] Figure 1 This is an exemplary product architecture block diagram of the present invention;

[0052] Figure 2 This is a structural diagram of the overall system device of the present invention;

[0053] Figure 3 This is one of the partial system device structure diagrams of the present invention;

[0054] Figure 4 This is a second partial system device structure diagram of the present invention;

[0055] Figure 5 This is the third partial system device structure diagram of the present invention;

[0056] Figure 6 This is the fourth partial system device structure diagram of the present invention;

[0057] Figure 7 This is the fifth partial system device structure diagram of the present invention;

[0058] Figure 8 This is the sixth partial system device structure diagram of the present invention;

[0059] Figure 9 This is a schematic diagram of an exemplary electronic computer product device according to the present invention.

[0060] The meanings of the labels in the diagram are as follows:

[0061] 100. Infrastructure Module; 101. Application Platform Module; 102. Data Source Module; 103. Technical Support Module; 104. Algorithm Management Module;

[0062] 200. Pipe scheduling unit; 201. Inter-process communication module; 202. Anonymous pipe module; 203. Named pipe module; 204. Round-robin scheduling module;

[0063] 300. Unified Service Unit; 301. Metadata Management Module; 302. Service Gateway Module; 303. Data Storage Module; 304. Identity Authentication Module;

[0064] 400. Data Management Unit; 401. Data Integration Module; 402. Data Exchange Module; 4021. Data Transmission Module; 4022. Node Management Module; 4023. Pipeline Matching Module; 4024. Exchange Approval Module; 403. Data Governance Module; 4031. Standards Management Module; 4032. Quality Management Module; 4033. Security Management Module; 4034. Lifecycle Module; 404. Master Data Management Module; 405. Data Application Module. Detailed Implementation

[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] Example 1

[0067] like Figures 1-9 As shown, this embodiment provides a distributed data governance system based on a novel pipeline scheduling algorithm, including...

[0068] The system comprises an infrastructure unit 100, a pipeline scheduling unit 200, a unified service unit 300, and a data management unit 400. These units are sequentially connected via network communication. The infrastructure unit 100 provides and manages the basic network topology supporting system operation. The pipeline scheduling unit 200 provides and manages different inter-process communication pipelines and performs allocation and scheduling. The unified service unit 300 provides multiple unified service functions to achieve standardized management and analysis of data for the enterprise. The data management unit 400 performs quality control and comprehensive governance of data.

[0069] The infrastructure unit 100 includes an application platform module 101, a data source module 102, a technical support module 103, and an algorithm management module 104;

[0070] The pipe scheduling unit 200 includes an inter-process communication module 201, an anonymous pipe module 202, a named pipe module 203, and a round-robin scheduling module 204;

[0071] The unified service unit 300 includes a metadata management module 301, a service gateway module 302, a data storage module 303, and an identity authentication module 304;

[0072] The data management unit 400 includes a data integration module 401, a data exchange module 402, a data governance module 403, a master data management module 404, and a data application module 405.

[0073] In this embodiment, the application platform module 101, data source module 102, technical support module 103, and algorithm management module 104 are sequentially connected via network communication and operate in parallel. The application platform module 101 is used to build a multi-functional data governance application platform based on big data and blockchain technology to achieve human-computer interaction. The data source module 102 is used to establish signal connection and data transmission channels between the system and various data sources and to manage and allocate the data source platform. The technical support module 103 is used to load various intelligent technologies to support and improve the functionality of the system. The algorithm management module 104 is used to encapsulate various intelligent algorithms that support system functions.

[0074] Among these, intelligent technologies include, but are not limited to, big data technology, big data analytics technology, blockchain technology, and front-end / back-end separation technology.

[0075] Among them, intelligent algorithms include, but are not limited to, intelligent scheduling algorithms, encryption algorithms, consensus algorithms, and matching algorithms.

[0076] In this embodiment, the signal output terminal of the inter-process communication module 201 is connected to the signal input terminals of the anonymous pipe module 202 and the named pipe module 203. The anonymous pipe module 202 and the named pipe module 203 run in parallel. The signal output terminals of the anonymous pipe module 202 and the named pipe module 203 are connected to the signal input terminal of the polling scheduling module 204. The inter-process communication module 201 is used to allocate a buffer in the kernel using memory as a medium to realize inter-process data exchange communication. The anonymous pipe module 202 is used to create anonymous pipes between related processes by calling the pipe function to provide one-way communication. The named pipe module 203 is used to create system-visible named pipes between any two processes by calling the mknod or mkfifo command to realize communication between the two processes. The polling scheduling module 204 is used to schedule pipes in a polling manner to achieve load balancing of pipe scheduling.

[0077] Specifically, in the round-robin scheduling module 204, the scheduling algorithm types include the round-robin scheduling algorithm and the improved weighted round-robin scheduling algorithm. The round-robin scheduling algorithm is suitable for situations where all servers in the server group have the same hardware and software configuration and the average service requests are relatively balanced. The weighted round-robin scheduling algorithm is suitable for situations where the configurations, installed business applications, and processing capabilities of the servers in the server group are different. Therefore, the calculation expression for weight allocation in the weighted round-robin scheduling algorithm is:

[0078]

[0079] In the formula, degree k Use load on the server.

[0080] In this embodiment, the metadata management module 301, service gateway module 302, data storage module 303, and identity authentication module 304 communicate sequentially via the network and operate in parallel. The metadata management module 301 provides unified management of information description metadata and information authorization metadata, as well as metadata description, metadata interfaces, and a high-performance metadata access mechanism. The service gateway module 302 provides a unified data service gateway to achieve secure and controllable connection and service between internal and external systems, and supports full lifecycle management of APIs. The data storage module 303 provides a unified distributed data storage service, enabling unified storage and access of various data types such as structured data, unstructured data, and files, and can achieve self-description and information authorization through metadata management information. The identity authentication module 304 provides unified user management and identity authentication services, supports multiple access methods, and supports the use of third-party ID cards to meet the identity authentication needs of different application systems.

[0081] The unified metadata management system includes a rich set of built-in acquisition adapters, enabling end-to-end automated acquisition and one-click metadata analysis. It can quickly clean up data resources, understand the origin and development of data, and build data maps to provide basic support for data standardization and data quality. The metadata management functions include, but are not limited to, metadata acquisition, metadata retrieval, data mapping, lineage analysis, and impact analysis.

[0082] During the unified identity authentication process, login methods include, but are not limited to, mobile phone, email, QQ, WeChat, Weibo, etc.

[0083] In this embodiment, the data integration module 401, data exchange module 402, data governance module 403, master data management module 404, and data application module 405 are sequentially connected and interleaved via network communication. The data integration module 401 is used to load, clean, transform, and integrate cross-block data, and supports custom scheduling and graphical monitoring, thereby achieving unified scheduling and unified monitoring to meet the needs of operation and maintenance visualization. The data exchange module 402 is used to realize the transmission and sharing of data or files between several business subsystems, and can integrate data collection, processing, distribution, and exchange transmission. The data governance module 403 is used to integrate and govern data from multiple aspects. The master data management module 404 is used to establish a unified view and centralized management of data that needs to be shared, thereby providing optimal data for business system data calls. The data application module 405 is used to provide users with high-quality, well-governed data for retrieval and application through deep mining, statistical reports, and other formats.

[0084] The data integration functions include, but are not limited to, real-time data collection, cleaning and transformation, encryption and desensitization.

[0085] The master data management functions include, but are not limited to, application identification, master data retrieval, master data storage, sharing and publishing, and master data monitoring.

[0086] Furthermore, the data exchange module 402 includes a data transmission module 4021, a node management module 4022, a pipeline matching module 4023, and an exchange approval module 4024. The data transmission module 4021, node management module 4022, pipeline matching module 4023, and exchange approval module 4024 are sequentially connected via communication. The data transmission module 4021 is used to implement secure data transmission in various ways through multiple data exchange components. The node management module 4022 is used for visual configuration of data transmission and exchange nodes, controlling the data transmission status of nodes, supporting the transmission of various data formats, and shielding the data type differences between systems. The pipeline matching module 4023 is used to select the inter-process pipeline to be used based on the correlation analysis results between business data and processes using metadata, through a matching algorithm. The exchange approval module 4024 is used to support the design of data exchange task methods and to realize the detection and approval of data exchange tasks through flexible extraction and exchange of data.

[0087] During the data exchange process, data exchange components include, but are not limited to, table exchange, file transfer, SFTP upload and download, and HTTP components; data transmission methods include, but are not limited to, encrypted data transmission, power outage resume transmission, de-identification algorithms, dual control of data permissions and function permissions, data partitioning, and parallel loading technology.

[0088] During node management, data formats include, but are not limited to, mainstream databases, text files, Excel files, API interfaces, and WebService services.

[0089] Specifically, in the pipeline matching module 4023, the correlation-based matching algorithm is used to determine the relationship between process relationships and pipeline types, and the NCC normalized cross-correlation coefficient is used to represent the degree of correlation between the two. The calculation expression is as follows:

[0090]

[0091] In the formula, X and Y are two random variables, μ X μ Y Let σ be the mean of two random variables. X σ Y Let be the standard deviation of two random variables;

[0092] In the above formula, the denominator is the standard deviation of the two random variables, which serves as a normalization function. The mean of the two random variables is also subtracted from the numerator, which is called the centering.

[0093] Normalization and centering are not the essence of the correlation coefficient. After separating these two, what remains is the expectation of the product of two random variables, that is, the inner product of two vectors, which is the essence of the correlation coefficient.

[0094] Specifically, it can be understood as follows: if random variables X and Y are sampled multiple times and the sampled samples are placed into two vectors, then the inner product of these two vectors is the expectation of the product of X and Y; conversely, it can be understood as follows: if the two vectors are regarded as the joint distribution of two random variables, then the inner product of the two vectors is the expectation of the product of the two random variables.

[0095] Furthermore, the data governance module 403 includes a standards management module 4031, a quality management module 4032, a security management module 4033, and a lifecycle module 4034; the standards management module 4031, quality management module 4032, security management module 4033, and lifecycle module 4034 are sequentially connected via network communication; the standards management module 4031 is used to provide a comprehensive and complete data standards management process and method to determine and establish a single, accurate, and authoritative source of facts, to achieve complete, effective, consistent, standardized, open, and shared management of data, and to provide standard basis for data quality inspection and data security management; the quality management module 4032 is used to guide... The system integrates the workflow of quality assessment, quality inspection, quality rectification, and quality reporting through methods such as visualization and automation, forming a complete data quality management closed loop based on data standards and metadata. The security management module 4033 provides various data security management measures throughout the data governance process, including encryption, desensitization, obfuscation, and database authorization monitoring of privacy data, to ensure comprehensive data security. The lifecycle module 4034 records the entire flow of data from its creation and initial storage to its deletion due to obsolescence, and performs near-line archiving, offline archiving, destruction, and full lifecycle monitoring of data.

[0096] The data standards management functions include, but are not limited to, standard publishing, standard mapping, and standard querying.

[0097] Data quality management includes, but is not limited to, rule management, quality reports, and problem rectification.

[0098] The data security management functions include, but are not limited to, change tracking, anomaly monitoring, data anonymization, secure transmission, and de-privacy protection.

[0099] The data lifecycle management functions include, but are not limited to, data policies, level agreements, classification and assessment, data archiving, and data destruction.

[0100] This embodiment also provides a method for operating a distributed data governance system based on a novel pipeline scheduling algorithm, including the following steps:

[0101] S1. Based on blockchain and big data technologies, build a distributed data governance system and connect the application platform with multiple data sources;

[0102] S2, Load the corresponding inter-process communication pipe creation method and round-robin scheduling algorithm;

[0103] S3. Implement unified management of the data governance system, pre-manage metadata, and monitor and adjust the standard quality of metadata;

[0104] S4. When exchanging data between the data source and the data integration module, based on the metadata's provisions for different processes and the results of lineage analysis, and according to the correlation-based matching algorithm, select and create the corresponding inter-process pipelines to achieve secure data transmission.

[0105] S5. Implement security management and lifecycle management for data, and manage data collection, transformation, cleaning, and archiving in a unified manner to provide high-quality data.

[0106] S6. Apply the well-managed high-quality data to front-end business operations;

[0107] S7. Users access the system's application platform through various methods and log in after unified identity authentication, obtain high-quality data from the system, and should call relevant data for application.

[0108] like Figure 9 As shown, this embodiment also provides a running device for a distributed data governance system based on a novel pipeline scheduling algorithm. The device includes a processor, a memory, and a computer program stored in the memory and running on the processor.

[0109] The processor includes one or more processing cores. The processor is connected to the memory via a bus. The memory is used to store program instructions. When the processor executes the program instructions in the memory, it implements the aforementioned distributed data governance system based on the novel pipeline-based scheduling algorithm.

[0110] Optionally, the memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0111] In addition, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned distributed data governance system based on a novel pipeline-based scheduling algorithm.

[0112] Optionally, the present invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the aforementioned aspects of the pipeline-based novel scheduling algorithm for distributed data governance system.

[0113] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0114] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distributed data governance system based on a novel scheduling algorithm for pipelines, characterized in that: Comprise Infrastructure unit (100), pipeline scheduling unit (200), unified service unit (300) and data management unit (400); the infrastructure unit (100), the pipeline scheduling unit (200), the unified service unit (300) and the data management unit (400) are sequentially connected by network communication; the infrastructure unit (100) is used to provide and manage the basic network topology architecture supporting system operation; the pipeline scheduling unit (200) is used to provide and manage different interprocess communication pipelines and scheduling; the unified service unit (300) is used to realize the unified standardized management and analysis of enterprise data through multiple unified service functions; the data management unit (400) is used for quality control and comprehensive management of data; The infrastructure unit (100) comprises an application platform module (101), a data source module (102), a technical support module (103) and an algorithm management module (104); The pipeline scheduling unit (200) comprises an interprocess communication module (201), an anonymous pipeline module (202), a named pipeline module (203) and a polling scheduling module (204); The unified service unit (300) comprises a metadata management module (301), a service gateway module (302), a data storage module (303) and an identity authentication module (304); The data management unit (400) comprises a data integration module (401), a data exchange module (402), a data governance module (403), a master data management module (404) and a data application module (405); The operation method of the system comprises the following steps: S1, based on the block chain and big data technology, build a distributed data governance system, and connect the application platform with multiple data sources; S2, load the creation mode of the interprocess communication pipeline and the polling scheduling algorithm; S3, unified management of data governance system, preset management of metadata, and monitoring and adjusting the standard quality of metadata; S4, when data exchange between data source and data integration module, based on the regulation and blood relationship analysis result of different interprocesses, according to the matching algorithm based on correlation, select and create the corresponding interprocess pipeline, realize the safe transmission of data; S5, safe management and life cycle management of data, unified collection, conversion, cleaning and archiving management of data, to provide high-quality gold data; S6, apply the well governed high-quality data to the front-end business; S7, users access through multiple ways, and log in to the application platform of the system after unified identity authentication, obtain high-quality data from the system, and call related data for application; In the polling scheduling module (204), the type of scheduling algorithm includes a polling scheduling algorithm and a weight polling scheduling algorithm improved on the basis of the polling scheduling algorithm, wherein: the polling scheduling algorithm is applicable to the case that all servers in the server group have the same software and hardware configuration and the average service request is relatively balanced, and the weight polling scheduling algorithm is applicable to the case that the configurations, installed service applications and processing capacities of the servers in the server group are different; in the weight polling scheduling algorithm, the calculation expression of weight allocation is: ; In the formula, is the load used by the server.

2. The distributed data governance system based on pipeline new scheduling algorithm according to claim 1, characterized in that: The application platform module (101), the data source module (102), the technical support module (103) and the algorithm management module (104) are sequentially connected through network communication and run side by side; the application platform module (101) is used to build a multifunctional data governance application platform based on big data and blockchain technology to realize human-computer interaction; the data source module (102) is used to build a signal connection and data transmission channel between the system and each data source and manage and distribute the data source platform; the technical support module (103) is used to load various intelligent technologies to support and improve the functionality of the system; the algorithm management module (104) is used to encapsulate various intelligent algorithms supporting the functionality of the system.

3. The distributed data governance system based on pipeline-based novel scheduling algorithm of claim 2, wherein: The signal output end of the inter-process communication module (201) is connected with the signal input end of the anonymous pipe module (202) and the named pipe module (203), the anonymous pipe module (202) and the named pipe module (203) run side by side, and the signal output end of the anonymous pipe module (202) and the named pipe module (203) is connected with the signal input end of the polling scheduling module (204); the inter-process communication module (201) is used to open a buffer area in the kernel to realize the communication of inter-process data exchange by taking memory as a medium; the anonymous pipe module (202) is used to create an anonymous pipe between processes with blood relationship by calling the pipe function to provide one-way communication; the named pipe module (203) is used to create a system visible named pipe between any two processes by calling the mknod or mkfifo command to realize the communication between the two processes; the polling scheduling module (204) is used to schedule the pipe by polling to realize the load balancing of pipe scheduling.

4. The distributed data governance system based on pipeline new scheduling algorithm according to claim 3, characterized in that: The metadata management module (301), the service gateway module (302), the data storage module (303) and the identity authentication module (304) are sequentially connected through network communication and run side by side; the metadata management module (301) is used to provide unified management of information description metadata and information authorization metadata, metadata description, metadata interface and high-performance metadata access mechanism; The service gateway module (302) is configured to provide a unified data service gateway to realize safe and controllable connection and service between internal and external systems, and support full life cycle management of API; the data storage module (303) is configured to provide a unified distributed data storage service, which can realize unified storage and access of structured data, unstructured data and file data of various types, and can realize self-description and information authorization through metadata management information; the identity authentication module (304) is configured to provide unified user management and identity authentication service, and support multi-mode access and use of third-party identity cards to meet the identity authentication needs of different application systems.

5. The distributed data governance system based on pipeline-based novel scheduling algorithm of claim 4, wherein: The data integration module (401), the data exchange module (402), the data governance module (403), the master data management module (404) and the data application module (405) are sequentially connected by network communication and run alternately; the data integration module (401) is configured to realize loading, cleaning, conversion and integration of cross-block data, and support custom scheduling and graphical monitoring, so as to realize unified scheduling and unified monitoring to meet the visual operation and maintenance needs; the data exchange module (402) is configured to realize data or file transmission and sharing between a plurality of business subsystems, and can integrate data collection, processing and distribution, exchange and transmission; the data governance module (403) is configured to integrate and govern data from multiple aspects; the master data management module (404) is configured to establish a unified view and centralized management for data to be shared, so as to provide optimal data for business system data calling; the data application module (405) is configured to provide high-quality data to users for application by deep mining and statistical report format.

6. The distributed data governance system based on pipeline-based novel scheduling algorithm of claim 5, wherein: The data exchange module (402) includes a data transmission module (4021), a node management module (4022), a pipeline matching module (4023) and an exchange approval module (4024); the data transmission module (4021), the node management module (4022), the pipeline matching module (4023) and the exchange approval module (4024) are sequentially connected by communication; the data transmission module (4021) is configured to realize data safe transmission process in multiple modes through various data exchange components; the node management module (4022) is configured to visually configure data transmission exchange nodes, control data transmission state of nodes, and support transmission of various format data, while shielding data type differences between systems; the pipeline matching module (4023) is configured to select inter-process pipelines used according to the correlation analysis result of business data and processes by metadata; the exchange approval module (4024) is configured to support design of data exchange task mode and realize detection and approval of data exchange task by flexibly extracting and exchanging data.

7. The distributed data governance system based on pipeline-based novel scheduling algorithm of claim 6, wherein: In the pipeline matching module (4023), the correlation between the judgment process and the pipeline type is matched by using a correlation-based matching algorithm, and a normalized cross-correlation coefficient (NCC) is used to represent the correlation degree between them, and the calculation expression is: ; In the formula, X, Y are two random variables, , are the mean values of the two random variables, respectively, , are the standard deviations of the two random variables. In the above formula, the denominator is the standard deviation of two random variables, which plays a role of normalization, and the mean of two random variables is subtracted from the numerator, which is called centering.

8. The distributed data governance system based on pipeline-based novel scheduling algorithm of claim 7, wherein: The data governance module (403) comprises a standard management module (4031), a quality management module (4032), a security management module (4033) and a life cycle module (4034); the standard management module (4031), the quality management module (4032), the security management module (4033) and the life cycle module (4034) are sequentially connected by network communication; the standard management module (4031) is used to provide a complete and integrated data standard management process and method, to determine and establish a single, accurate and authoritative fact source, to realize complete, effective, consistent, standardized, open and shared management of data, and to provide a standard basis for data quality inspection and data security management; the quality management module (4032) is used to integrate the quality evaluation, quality inspection, quality improvement and quality report working environment through guided and visualized operation means, to form a complete data quality management closed loop taking data standards as data inspection basis and metadata as data inspection object; the security management module (4033) is used to provide encryption, desensitization, fuzzing processing and database authorization monitoring of private data in the whole data governance process, to realize all-round protection of data security operation process; the life cycle module (4034) is used to record the whole flow process of data from creation and initial storage to deletion, to perform near-line archiving, off-line archiving, destruction and full life cycle monitoring of data during the survival period.

Citation Information

Patent Citations

  • Systems and methods for managing access control between processes in a computing device

    CN111357256A

  • Big data acquisition and governance quick retrieval system based on data lake

    CN111460236A