A multi-machine multi-process synchronous data management system and method based on block tasks

The multi-machine, multi-process synchronous data management system based on blockchain and big data technologies solves the problem of low data synchronization efficiency in multi-machine, multi-process systems, enables fast and accurate synchronization and management of massive amounts of data, simplifies user operations, and improves system efficiency and data security.

CN113672399BActive Publication Date: 2026-04-21CHINA NAT BUILDING MATERIALS TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA NAT BUILDING MATERIALS TECH CO LTD
Filing Date
2021-08-13
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, multi-machine, multi-process synchronous data management systems are inefficient, making it difficult to synchronize data quickly and accurately, increasing the difficulty of data management. Furthermore, users need to manually operate these systems, which consumes a lot of manpower and resources, and data duplication, corruption, and chaos are likely to occur.

Method used

A multi-machine, multi-process synchronous data management system based on block tasks is adopted. It utilizes blockchain and big data technologies to build a network architecture, combines inter-process communication and synchronization methods, and achieves the synchronization and management of massive amounts of data through steps such as synchronization software, deduplication and cleaning, and classified storage. A user management mechanism is also introduced.

Benefits of technology

It improves system parallel efficiency, reduces data transmission workload and probability of loss, lowers the probability of data corruption, improves resource utilization and system throughput, simplifies user operations, and ensures smooth business operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113672399B_ABST
    Figure CN113672399B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data management technology, specifically to a multi-machine, multi-process synchronous data management system and method based on block tasks. It includes a technical support unit, a multi-process synchronization unit, a data synchronization unit, and a data management unit. The technical support unit supports system connection and operation; the multi-process synchronization unit manages the synchronization of multiple concurrent processes; the data synchronization unit implements data synchronization within the system; and the data management unit manages the synchronized data. The system designed in this invention can improve system parallel efficiency, quickly achieve massive data synchronization, and then process and provide the data for user access and application. Its method can reduce the workload of data transmission, lower the probability of data loss, improve data synchronization efficiency, reduce user workload, lower the probability of data corruption, ensure business operation progress, effectively improve resource utilization, and increase system throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and more specifically, to a multi-machine, multi-process synchronous data management system and method based on block tasks. Background Technology

[0002] With the continuous development of computer technology, many individuals and businesses can operate and manage their operations online, involving large amounts of data. In daily applications, this data is often distributed across different hardware devices, software programs, and network platforms. When this data needs to be synchronized, users often need to process it manually, which consumes significant manpower, resources, and time. Furthermore, the process is prone to data duplication, corruption, and chaos, greatly impacting business operations. Windows applications have two message delivery methods—direct and queued. In traditional operating systems, multiprogramming is typically used to improve resource utilization and system throughput, loading multiple programs into memory simultaneously and executing them concurrently. Furthermore, the introduction of processes into the OS allows for concurrent execution of multiple programs, improving resource utilization and significantly increasing system throughput. However, it also makes the system more complex, especially in multi-machine, multi-process systems where the disorderly competition for system resources leads to uncertainty in the results of each processing step. This makes it difficult to synchronize data quickly and accurately, increasing the difficulty of data management. Currently, however, there is no efficient multi-machine, multi-process synchronized data management system or method. Summary of the Invention

[0003] The purpose of this invention is to provide a multi-machine, multi-process synchronous data management system and method based on block tasks, so as to solve the problems mentioned in the background art.

[0004] To address the aforementioned technical problems, one objective of this invention is to provide a multi-machine, multi-process synchronous data management system based on block tasks, including...

[0005] The system comprises a technical support unit, a multi-process synchronization unit, a data synchronization unit, and a data management unit. These units are sequentially connected via network communication. The technical support unit loads various intelligent technologies to support the system's connection and operation. The multi-process synchronization unit manages the synchronization of multiple processes running concurrently on multiple processors and within a single processor. The data synchronization unit selects appropriate software or methods to achieve data synchronization within the system. The data management unit processes, applies, and manages the synchronized data.

[0006] The technical support unit includes a block connection module, a synchronization software module, a big data module, and an inter-process communication module;

[0007] The multi-process synchronization unit includes a hardware synchronization module, a semaphore module, a monitor mechanism module, and a Petri net module.

[0008] The data synchronization unit includes a publish-subscribe module, an SQL-JOB module, a message queue module, and a select-call module;

[0009] The data management unit includes a deduplication and cleaning module, a summarization and classification module, a secure storage module, and a scheduling and application module;

[0010] This multi-machine, multi-process synchronous data management system based on block tasks first establishes a basic network architecture supported by technologies such as blockchain and big data. On this basis, it builds and loads data synchronization software, introduces inter-process communication functions and process synchronization solutions, realizes inter-process communication and process synchronization within a single machine, and achieves synchronization between multiple machines and multiple processes. Then, based on multi-process synchronization, relevant methods are selected to synchronize massive amounts of data. Next, the massive amounts of data are deduplicated, cleaned, classified, stored and managed for security. Finally, a user management mechanism is introduced to allow users to schedule and apply the data.

[0011] As a further improvement to this technical solution, the block connection module, the synchronization software module, the big data module, and the inter-process communication module are sequentially connected via network communication. The block connection module is used to connect multiple processors in the system as nodes in a loop based on blockchain. The synchronization software module is used to load various data synchronization software into each processor to quickly achieve data synchronization within the machine. The big data module is used to manage and allocate massive amounts of synchronized data through big data technology and big data processing technology. The inter-process communication module is used to provide and manage different inter-process communication channels and allocate and schedule them for multiple concurrent processes within a single processor. Specifically, it uses memory as a medium to open a buffer in the kernel to realize inter-process data exchange communication.

[0012] Data synchronization software includes, but is not limited to, SyncToy v2.0, Allway Sync, Live Sync, Dropbox, and Data Sync Wizard.

[0013] As a further improvement to this technical solution, the inter-process communication module includes an anonymous pipe module, a named pipe module, and a round-robin scheduling module. The anonymous pipe module runs in parallel with the named pipe module, and the signal output terminals of the anonymous pipe module and the named pipe module are connected to the signal input terminal of the round-robin scheduling module. The anonymous pipe module is used to create anonymous pipes between related processes by calling the pipe function to provide one-way communication. The named pipe module is used to create system-visible named pipes between any two processes by calling the mknod or mkfifo command to achieve communication between the two processes. The round-robin scheduling module is used to schedule pipes in a round-robin manner to achieve load balancing of pipe scheduling.

[0014] As a further improvement to this technical solution, the round-robin scheduling module includes a round-robin scheduling algorithm and an improved weighted round-robin scheduling algorithm. The round-robin scheduling algorithm is suitable for situations where all servers in a server group have the same hardware and software configuration and the average service requests are relatively balanced. The weighted round-robin scheduling algorithm is suitable for situations where the configurations, installed business applications, and processing capabilities of the servers in the server group are different. Therefore, the calculation expression for weight allocation in the weighted round-robin scheduling algorithm is:

[0015] ;

[0016] In the formula, Use load on the server.

[0017] As a further improvement to this technical solution, the hardware synchronization module, the semaphore module, the monitor mechanism module, and the Petri net module are sequentially connected via network communication and operate in parallel. The hardware synchronization module is used to provide special hardware instructions from the computer to allow the detection and correction of the content in a word or the swapping of two words, thereby solving simple process synchronization problems. The semaphore module is used to evolve from integer semaphores to record-type semaphores to semaphore sets, and to widely apply the semaphore mechanism in single-processor and multi-processor systems as well as computer networks to achieve more complex process synchronization. The monitor mechanism module is used to accept or block multiple concurrent processes requesting access to shared resources based on resource availability, ensuring that only one process enters the monitor at a time, executes the monitor, uses the shared resources, and selects unified management of all access to shared resources to effectively achieve process mutual exclusion. The Petri net module is used in parallel computing and distributed database design to construct system models and analyze dynamic characteristics, and is suitable for describing and analyzing problems in asynchronous concurrent systems.

[0018] The physical description of Petri net structural elements includes: a) Locations: describing possible local states of the system (conditions or states, such as queues, resources, etc.); b) Transitions: describing events that modify the system state, such as information processing, sending, resource access, etc.; c) Arcs: defining the relationship between local states and events, connecting locations and transitions; d) Markers or tags: active elements contained in the location set. If a location describes a condition, a tag indicates that the condition is true; otherwise, the condition is false. They can also be used to represent processed information units, data resource units, and object entities such as customers and users.

[0019] As a further improvement to this technical solution, the semaphore module includes an integer information module, a record-type information module, an AND-type information module, and an information set module. The integer information module, the record-type information module, the AND-type information module, and the information set module are sequentially connected via network communication and run in parallel. The integer information module, except for initialization, accesses data information through only two standard atomic operations, and is used to handle processes in a "busy waiting" state. The record-type information module uses a record-type data structure and adds a process linked list pointer to the semaphore mechanism to link all waiting processes accessing the same critical resource. The AND-type information module allocates all resources needed by a process during its entire operation to the process at once, and releases them all together after the process has finished using them, thereby achieving AND synchronization and avoiding deadlock. The information set module completes the allocation or release of all resources requested by a process and the different resource requirements for each type of resource in a single atomic operation.

[0020] As a further improvement to this technical solution, the publish / subscribe module, the SQL-JOB module, and the message queue module are connected sequentially via network communication and run in parallel. The signal output terminals of the publish / subscribe module, the SQL-JOB module, and the message queue module are connected to the signal input terminal of the selection and invocation module. The publish / subscribe module is used to quickly achieve data backup and synchronization without writing any code, through the publish / subscribe database backup mechanism built into SQL Server. The SQL-JOB module achieves data synchronization through SQL Job scheduled jobs, that is, it is used to read data from the source server and update the target server by writing SQL statements through the connection between the target server and the source server. The message queue module is used to provide SQL Server with queues, reliable message passing, and a powerful asynchronous programming model through SQL Server Service Broker, thereby providing reliable message passing services, shortening interactive response time to increase the overall throughput of the application, and thus achieving data synchronization. The selection and invocation module is used to select the optimal and applicable data synchronization method according to the source and type of the data.

[0021] As a further improvement to this technical solution, the signal output terminal of the deduplication and cleaning module is connected to the signal input terminal of the summarization and classification module, the signal output terminal of the summarization and classification module is connected to the signal input terminal of the secure storage module, and the signal output terminal of the secure storage module is connected to the signal input terminal of the scheduling and application module. The deduplication and cleaning module is used to compare and deduplicate massive amounts of data and remove duplicate, damaged, invalid, and expired data. The summarization and classification module is used to classify the synchronized and cleaned massive amounts of data using clustering techniques. The secure storage module is used to securely manage the data and distribute its storage according to the classified items. The scheduling and application module is used to provide users with a way to schedule data for application through querying.

[0022] Data applications include viewing, downloading, and forwarding.

[0023] As a further improvement to this technical solution, the inductive classification module adopts the K-Means clustering algorithm.

[0024] The second objective of this invention is to provide a multi-machine, multi-process synchronous data management method based on block tasks, including the aforementioned multi-machine, multi-process synchronous data management system based on block tasks, comprising the following steps:

[0025] S1. Based on blockchain, build a system architecture that includes multiple processors, and install multiple data synchronization software in each processor within the system;

[0026] S2, the processors distributed across the nodes of the block, collect, calculate and process various types of data near the node. In the process of handling a large number of complex business processes, the operating system schedules and manages the execution of multiple processes.

[0027] S3. When a single processor performs multiple processes concurrently, it allows multiple instruction streams to be executed concurrently in a program. Each instruction stream is called a process, and each process is independent of the others.

[0028] S4. During the operation of a single processor, a communication channel is continuously established between multiple concurrent processes through a polling scheduling algorithm to achieve initial data connection between multiple processes on a single processor, and the optimal and suitable data synchronization software is selected to achieve initial data synchronization.

[0029] S5. Among multiple concurrent processes on multiple processors, the optimal and applicable process synchronization method is selected to achieve process synchronization according to the cooperation conditions between processes and the corresponding constraint relationship.

[0030] S6. Combine several concurrent processes after synchronization, select the optimal and applicable data synchronization software or method to synchronize and integrate the massive data involved in multiple machines and multiple processes.

[0031] S7. Compare and check massive amounts of data for duplicates, clean and filter out invalid, expired, damaged and duplicate data, classify and summarize the data through clustering methods, distribute and store the classified data in a cloud database, and perform certain security management on the data.

[0032] S8. Add user management function to the system and assign corresponding operation permissions to different users. Users can log in to the system with legitimate identity, perform data query operations, and call, forward, and download relevant data for application within their permissions.

[0033] The third objective of this invention is to provide an operating device for a multi-machine, multi-process synchronous data management system based on block tasks, including a processor, a memory, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the aforementioned multi-machine, multi-process synchronous data management system and method based on block tasks.

[0034] The fourth objective of this invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned multi-machine, multi-process synchronous data management system and method based on block tasks.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] 1. This multi-machine, multi-process synchronous data management system based on block tasks builds a multi-processor system on the basis of blockchain technology, which can improve the parallel efficiency of the system. A single processor can perform multiple processes concurrently, and multiple processes between multiple processors can be synchronized, thereby quickly realizing the synchronization of massive amounts of data. Then, the massive amounts of data are processed by deduplication, cleaning, classification, and distributed storage, and made available for users to access and apply.

[0037] 2. This multi-machine, multi-process synchronous data management method based on block tasks can perform business processing simultaneously and efficiently at each block node, reducing the workload of data transmission and the probability of data loss. At the same time, through concurrent multi-processes, multiple data connection and synchronization operations are performed to improve the efficiency of data synchronization, reduce the workload of users manually performing data synchronization operations, reduce the probability of data corruption, ensure the progress of business operations, effectively improve resource utilization, and increase system throughput. Attached Figure Description

[0038] Figure 1 This is an exemplary product architecture diagram of the present invention;

[0039] Figure 2 This is a structural diagram of the overall system device of the present invention;

[0040] Figure 3 This is one of the partial overall device structure diagrams of the present invention;

[0041] Figure 4 This is the second partial overall device structure diagram of the present invention;

[0042] Figure 5 This is the third partial overall device structure diagram of the present invention;

[0043] Figure 6 This is the fourth partial overall device structure diagram of the present invention;

[0044] Figure 7 This is the fifth partial overall device structure diagram of the present invention;

[0045] Figure 8 This is the sixth partial overall device structure diagram of the present invention;

[0046] Figure 9 This is a flowchart of the method of the present invention;

[0047] Figure 10 This is a structural diagram of an exemplary electronic computer product device according to the present invention.

[0048] In the picture:

[0049] 100. Technical Support Unit; 101. Block Chaining Module; 102. Synchronization Software Module; 103. Big Data Module; 104. Inter-Process Communication Module; 1041. Anonymous Pipe Module; 1042. Named Pipe Module; 1043. Round Robin Scheduling Module;

[0050] 200. Multi-process synchronization unit; 201. Hardware synchronization module; 202. Semaphore module; 2021. Integer information module; 2022. Record-type information module; 2023. AND-type information module; 2024. Information set module; 203. Monitor mechanism module; 204. Petri net module;

[0051] 300. Data Synchronization Unit; 301. Publish / Subscribe Module; 302. SQL-JOB Module; 303. Message Queue Module; 304. Select and Call Module;

[0052] 400. Data Management Unit; 401. Deduplication and Cleaning Module; 402. Summarization and Classification Module; 403. Secure Storage Module; 404. Scheduling and Application Module. Detailed Implementation

[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] Example 1

[0055] like Figures 1-8 As shown, this embodiment provides a multi-machine, multi-process synchronous data management system based on block tasks, including...

[0056] The system comprises a technical support unit 100, a multi-process synchronization unit 200, a data synchronization unit 300, and a data management unit 400. These units are sequentially connected via network communication. The technical support unit 100 loads various intelligent technologies to support the system's connection and operation. The multi-process synchronization unit 200 manages the synchronization of multiple processes running concurrently on multiple processors or within a single processor. The data synchronization unit 300 selects appropriate software or methods to achieve data synchronization within the system. The data management unit 400 manages the synchronized data, including processing and application.

[0057] The technical support unit 100 includes a block connection module 101, a synchronization software module 102, a big data module 103, and an inter-process communication module 104.

[0058] The multi-process synchronization unit 200 includes a hardware synchronization module 201, a semaphore module 202, a monitor mechanism module 203, and a Petri net module 204.

[0059] The data synchronization unit 300 includes a publish-subscribe module 301, an SQL-JOB module 302, a message queue module 303, and a select-call module 304;

[0060] The data management unit 400 includes a deduplication and cleaning module 401, a summarization and classification module 402, a secure storage module 403, and a scheduling and application module 404;

[0061] This multi-machine, multi-process synchronous data management system based on block tasks first establishes a basic network architecture supported by technologies such as blockchain and big data. On this basis, it builds and loads data synchronization software, introduces inter-process communication functions and process synchronization solutions, realizes inter-process communication and process synchronization within a single machine, and achieves synchronization between multiple machines and multiple processes. Then, based on multi-process synchronization, relevant methods are selected to synchronize massive amounts of data. Next, the massive amounts of data are deduplicated, cleaned, classified, stored and managed for security. Finally, a user management mechanism is introduced to allow users to schedule and apply the data.

[0062] In this embodiment, the block connection module 101, the synchronization software module 102, the big data module 103, and the inter-process communication module 104 are sequentially connected via network communication. The block connection module 101 is used to connect multiple processors in the system as nodes in a loop based on blockchain. The synchronization software module 102 is used to load various data synchronization software in each processor to quickly realize data synchronization within the machine. The big data module 103 is used to manage and allocate massive amounts of synchronized data through big data technology and big data processing technology. The inter-process communication module 104 is used to provide and manage different inter-process communication channels for multiple concurrent processes in a single processor and to allocate and schedule them. Specifically, it uses memory as a medium to open a buffer in the kernel to realize inter-process data exchange communication.

[0063] Data synchronization software includes, but is not limited to, SyncToy v2.0, Allway Sync, Live Sync, Dropbox, and Data Sync Wizard.

[0064] Furthermore, the inter-process communication module 104 includes an anonymous pipe module 1041, a named pipe module 1042, and a round-robin scheduling module 1043. The anonymous pipe module 1041 and the named pipe module 1042 operate in parallel, and the signal output terminals of the anonymous pipe module 1041 and the named pipe module 1042 are connected to the signal input terminals of the round-robin scheduling module 1043. The anonymous pipe module 1041 is used to create anonymous pipes between related processes by calling the pipe function to provide one-way communication. The named pipe module 1042 is used to create system-visible named pipes between any two processes by calling the mknod or mkfifo command to achieve communication between the two processes. The round-robin scheduling module 1043 is used to schedule pipes in a round-robin manner to achieve load balancing of pipe scheduling.

[0065] Specifically, in the round-robin scheduling module 1043, the scheduling algorithm types include the round-robin scheduling algorithm and the improved weighted round-robin scheduling algorithm. The round-robin scheduling algorithm is suitable for situations where all servers in a server group have the same hardware and software configuration and the average service requests are relatively balanced. The weighted round-robin scheduling algorithm is suitable for situations where the configurations, installed business applications, and processing capabilities of the servers in the server group are different. Therefore, the calculation expression for weight allocation in the weighted round-robin scheduling algorithm is:

[0066] ;

[0067] In the formula, Use load on the server.

[0068] In this embodiment, the hardware synchronization module 201, semaphore module 202, monitor mechanism module 203, and Petri net module 204 are connected sequentially via network communication and operate in parallel. The hardware synchronization module 201 is used to provide special hardware instructions from the computer to allow the detection and correction of the content in a word or the exchange of two words, thereby solving simple process synchronization problems. The semaphore module 202 is used to evolve from integer semaphores to record-type semaphores to semaphore sets, and to widely apply the semaphore mechanism in single-processor and multi-processor systems as well as computer networks to achieve more complex process synchronization. The monitor mechanism module 203 is used to accept or block multiple concurrent processes requesting access to shared resources according to the resource situation, to ensure that only one process enters the monitor at a time, executes the monitor, uses the shared resources, and selects unified management of all access to shared resources to effectively achieve process mutual exclusion. The Petri net module 204 is used in computer parallel computing and distributed database design to construct system models and analyze dynamic characteristics, and is suitable for describing and analyzing problems in asynchronous concurrent systems.

[0069] The physical description of Petri net structural elements includes: a) Locations: describing possible local states of the system (conditions or states, such as queues, resources, etc.); b) Transitions: describing events that modify the system state, such as information processing, sending, resource access, etc.; c) Arcs: defining the relationship between local states and events, connecting locations and transitions; d) Markers or tags: active elements contained in the location set. If a location describes a condition, a tag indicates that the condition is true; otherwise, the condition is false. They can also be used to represent processed information units, data resource units, and object entities such as customers and users.

[0070] Further, the semaphore module 202 includes an integer information module 2021, a record-type information module 2022, an AND-type information module 2023, and an information set module 2024; the integer information module 2021, the record-type information module 2022, the AND-type information module 2023, and the information set module 2024 are connected sequentially via network communication and operate in parallel; the integer information module 2021 is used to access data information through only two standard atomic operations, except for initialization, and is used to handle processes in a "busy waiting" state; the record-type... The information volume module 2022 uses a record-type data structure to add a process linked list pointer to the semaphore mechanism to link all waiting processes accessing the same critical resource; the AND-type information volume module 2023 is used to allocate all the resources needed by a process during its entire operation to the process at once, and then release them all at once after the process has finished using them, thereby achieving AND synchronization and avoiding deadlock; the information set module 2024 is used to complete the allocation or release of all resources requested by a process and the different resource requirements of each type of resource in a single atomic operation.

[0071] In this embodiment, the publish / subscribe module 301, the SQL-JOB module 302, and the message queue module 303 are connected sequentially via network communication and run in parallel. The signal output terminals of the publish / subscribe module 301, the SQL-JOB module 302, and the message queue module 303 are connected to the signal input terminals of the selection and invocation module 304. The publish / subscribe module 301 is used to quickly achieve data backup and synchronization without writing any code, through the publish / subscribe database backup mechanism built into SQL Server. The SQL-JOB module 302 achieves data synchronization through SQL Job scheduled jobs, that is, it is used to read data from the source server and update the target server by writing SQL statements through the connection between the target server and the source server. The message queue module 303 is used to provide SQL Server with queues, reliable message passing, and a powerful asynchronous programming model through SQL Server Service Broker, thereby providing reliable message passing services, shortening interactive response time to increase the total throughput of the application, and thus achieving data synchronization. The selection and invocation module 304 is used to select the optimal and applicable data synchronization method according to the source and type of the data.

[0072] In this embodiment, the signal output terminal of the deduplication cleaning module 401 is connected to the signal input terminal of the summarization and classification module 402, the signal output terminal of the summarization and classification module 402 is connected to the signal input terminal of the secure storage module 403, and the signal output terminal of the secure storage module 403 is connected to the signal input terminal of the scheduling application module 404. The deduplication cleaning module 401 is used to compare and check massive amounts of data for deduplication and to remove duplicate, damaged, invalid, and expired data. The summarization and classification module 402 is used to classify the synchronized and cleaned massive amounts of data using clustering techniques. The secure storage module 403 is used to manage the data securely and distribute the data according to the classified items. The scheduling application module 404 is used to provide users with a way to schedule data for application through querying.

[0073] Data applications include viewing, downloading, and forwarding.

[0074] Specifically, the inductive classification module 402 uses the K-Means clustering algorithm.

[0075] like Figure 9 As shown, this embodiment provides a multi-machine, multi-process synchronous data management method based on block tasks, including the aforementioned multi-machine, multi-process synchronous data management system based on block tasks, comprising the following steps:

[0076] S1. Based on blockchain, build a system architecture that includes multiple processors, and install multiple data synchronization software in each processor within the system;

[0077] S2, the processors distributed across the nodes of the block, collect, calculate and process various types of data near the node. In the process of handling a large number of complex business processes, the operating system schedules and manages the execution of multiple processes.

[0078] S3. When a single processor performs multiple processes concurrently, it allows multiple instruction streams to be executed concurrently in a program. Each instruction stream is called a process, and each process is independent of the others.

[0079] S4. During the operation of a single processor, a communication channel is continuously established between multiple concurrent processes through a polling scheduling algorithm to achieve initial data connection between multiple processes on a single processor, and the optimal and suitable data synchronization software is selected to achieve initial data synchronization.

[0080] S5. Among multiple concurrent processes on multiple processors, the optimal and applicable process synchronization method is selected to achieve process synchronization according to the cooperation conditions between processes and the corresponding constraint relationship.

[0081] S6. Combine several concurrent processes after synchronization, select the optimal and applicable data synchronization software or method to synchronize and integrate the massive data involved in multiple machines and multiple processes.

[0082] S7. Compare and check massive amounts of data for duplicates, clean and filter out invalid, expired, damaged and duplicate data, classify and summarize the data through clustering methods, distribute and store the classified data in a cloud database, and perform certain security management on the data.

[0083] S8. Add user management function to the system and assign corresponding operation permissions to different users. Users can log in to the system with legitimate identity, perform data query operations, and call, forward, and download relevant data for application within their permissions.

[0084] like Figure 10 As shown, this embodiment also provides an operating device for a multi-machine, multi-process synchronous data management system based on block tasks. The device includes a processor, a memory, and a computer program stored in the memory and running on the processor.

[0085] The processor includes one or more processing cores. The processor is connected to the memory via a bus. The memory is used to store program instructions. When the processor executes the program instructions in the memory, it implements the above-mentioned multi-machine multi-process synchronous data management system and method based on block tasks.

[0086] Optionally, the memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0087] In addition, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described block-task-based multi-machine multi-process synchronous data management system and method.

[0088] Optionally, the present invention also provides a computer program product containing instructions that, when run on a computer, causes the computer to execute the steps of the above-described multi-machine, multi-process synchronous data management system and method based on block tasks.

[0089] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0090] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-machine, multi-process synchronous data management system based on block tasks, characterized in that: include The system comprises a technical support unit, a multi-process synchronization unit, a data synchronization unit, and a data management unit, which are sequentially connected via network communication. The technical support unit is used to load various intelligent technologies to support the connection and operation of the system. The multi-process synchronization unit is used to manage the synchronization of multiple processes running concurrently on multiple processors and on a single processor. The data synchronization unit is used to select appropriate software or methods to achieve data synchronization within the system. The data management unit is used to process and manage the synchronized data. The technical support unit includes a block connection module, a synchronization software module, a big data module, and an inter-process communication module. The inter-process communication module is used to provide and manage different inter-process communication pipes and allocate and schedule them for multiple concurrent processes within a single processor. Specifically, it uses memory as a medium to open a buffer in the kernel to realize inter-process data exchange communication. The inter-process communication module includes an anonymous pipe module, a named pipe module, and a polling scheduling module. The multi-process synchronization unit includes a hardware synchronization module, a semaphore module, a monitor mechanism module, and a Petri net module. The data synchronization unit includes a publish / subscribe module, an SQL-JOB module, a message queue module, and a select-to-invoke module. The publish / subscribe module enables rapid data backup and synchronization without writing any code, utilizing SQL Server's built-in publish / subscribe database backup mechanism. The SQL-JOB module achieves data synchronization through scheduled SQL Jobs, connecting the target and source servers and using SQL statements to read data from the source server and update it on the target server. The message queue module provides SQL Server with queues, message passing, and an asynchronous programming model via SQL Server Service Broker, offering message passing services that shorten interactive response time, increase overall application throughput, and ultimately achieve data synchronization. The select-to-invoke module selects the optimal data synchronization method based on the data's source and type. The data management unit includes a deduplication and cleaning module, a summarization and classification module, a secure storage module, and a scheduling and application module; This multi-machine, multi-process synchronous data management system based on block tasks first establishes a basic network architecture supported by blockchain and big data technologies. On this basis, it builds and loads data synchronization software, introduces inter-process communication functions and process synchronization solutions, realizes inter-process communication and process synchronization within a single machine, and achieves synchronization between multiple machines and multiple processes. Then, based on multi-process synchronization, relevant methods are selected to synchronize massive amounts of data. Next, the massive amounts of data are deduplicated, cleaned, classified, stored and managed for security. Finally, a user management mechanism is introduced to allow users to schedule and apply the data.

2. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 1, characterized in that: The block connection module, the synchronization software module, the big data module, and the inter-process communication module are sequentially connected via network communication. The block connection module is used to connect multiple processors in the system as nodes in a loop based on blockchain. The synchronization software module is used to load various data synchronization software into each processor to quickly realize data synchronization within the machine. The big data module is used to manage and distribute massive amounts of synchronized data through big data technology and big data processing technology.

3. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 2, characterized in that: The anonymous pipe module runs in parallel with the named pipe module, and the signal output terminals of the anonymous pipe module and the named pipe module are connected to the signal input terminal of the polling scheduling module. The anonymous pipe module is used to create anonymous pipes between related processes by calling the pipe function to provide one-way communication. The named pipe module is used to create system-visible named pipes between any two processes by calling the mknod or mkfifo command to achieve communication between the two processes. The polling scheduling module is used to schedule pipes in a round-robin manner to achieve load balancing of pipe scheduling.

4. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 3, characterized in that: The round-robin scheduling module includes round-robin scheduling and an improved weighted round-robin scheduling algorithm. The round-robin scheduling algorithm is suitable when all servers in the server group have the same hardware and software configuration and the average service requests are relatively balanced. The weighted round-robin scheduling algorithm is suitable when the configurations, installed business applications, and processing capabilities of the servers in the server group are different. Therefore, the weight allocation calculation expression in the weighted round-robin scheduling algorithm is: ; In the formula, Use load on the server.

5. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 4, characterized in that: The hardware synchronization module, semaphore module, monitor mechanism module, and Petri net module are sequentially connected via network communication and operate in parallel. The hardware synchronization module provides special hardware instructions from the computer to allow the detection and correction of the content of a word or the swapping of two words, thereby solving simple process synchronization problems. The semaphore module evolves from integer semaphores to record-type semaphores and then to semaphore sets, applying the semaphore mechanism to single-processor and multi-processor systems as well as computer networks to achieve process synchronization. The monitor mechanism module accepts or blocks multiple concurrent processes requesting access to shared resources based on resource availability, ensuring that only one process enters the monitor at a time to execute the monitor, use the shared resources, and selects unified management of all access to shared resources to effectively achieve process mutual exclusion. The Petri net module is used in parallel computing and distributed database design to construct system models and analyze dynamic characteristics, and is suitable for describing and analyzing problems in asynchronous concurrent systems.

6. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 5, characterized in that: The semaphore module includes an integer information module, a record-type information module, an AND-type information module, and an information set module. These modules are sequentially connected via network communication and operate in parallel. The integer information module, except for initialization, accesses data information through only two standard atomic operations, handling processes in a "busy waiting" state. The record-type information module uses a record-type data structure and adds a process linked list pointer to the semaphore mechanism to link all waiting processes accessing the same critical resource. The AND-type information module allocates all resources needed by a process during its entire operation to the process at once, releasing them all at once after the process has finished using them, thus achieving AND synchronization and preventing deadlock. The information set module completes the allocation or release of all resources requested by a process and the different resource requirements for each type of resource in a single atomic operation.

7. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 6, characterized in that: The publish / subscribe module, the SQL-JOB module, and the message queue module are connected in sequence via network communication and run in parallel. The signal output terminals of the publish / subscribe module, the SQL-JOB module, and the message queue module are connected to the signal input terminal of the select / call module.

8. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 7, characterized in that: The signal output terminal of the deduplication cleaning module is connected to the signal input terminal of the summarization and classification module, the signal output terminal of the summarization and classification module is connected to the signal input terminal of the secure storage module, and the signal output terminal of the secure storage module is connected to the signal input terminal of the scheduling application module. The deduplication cleaning module is used to compare and check massive amounts of data for deduplication and to remove duplicate, damaged, invalid, and expired data. The summarization and classification module is used to classify the synchronized and cleaned massive amounts of data using clustering techniques. The secure storage module is used to manage the data securely and distribute its storage according to the classified items. The scheduling application module is used to provide users with a way to schedule data for application through querying.

9. The multi-machine, multi-process synchronous data management system based on block tasks according to claim 8, characterized in that: The inductive classification module uses the K-Means clustering algorithm.

10. A method for managing synchronous data across multiple machines and processes based on block tasks, comprising the synchronous data management system for multiple machines and processes based on block tasks as described in claim 9, characterized in that: Includes the following steps: S1. Based on blockchain, build a system architecture that includes multiple processors, and install multiple data synchronization software in each processor within the system; S2, the processors distributed across the nodes of the block collect, calculate and process various types of data near the node. In the process of handling a large number of complex business processes, the operating system schedules and manages the execution of multiple processes. S3. When a single processor performs multiple processes concurrently, it allows multiple instruction streams to be executed concurrently in a program. Each instruction stream is called a process, and each process is independent of the others. S4. During the operation of a single processor, a communication channel is continuously established between multiple concurrent processes through a polling scheduling algorithm to achieve initial data connection between multiple processes on a single processor, and the optimal and suitable data synchronization software is selected to achieve initial data synchronization. S5. Among multiple concurrent processes on multiple processors, the optimal and applicable process synchronization method is selected to achieve process synchronization according to the cooperation conditions between processes and the corresponding constraint relationship. S6. Combine several concurrent processes after synchronization, select the optimal and applicable data synchronization software or method to synchronize and integrate the massive data involved in multiple machines and multiple processes. S7. Compare and check massive amounts of data for duplicates, clean and filter out invalid, expired, damaged and duplicate data, classify and summarize the data through clustering methods, distribute and store the classified data in a cloud database, and manage the data securely. S8. Add user management function to the system and assign corresponding operation permissions to different users. Users can log in to the system with legitimate identity, perform data query operations, and call, forward, and download relevant data for application within their permissions.

Citation Information

Patent Citations

  • Key data synchronizing method for multi-processor

    CN101562515A

  • Online data synchronization method and device based on block chain, terminal and storage medium

    CN110471981A