Multi-system redundant computer processing method and system and electronic equipment
Through multi-system redundant architecture and dynamic resource scheduling, the problems of unreasonable resource scheduling and poor stability in traditional computer processing methods are solved, and faster data processing speed and higher system reliability are achieved.
Patent Information
- Application Number
- CN202510278541.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional computer processing methods use single nodes to process data, the resource scheduling is unreasonable, the data processing speed is slow, and the redundant function is lacking, and the backup nodes or additional redundant space cannot be provided, resulting in poor long-term stability.
A multi-system redundancy architecture is adopted, including computer nodes, computer computing nodes, thermal redundancy one node, thermal redundancy two nodes and temperature redundancy nodes. Communication channels are established and encrypted through the TCP/IP protocol, dynamic cloning and select resource scheduling, real-time monitoring of load and starting redundancy switching mechanism, and using temperature redundancy nodes to share the load.
It achieves more reasonable resource allocation, improves data processing speed and system reliability, ensures fast switching and load balancing of nodes in the event of failure, and improves the long-term stability and security of the computer system.
Smart Images

Figure CN120386667A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and more specifically, particularly relates to a multi-system redundant computer processing method, system, and electronic device. Background Art
[0002] In today's digital age, computer systems have been widely penetrated into all fields of social life. From critical infrastructures such as power, transportation, and finance to daily office work, entertainment, etc., they all rely on the stable operation of computer systems. With the continuous expansion and deepening of computer application scenarios, the requirements for their reliability and stability are becoming increasingly stringent. However, traditional computer processing methods often use a single node to process data, with relatively unreasonable resource scheduling, slow data processing speed. Moreover, when a single node processes data, it is often in a high-load state, which is not conducive to long-term stable operation. Secondly, most traditional computer processing methods do not have a redundancy function and cannot provide a backup node for the node or provide additional redundant space for data processing. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention provides a multi-system redundant computer processing method, system, and electronic device to solve the technical problems in the prior art that traditional computer processing methods often use a single node to process data, with relatively unreasonable resource scheduling, slow data processing speed, and most do not have a redundancy function and cannot provide a backup node for the node or provide additional redundant space for data processing.
[0004] The purpose and efficacy of a multi-system redundant computer processing method, system, and electronic device of the present invention are achieved by the following specific technical means: A multi-system redundant computer processing method includes the following steps: S101: Obtain eight different nodes, where one node is the computer master node for controlling the entire data processing system, three are computer computing nodes for processing data, two are hot redundancy one nodes as backup computer computing nodes, one is a warm redundancy node for taking over local load data in the computer computing nodes. The warm redundancy node includes two simultaneously running load data processing systems, and one is a hot redundancy two node as a backup computer master node; S102: Establish a communication channel using the TCP / IP protocol between eight different nodes, and encrypt the communication channel using public-private key pairs. Based on the communication channel, establish an on-off channel between the computer general node and the hot redundant two nodes. Based on the communication channel, establish three on-off channels between three computer computing nodes and two hot redundant first nodes and the warm redundant node. Conduct data communication based on the communication channel between the eight nodes, and establish a backup database, which is used for real-time storage and backup of data. S103: Obtain a task feature vector, which represents the task features of the data processing that the computer computing node needs to execute. Obtain a resource feature vector, which represents the resource features of the computer computing node. Conduct a dynamic clone selection algorithm based on the task feature vector and the resource feature vector. Obtain an optimal resource scheduling plan based on the dynamic clone selection algorithm. The computer general node schedules resources for the three computer computing nodes based on the optimal resource scheduling plan, and generate computing node resource scheduling information based on the resource scheduling. S104: Conduct real-time detection on the three computer computing nodes and the computer general node based on heartbeat detection and instruction response delay. If any one of the computer general node or the three computer computing nodes fails, start the redundancy switching mechanism to ensure that the data processing within the computer computing node continues and the computer general node continues to control the data processing system, and generate fault information based on the redundancy switching mechanism. S105: The computer general node monitors the real-time resource usage of each computer computing node in real time. When the load of at least one computer computing node reaches 85%, start the redundant resource scheduling mechanism. Based on the redundant resource scheduling mechanism, transfer the data of the computer node with too high a load to the warm redundant node, and the warm redundant node takes over the local data exceeding the load. Generate warm redundant node resource scheduling information based on the redundant resource scheduling mechanism. S106: During operation, the backup database records the fault information, the computer computing node resource scheduling information, and the warm redundant node resource scheduling information in real time. The user can call the backup database to view the records of the fault information, the computer computing node resource scheduling information, and the warm redundant node resource scheduling information.
[0005] As a further solution of the present invention, the resource scheduling specifically includes: Obtain the task type, which represents the task type of the data processing that the computer computing node needs to execute. Obtain the amount of computation, which represents the complexity of the data processing that the computer computing node needs to execute. Obtain the latency requirement, which represents the time requirement of the data processing that the computer computing node needs to execute. Establish a task feature vector based on the task type, the amount of computation, and the latency requirement. Obtain the number of CPU cores of the computer computing node. The number of CPU cores reflects the data processing ability. Obtain the cache hit rate, which reflects the data access speed. Obtain the CPU energy consumption index, which reflects the energy efficiency of the computer computing node. Establish a resource feature vector based on the number of CPU cores, cache hit rate, and CPU energy consumption index; The computer total node monitors the real-time resource usage of each computer computing node in real time. When the load of one or more computer computing nodes reaches 70%, trigger the resource scheduling mechanism; Generate multiple candidate scheduling plans based on the dynamic cloning selection algorithm. The multiple candidate scheduling plans are represented as different combinations of resource scheduling. Calculate the affinity of each candidate scheduling plan through the affinity function, and obtain the optimal resource scheduling plan based on the affinity. The computer total node performs resource scheduling on three computer computing nodes based on the optimal resource scheduling plan; Establish a memory cell call and update mechanism. The memory cell call and update mechanism specifically includes: Establish a historical scheduling case library. Store the optimal resource scheduling plan, task feature vector, and resource feature vector of each scheduling into the historical scheduling case library. Establish an LSTM neural network based on the historical scheduling case library. Predict the affinity of the computer computing node currently performing the data processing task based on the LSTM neural network, and query in the historical scheduling case library based on the affinity; If the affinity conforms to an optimal resource scheduling plan in the historical scheduling case library, the computer total node directly calls the optimal resource scheduling plan to perform resource scheduling on three computer computing nodes; If the affinity does not conform to an optimal resource scheduling plan in the historical scheduling case library, trigger the resource scheduling mechanism, and continuously import the newly generated optimal resource scheduling plan into the historical scheduling case library.
[0006] As a further solution of the present invention, the task feature vector can be expressed as: ; Among them, is expressed as the task feature vector, is expressed as the task type, is expressed as the amount of computation, is expressed as the delay requirement; The resource feature vector can be expressed as: ; Among them, is expressed as the resource feature vector, is expressed as the number of CPU cores, is expressed as the cache hit rate, is expressed as the CPU energy consumption index; The affinity function can be expressed as: ; Among them, is expressed as an affinity function, , , are expressed as weight factors, is expressed as a load, is expressed as a cache matching degree, is expressed as a migration cost.
[0007] As a further solution of the present invention, step S104 specifically includes: Two hot redundant one nodes pre-load the mirror system of the computer computing node in advance, and two hot redundant two nodes pre-load the mirror system of the computer total node in advance; When performing heartbeat detection and instruction response delay detection, three computer computing nodes and the computing total node detect each other. When a node fails to respond to the heartbeat for more than 3 cycles and the instruction response delay exceeds 2 times the threshold, it is determined that the node is faulty, and the redundant switching mechanism is immediately started: If at least one of the three computer computing nodes fails, the on-off channels between the three computer computing nodes and the two hot redundant one nodes are immediately closed, and the hot redundant one node starts and takes over the data processing of the faulty computer computing node within 200 ms. The takeover mode of the hot redundant one node is one-to-one takeover; If the computer total node fails, the on-off channel between the computer total node and the hot redundant two node is immediately closed, and the hot redundant two node starts and takes over the control operation of the faulty computer total node within 100 ms; When the response heartbeat and instruction response delay of the faulty node are restored, based on the sandbox verification mode, a progressive recovery strategy is adopted to gradually reload to the original node. After the recovery is completed, the on-off channel between the recovery node and the redundant node is disconnected.
[0008] As a further solution of the present invention, step S105 specifically includes: The two load data processing systems included in the warm redundant node are the x86 system and the ARM system respectively. The x86 system is used to process I / O intensive data, and the ARM system is used to process compute intensive data; When the load of at least one computer computing node exceeds 80%, the warm redundant node starts preheating. The preheating time of the warm redundant system is 30s. When the load of at least one computer computing node exceeds 85%, the on-off channels between the warm redundant node and the three computer computing nodes are immediately closed. The preheated warm redundant node takes over the local data exceeding the load of the computer computing node. When the load of the computer computing node drops back to 70%, the computer computing node gradually takes over the data previously migrated to the warm redundant node, and rationally distributes the data to the three computer computing nodes based on the dynamic load balancing algorithm. After the data migration is completed, the warm redundant node cools down, and the on-off channels between the warm redundant node and the three computer computing nodes are disconnected.
[0009] As a further solution of the present invention, a communication channel using the TCP / IP protocol is established between eight different nodes, and the communication channel is encrypted using public and private key pairs, including: Encrypt the data communication in the communication channel based on TLS; Configure a blockchain network node for each of the eight nodes, and the nodes perform data transmission based on the P2P communication protocol. Configure a smart contract agent for each node, and the smart contract agent is used to verify all the data transmitted; During each data transmission, the sending node will perform a hash calculation on the data based on the hash function to generate a hash value of the data, and sign it with the private key. The receiving node verifies the legality of the signature through the public key. All changes, transmission processes, and reception confirmations of the data are recorded based on the blockchain. After the three computer computing nodes receive the data, the received data is verified based on the voting mechanism of the PBFT consensus algorithm. When at least two computer computing nodes confirm that the data is valid, the data will be accepted, and the blockchain state is updated synchronously; When data conflicts occur, that is, when two computer nodes modify the same data, the blockchain record and hash value are used to help determine the correct data version, so as to resolve the conflict; When a node fails or data is lost, the data is repaired based on the PBFT consensus algorithm and the backup database.
[0010] As a further solution of the present invention, during each data transmission, the sending node will perform a hash calculation on the data, and at the same time append a timestamp, node fingerprint, and signature information to generate a data packet, which can be expressed as: ; Among them, represents the data packet, represents the data, represents the timestamp, represents the node fingerprint, Represented as signature information, the data packet remains encrypted during transmission to prevent tampering; After receiving the data packet, the receiving node verifies the integrity of the data packet through the intelligent contract proxy verification. If the verification passes, the receiving node updates the record in the blockchain. If the verification fails, the data is discarded; Regularly check the blockchain status of all nodes, and ensure the consistency of data of each node based on the distributed ledger mechanism of the blockchain.
[0011] As a further solution of the present invention, establishing a communication channel specifically includes: The computer general node is connected to three computer computing nodes based on a one-way communication channel. The three computer computing nodes are respectively connected to each other based on a two-way communication channel. The computer general node is connected to the hot redundant two-node based on a two-way communication channel. The three computer computing nodes are respectively connected to two hot redundant one-nodes and a warm redundant node based on a two-way communication channel. The three computer computing nodes are connected to the backup database based on a two-way communication channel. The two hot redundant one-nodes and the warm redundant node are connected to the computer general node based on a communication channel. Between the computer general node and the hot redundant two-node, the on-off of the communication channel is controlled by an on-off channel. Between the three computer computing nodes and the two hot redundant one-nodes and the warm redundant node, the on-off of the communication channel is controlled by three on-off channels. The three on-off channels between the three computer computing nodes and the two hot redundant one-nodes and the warm redundant node are independently controlled.
[0012] A multi-system redundant computer processing system includes: A heterogeneous node cluster, adopting a 3-2-1-1-1 heterogeneous redundancy architecture, is used for processing and computing data. This architecture includes three computer computing nodes, two hot redundant one-nodes, one computer general node, one warm redundant node and one hot redundant two-node; A monitoring module, which is used to continuously monitor the load of the computer computing nodes in real time and obtain the load value. At the same time, it is also used to obtain the task feature vector, which represents the task characteristics of the data processing that the computer computing node needs to execute, and obtain the resource feature vector, which represents the resource characteristics of the computer computing node; A communication module, which is used to establish communication channels between different nodes in the heterogeneous node cluster and establish a communication channel between the backup database and the heterogeneous node cluster; An on-off module, which is used to establish on-off channels between different nodes in the heterogeneous node cluster and independently control the on-off channels; A scheduling module, which is used to perform resource scheduling on the three computer computing nodes; A fault module, which is used to detect the three computer computing nodes and the computer general node in real time, and start the redundancy switching mechanism based on heartbeat detection and instruction response delay; A redundancy module, used to start the redundancy resource scheduling mechanism; A recording module, used to generate fault information, computer computing node resource scheduling information, and warm redundancy node resource scheduling information, and record them in real time. At the same time, import the above information into the backup database; A backup module, used to establish a backup database. The backup database is used to store and back up data in real time, and is also used to store fault information, computer computing node resource scheduling information, and warm redundancy node resource scheduling information.
[0013] An electronic device, comprising: At least one processor; and a memory communicatively connected to at least one processor; wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can execute the method proposed in Embodiment 1 of the present invention.
[0014] Compared with the prior art, the present invention has the following beneficial effects: First, obtain eight different nodes, which respectively include three groups of computer computing nodes, two groups of hot redundancy nodes, one group of warm redundancy nodes, one group of hot redundancy nodes, and one group of computer summary nodes. Establish a communication channel using the TCP / IP protocol between the eight different nodes, and encrypt the communication channel. Establish an on-off channel between the eight different nodes, so as to form a heterogeneous node cluster adopting a 3-2-1-1-1 heterogeneous redundancy architecture, and connect the backup database to perform real-time backup of data. Subsequently, obtain the task feature vector and the resource feature vector, perform a dynamic clone selection algorithm based on the task feature vector and the resource feature vector, and obtain the optimal resource scheduling scheme based on the dynamic clone selection algorithm. Thus, schedule the resources of the three groups of computer computing nodes based on the optimal resource scheduling scheme, making the resource allocation more reasonable, improving the availability of resources, and enhancing the data processing speed. When the nodes are running, real-time detection is performed on the three computer computing nodes and the computer summary node. If any one of the nodes fails, the redundancy switching mechanism is started. Through the redundancy switching mechanism, the three computer computing nodes and the computer summary node can be taken over by the hot redundancy nodes and the hot redundancy nodes, so as to ensure the continuous operation of the nodes and ensure the reliability and security during operation. At the same time, real-time monitor the real-time resource usage of each computer computing node. When the load of at least one computer computing node reaches 85%, start the redundancy resource scheduling mechanism. The warm redundancy node takes over the local data exceeding the load. The warm redundancy node and the redundancy resource scheduling mechanism can call redundant resources to share the resource load of the computer computing node, improving the long-term operation stability of the computer computing node. And two independent systems are built in the warm redundancy system, which can improve the data processing ability. Brief Description of the Drawings
[0015] Figure 1 is a flowchart of the steps of a multi-system redundant computer processing method of the present invention; Figure 2 is a schematic diagram of a heterogeneous node cluster in a multi-system redundant computer processing system of the present invention. Detailed Embodiments
[0016] The following further describes in detail the embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the technical solutions of the present invention, but cannot be used to limit the protection scope of the present invention.
[0017] Embodiment 1: As shown in the attached Figure 1 to the attached Figure 2 figures: The present invention provides a multi-system redundant computer processing method, including the following steps: S101: Obtain eight different nodes, where one node is the computer total node for controlling the entire data processing system, three are computer computing nodes for processing data, two are hot redundant one nodes as standby computer computing nodes, one is a warm redundant node for taking over local load data in the computer computing nodes, the warm redundant node includes two simultaneously running load data processing systems, and one is a hot redundant two node as a standby computer total node.
[0018] It can be understood that the hot redundant one node and the hot redundant two node run in real-time synchronization as standby nodes. When a computer computing node or a computer total node fails, the hot redundant one node can quickly take over the data processing task of the computer computing node, and the hot redundant two node can quickly take over the control task of the computer total node to ensure the continuous and stable operation of the entire system. The warm redundant node requires a short time to start when starting, and is suitable for sharing the data processing task of the computer computing node. In practical applications, such as trading platforms and payment systems in the financial system that need to be 24 / 7 without interruption, or automated equipment in nuclear power plants and chemical plants in industrial control that need to prevent downtime from causing accidents, this method and system can be applied to the above scenarios to avoid failure downtime and improve the reliability and security of the system.
[0019] S102: Establish a communication channel using the TCP / IP protocol between eight different nodes, and encrypt the communication channel using public-private key pairs. Based on the communication channel, establish an on-off channel between the computer general node and the hot redundant two-node. Based on the communication channel, establish three on-off channels between three computer computing nodes and two hot redundant first nodes and a warm redundant node. Conduct data communication based on the communication channel between the eight nodes, and establish a backup database, which is used for real-time storage and backup of data.
[0020] As shown in the Figure 2 accompanying figure, in the figure, "Summary node" represents the computer general node, "Hot node 2" represents the hot redundant two-node, "Computing node" represents the computer computing node, "Hot node 1" represents the hot redundant first node. "Warm node" represents the warm redundant node, and "Backup database" represents the backup database. When establishing the channels, the computer general node is connected to three computer computing nodes based on a unidirectional communication channel, and the three computer computing nodes are respectively connected to each other based on a bidirectional communication channel. The computer general node is connected to the hot redundant two-node based on a bidirectional communication channel. The three computer computing nodes are respectively connected to two hot redundant first nodes and a warm redundant node based on a bidirectional communication channel. The three computer computing nodes are connected to the backup database based on a bidirectional communication channel. The two hot redundant first nodes and the warm redundant node are connected to the computer general node based on the communication channel. The on-off of the communication channel between the computer general node and the hot redundant two-node is controlled by an on-off channel. The on-off of the communication channel between the three computer computing nodes and the two hot redundant first nodes and the warm redundant node is controlled by three on-off channels. The three on-off channels between the three computer computing nodes and the two hot redundant first nodes and the warm redundant node are independently controlled.
[0021] Furthermore, establish a heterogeneous node cluster with a 3-2-1-1-1 heterogeneous redundancy architecture through the above steps. The two hot redundant first nodes can simultaneously take over the failures of any two of the three computer computing nodes, achieving N+2 fault tolerance, while the traditional architecture method is mostly N+1. When the load of the three computer computing nodes exceeds the limit, the warm redundant node can take over the local data processing tasks to avoid the hot redundant nodes being overly occupied by non-fatal failures. The computer general node is set with independent redundancy to prevent the control layer from becoming a single point of failure source.
[0022] Furthermore, in practical applications, users can combine the heterogeneous node clusters with a 3-2-1-1-1 heterogeneous redundancy architecture in series or parallel.
[0023] If a parallel heterogeneous node cluster is adopted, multiple nodes within the parallel heterogeneous node cluster can work in parallel. If a certain heterogeneous node cluster fails, other heterogeneous node clusters can still work normally without affecting the operation of the overall system. Moreover, in the parallel heterogeneous node cluster, the data processing capacity of the system can be expanded by adding heterogeneous node clusters, enabling the system to handle higher loads and larger-scale tasks. In the parallel heterogeneous node cluster, each heterogeneous node cluster can handle independent data processing tasks, enhancing the parallel processing capacity and processing efficiency of the system.
[0024] Specifically, the parallel heterogeneous node cluster is suitable for large-scale parallel processing tasks (such as big data analysis, AI model training, etc.), tasks with fault tolerance requirements (such as financial trading systems, cloud computing platforms, etc.), and tasks of task load distribution and elastic expansion (such as distributed databases, distributed computing systems, etc.).
[0025] In a series heterogeneous node cluster, the heterogeneous node clusters are connected in sequence, with each heterogeneous node cluster serving as the input for the next one. The series heterogeneous node cluster is suitable for some scenarios with a clear task chain. Each heterogeneous node cluster performs a specific processing task, and the subsequent heterogeneous node clusters rely on the processing results of the previous node. In the series heterogeneous node cluster, the tasks of each heterogeneous node cluster are clear, there are clear inputs and outputs between the heterogeneous node clusters, and the data flow is unidirectional, making it easy to control and debug. It is suitable for scenarios that require sequential processing of tasks, such as data preprocessing, calculation, postprocessing, etc. Moreover, data synchronization and coordination in the series heterogeneous node cluster are relatively simple because the data only continues to flow to the next heterogeneous node cluster after being processed by the next one, avoiding data consistency issues between multiple heterogeneous node clusters.
[0026] Specifically, the parallel heterogeneous node cluster is suitable for task chain processing (image processing, pipeline operations, etc.), complex data conversion or format conversion tasks (data cleaning, feature engineering, etc.).
[0027] In some complex systems, in addition to the parallel and series heterogeneous node clusters, users can adopt a combined method of parallel and series to use the heterogeneous node clusters. For example, multiple heterogeneous node clusters are configured in parallel to form a parallel cluster, and the heterogeneous node clusters within each cluster are combined in series to handle specific tasks. In the combined method of parallel and series, the parallel part can be used for parallel working data processing tasks, while the series part can be used to handle some complex tasks or data processing links, such as data preprocessing, feature extraction, postprocessing, etc., ensuring the consistency and orderliness of the data flow.
[0028] Specifically, when establishing a communication channel using the TCP / IP protocol between eight different nodes and encrypting the communication channel using public-private key pairs, it includes: Encrypting the data communication within the communication channel based on TLS; It can be understood that a secure communication channel is established between all eight nodes, data exchange is carried out based on the transport layer protocol (TCP / IP), and each node establishes an encrypted end-to-end communication channel through public-private key pairs to ensure the security and integrity of the communication content. During the transmission process, TLS (Transport Layer Security) is used to encrypt the communication to prevent man-in-the-middle attacks or data tampering.
[0029] Configure a blockchain network node for each of the eight nodes, and data is transmitted between the nodes based on the P2P communication protocol. Configure a smart contract agent for each node, and the smart contract agent is used to verify all the data transmitted; It can be understood that each node is configured with a blockchain network node to form a fully distributed consensus system. The nodes exchange data and verify transactions through the P2P (peer-to-peer) communication protocol to ensure that the system has high availability and fault tolerance. And a smart contract agent is configured for each node, which can be used to specify the rules and verification process of data exchange. The smart contract agent ensures that each data exchange between nodes complies with the predetermined protocol to ensure the compliance of the data. All communication metadata (such as data packets, timestamps, node fingerprints, etc.) need to be verified through the smart contract agent.
[0030] During each data transmission, the sending node will perform a hash calculation on the data based on the hash function to generate a hash value of the data, and sign it with the private key. The receiving node verifies the legality of the signature through the public key. All changes, transmission processes, and receipt confirmations of the data are recorded based on the blockchain. After three computer computing nodes receive the data, the received data is verified based on the voting mechanism of the PBFT consensus algorithm. When at least two computer computing nodes confirm that the data is valid, the data will be accepted and the blockchain state will be updated synchronously.
[0031] It can be understood that to ensure the communication consistency between multiple nodes, the PBFT consensus algorithm is used to ensure the consistency of data transmission. PBFT reaches a consensus to verify the validity of data packets through the voting mechanism between nodes during each data transmission.
[0032] The specific process is that each node will perform verification after receiving the data and execute the verification based on the current state. If most nodes (at least 2 / 3 nodes) confirm that the data is valid, the data will be accepted and the blockchain state will be updated synchronously. During the process of transmitting data, the PBFT consensus algorithm can prevent data loss or tampering.
[0033] When a node failure occurs or data is lost, the data is repaired based on the PBFT consensus algorithm and the backup database.
[0034] When data conflicts occur, that is, when two computer nodes modify the same data, the blockchain-based records and hash values help determine the correct data version for conflict resolution.
[0035] Specifically, each time data is transmitted, the sending node calculates the hash of the data and attaches the timestamp, node fingerprint, and signature information at the same time to generate a data packet, which can be expressed as: ; where represents the data packet, represents the data, represents the timestamp, represents the node fingerprint, represents the signature information. The data packet remains encrypted during transmission to prevent tampering; After receiving the data packet, the receiving node verifies the integrity of the data packet through the intelligent contract proxy verification. If the verification passes, the receiving node updates the records in the blockchain. If the verification fails, the data is discarded; Regularly check the blockchain status of all nodes, and ensure the consistency of data for each node based on the blockchain's distributed ledger mechanism.
[0036] S103: Obtain the task feature vector, which represents the task characteristics of the data processing that the computer computing node needs to execute, obtain the resource feature vector, which represents the resource characteristics of the computer computing node, perform a dynamic cloning selection algorithm based on the task feature vector and the resource feature vector, obtain the optimal resource scheduling scheme based on the dynamic cloning selection algorithm, and the computer general node performs resource scheduling on the three computer computing nodes based on the optimal resource scheduling scheme, and generates computing node resource scheduling information based on the resource scheduling.
[0037] Specifically, the resource scheduling specifically includes: Obtain the task type, which represents the task type of the data processing that the computer computing node needs to execute, obtain the computing amount, which represents the complexity of the data processing that the computer computing node needs to execute, obtain the latency requirement, which represents the time requirement of the data processing that the computer computing node needs to execute, and establish a task feature vector based on the task type, computing amount, and latency requirement.
[0038] The task feature vector can be expressed as: ; where Represented as a task feature vector, Represented as a task type, Represented as the computational workload, Represented as the latency requirement.
[0039] Specifically, when obtaining the task type, determine whether the task is a compute-intensive, I / O-intensive, or hybrid task, which helps with resource allocation and scheduling. When obtaining the computational workload, based on the complexity of the task, quantify the computing requirements. For example, the number of floating-point operations or the amount of data processing of the task, etc. When obtaining the latency requirement, obtain it based on the real-time requirement of the task. For example, low-latency tasks need to be scheduled first to avoid long queues.
[0040] Obtain the number of CPU cores of the computer computing node. The number of CPU cores reflects the data processing ability. Obtain the cache hit rate, which reflects the data access speed. Obtain the CPU energy consumption index, which reflects the energy efficiency of the computer computing node. Establish a resource feature vector based on the number of CPU cores, cache hit rate, and CPU energy consumption index; The resource feature vector can be represented as: ; Among them, Represented as the resource feature vector, Represented as the number of CPU cores, Represented as the cache hit rate, Represented as the CPU energy consumption index.
[0041] The computer total node monitors the real-time resource usage of each computer computing node in real time. When the load of one or more computer computing nodes reaches 70%, trigger the resource scheduling mechanism.
[0042] Furthermore, in addition to the load of the computer computing node, users can also monitor parameters such as the temperature of the node, memory usage, cache hit rate, etc. to comprehensively evaluate the health status of the node.
[0043] The resource scheduling mechanism specifically includes generating multiple candidate scheduling schemes based on the dynamic cloning selection algorithm. The multiple candidate scheduling schemes are represented as different combinations of resource scheduling. The candidate schemes will consider the type of task, computational workload, latency requirement, and the load situation of the node, and generate different combinations of resource scheduling. Calculate the affinity of each candidate scheduling scheme through the affinity function. The affinity function combines the matching degree of task requirements and node resources, and preferentially selects the scheme with the lightest load and the shortest response time.
[0044] Specifically, the affinity function can be represented as: ; Among them, Expressed as an affinity function, , , expressed as a weight factor, and the weight factor is dynamically adjusted according to different task requirements and node conditions. Expressed as a load, the lower the load, the higher the affinity. Expressed as a cache matching degree, which is the matching degree between the data access pattern of the task and the node cache. The higher the cache hit rate, the higher the affinity. Expressed as a migration cost, and the migration cost represents the resource consumption caused by migrating the task to other nodes. The lower the migration cost, the higher the affinity.
[0045] Based on the affinity, obtain the optimal resource scheduling plan. The set of candidate scheduling plans with the highest affinity is the optimal resource scheduling plan. The computer's total node performs resource scheduling on the three computer computing nodes based on the optimal resource scheduling plan.
[0046] Furthermore, establish a memory cell call and update mechanism. The memory cell call and update mechanism specifically includes: Establish a historical scheduling case library, and store the optimal resource scheduling plan, task feature vector, and resource feature vector of each scheduling into the historical scheduling case library. In addition, when the user is actually using it, the user can also store the type of the task, the scheduled node, the resource usage situation during task execution, the execution time, the task success rate, etc. into the historical scheduling case library. The scheduling case library provides empirical data for future scheduling decisions and can be used to discover potential scheduling patterns.
[0047] Based on the historical scheduling case library, establish an LSTM neural network. The LSTM neural network is suitable for processing time series data and can capture the changing rules of task loads. The LSTM neural network learns the changing trends and execution patterns of task loads according to the relevant data in the historical scheduling case library, so as to provide prediction information for resource scheduling tasks. Based on the LSTM neural network, predict the affinity of the computer computing nodes currently performing data processing tasks, and query in the historical scheduling case library based on the affinity.
[0048] If the affinity conforms to an optimal resource scheduling plan in the historical scheduling case library, the computer's total node directly calls the optimal resource scheduling plan to perform resource scheduling on the three computer computing nodes.
[0049] If the affinity does not conform to an optimal resource scheduling plan in the historical scheduling case library, trigger the resource scheduling mechanism and continuously import the newly generated optimal resource scheduling plan into the historical scheduling case library.
[0050] S104: Based on heartbeat detection and instruction response delay, real-time detection is performed on three computer computing nodes and the total computer nodes. If any one of the total computer nodes or the three computer computing nodes fails, the redundant switching mechanism is activated to ensure the continuous data processing within the computer computing nodes and the continuous control of the data processing system by the total computer nodes, and fault information is generated based on the redundant switching mechanism.
[0051] Specifically, two hot-standby nodes pre-load the mirror system of the computer computing nodes, and the second hot-standby node pre-loads the mirror system of the total computer nodes. By pre-loading the mirror systems of the computer computing nodes and the total computer nodes, the data and state consistency during switching are ensured. And during switching, the first hot-standby node immediately takes over the data processing tasks of the computer computing nodes, and the second hot-standby node immediately takes over the control tasks of the total computer nodes, keeping the computing power of the system and the control tasks of the total computer nodes unaffected.
[0052] When performing heartbeat detection and instruction response delay detection, the three computer computing nodes and the total computing node detect each other. When a node fails to respond to the heartbeat for more than 3 cycles and the instruction response delay exceeds 2 times the threshold, it is determined that the node has failed, and the redundant switching mechanism is immediately activated. And when the system starts, the first hot-standby node synchronizes the data and task status from the total computer node to establish a mirror system, which stores task information, resource status, and data processing history records. The second hot-standby node synchronizes the status of the control task and subsequent control tasks from the total computer node to ensure seamless takeover.
[0053] The redundant switching mechanism is specifically that if at least one of the three computer computing nodes fails, the on-off channels between the three computer computing nodes and the two first hot-standby nodes are immediately closed, and the first hot-standby node starts and takes over the data processing of the failed computer computing node within 200 ms. The takeover mode of the first hot-standby node is one-to-one takeover; If the total computer node fails, the on-off channel between the total computer node and the second hot-standby node is immediately closed, and the second hot-standby node starts and takes over the control operation of the failed total computer node within 100 ms; When the second hot-standby node takes over the control task of the total computer node, the permission migration is completed by automatically generating a digital certificate. The digital certificate is used to confirm that the second hot-standby node, as the new total node, has the control right over the entire system. During the process of generating the digital certificate, the system authenticates the control permission of the original total node and safely transfers the control right to the second hot-standby node.
[0054] It is understandable that the digital certificate verifies its validity through the blockchain network to ensure the correct migration of the total node control authority. This process ensures data consistency and system integrity, and avoids operation anomalies caused by permission switching.
[0055] Furthermore, when the response heartbeat and instruction response delay of the faulty node are restored, based on the sandbox verification mode, a progressive recovery strategy is adopted to gradually reload to the original node. After the recovery is completed, the on-off channel between the recovered node and the redundant node is disconnected.
[0056] Specifically, when the faulty node is repaired, the system does not immediately reconnect it to the working environment, but guides it to the sandbox verification mode. In the sandbox mode, the faulty node will synchronize with other healthy nodes to verify whether its computing and data states are consistent. At this time, the node will not participate in actual computing and only serves as a part of the verification environment. In the sandbox mode, the system will ensure that the repair results and data synchronization of the faulty node meet the requirements through the status logs recorded by the blockchain and consistency checks. The system will use means such as data verification, version control, and task consistency verification to ensure that the repaired node can return to the normal state. In the sandbox mode, the system checks each item of the node to ensure that the data, computing power, and load status of this node are consistent with other nodes.
[0057] Specifically, the consistency verification uses the Paxos algorithm for state synchronization and verification to ensure that the results of the faulty node recovery will not affect the overall stability of the system. Once the node passes the consistency verification, the system will gradually introduce it into the production environment. First, it will be restored to a lower load state, and then the computing and resource consumption will be gradually increased to ensure that the system will not have too much impact on other nodes during the recovery process. The node will continue to receive task scheduling during the gradual reloading process and gradually return to the fully available state during the computing. After the recovery is completed, the on-off channel between the recovered node and the redundant node is disconnected to prevent the system from misoperating and causing the redundant node to take over the other nodes when there is no fault. Through this redundant switching mechanism, the system can quickly respond when a node fails, minimize the interruption time of data processing tasks, and at the same time ensure the consistency and security of data and computing resources.
[0058] S105: The computer total node monitors the real-time resource usage of each computer computing node in real time. When the load of at least one computer computing node reaches 85%, start the redundant resource scheduling mechanism. Through the redundant resource scheduling mechanism, share the data processing tasks for the computer computing node to ensure that the load of the high-load node will not be too high. Based on the redundant resource scheduling mechanism, transfer the data of the computer node with too high a load to the warm redundant node, and the warm redundant node takes over the local data exceeding the load. Generate the warm redundant node resource scheduling information based on the redundant resource scheduling mechanism.
[0059] Specifically, the two load data processing systems included in the warm redundant node are the x86 system and the ARM system respectively. The dual systems will coordinate according to the load conditions and are responsible for the data processing tasks of different computer computing nodes respectively. The x86 system is used to process I / O-intensive data, and the ARM system is used to process compute-intensive data. The x86 system and the ARM system of the warm redundant node determine which data processing tasks to take over respectively.
[0060] When the load of at least one computer computing node exceeds 80%, the warm redundant node starts to preheat. The preheating time of the warm redundant system is 30s. When the load of at least one computer computing node exceeds 85%, the on-off channels between the warm redundant node and the three computer computing nodes are immediately closed. The preheated warm redundant node takes over the local data exceeding the load of the computer computing node. By analyzing the task load on each computer computing node in real time, it is judged which data processing tasks are compute-intensive and which are I / O-intensive, and then it is determined which local data exceeding the load is migrated to the warm redundant node. When the local data is migrated to the warm redundant node, the x86 system and the ARM system in the warm redundant node will take over different local data on the computer computing node according to the characteristics of the data processing tasks. When the load of the computer computing node drops back to 70%, the computer computing node gradually takes over the data previously migrated to the warm redundant node, and reasonably distributes the data to the three computer computing nodes based on the dynamic load balancing algorithm to ensure that the computer computing node will not be overloaded again when taking over the data previously migrated to the warm redundant node, and can avoid over-concentration of tasks and ensure the efficient utilization of system resources. After the data migration is completed, the warm redundant node cools down, and the on-off channels between the warm redundant node and the three computer computing nodes are disconnected to prevent the warm redundant node from taking over the other nodes without overloading due to system misoperation.
[0061] S106: During operation, the backup database records the fault information, computer computing node resource scheduling information, and warm redundant node resource scheduling information in real time. Users can call the backup database to view the records of the fault information, computer computing node resource scheduling information, and warm redundant node resource scheduling information.
[0062] Specifically, the system will monitor the operating status of computer computing nodes, warm redundant nodes, hot redundant node 1, and hot redundant node 2 in real time, such as indicators like node load, memory usage rate, response latency, etc. When a certain node fails, such as a heartbeat loss or response timeout, the system will immediately detect and determine the type of failure. Once the failure is determined, the system will automatically record the failure information in the backup database. The recorded content includes the timestamp of the failure occurrence, the identifier of the failed node, the type of failure, the failure repair time, etc. The resource scheduling information of computer computing nodes and warm redundant nodes uses a similar method to record the above failure information. Based on the recorded failure information, computer computing node resource scheduling information, and warm redundant node resource scheduling information in the backup database, users can conveniently query and view these records to obtain the operating status of the system and the task execution situation, providing strong data support for subsequent optimization, problem-solving, and performance improvement.
[0063] A multi-system redundant computer processing system includes: A heterogeneous node cluster, adopting a 3-2-1-1-1 heterogeneous redundant architecture, is used for processing and computing data. This architecture includes three computer computing nodes, two hot redundant nodes 1, one computer general node, one warm redundant node, and one hot redundant node 2; A monitoring module, which is used to continuously monitor the load of computer computing nodes in real time and obtain the load value. At the same time, it is also used to obtain the task feature vector, which represents the task characteristics of the data processing that the computer computing node needs to execute, and obtain the resource feature vector, which represents the resource characteristics of the computer computing node; A communication module, which is used to establish communication channels between different nodes in the heterogeneous node cluster and establish a communication channel between the backup database and the heterogeneous node cluster; A connection and disconnection module, which is used to establish connection and disconnection channels between different nodes in the heterogeneous node cluster and independently control the connection and disconnection channels; A scheduling module, which is used to perform resource scheduling on the three computer computing nodes; A failure module, which is used to detect the three computer computing nodes and the computer general node in real time and start the redundant switching mechanism based on heartbeat detection and instruction response delay; A redundancy module, which is used to start the redundant resource scheduling mechanism; A recording module, which is used to generate failure information, computer computing node resource scheduling information, and warm redundant node resource scheduling information, and record them in real time. At the same time, the above information is imported into the backup database; A backup module, which is used to establish a backup database. The backup database is used to perform real-time storage backup of data. At the same time, it is also used to store failure information, computer computing node resource scheduling information, and warm redundant node resource scheduling information.
[0064] Specific usage mode and function of the first embodiment: First, obtain eight different nodes, which respectively include three groups of computer computing nodes, two groups of hot redundancy one nodes, one group of warm redundancy nodes, one group of hot redundancy two nodes, and one group of computer summary nodes. Establish a communication channel using the TCP / IP protocol between the eight different nodes, and encrypt the communication channel. Establish an on-off channel between the eight different nodes, thereby forming a heterogeneous node cluster with a 3-2-1-1-1 heterogeneous redundancy architecture, and connect to the backup database to perform real-time backup of data. Subsequently, obtain the task feature vector and the resource feature vector, perform a dynamic clone selection algorithm based on the task feature vector and the resource feature vector, and obtain the optimal resource scheduling scheme based on the dynamic clone selection algorithm. Then, schedule the resources of the three groups of computer computing nodes based on the optimal resource scheduling scheme, making the resource allocation more reasonable, improving the availability of resources, and enhancing the data processing speed. When the nodes are running, real-time detection is performed on the three computer computing nodes and the computer summary node. If any one of the nodes fails, the redundancy switching mechanism is activated. Through the redundancy switching mechanism, the three computer computing nodes and the computer summary node can be taken over by the hot redundancy one node and the hot redundancy two nodes, thereby ensuring the continuous operation of the nodes and ensuring the reliability and security during operation. At the same time, the real-time resource usage situation of each computer computing node is monitored in real time. When the load of at least one computer computing node reaches 85%, the redundant resource scheduling mechanism is activated, and the warm redundancy node takes over the local data exceeding the load. The warm redundancy node and the redundant resource scheduling mechanism can call redundant resources to share the resource load of the computer computing nodes, improving the long-term operation stability of the computer computing nodes. Moreover, two independent systems are built into the warm redundancy system, which can improve the data processing ability.
[0065] An electronic device, comprising: At least one processor; and a memory communicatively connected to at least one processor; wherein the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable at least one processor to execute the method proposed in the first embodiment of the present invention.
[0066] The following is a specific introduction to the various components of the electronic device: Among them, the processor is the control center of the electronic device, which can be a single processor or a collective term for multiple processing elements. For example, the processor can be one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement Embodiment 1 of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0067] Among them, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0068] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0069] The memory can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device. The embodiments of the present invention do not make specific limitations on this.
[0070] The above-described embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above-described embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by means of a wired (such as infrared, wireless, microwave, etc.) connection. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0071] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can mean that A exists alone, A and B exist simultaneously, or B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally indicates an "or" relationship between the associated objects before and after, but it may also indicate an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0072] It should be understood that in the embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0073] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A multi-system redundant computer processing method, characterized in that It includes the following steps: S101: Obtain eight different nodes, where one node is the computer general node for controlling the entire data processing system, three are computer computing nodes for processing data, two are hot redundancy one nodes as standby computer computing nodes, one is the warm redundancy node for taking over local load data in the computer computing nodes. The warm redundancy node includes two simultaneously running load data processing systems, and one is the hot redundancy two node as the standby computer general node; S102: Establish a communication channel using the TCP / IP protocol among the eight different nodes, and encrypt the communication channel using public-private key pairs. Based on the communication channel, establish an on-off channel between the computer general node and the hot redundancy two node, and establish three on-off channels between the three computer computing nodes and the two hot redundancy first nodes and the warm redundancy node. Conduct data communication based on the communication channel among the eight nodes, and establish a backup database for real-time data storage and backup; S103: Obtain a task feature vector representing the task features of the data processing that the computer computing nodes need to execute, and obtain a resource feature vector representing the resource features of the computer computing nodes. Conduct a dynamic cloning selection algorithm based on the task feature vector and the resource feature vector, and obtain an optimal resource scheduling scheme based on the dynamic cloning selection algorithm. The computer general node schedules resources for the three computer computing nodes based on the optimal resource scheduling scheme, and generates computing node resource scheduling information based on the resource scheduling; S104: Conduct real-time detection on the three computer computing nodes and the computer general node based on heartbeat detection and instruction response delay. If any one of the computer general node or the three computer computing nodes fails, start the redundancy switching mechanism to ensure the continuous data processing within the computer computing nodes and the continuous control of the data processing system by the computer general node, and generate fault information based on the redundancy switching mechanism; S105: The computer general node monitors the real-time resource usage of each computer computing node in real time. When the load of at least one computer computing node reaches 85%, start the redundant resource scheduling mechanism. Based on the redundant resource scheduling mechanism, transfer the data of the computer node with too high load to the warm redundancy node, and the warm redundancy node takes over the local data exceeding the load. Generate warm redundancy node resource scheduling information based on the redundant resource scheduling mechanism; S106: During operation, the backup database records the fault information, computer computing node resource scheduling information, and warm redundancy node resource scheduling information in real time. Users can call the backup database to view the records of the fault information, computer computing node resource scheduling information, and warm redundancy node resource scheduling information.
2. A multi-system redundant computer processing method according to claim 1, characterized in that, The resource scheduling specifically includes: Obtain the task type, which represents the type of data processing task that the computer computing node needs to execute. Obtain the computing volume, which represents the complexity of the data processing that the computer computing node needs to execute. Obtain the latency requirement, which represents the time requirement for the data processing that the computer computing node needs to execute. Establish a task feature vector based on the task type, computing volume, and latency requirement; Obtain the number of CPU cores of the computer computing node, which reflects the data processing ability. Obtain the cache hit rate, which reflects the access speed of the data. Obtain the CPU energy consumption index, which reflects the energy efficiency of the computer computing node. Establish a resource feature vector based on the number of CPU cores, cache hit rate, and CPU energy consumption index; The computer master node monitors the real-time resource usage of each computer computing node in real time. When the load of one or more computer computing nodes reaches 70%, trigger the resource scheduling mechanism; Generate multiple candidate scheduling schemes based on the dynamic cloning selection algorithm. The multiple candidate scheduling schemes represent different combinations of resource scheduling. Calculate the affinity of each candidate scheduling scheme through the affinity function. Obtain the optimal resource scheduling scheme based on the affinity. The computer master node performs resource scheduling on the three computer computing nodes based on the optimal resource scheduling scheme; Establish a memory cell call and update mechanism, which specifically includes: Establish a historical scheduling case library. Store the optimal resource scheduling scheme, task feature vector, and resource feature vector of each scheduling in the historical scheduling case library. Establish an LSTM neural network based on the historical scheduling case library. Predict the affinity of the computer computing node currently executing the data processing task based on the LSTM neural network. Query in the historical scheduling case library based on the affinity; If the affinity conforms to an optimal resource scheduling scheme in the historical scheduling case library, the computer master node directly calls the optimal resource scheduling scheme to perform resource scheduling on the three computer computing nodes; If the affinity does not conform to an optimal resource scheduling scheme in the historical scheduling case library, trigger the resource scheduling mechanism and continuously import the newly generated optimal resource scheduling scheme into the historical scheduling case library.
3. A multi-system redundant computer processing method according to claim 2, wherein, The task feature vector can be expressed as: ; Among them, is represented as a task feature vector, is represented as a task type, is represented as the amount of computation, is represented as the latency requirement; The resource feature vector can be expressed as: ; Among them, is represented as a resource feature vector, is represented as the number of CPU cores, is represented as the cache hit rate, is represented as the CPU energy consumption index; The affinity function can be expressed as: ; Among them, is expressed as an affinity function, , , are expressed as weight factors, is expressed as a load, is expressed as a cache matching degree, is expressed as a migration cost.
4. A multi-system redundant computer processing method according to claim 1, characterized in that Step S104 specifically includes: Two hot-standby nodes pre-load the image system of the computer computing node in advance, and the second hot-standby node pre-loads the image system of the computer master node in advance; When performing heartbeat detection and instruction response latency detection, the three computer computing nodes and the computing master node detect each other. When a node fails to respond to the heartbeat for more than 3 cycles and the instruction response latency exceeds 2 times the threshold, it is determined that the node is faulty, and the redundant switching mechanism is immediately started: If at least one of the three computer computing nodes fails, the on-off channels between the three computer computing nodes and the two hot-standby nodes are immediately closed. The hot-standby node starts within 200 ms and takes over the data processing of the faulty computer computing node. The takeover mode of the hot-standby node is one-to-one takeover; If the computer total node fails, the on-off channel between the computer total node and the hot redundant dual node closes immediately, and the hot redundant dual node starts up within 100 ms and takes over the control operation of the failed computer total node; When the response heartbeat and instruction response delay of the failed node recover, a progressive recovery strategy is adopted based on the sandbox verification mode to gradually reload to the original node. After the recovery ends, the on-off channel between the recovered node and the redundant node disconnects.
5. A multi-system redundant computer processing method according to claim 1, characterized in that, Step S105 specifically includes: The two load data processing systems included in the warm redundant node are the x86 system and the ARM system respectively. The x86 system is used to process I / O-intensive data, and the ARM system is used to process compute-intensive data; When the load of at least one computer computing node exceeds 80%, the warm redundant node starts preheating. The preheating time of the warm redundant system is 30 s. When the load of at least one computer computing node exceeds 85%, the on-off channel between the warm redundant node and the three computer computing nodes closes immediately. The preheated warm redundant node takes over the local data of the computer computing node that exceeds the load. When the load of the computer computing node drops to 70%, the computer computing node gradually takes over the data previously migrated to the warm redundant node, and based on the dynamic load balancing algorithm, the data is reasonably allocated to the three computer computing nodes. After the data migration is completed, the warm redundant node cools down, and the on-off channel between the warm redundant node and the three computer computing nodes disconnects.
6. A multi-system redundant computer processing method according to claim 1, characterized in that, A communication channel using the TCP / IP protocol is established between eight different nodes, and the communication channel is encrypted using a public-private key pair, including: Encrypting the data communication in the communication channel based on TLS; Configuring a blockchain network node for each of the eight nodes, and data is transmitted between the nodes based on the P2P communication protocol. A smart contract agent is configured for each node, and the smart contract agent is used to verify all the data transmitted; During each data transmission, the sending node will perform a hash calculation on the data based on the hash function to generate the hash value of the data, and sign it with the private key. The receiving node verifies the legality of the signature through the public key. All changes, transmission processes, and receipt confirmations of the data are recorded based on the blockchain. After the three computer computing nodes receive the data, the received data is verified based on the voting mechanism of the PBFT consensus algorithm. When at least two computer computing nodes confirm that the data is valid, the data will be accepted and the blockchain state will be updated synchronously; When data conflicts occur, that is, when two computer nodes modify the same data, the blockchain record and hash value are used to help determine the correct data version for conflict resolution; When a node fails or data is lost, the data is repaired based on the PBFT consensus algorithm and the backup database.
7. A multi-system redundant computer processing method according to claim 6, characterized in that, During each data transmission, the sending node will perform a hash calculation on the data and attach the timestamp, node fingerprint, and signature information at the same time to generate a data packet. The data packet can be expressed as: ; Among them, is represented as a data packet, is represented as data, is represented as a timestamp, is represented as a node fingerprint, is represented as signature information. During the transmission process, the data packet remains encrypted to prevent tampering; After receiving a data packet, the receiving node verifies the integrity of the data packet through the intelligent contract proxy. If the verification passes, the receiving node updates the record in the blockchain. If the verification fails, the data is discarded; Regularly check the blockchain status of all nodes, and ensure the consistency of data of each node based on the distributed ledger mechanism of the blockchain.
8. A multi-system redundant computer processing method according to claim 1, characterized in that, Establishing a communication channel specifically includes: The computer general node is connected to three computer computing nodes based on a one-way communication channel. The three computer computing nodes are respectively connected to each other based on a two-way communication channel. The computer general node is connected to the hot redundant two-node based on a two-way communication channel. The three computer computing nodes are respectively connected to two hot redundant one-nodes and a warm redundant node based on a two-way communication channel. The three computer computing nodes are connected to the backup database based on a two-way communication channel. The two hot redundant one-nodes and the warm redundant node are connected to the computer general node based on a communication channel. The on-off of the communication channel between the computer general node and the hot redundant two-node is controlled by an on-off channel. The on-off of the communication channel between the three computer computing nodes and the two hot redundant one-nodes and the warm redundant node is controlled by three on-off channels. The three on-off channels between the three computer computing nodes and the two hot redundant one-nodes and the warm redundant node are independently controlled.
9. A multi-system redundant computer processing system, characterized in that, Including: A heterogeneous node cluster, adopting a 3-2-1-1-1 heterogeneous redundant architecture, is used for processing and computing data. This architecture includes three computer computing nodes, two hot redundant one-nodes, one computer general node, one warm redundant node, and one hot redundant two-node; A monitoring module is used to continuously monitor the load of the computer computing nodes in real time and obtain the load value. At the same time, it is also used to obtain the task feature vector, which represents the task characteristics of the data processing that the computer computing node needs to execute, and obtain the resource feature vector, which represents the resource characteristics of the computer computing node; A communication module is used to establish communication channels between different nodes in the heterogeneous node cluster, and establish a communication channel between the backup database and the heterogeneous node cluster; An on-off module is used to establish on-off channels between different nodes in the heterogeneous node cluster and independently control the on-off channels; A scheduling module is used to perform resource scheduling for the three computer computing nodes; A fault module is used to detect the three computer computing nodes and the computer general node in real time, and start the redundant switching mechanism based on heartbeat detection and instruction response delay; A redundancy module is used to start the redundant resource scheduling mechanism; A recording module is used to generate fault information, computer computing node resource scheduling information, and warm redundant node resource scheduling information, and record them in real time. At the same time, the above information is imported into the backup database; A backup module is used to establish a backup database. The backup database is used to store and backup data in real time. At the same time, it is also used to store fault information, computer computing node resource scheduling information, and warm redundant node resource scheduling information.
10. An electronic device, characterized in that, Including: At least one processor; and, a memory communicatively connected to at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.
Citation Information
Cited By
Distributed PLC (Programmable Logic Controller) double-CPU (Central Processing Unit) hot standby redundant system
CN121209408A
Backup node scheduling method and device, computer equipment, medium and program product
CN121542053A