Highway cloud toll collection system high availability realization method
By establishing multi-node cloud services in the highway cloud toll system, configuring lane load balancing, real-time monitoring of node status and automatic fault switching, the system's long problem of abnormal monitoring, load balancing and abnormal recovery time is solved, and the system's availability and vehicle traffic efficiency are improved.
Patent Information
- Application Number
- CN202510210096.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-23
AI Technical Summary
The existing highway cloud charging system has low coupling between the abnormality monitoring mechanism and the station-level trading system, insufficient load balancing and network status detection during multi-host deployment, and excessive abnormal recovery time, which affects vehicle traffic efficiency.
High availability guarantee is achieved through the establishment of multi-node cloud services, configuring lane load balancing, real-time monitoring of node status and automatic fault switching. Specific steps include: establishing multi-node cloud services, configuring lane load balancing, monitoring node status in real time, and automatically switching load and recovery nodes when a failure occurs.
It improves the load capacity and processing speed of the system, enhances the availability and performance of the system, and ensures the stable operation of the system and vehicle traffic efficiency.
Smart Images

Figure CN120034537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a method for implementing high availability of a highway cloud toll collection system. Background Art
[0002] The highway toll collection system is an important part of my country's transportation sector. At present, it has achieved many technological innovations such as automation, intelligence and informatization. The traditional highway toll collection system adopts a centralized architecture, which has problems such as single point failure, high load pressure and limited performance. These problems will lead to the risk of poor reliability, availability and scalability of the system, which will bring troubles to the management and operation of highway toll collection. After the cancellation of provincial boundary toll stations, some provinces and cities have innovated the charging form of highway station-level transaction systems and adopted unmanned cloud charging methods. The cloud charging method refers to the transformation of the original toll collection system deployed in a single lane into a set of station-level transaction systems uniformly deployed at the toll station, so as to realize the intensive deployment and unified management of system resources. However, in the current construction of intelligent cloud toll stations, more attention is paid to the functional realization of cloud toll stations, and the high availability guarantee system of cloud toll systems is rarely mentioned. The guarantee mechanism adopted by most provinces and cities relies on the redundant backup mechanism and abnormal pull-up mechanism provided by the cloud platform itself. Although the cloud platform provides a recovery mechanism after system anomalies, it still has the following shortcomings: First, the anomaly monitoring mechanism is poorly coupled with the station-level transaction system, and it is impossible to accurately and promptly judge system anomalies; second, for multi-host deployment, it is impossible to achieve load balancing and multi-dimensional detection of network status; third, the anomaly recovery time cannot meet the vehicle traffic requirements of the station-level transaction system. Summary of the invention
[0003] Aiming at the problems that the abnormality monitoring mechanism and the station-level transaction system in the existing technology are poorly coupled, resulting in the inability to accurately and timely judge the abnormality of the highway toll system, and the load balancing and multi-dimensional detection of the network status cannot be achieved when multiple hosts are deployed, and the abnormality recovery time is too long, affecting the vehicle traffic efficiency, a high-availability implementation method of the highway cloud toll system is proposed, which enables the highway toll system to make full use of the computing resources of each node, improve the system's load capacity and processing speed, and thus improve the system's availability and performance; and can promptly discover and repair cloud service node failures to ensure the stable operation of the system; realize real-time fault switching and improve the system's availability.
[0004] A method for realizing high availability of a highway cloud toll collection system.
[0005] S1: Establishing a multi-node cloud service: Based on the multi-node cloud service, a station-level transaction system cluster is formed; the station-level transaction system cluster obtains a master node, a slave node and a backup node through cloud resource division, and establishes a mutual access channel; through cloud service network access, the master node and the slave node are connected to the network environment of the toll station level, and the backup node is connected to the toll plaza network environment;
[0006] S2: Configure lane load balancing: Configure the basic information of all lanes under the toll station and cloud service node parameters for each node, and evenly distribute the lanes under the toll station to all nodes for lane task bearing;
[0007] S3: Real-time monitoring of node status: Deploy a monitoring program inside the container where the station-level transaction system of the station-level transaction system cluster described in S1 is located, and implement real-time process abnormality monitoring of each node by running the monitoring program. The monitoring content includes: cloud charging process monitoring, network status monitoring and cloud service node status monitoring;
[0008] S4: Automatic fault switching: When S3 detects an abnormal process in a node, it executes the process pull-up operation; when a node failure or network disconnection is detected, the lane carrying task pre-assigned to the failed node is terminated; the remaining cloud service nodes can automatically balance the load of the running tasks of the lane to which the failed node belongs; when the fault is restored, each node returns to the state of the pre-assigned lane carrying task.
[0009] Preferably, the specific method for establishing a multi-node cloud service is:
[0010] S11: Building a virtualization platform and a downgraded standby server: Building a virtualization platform on a hyper-converged server; the downgraded standby server is a bare metal server;
[0011] S12: Cloud service resource division: The hyper-converged server is divided into resources through the virtualization platform to obtain cloud service nodes; the cloud service nodes include: a master service virtual machine node, i.e., a master node, and a slave service virtual machine node, i.e., a slave node; the downgraded standby server serves as a standby node;
[0012] S13: Cloud service network access: connect the network of the master node and the slave node to the network environment at the toll station level; connect the network of the standby node to the network environment of the toll station square; establish mutual access channels between the master node, the slave node and the standby node through the network access strategy; the standby node and the electromechanical equipment of the toll lane are in the same switch.
[0013] Preferably, the specific method for configuring lane load balancing in step S2 is:
[0014] S21: Configure basic information of all lanes: configure basic information of all lanes under the toll station for each node; the basic lane information includes toll station number, total number of lanes, lane number, lane type, basic operation parameters, rate parameters and connection parameters of various electromechanical equipment;
[0015] S22: Configure cloud service node parameters: configure cloud service node parameters for each node according to the node operation role; the cloud service node parameters include: communication address, heartbeat monitoring IP and port; the node operation roles include: master node, slave node and standby node;
[0016] S23: Configure the cloud service initialization bearer task: According to the total number of lanes under the toll station, the lanes are evenly distributed to all nodes for bearing; the lanes of each node must have both entry lane type and exit lane type.
[0017] Preferably, the specific method for real-time monitoring of node status in step S3 is:
[0018] S31: Cloud charging process monitoring, including heartbeat detection and process existence detection; the heartbeat detection is that the monitoring program determines whether the cloud charging process is running normally by judging the heartbeat interval and heartbeat status sent by the cloud charging process; the process existence detection is to determine whether the cloud charging process is still running by judging the existence of the cloud charging process;
[0019] S32: Network status monitoring: The monitoring program of each cloud service node receives the controller heartbeats and electromechanical equipment heartbeats of all lanes in real time, and obtains the operation status, network status and load status of each cloud service node in the station-level transaction system cluster in real time; when the monitoring program detects that the heartbeat has timed out and still has not received the heartbeat of any lane controller, electromechanical equipment heartbeat and other node heartbeat, the network disconnection and offline operation mechanism is triggered, and the existing lane operation tasks are no longer carried;
[0020] S33: Cloud service node status monitoring: The current node receives heartbeat messages from other nodes, and determines the heartbeat interval and heartbeat status of other nodes to implement status monitoring of other nodes; the current node's own network must be in a normal state.
[0021] Preferably, the specific method of the process existence detection is:
[0022] A1: Load the name of the pre-configured cloud charging process into memory through the monitoring program;
[0023] A2: The monitoring program starts a thread to poll and determine whether the cloud charging process name is still running;
[0024] A3: If it is determined that the cloud charging process is not running, the monitoring program obtains the result that the cloud charging process is running abnormally.
[0025] Preferably, the heartbeat format sent by the cloud charging process to the monitoring program is: {type|category code|operating status}; the format in which the monitoring program receives the controller heartbeat and electromechanical equipment heartbeat is: {type|category code|operating status}; the format in which the current node receives the heartbeat information of other nodes is: {type|category code|operating status|timestamp}.
[0026] Preferably, the specific method of executing the process pull-up operation is:
[0027] B1: The monitoring program pre-reads the configured process name;
[0028] B2: Obtain the application directory corresponding to the process according to the running directory and process name of the current monitoring program;
[0029] B3: Call the startDetached method instruction of the system QProcess to start running the application; under the Linux platform, you need to use ShellCommand to execute the chmod instruction to obtain the application running permission, and then start running the application.
[0030] Preferably, when a node failure or network disconnection is detected, the specific method of automatically switching the fault is:
[0031] C1: Failure switching: End the pre-assigned lane load task of the faulty node, making it run without load; assign the lane task of the faulty node to the normal node;
[0032] C2: Fault recovery: When the faulty node network is restored or the fault is eliminated, each node returns to the state of pre-assigned lane carrying tasks.
[0033] Preferably, the state of recovering to the pre-allocated lane load-bearing task includes: the monitoring program running in the normal node receives the heartbeat message sent by the faulty node again, and the normal node releases the lane task of carrying the faulty node; after the faulty node receives the load recovery message from the normal node, it runs the pre-allocated lane load-bearing task again.
[0034] Beneficial effects:
[0035] The present invention proposes a method for realizing high availability of a highway cloud toll collection system. By establishing a multi-node cloud service, configuring lane load balancing, real-time monitoring of node status and automatic fault switching, a high availability protection measure is provided for coping with software and hardware failures, thereby avoiding the situation in which a single cloud service node fails and causes the unavailability of the production system, and providing a high availability protection mechanism for highway station-level transaction systems.
[0036] By deploying the station-level transaction system on each cloud service node, the problem of the station-level transaction system being unavailable due to a single node downtime is avoided; by connecting the standby node, i.e. the downgraded standby server, to the toll plaza network environment and placing it on the same switch as the electromechanical equipment of the toll lane, the problem of being unable to access the station-level transaction system and control the electromechanical equipment when the toll station-level network fails is avoided. By balancing the lane task load and loading the lane as a unit, the resources of each node can be effectively used to carry out business while ensuring the normal transaction of the lane business, thus achieving efficient use of resources.
[0037] By deploying a monitoring program inside the container where the station-level transaction system is located, the problem of being unable to effectively monitor the status of the station-level transaction system in the container is solved, and the inaccurate problem of relying solely on the container running status and port monitoring to judge abnormalities is also solved; network status monitoring is used to solve the problem of only being able to judge the running status of other nodes but not whether the node itself is disconnected from the network; through cloud charging process monitoring, network status monitoring and cloud service node status monitoring, a composite judgment of whether a node has a fault is achieved, thereby achieving accurate identification of faults.
[0038] By starting the process, the station-level transaction system can be restored to normal function as soon as possible, reducing the business downtime caused by faults, ensuring timely data processing and transmission, and maintaining the stable operation of the system; by assigning the lane tasks of the faulty node to the normal node, when the faulty node returns to normal, each lane resumes executing the original lane task, which solves the problem that the traditional switching mechanism requires restarting the virtual machine or restarting the container, resulting in a long switching time, and realizes the function of quickly switching the business from the faulty node to the normal node.
[0039] In summary, by establishing multi-node cloud services, configuring lane load balancing, real-time monitoring of node status and automatic fault switching technology, the stability and availability of the system are improved, the continuity of highway toll collection services can be effectively guaranteed, and the impact of software and hardware system failures and network failures that occur during daily operations can be effectively reduced. Compared with the existing high-availability guarantee methods, the present invention further makes up for the shortcomings of the existing guarantee methods, so that the highway cloud toll collection system has higher availability and better practicality on the basis of realizing smart transportation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 The present invention is a flowchart of a method for realizing high availability of a highway cloud toll collection system.
[0041] Figure 2 It is a flowchart for establishing a multi-node cloud service.
[0042] Figure 3It is a schematic diagram of establishing a multi-node cloud service network environment.
[0043] Figure 4 This is a flowchart for configuring lane load balancing.
[0044] Figure 5 It is a flow chart for real-time monitoring of node status.
[0045] Figure 6 This is a failover flowchart.
[0046] Figure 7 This is a fault recovery flowchart. DETAILED DESCRIPTION
[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0048] like Figure 1 The process of a method for implementing high availability of a highway cloud toll collection system is shown below:
[0049] S1: Establishing a multi-node cloud service: Based on the multi-node cloud service, a station-level transaction system cluster is formed; the station-level transaction system cluster obtains a master node, a slave node and a backup node through cloud resource division, and establishes a mutual access channel; through cloud service network access, the master node and the slave node are connected to the network environment of the toll station level, and the backup node is connected to the toll plaza network environment;
[0050] S2: Configure lane load balancing: Configure the basic information of all lanes under the toll station and cloud service node parameters for each node, and evenly distribute the lanes under the toll station to all nodes for lane task bearing;
[0051] S3: Real-time monitoring of node status: Deploy a monitoring program inside the container where the station-level transaction system of the station-level transaction system cluster described in S1 is located, and implement real-time process abnormality monitoring of each node by running the monitoring program. The monitoring content includes: cloud charging process monitoring, network status monitoring and cloud service node status monitoring;
[0052] S4: Automatic fault switching: When S3 detects an abnormal process in a node, it executes the process pull-up operation; when a node failure or network disconnection is detected, the lane carrying task pre-assigned to the failed node is terminated; the remaining cloud service nodes can automatically balance the load of the running tasks of the lane to which the failed node belongs; when the fault is restored, each node returns to the state of the pre-assigned lane carrying task.
[0053] The specific implementation process is:
[0054] (I) Establishing a multi-node cloud service
[0055] The multi-node cloud service consists of 3 cloud service nodes, forming a station-level transaction system cluster. Two of the cloud service nodes are deployed in the station-level network environment of the toll station, assuming the functions of the master node and the slave node respectively, and the remaining 1 cloud service node is deployed in the toll plaza network environment, assuming the downgraded standby function. The 3 cloud service nodes carry the same type of business. When the business is running, the 3 cloud service nodes distribute the load tasks according to the priority corresponding to the role. The station-level transaction system is deployed on 3 cloud service nodes to form a station-level transaction system cluster, avoiding the problem of the station-level transaction system being unusable due to the downtime of a single node. The standby node is in the toll plaza network environment and is in the same switch as the electromechanical equipment of the toll lane, avoiding the problem of the station-level transaction system being inaccessible and the electromechanical equipment being uncontrollable due to network failure.
[0056] The hyper-converged server and the downgraded standby server both use the ARM processor architecture, and the three cloud service nodes all use the Kylin V10 server operating system as the basic environment. This toll station has 8 lanes, with 3 entrance lanes and 5 exit lanes. The CPU resources of the master node and the slave node are divided into 32 cores, 32G memory resources, and 6T hard disk resources. The CPU resources of the standby node are 32 cores, 128G memory resources, and 8T hard disk resources. The specific implementation steps for establishing a multi-node cloud service are as follows: Figure 2 :
[0057] 1. Build a virtualization platform and downgrade the standby server. The virtualization platform runs on a hyper-converged server. The virtualization platform uses Kernel-Based Virtual Machines (KVM) to achieve unified interface management and scheduling of virtualized resources, and downgrade the standby server to a single bare metal server.
[0058] 2. Division of cloud service resources. Use the virtualization platform to divide the hyper-converged server into two cloud service nodes, one for the master node and the other for the station-level transaction business. In addition, according to business needs, the virtualization platform also divides the cloud service node resources required for station-level management, station-level storage, and station-level database to support the networked toll collection business of the entire toll station. The downgraded standby server assumes the standby node required for the station-level transaction business, and the hardware resources are directly allocated by the operating system, and all resources are directly used for the station-level transaction business.
[0059] 3. Cloud service network access. Connect the master node and slave node networks to the toll station network environment, and connect the downgraded standby server network to the toll station square network environment. At the same time, establish a network access strategy so that the three cloud service nodes have the ability to access each other. The network environment diagram is as follows: Figure 3As shown. After the physical network is connected, the IP of the master node is 192.168.1.226, the IP of the slave node is 192.168.1.227, and the IP of the standby node is 192.168.1.189.
[0060] (II) Configuring Lane Load Balancing
[0061] By configuring lane load balancing, the operation task load is performed on a lane-by-lane basis. In the initial state, the 8 lanes of the current toll station are evenly distributed to 3 cloud service nodes for load. The master node and slave node load 3 lanes respectively, and the backup node loads 2 lanes. When one of the nodes in the multi-node cloud service goes down or has a network failure, the remaining cloud service nodes can automatically balance the operation tasks of the lanes to which the failed node belongs. The specific operation steps are as follows: Figure 4 :
[0062] 1. Configure basic information of all lanes. Configure the basic information of the 8 lanes under the current toll station to the 3 cloud service nodes to ensure that the 3 cloud service nodes can independently and completely carry all lane passing tasks of this station. The basic information of the lane includes the toll station number, the total number of lanes, the lane number, the lane type, the basic operating parameters, the rate parameters and the connection parameters of various electromechanical equipment. The electromechanical equipment includes ceiling lights, barrier machines, fee displays, lane cameras, license plate recognition cameras, antennas, coils, IC card readers, ticket printers, self-service card receiving and sending machines and mobile payment devices.
[0063] 2. Configure cloud service node parameters. Configure roles and communication addresses on the three cloud service nodes. Configure the roles of the three cloud service nodes as master node, slave node, and standby node, and configure the heartbeat monitoring IP and port of the three cloud service nodes. The configuration of roles and communication addresses provides a way to implement subsequent node status monitoring. The parameter configuration table of the three cloud service nodes is as follows:
[0064] Table 1 Cloud service node parameter configuration table
[0065]
[0066] ●nodeCount: The total number of cloud service nodes. This implementation scenario has 3 nodes.
[0067] ●otherServerIp: corresponds to other cloud service node IP, used for the interaction of heartbeat information. The IP of other cloud service nodes corresponding to the master-slave node is the IP of the standby node, and the IP of other cloud service nodes corresponding to the standby node is the IP of the master node.
[0068] laneNoLists: Lane information carried by the initial state of the current node. The entry lane number starts from 1, and the exit lane number starts from 101.
[0069] ●isFirtNode: whether it is a spare node.
[0070] 3. Configure the node to initialize the task. According to the total number of lanes under the current toll station, divide the lanes evenly to 3 cloud service nodes for task carrying to ensure that all cloud service node resources are effectively utilized under normal circumstances. The lane type configured for each cloud service node must have both entrance lane and exit lane types to ensure that at least one entrance and exit lane is available in any case. The master node is initially configured with entrance 1, exit 1, and exit 2 lanes, the slave node is initially configured with entrance 2, exit 3, and exit 4 lanes, and the standby node is initially configured with entrance 3 and exit 5 lanes.
[0071] (III) Real-time monitoring of node status
[0072] Highway cloud toll collection systems are usually deployed and operated in a containerized manner, that is, the cloud toll collection system runs inside a container, and the container runs on a cloud node. Containerized operation has the advantages of fast and lightweight deployment, isolation of operating resources, and consistency of operating environment. The station-level transaction system consists of a cloud toll collection process, other service processes, and a monitoring program. The node operation status monitoring function is performed by a monitoring program running in a container. First, the monitoring program obtains the operation status of the station-level transaction system in real time by exchanging data with the cloud toll collection process in the current container; second, the monitoring programs of each cloud service node interact with each other by heartbeat, and obtain the operation status, network status, and load status of each cloud service node in the station-level transaction system cluster in real time. The data interaction protocol uses the User Datagram Protocol (UDP), and the encoding format is UTF8. The monitoring program listening port is set to 10002, and the cloud toll collection process listening port is 34006. The specific operation steps are as follows: Figure 5 :
[0073] 1. Cloud charging process monitoring
[0074] The cloud charging process running status monitoring is divided into heartbeat detection and process existence detection. On the one hand, by judging the cloud charging heartbeat interval and heartbeat status, it is determined whether the cloud charging process is running normally; on the other hand, by judging the existence of the cloud charging process, it is determined whether the cloud charging process has abnormal exit. When the monitoring program in the container detects that the cloud charging process is running abnormally, it triggers the abnormal pull-up mechanism to ensure the stable operation of the system.
[0075] The heartbeat format sent by the cloud charging process to the monitoring program is: {laneupdate|12|1}. The heartbeat content consists of type, category code and running status. The contents are connected by vertical lines, and the entire content is contained in a pair of curly brackets {}. The cloud charging process sends a heartbeat message to the monitoring program every 5 seconds. If the monitoring program does not receive a heartbeat message within 3 cycles (15 seconds), the cloud charging process is determined to be in an abnormal state; in addition, if the running status in the heartbeat message received by the monitoring program is 0, the cloud charging process is also determined to be in an abnormal state. Based on the abnormal state, the monitoring program will restart the cloud charging process to ensure that the cloud charging process is in a normal running state. If the heartbeat interval and monitoring cycle are too short, it is more sensitive to network fluctuations; if the heartbeat interval and monitoring cycle are too long, the switching timeliness will be reduced. This implementation scenario comprehensively considers the factors of network fluctuations and switching timeliness, and sets the heartbeat interval to 5 seconds and the monitoring cycle to 3 times.
[0076] Considering that the cloud charging process may be terminated due to unpredictable reasons, the monitoring program also has the function of judging the existence of the cloud charging process. When the monitoring program is running, first, the name of the pre-configured cloud charging process is loaded into the memory; secondly, a thread is started to poll and determine whether the process name is still running; finally, the terminated process is restarted. The cloud charging process includes ETC transaction service (etcserver), MTC transaction service (mtcserver), field control service (devicecontrolser ver), billing service (feeserver), namelist service (namelistserver), remote duty service (remotecontrolserver), data processing and transmission service (datatransserver), log service (logserver) and parameter management service (paramserver).
[0077] 2. Network status monitoring
[0078] The monitoring program of each cloud service node receives the controller heartbeats and electromechanical equipment heartbeats of all lanes in real time. When the monitoring program detects that the preset timeout period has expired and no lane controller heartbeat, electromechanical equipment heartbeat, or other cloud service node heartbeat has been received, the offline operation mechanism is triggered and the existing lane operation business is no longer carried.
[0079] The heartbeat format of the controller and electromechanical equipment is: {laneupdate|79|1}. The heartbeat content consists of the type, category code and operating status. The contents are connected by vertical lines, and the entire content is contained in a pair of curly brackets {}. The controller and electromechanical equipment send a heartbeat message to the three cloud service node monitoring programs every 5 seconds. If the monitoring program does not receive a heartbeat message within 3 cycles (15 seconds), it is determined that the current node cloud service itself is in a disconnected state. In the disconnected state, in order to avoid the situation where multiple cloud service nodes occupy the same lane controller and electromechanical equipment at the same time, the monitoring program will end the MTC transaction service and ETC transaction service processes. At the same time, the cloud charging process monitoring function will no longer determine the existence of the MTC transaction service and ETC transaction service processes.
[0080] 3. Cloud service node status monitoring
[0081] Each cloud service node has the function of sending the heartbeat information of its own operation status, and also has the function of receiving the heartbeat information of other cloud service nodes. By judging the heartbeat interval and heartbeat status of other cloud service nodes, the operation status of other cloud service nodes can be monitored. When other cloud service nodes are abnormal, the automatic switching mechanism is triggered.
[0082] The heartbeat format between cloud service nodes is: {laneupdate|47|1|14:18:07}. The heartbeat content consists of type, category code, running status and timestamp. The contents are connected by vertical lines and the entire content is contained in a pair of curly brackets {}. Each cloud service node sends a heartbeat message to other cloud service node monitoring programs every 5 seconds. If the current cloud service node monitoring program does not receive the heartbeat sent by other cloud service node monitoring programs within 3 cycles (15 seconds), and the current cloud service node’s own network is normal, it is determined that other cloud service nodes are disconnected or faulty.
[0083] (IV) Automatic Failure Switching
[0084] Based on the function of real-time monitoring of cloud service node status, automatic fault switching has the composite judgment function of software running status judgment, current node disconnection judgment and other node running status judgment. It realizes the accurate identification of cloud service node failure. In addition, fault switching does not require the restart of virtual machines or containers. It only needs to end or add lane running tasks based on existing processes, which shortens the fault switching time.
[0085] When the monitoring program in the container detects that the cloud charging process is running abnormally, it executes the process pull-up operation. The specific steps are as follows:
[0086] ① Get the monitoring program to pre-read the configured process name.
[0087] ② Get the application directory corresponding to the process based on the current running directory and process name.
[0088] ③Call the startDetached method of the system QProcess to run the application. On the Linux platform, you must first use ShellCommand to execute the chmod command to obtain the application's running permissions.
[0089] When a node fails or loses power, perform failover and failback operations:
[0090] 1. Failover.
[0091] The failover process is as follows Figure 6 As shown, fault switching refers to the process of switching the lane carrying task of the faulty node to the normal node. It occurs on two or more nodes. When one or more of the nodes are disconnected from the network or fail, the monitoring program running in the faulty node ends the pre-allocated lane carrying task, making it run at no load, to prevent multiple cloud service nodes from running the same lane task at the same time when the fault is recovered, resulting in resource occupation; the monitoring program running in the normal node assigns the lane task of the faulty node to the station-level transaction system. Therefore, in addition to carrying the original pre-allocated lane task, the station-level transaction system of the normal node also needs to carry the lane task of the faulty node. Take the network failure of the backup node 192.168.1.189 as an example:
[0092] ① After the standby node monitoring program determines that its own node is in a disconnected state through the controller heartbeat and the electromechanical equipment heartbeat, it ends the pre-assigned entrance 3 lane task and exit 5 lane task.
[0093] ② The master node monitoring program determines that the standby node has a network failure by not receiving the heartbeat information between the monitoring programs for three consecutive cycles. The master node monitoring program sends load switching instructions to the cloud charging processes of the master node and the slave node respectively: in addition to the load entry 1, exit 1 and exit 2 lane tasks, the master node also needs a load entry 3 lane task, and in addition to the load entry 2, exit 3 and exit 4 lane tasks, the slave node also needs a load exit 5 lane task.
[0094] 2. Fault recovery
[0095] Fault recovery process Figure 7 As shown, fault recovery means that after the faulty node network is restored or the fault is eliminated, each cloud service node is restored to the state of pre-allocated lane carrying tasks. When the faulty node is restored, the monitoring program running in the normal node receives the heartbeat message sent by the faulty node again, and the normal node no longer carries the lane task of the faulty node. After the faulty node receives the load recovery message from the normal node, it runs the pre-allocated lane carrying task again. Take the network fault recovery of the standby node 192.168.1.189 as an example:
[0096] ① The main node monitoring program receives the heartbeat information sent by the standby node monitoring program and determines that the failure of the standby node has been resolved. The main node monitoring program sends load recovery instructions to the cloud charging processes of the main node and the slave node respectively, canceling the loads of lanes 3 at the entrance and 5 at the exit.
[0097] After the loads of the master and slave nodes are restored to the initial state, the main node monitoring program sends a load recovery instruction to the standby node monitoring program. The standby node monitoring program restores the monitoring function for the cloud charging process and resumes the operation tasks of lanes 3 at the entrance and 5 at the exit.
[0098] Finally, it should be noted that the above is only used to illustrate the technical solution of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred arrangement, those of ordinary skill in the art should understand that the technical solution of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for implementing high availability of a highway cloud toll collection system, characterized in that: S1: Establishing a multi-node cloud service: Based on the multi-node cloud service, a station-level transaction system cluster is formed; the station-level transaction system cluster obtains a master node, a slave node and a backup node through cloud resource division, and establishes a mutual access channel; through cloud service network access, the master node and the slave node are connected to the network environment of the toll station level, and the backup node is connected to the toll plaza network environment; S2: Configure lane load balancing: Configure the basic information of all lanes under the toll station and cloud service node parameters for each node, and evenly distribute the lanes under the toll station to all nodes for lane task bearing; S3: Real-time monitoring of node status: Deploy a monitoring program inside the container where the station-level transaction system of the station-level transaction system cluster described in S1 is located, and implement real-time process abnormality monitoring of each node by running the monitoring program. The monitoring content includes: cloud charging process monitoring, network status monitoring and cloud service node status monitoring; S4: Automatic fault switching: When S3 detects an abnormal process in a node, it executes the process pull-up operation; when a node failure or network disconnection is detected, the lane carrying task pre-assigned to the failed node is terminated; the remaining cloud service nodes can automatically balance the load of the running tasks of the lane to which the failed node belongs; when the fault is restored, each node returns to the state of the pre-assigned lane carrying task.
2. A method for realizing high availability of a highway cloud toll collection system according to claim 1, characterized in that: The specific method for establishing a multi-node cloud service is: S11: Building a virtualization platform and downgraded standby server: Building a virtualization platform on a hyper-converged server; the downgraded standby server is a bare metal server; S12: Cloud service resource division: The hyper-converged server is divided into resources through the virtualization platform to obtain cloud service nodes; the cloud service nodes include: a master service virtual machine node, i.e., a master node, and a slave service virtual machine node, i.e., a slave node; the downgraded standby server serves as a standby node; S13: Cloud service network access: connect the network of the master node and the slave node to the network environment at the toll station level; connect the network of the standby node to the network environment of the toll station square; establish mutual access channels between the master node, the slave node and the standby node through the network access strategy; the standby node and the electromechanical equipment of the toll lane are in the same switch.
3. A method for realizing high availability of a highway cloud toll collection system according to claim 1, characterized in that: The specific method for configuring lane load balancing in step S2 is: S21: Configure basic information of all lanes: configure basic information of all lanes under the toll station for each node; the basic lane information includes toll station number, total number of lanes, lane number, lane type, basic operation parameters, rate parameters and connection parameters of various electromechanical equipment; S22: Configure cloud service node parameters: configure cloud service node parameters for each node according to the node operation role; The cloud service node parameters include: communication address, heartbeat monitoring IP and port; the node operation roles include: master node, slave node and standby node; S23: Configure the cloud service initialization bearer task: According to the total number of lanes under the toll station, the lanes are evenly distributed to all nodes for bearing; the lanes of each node must have both entry lane type and exit lane type.
4. A method for realizing high availability of a highway cloud toll collection system according to claim 1, characterized in that: The specific method of real-time monitoring of node status described in step S3 is: S31: Cloud charging process monitoring, including heartbeat detection and process existence detection; the heartbeat detection is that the monitoring program determines whether the cloud charging process is running normally by judging the heartbeat interval and heartbeat status sent by the cloud charging process; the process existence detection is to determine whether the cloud charging process is still running by judging the existence of the cloud charging process; S32: Network status monitoring: The monitoring program of each cloud service node receives the controller heartbeat and electromechanical equipment heartbeat of all lanes in real time, and obtains the operation status, network status and load status of each cloud service node of the station-level transaction system cluster in real time; When the monitoring program detects that the heartbeat has timed out and still has not received any lane controller heartbeat, electromechanical equipment heartbeat, or other node heartbeat, the offline operation mechanism is triggered and the existing lane operation tasks are no longer carried out; S33: Cloud service node status monitoring: The current node receives heartbeat messages from other nodes, and determines the heartbeat interval and heartbeat status of other nodes to implement status monitoring of other nodes; the current node's own network must be in a normal state.
5. A method for realizing high availability of a highway cloud toll collection system according to claim 4, characterized in that: The specific method of the process existence detection is: A1: Load the name of the pre-configured cloud charging process into memory through the monitoring program; A2: The monitoring program starts a thread to poll and determine whether the cloud charging process name is still running; A3: If it is determined that the cloud charging process is not running, the monitoring program obtains the result that the cloud charging process is running abnormally.
6. A method for realizing high availability of a highway cloud toll collection system according to claim 4, characterized in that: The heartbeat format sent by the cloud charging process to the monitoring program is: {type|category code|operating status}; the format in which the monitoring program receives the controller heartbeat and electromechanical equipment heartbeat is: {type|category code|operating status}; the format in which the current node receives the heartbeat information of other nodes is: {type|category code|operating status|timestamp}.
7. A method for realizing high availability of a highway cloud toll collection system according to claim 1, characterized in that: The specific method of executing the process pull-up operation is: B1: The monitoring program pre-reads the configured process name; B2: Obtain the application directory corresponding to the process according to the running directory and process name of the current monitoring program; B3: Call the startDetached method instruction of the system QProcess to start running the application; under the Linux platform, you need to use ShellCommand to execute the chmod instruction to obtain the application running permission, and then start running the application.
8. A method for realizing high availability of a highway cloud toll collection system according to claim 1, characterized in that: When a node failure or network disconnection is detected, the specific method of automatic fault switching is as follows: C1: Failure switching: End the pre-assigned lane load task of the faulty node, making it run without load; assign the lane task of the faulty node to the normal node; C2: Fault recovery: When the faulty node network is restored or the fault is eliminated, each node returns to the state of pre-assigned lane carrying tasks.
9. A method for realizing high availability of a highway cloud toll collection system according to claim 8, characterized in that: The state of recovering to the pre-allocated lane load-bearing task includes: the monitoring program running in the normal node receives the heartbeat message sent by the faulty node again, and the normal node releases the lane task of carrying the faulty node; after the faulty node receives the load recovery message from the normal node, it runs the pre-allocated lane load-bearing task again.
Citation Information
Patent Citations
Highway toll passage management and control system based on cloud toll architecture
CN115529325A
Charging system and method based on Internet of Things
CN117058775A
Road toll collection system and method, electronic equipment and storage medium
CN118262425A
Toll collection system
JP2022113940A
Cited By
Road networking charging management and information resource wide-area fusion scheduling system
CN121436603A
A highway networking toll management and information resource wide-area fusion scheduling system
CN121436603B
Global resource allocation method for cross-regional road charging management system
CN121441916A
Disaster recovery processing method and device for charge management system based on distributed architecture
CN122476101A