Oracle RAC heartbeat network detection switching method and device and storage medium

By using network card hardware counters and hardware timestamps on the computing nodes of the Oracle RAC cluster in real time to monitor the heartbeat packet situation, judge and switch network card or virtual IP, the business pause caused by network jump failure in Oracle RAC cluster center is solved, and the reliability and stability of the system is improved.

CN120128504AInactive Publication Date: 2025-06-10POWERLEADER COMPUTER SYST CO LTD

Patent Information

Application Number
CN202510614666.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In Oracle RAC technology application scenarios, a single server downtime causes a delay in the migration of the heartbeat network, which may cause database service suspension and reduce the reliability of the Oracle RAC cluster.

Method used

By binding the network card on the computing node of the Oracle RAC cluster, using the network card hardware counter and hardware timestamp, the sending and receiving of heartbeat packets is collected and analyzed in real time, the network delay, packet loss rate and bandwidth utilization rate are calculated, the heartbeat network failure is judged, and the network card or virtual IP is switched to restore services.

Benefits of technology

It realizes detection of heartbeat network failures in milliseconds, reduces the impact of failures on Oracle RAC cluster, ensures timely traffic monitoring, and improves system reliability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128504A_ABST
    Figure CN120128504A_ABST
Patent Text Reader

Abstract

The invention relates to an Oracle RAC heartbeat network detection switching method and device and a storage medium, and the method comprises the steps: obtaining the number of sent and received heartbeat packets, the number of received and sent bytes and the round-trip time of the heartbeat packets in a network card hardware counter, calculating the network delay according to the round-trip time of the heartbeat packets through employing a hardware timestamp, and carrying out the detection switching of the Oracle RAC heartbeat network. And calculating a packet loss rate according to the number of sent and received heartbeat packets, calculating a bandwidth utilization rate according to the number of received and sent bytes, judging whether a heartbeat network fault is detected according to the network delay, the packet loss rate and the bandwidth utilization rate, and if the heartbeat network fault is detected, switching a network card or a virtual IP. The network flow data comprises the number of sent and received heartbeat packets and the number of received and sent bytes, the network flow data can be collected in real time through a network card hardware counter and a hardware timestamp, the timeliness of flow monitoring can be ensured, the accuracy can be up to milliseconds, heartbeat network faults can be detected within the millisecond level, and the service life of the network is prolonged. And the influence of the heartbeat network fault on the reliability of the Oracle RAC cluster is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and particularly to an Oracle RAC heartbeat network detection and switching method, device, and storage medium. Background Art

[0002] In the application scenario of Oracle RAC technology, when a single server (compute node or storage node) fails, there is a possibility of service interruption. For example, when one of the servers in the compute node fails abnormally due to internal and external factors, if the heartbeat network is on the faulty compute node at this time, the cluster heartbeat network will experience a process of migrating to another compute node. During this migration process, there will be a certain delay, which will cause the database service to be temporarily stopped and unable to be normally written to disk, thereby reducing the reliability of the Oracle RAC cluster. Summary of the Invention

[0003] The present invention provides an Oracle RAC heartbeat network detection and switching method, device, and storage medium, aiming to solve at least one of the technical problems existing in the prior art.

[0004] The technical solution of the present invention is an Oracle RAC heartbeat network detection and switching method, which is applied to an Oracle RAC heartbeat network detection and switching device for detecting the status of the heartbeat network of an Oracle RAC cluster, and includes: Binding a network card to the compute nodes in the Oracle RAC cluster, where the network card has a virtual IP setting, and obtaining the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter through the network card driver or API; Using a hardware timestamp to calculate the network delay according to the round-trip time of the heartbeat packets; Calculating the packet loss rate according to the number of heartbeat packets sent and received; Calculating the bandwidth utilization rate according to the number of bytes received and sent; Judging whether a heartbeat network failure is detected according to the network delay, the packet loss rate, and the bandwidth utilization rate; If the heartbeat network failure is detected, switching the network card or the virtual IP.

[0005] According to some embodiments of the present invention, the using a hardware timestamp to calculate the network delay according to the round-trip time of the heartbeat packets includes: Using the ping command to detect the network connectivity by sending and receiving heartbeat packets; Using the hardware timestamp, calculate the network latency according to the round-trip time of the heartbeat packet.

[0006] According to some embodiments of the present invention, determining whether a heartbeat network failure is detected based on the network latency, the packet loss rate, and the bandwidth utilization rate includes: Set a first threshold, a second threshold, and a third threshold for the network latency, the packet loss rate, and the bandwidth utilization rate respectively; Determine whether the network latency, the packet loss rate, and the bandwidth utilization rate exceed the first threshold, the second threshold, and the third threshold respectively; If one or more combinations of the network latency exceeding the first threshold, the packet loss rate exceeding the second threshold, and the bandwidth utilization rate exceeding the third threshold occur, then the heartbeat network failure is detected; otherwise, the heartbeat network failure is not detected.

[0007] According to some embodiments of the present invention, if the heartbeat network failure is detected, switching the network card includes: Use ifenslave to bind multiple network cards to the computing node; Deploy the heartbeat network on the network card; If the heartbeat network failure is detected, determine the network card of the computing node corresponding to the faulty heartbeat network as the faulty network card; Switch the computing node from the faulty network card to the standby network card.

[0008] According to some embodiments of the present invention, if the heartbeat network failure is detected, switching the virtual IP includes: Deploy the heartbeat network on the computing node; If the heartbeat network failure is detected, determine the computing node corresponding to the faulty heartbeat network as the faulty computing node; Evict the faulty computing node from the Oracle RAC cluster; Migrate the virtual IP of the faulty computing node to a healthy computing node; Send an ARP broadcast through Oracle Clusterware to update the MAC address table of the network device to ensure that traffic is routed to the healthy computing node; The client redirects the connection to the virtual IP on the healthy computing node.

[0009] According to some embodiments of the present invention, after the client redirects the connection to the virtual IP on the healthy computing node, if the heartbeat network failure is detected, switching the virtual IP further includes: Migrate the client session to the healthy computing node through TAF.

[0010] According to some embodiments of the present invention, it further includes: Create a dashboard of the network latency, the packet loss rate, and the bandwidth utilization rate using Grafana to display the network latency, the packet loss rate, and the bandwidth utilization rate.

[0011] The technical solution of the present invention further relates to an Oracle RAC heartbeat network detection and switching device for performing an Oracle RAC heartbeat network detection and switching method as described above. The Oracle RAC heartbeat network detection and switching device includes: A data acquisition module for obtaining the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter through the network card driver or the API; A data analysis module for calculating the network latency, the packet loss rate, and the bandwidth utilization rate; the data acquisition module is connected to the data analysis module; A fault detection module for determining whether the heartbeat network fault is detected according to the network latency, the packet loss rate, and the bandwidth utilization rate; the data analysis module is connected to the fault detection module; An automatic recovery module for switching the network card or the virtual IP; the fault detection module is connected to the automatic recovery module.

[0012] According to some embodiments of the present invention, it further includes: A visualization module for creating a dashboard of the network latency, the packet loss rate, and the bandwidth utilization rate using Grafana to display the network latency, the packet loss rate, and the bandwidth utilization rate; the automatic recovery module is connected to the visualization module.

[0013] The technical solution of the present invention further relates to an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements an Oracle RAC heartbeat network detection and switching method as described above.

[0014] The technical solution of the present invention further relates to a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements an Oracle RAC heartbeat network detection and switching method as described above.

[0015] The beneficial effects of the present invention include: binding the computing nodes in the Oracle RAC cluster to network cards, where virtual IP settings exist for the network cards. Obtain the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter through the network card driver or API. Then, use the hardware timestamp to calculate the network latency based on the round-trip time of the heartbeat packets, calculate the packet loss rate based on the number of heartbeat packets sent and received, calculate the bandwidth utilization rate based on the number of bytes received and sent, and determine whether a heartbeat network failure is detected based on the network latency, packet loss rate, and bandwidth utilization rate. If a heartbeat network failure is detected, switch the network card or virtual IP. Through the network card hardware counter and hardware timestamp, it can be accurate to milliseconds, detect heartbeat network failures within milliseconds, and reduce the impact of heartbeat network failures on the reliability of the Oracle RAC cluster. The network traffic data includes the number of heartbeat packets sent and received, and the number of bytes received and sent. Through the network card hardware counter and hardware timestamp, the network traffic data can be collected in real time, which is beneficial to ensuring the timeliness of traffic monitoring.

[0016] In addition, some of the additional aspects and advantages of the present invention will be given in the following description, some will become apparent from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an optional flowchart of a method for detecting and switching the Oracle RAC heartbeat network in an embodiment of the present invention.

[0018] Figure 2 is a schematic diagram of a device for detecting and switching the Oracle RAC heartbeat network in an embodiment of the present invention.

[0019] Figure 3 is an optional network topology diagram of the Oracle RAC cluster in an embodiment of the present invention.

[0020] Figure 4 is an optional schematic diagram of the Oracle RAC cluster and the device for detecting and switching the Oracle RAC heartbeat network in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with the embodiments and the drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0022] It should be noted that, unless otherwise specified, when a certain feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. In addition, the descriptions such as upper, lower, left, right, top, bottom, etc. used in the present invention are only relative to the mutual positional relationship of the components of the present invention in the drawings.

[0023] In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this technology belongs. The terms used in the specification of this article are only for describing specific embodiments, rather than for limiting the present invention. The term "and / or" used herein includes any combination of one or more of the related listed items.

[0024] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of the present invention, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.

[0025] Refer to Figures 1 to 4 , in some embodiments, the technical solution of the present invention is an Oracle RAC heartbeat network detection and switching method, which is applied to an Oracle RAC heartbeat network detection and switching device. The Oracle RAC heartbeat network detection and switching device is used to detect the status of the heartbeat network of the Oracle RAC cluster, including but not limited to steps S101 to S106. The following will introduce each step in turn.

[0026] Step S101: Bind the network card to the computing nodes in the Oracle RAC cluster. The network card has a virtual IP setting. Obtain the number of heartbeat packets sent and received, the number of bytes sent and received, and the round-trip time of the heartbeat packets in the network card hardware counter through the network card driver or API.

[0027] Note that Oracle RAC (Real Application Clusters) is a high-availability and scalable clustering solution for Oracle databases, allowing multiple database instances to access the same database simultaneously. By running Oracle database instances on multiple servers (compute nodes), RAC provides high availability, load balancing, and failover capabilities, ensuring the continuous availability of database services. An Application Programming Interface (API) is a set of definitions and protocol specifications that allows software applications or components to interact and exchange data. An Oracle RAC (Real Application Clusters) cluster is a database cluster architecture based on Oracle RAC technology that allows multiple physical servers (compute nodes) to work together to access and manage the same database instance. This architecture achieves high availability, load balancing, and failover capabilities through shared storage and cluster software (such as Oracle Clusterware), ensuring the continuous availability and high performance of database services. Network Interface Card Hardware Counters are registers or storage units in network card hardware used to record network communication-related statistical information.

[0028] Referring to Figure 3 , in a specific embodiment, the Oracle RAC heartbeat network detection and switching method is applied to a 2+3 architecture, that is, an Oracle RAC cluster built with 2 compute nodes and 3 storage nodes.

[0029] Specifically, a compute node is a physical server or virtual machine that runs an Oracle database instance. Each compute node is installed with Oracle database software and Oracle Clusterware to manage cluster resources and database instances. A storage node is a device that provides shared storage resources, and all compute nodes access database files through the storage node. A storage node is usually a Storage Area Network (SAN) or Network Attached Storage (NAS) that allows multiple compute nodes to access the same storage resource simultaneously. The heartbeat network is a dedicated network for communication between compute nodes, which is used to detect the connectivity between nodes and ensure that each node can operate normally. Heartbeat Packets are a communication mechanism used to monitor the status of each node in the cluster. Heartbeat Packets send signals between nodes periodically to ensure that each node can respond to requests from other nodes in a timely manner, thus maintaining the high availability and stability of the cluster.

[0030] In a specific embodiment, the heartbeat network of the Oracle RAC cluster architecture is deployed on 10 / 25GE network cards. Using dedicated network cards avoids sharing bandwidth with other network traffic, reduces network congestion and interference, and achieves the effect of physical isolation.

[0031] In a possible implementation, data such as network traffic, packet loss, and errors in the network card hardware counter is obtained through the network card driver or API.

[0032] In a specific embodiment, ethtool is used to obtain the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets. Among them, ethtool is used to obtain the detailed statistics of network card eth0. The detailed statistics include the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets, etc. The example is as follows: ethtool -S eth0 In a specific embodiment, network traffic data and the round-trip time of heartbeat packets in the network card hardware counter are obtained through the network card driver or API. Among them, the network traffic data includes the number of heartbeat packets sent and received, and the number of bytes received and sent.

[0033] Step S102: Use the hardware timestamp to calculate the network latency based on the round-trip time of the heartbeat packet.

[0034] It should be noted that the hardware timestamp (Hardware Timestamping) is a timestamp directly generated by hardware devices (such as network cards, FPGAs, etc.) and is used to record the sending and receiving times of data packets.

[0035] In some embodiments, using the hardware timestamp to calculate the network latency based on the round-trip time of the heartbeat packet includes: Use the ping command to detect the network connectivity by sending and receiving heartbeat packets; Use the hardware timestamp to calculate the network latency based on the round-trip time of the heartbeat packet.

[0036] In some embodiments, calculating the network latency based on the round-trip time of the heartbeat packet includes: Divide the round-trip time (RTT) of the heartbeat packet by 2 to obtain the one-way latency (i.e., the aforementioned network latency).

[0037] Specifically, the round-trip time (RTT) refers to the total time for the heartbeat packet to be sent and returned. The one-way latency refers to the time for data to travel from the sender to the receiver or from the receiver to the sender. The round-trip time includes the latency in both directions. Divide the round-trip time (RTT) of the heartbeat packet by 2 to obtain the one-way latency.

[0038] In a possible implementation, the network card hardware timestamp function is used to accurately measure network latency.

[0039] In a possible implementation, PTP (Precision Time Protocol) is used to accurately measure network latency. Specifically, PTP (Precision Time Protocol) is a protocol based on the IEEE 1588 standard and is used to achieve high-precision time synchronization in a computer network. PTP can achieve synchronization accuracy at the sub-microsecond level. The example is as follows: ptp4l -i eth0 -m In a possible implementation, the network connectivity and latency are detected by sending and receiving custom heartbeat packets. Among them, the ping command is used to detect the network connectivity and latency by sending and receiving custom heartbeat packets. The example is as follows: ping -I eth0 192.168.1.1 Step S103: Calculate the packet loss rate based on the number of heartbeat packets sent and received.

[0040] Specifically, the packet loss rate refers to the proportion of data packets lost due to various reasons (such as network congestion, errors, device failures, etc.) in network communication.

[0041] In a specific embodiment, calculating the packet loss rate based on the number of heartbeat packets sent and received includes: Dividing the difference between the number of heartbeat packets sent and the number of heartbeat packets received by the number of heartbeat packets sent to obtain the packet loss rate.

[0042] Step S104: Calculate the bandwidth utilization rate based on the number of bytes received and sent.

[0043] In a specific embodiment, calculating the bandwidth utilization rate based on the number of bytes received and sent includes: Subtracting the number of bytes received for the first time from the number of bytes received for the second time to obtain the difference in the number of received bytes; Subtracting the number of bytes sent for the first time from the number of bytes sent for the second time to obtain the difference in the number of sent bytes; Adding the difference in the number of received bytes and the difference in the number of sent bytes to obtain a first intermediate value; Multiplying the sampling time interval by the network card bandwidth to obtain a second intermediate value; Dividing the first intermediate value by the second intermediate value and then multiplying by one hundred percent to obtain the bandwidth utilization rate.

[0044] Specifically, the sampling time interval is the time interval between two data readings (i.e., the number of bytes received and sent), in seconds. The network card bandwidth is the actual bandwidth of the network card, in bytes per second.

[0045] In a possible implementation, the bandwidth utilization rate is calculated based on the network card counter.

[0046] Step S105: Determine whether a heartbeat network failure is detected based on the network latency, packet loss rate, and bandwidth utilization rate.

[0047] In some embodiments, determining whether a heartbeat network failure is detected based on the network latency, packet loss rate, and bandwidth utilization rate includes: Set a first threshold, a second threshold, and a third threshold for the network latency, packet loss rate, and bandwidth utilization rate respectively; Determine whether the network latency, packet loss rate, and bandwidth utilization rate exceed the first threshold, the second threshold, and the third threshold respectively; If one or more combinations of the network latency exceeding the first threshold, the packet loss rate exceeding the second threshold, and the bandwidth utilization rate exceeding the third threshold occur, a heartbeat network failure is detected; otherwise, no heartbeat network failure is detected.

[0048] In a possible implementation, thresholds are set for metrics such as network latency, packet loss rate, and bandwidth utilization rate. When these metrics exceed the thresholds, it is determined as a network failure.

[0049] Step S106: If a heartbeat network failure is detected, switch the network card or virtual IP.

[0050] It should be noted that the Cluster IP, also known as the Virtual IP (VIP), is a technology used to balance the network communication load among multiple computer nodes. It creates a logical address to represent a group of nodes and distributes the received requests to multiple nodes associated with this address. The Internet Protocol (IP) is the basic communication protocol for data transmission between devices in a computer network.

[0051] In some embodiments, if a heartbeat network failure is detected, switching the network card includes: Use ifenslave to bind multiple network cards to the computing node; Deploy the heartbeat network on the network card; If a heartbeat network failure is detected, determine the network card of the computing node corresponding to the failed heartbeat network as the failed network card; Switch the computing node from the failed network card to the standby network card.

[0052] Specifically, ifenslave is a tool in the Linux system used to bind and manage network interfaces, achieving load balancing and redundant backup of network cards. Ifenslave allows users to bind multiple physical network cards to a virtual network card (bonding interface), thereby improving network bandwidth and reliability.

[0053] The example of using ifenslave to bind multiple network cards is as follows: ifenslave bond0 eth1 eth2 In a specific embodiment, the computing node adopts two 10 / 25GE network cards for deploying the heartbeat network. The heartbeat network adopts the deployment form of mode4: 802.3ad (LACP). LACP (Link Aggregation Control Protocol) is a protocol based on the IEEE 802.3ad standard, used to dynamically bundle multiple physical links together to form a logical link (link aggregation group, LAG).

[0054] It can be understood that using two network cards can achieve redundancy through Bonding, avoiding single point of failure. If one network card fails, the other network card can continue to provide services, which is beneficial to ensuring the stability of the Oracle RAC cluster. Through the load balancing mode of Bonding (such as Mode 4: 802.3ad), the bandwidth of the two network cards can be fully utilized to improve the performance of the heartbeat network. The IP corresponding to the heartbeat network in the Oracle database node here is the cluster IP (VIP).

[0055] In some embodiments, if a heartbeat network failure is detected, then switch the virtual IP, including: Deploy the heartbeat network on the computing node; If a heartbeat network failure is detected, then determine the computing node corresponding to the failed heartbeat network as the failed computing node; Evict the failed computing node from the Oracle RAC cluster; Migrate the virtual IP of the failed computing node to a healthy computing node; Send an ARP broadcast through Oracle Clusterware to update the MAC address table of the network device to ensure that traffic is routed to the healthy computing node; The client redirects the connection to the virtual IP on the healthy computing node.

[0056] In some embodiments, after the client redirects the connection to the virtual IP on the healthy computing node, if a heartbeat network failure is detected, then switching the virtual IP further includes: Migrate the client session to a healthy computing node through TAF.

[0057] Specifically, Oracle Clusterware is cross-platform cluster software provided by Oracle and is an essential component for running Oracle Real Application Clusters (Oracle RAC). Oracle Clusterware provides basic cluster services at the operating system level for the cluster mode of the Oracle database, enabling multiple independent servers to work together to form a highly available (HA) cluster environment. TAF (Transparent Application Failover) is a high-availability feature provided by the Oracle database, aiming to automatically redirect client connections to other available database instances when a database instance fails, thus minimizing service interruptions. ARP (Address Resolution Protocol) is a network protocol used to resolve IP addresses at the network layer to MAC addresses at the data link layer. The MAC (Media Access Control) address (media access control address) is the unique identifier of a network device and is used for communication at the data link layer. The MAC address is a 48-bit address represented in hexadecimal form.

[0058] It should be noted that the VIP is the highly available IP address for client connections in Oracle RAC. Each computing node has a VIP, usually in the same subnet as the physical IP address of the computing node. When a computing node fails, the VIP will automatically migrate to other healthy computing nodes to ensure that client connections are not interrupted.

[0059] Specifically, in an Oracle RAC environment, if one of the computing nodes fails, the cluster's IP address and resources will automatically switch to other healthy computing nodes. This is achieved through the Oracle Clusterware and VIP (Virtual IP) mechanisms.

[0060] In a possible implementation, Oracle Clusterware detects the status of compute nodes through the heartbeat network (Interconnect) and the disk heartbeat (Voting Disk). If a heartbeat network failure is detected, the compute node corresponding to the failed heartbeat network is determined as a failed compute node. The failed compute node cannot respond to the heartbeat, and Oracle Clusterware will evict it from the Oracle RAC cluster. The VIP of the failed compute node will automatically migrate to other healthy compute nodes. After migration, Oracle Clusterware sends an ARP broadcast to update the MAC address table of the network device to ensure that traffic is routed to the new healthy compute node. The client connection will be automatically redirected to the new VIP, and usually, the client will not perceive the compute node failure. If TAF is configured, the client session will automatically migrate to the new compute node.

[0061] The switching example is as follows: Suppose there is an Oracle RAC cluster with two compute nodes: Compute Node 1: VIP = 192.168.1.101 Compute Node 2: VIP = 192.168.1.102 When Compute Node 1 fails: Oracle Clusterware detects the failure of Compute Node 1.

[0062] The VIP of Compute Node 1 (192.168.1.101) migrates to Compute Node 2.

[0063] The client connection is automatically redirected to the VIP of Compute Node 2 (192.168.1.101).

[0064] In some embodiments, the Oracle RAC heartbeat network detection and switching method further includes: if a heartbeat network failure is detected, an alarm notification is sent via email, text message, or the monitoring system.

[0065] In some embodiments, the Oracle RAC heartbeat network detection and switching method further includes: creating dashboards for network latency, packet loss rate, and bandwidth utilization using Grafana to display the network latency, packet loss rate, and bandwidth utilization.

[0066] In a possible implementation, data such as network latency, packet loss rate, and bandwidth utilization are collected using Prometheus, and dashboards are created using Grafana to display metrics such as network latency, packet loss rate, and bandwidth utilization.

[0067] Specifically, Grafana is an open-source data visualization and monitoring tool. Prometheus is an open-source systems monitoring and alerting toolkit.

[0068] In a specific embodiment, configure the network card and enable its advanced functions, such as hardware timestamp and traffic monitoring. Among them, traffic monitoring is used to monitor network traffic to obtain network traffic data. Use Bonding or Teaming to configure multi-network card bonding to ensure network redundancy. Develop a monitoring program using Python, C, or other languages to collect and analyze network traffic data. Deploy the monitoring program on each server to regularly collect network traffic data and send it to the central monitoring system. Use Prometheus to collect data and use Grafana to display network traffic data.

[0069] Specifically, Bonding is a technology that binds multiple physical network cards into a virtual network card, which is used to increase network bandwidth, achieve load balancing, and redundant backup. Teaming is a technology similar to Bonding, mainly used in RHEL 7 and its derivatives. Teaming realizes network card bonding through the user-space teamd daemon process.

[0070] It can be understood that using the hardware functions of the network card, such as hardware timestamp, for data collection and processing can reduce the CPU load. Using the network card hardware timestamp function can accurately measure network latency and jitter. Through multi-network card bonding and alternate network paths, it is beneficial to ensure the high availability of the network. When a network failure is detected, automatically switch to the alternate network or path to reduce manual intervention. Send alarm notifications via email, SMS, or the monitoring system to facilitate timely handling by operation and maintenance personnel. Display the network health status through tools such as Prometheus and Grafana to facilitate monitoring and analysis by operation and maintenance personnel. Centralize the monitoring data of multiple servers to the central monitoring system for unified management. By monitoring network traffic and latency, detect abnormal behaviors (such as DDoS attacks) in a timely manner to enhance security. When a security threat is detected, automatically trigger protection measures. Specifically, a DDoS (Distributed Denial of Service) attack is a malicious network attack where attackers control a large number of computers (usually called a "botnet") to send a large amount of traffic to the target server or network, causing the target system to be unable to respond to the requests of legitimate users normally, thus achieving the purpose of denial of service.

[0071] In a specific embodiment, simulate network failures (such as disconnecting the network cable, increasing latency) to test the detection and recovery capabilities of the system. Optimize the thresholds and switching logic based on the test results.

[0072] Enable the hardware timestamp, the example is as follows: ethtool -T eth0 Python uses psutil to obtain network traffic data. The example is as follows: import psutil net_io = psutil.net_io_counters(pernic=True) print(net_io['eth0']) The following is an example Python code for heartbeat packet detection: import subprocess def check_latency(interface, target_ip): result = subprocess.run(['ping', '-I', interface, '-c', '1', target_ip], stdout=subprocess.PIPE) output = result.stdout.decode('utf-8') if 'time=' in output: latency = float(output.split('time=')[1].split(' ms')[0]) return latency return None if __name__ == '__main__': latency = check_latency('eth0', '192.168.1.1') if latency and latency > 10: # Latency exceeds 10ms print(f"High latency detected: {latency} ms") The following is an example Python code for network card traffic monitoring: import psutil import time def monitor_network(interface): while True: net_io = psutil.net_io_counters(pernic=True) sent = net_io[interface].bytes_sent recv = net_io[interface].bytes_recv print(f"Sent: {sent}, Received: {recv}") time.sleep(1) if __name__ == '__main__': monitor_network('eth0') In a specific embodiment, after the Oracle RAC cluster is set up, use the crsctl status crs command to monitor the status of all nodes in the cluster. A visual interface is used to display various types of logs in the Oracle RAC cluster to the user. Among them, it includes instance logs, cluster logs, and ASM logs.

[0073] Database instance logs are important log files generated during the operation of the database, which are used to record the status, events, and error information of the database instance. The purpose of monitoring instance logs is to promptly discover and solve potential problems to ensure the stability, performance, and security of the database.

[0074] Cluster logs record the running status, events, and error information of cluster components such as Oracle Clusterware, database instances, ASM, etc.

[0075] ASM logs: Oracle ASM (Automatic Storage Management) is a component of the Oracle database used to manage storage. It manages database files, disk groups, and storage resources in an automated manner. ASM logs record the running status, events, and error information of the ASM instance and are an important basis for diagnosing and solving ASM-related problems.

[0076] Focus on monitoring cluster logs based on the heartbeat network. Among the cluster logs, mainly monitor CSSD logs. CSSD (Cluster Synchronization Services Daemon) is one of the core components of Oracle Clusterware and is responsible for managing heartbeat communication and cluster synchronization between nodes. CSSD logs are the main logs related to the heartbeat network. Log location: $GRID_HOME / log / <hostname> / cssd / . The main contents include: Node heartbeat status: Records the heartbeat status between nodes.

[0077] Network heartbeat status: Records the connection status and communication quality of the heartbeat network.

[0078] Node eviction events: Records node eviction events caused by heartbeat loss or network failures.

[0079] Disk heartbeat status: Records the read / write status and heartbeat information of the disk heartbeat (Voting Disk).

[0080] Error and warning messages: Records errors and warnings related to the heartbeat network.

[0081] Through the above logs, the running status of the database is monitored from the bottom layer, and the cluster status information of the database is detailedly displayed to the operation and maintenance personnel through a visualization interface. When a single computing node fails, the visualization interface alarms the operation and maintenance personnel, and the VIP is quickly switched to a normal computing node using the switching strategy to ensure the normal operation of the service.

[0082] Considering cost and architecture, 10 / 25GE network cards are sufficient to meet the requirements of the heartbeat network. The heartbeat network switching can only be from the faulty computing node A to the healthy computing node B and cannot be switched to the storage node.

[0083] Specifically, Oracle ASM (Automatic Storage Management) is an intelligent and automated storage management solution provided by the Oracle database, specifically used to manage Oracle database files. The Oracle database is a relational database management system (RDBMS) developed and sold by Oracle Corporation.

[0084] It can be seen that by binding the network cards to the computing nodes in the Oracle RAC cluster, the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter are obtained through the network card driver or API. Using the hardware timestamp, the network latency is calculated based on the round-trip time of the heartbeat packets, the packet loss rate is calculated based on the number of heartbeat packets sent and received, the bandwidth utilization rate is calculated based on the number of bytes received and sent, and it is determined whether a heartbeat network failure is detected based on the network latency, packet loss rate, and bandwidth utilization rate. If a heartbeat network failure is detected, the network card or virtual IP is switched. Through the network card hardware counter and hardware timestamp, it can be accurate to milliseconds, and a heartbeat network failure can be detected within milliseconds, reducing the impact of the heartbeat network failure on the system. The network traffic data includes the number of heartbeat packets sent and received, and the number of bytes received and sent. Through the network card hardware counter and hardware timestamp, the network traffic data can be collected in real time, which is beneficial to ensuring the timeliness of traffic monitoring.

[0085] Referring to Figure 2 , the embodiment of the present invention also provides an Oracle RAC heartbeat network detection and switching device for executing the above Oracle RAC heartbeat network detection and switching method. The Oracle RAC heartbeat network detection and switching device includes: A data acquisition module for obtaining the number of heartbeat packets sent and received, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter through the network card driver or API; A data analysis module for calculating the network latency, packet loss rate, and bandwidth utilization rate; the data acquisition module is connected to the data analysis module; A fault detection module for determining whether a heartbeat network failure is detected based on the network latency, packet loss rate, and bandwidth utilization rate; the data analysis module is connected to the fault detection module; An automatic recovery module for switching the network card or virtual IP; the fault detection module is connected to the automatic recovery module.

[0086] In some embodiments, the Oracle RAC heartbeat network detection and switching device further includes: a visualization module for creating dashboards of the network latency, packet loss rate, and bandwidth utilization rate using Grafana to display the network latency, packet loss rate, and bandwidth utilization rate; the automatic recovery module is connected to the visualization module.

[0087] In a possible implementation, a heartbeat network health detection system is built into the 10 / 25GE network card. The heartbeat network health detection system includes a data collection module, a data analysis module, a fault detection module, an automatic recovery module, and a visualization module. Among them, the data collection module and the data analysis module monitor the internal logs of the database in real time, such as real-time monitoring of indicators such as network latency, packet loss rate, and bandwidth utilization. The fault detection module detects network faults within milliseconds and triggers switching and alarms. The automatic recovery module automatically switches to a standby network or path when a fault is detected. The visualization module provides a visual interface for the network health status, facilitating operation and maintenance personnel to monitor and manage.

[0088] Referring to Figure 4 , the Oracle RAC heartbeat network detection and switching method is applied to the Oracle RAC heartbeat network detection and switching device, and the upper computer controls the Oracle RAC heartbeat network detection and switching device to detect the link data of the Oracle RAC cluster.

[0089] An embodiment of the present invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned Oracle RAC heartbeat network detection and switching method is implemented. The electronic device can be any intelligent terminal including a computer, etc.

[0090] An embodiment of the present invention also provides a storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned Oracle RAC heartbeat network detection and switching method is implemented.

[0091] It should be recognized that the method steps in the embodiments of the present invention can be implemented or implemented by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.

[0092] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) collectively executed on one or more processors, by hardware, or by a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.

[0093] Further, the method can be implemented in operable connection with any suitable type of computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or in communication with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and can be used to configure and operate the computer to perform the processes described herein when the storage medium or device is read by the computer. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the inventions described herein include these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0094] The computer program is capable of being applied to input data to perform the functions described herein, thereby transforming the input data to generate output data stored to non-volatile memory. The output information can also be applied to one or more output devices, such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a particular visual depiction of the physical and tangible objects produced on a display.

[0095] As described above, only the preferred embodiments of the present invention are given, and the present invention is not limited to the above-described embodiments. As long as the same means achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and variations.< / hostname>

Claims

1. An Oracle RAC heartbeat network detection switching method, applied to an Oracle RAC heartbeat network detection switching device, wherein the Oracle RAC heartbeat network detection switching device is used to detect the state of the heartbeat network of the Oracle RAC cluster, characterized in that: include: Bind a computing node in the Oracle RAC cluster to a network card, where a virtual IP is set, and obtain the number of sent and received heartbeat packets, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter through a network card driver or an API; Using hardware timestamps, network latency is calculated based on the round-trip time of the heartbeat packet; Calculate the packet loss rate according to the number of heartbeat packets sent and received; Calculating bandwidth utilization based on the number of bytes received and sent; Determining whether a heartbeat network failure is detected according to the network delay, the packet loss rate, and the bandwidth utilization rate; If the heartbeat network failure is detected, the network card or the virtual IP is switched.

2. An Oracle RAC heartbeat network detection switching method according to claim 1, characterized in that: The using of hardware timestamp to calculate network delay according to the round trip time of the heartbeat packet includes: Use the ping command to detect network connectivity by sending and receiving heartbeat packets; The network delay is calculated based on the round-trip time of the heartbeat packet using the hardware timestamp.

3. An Oracle RAC heartbeat network detection switching method according to claim 1, characterized in that: Determining whether a heartbeat network failure is detected according to the network delay, the packet loss rate, and the bandwidth utilization rate includes: Setting a first threshold, a second threshold and a third threshold for the network delay, the packet loss rate and the bandwidth utilization rate respectively; Determining whether the network delay, the packet loss rate, and the bandwidth utilization rate exceed the first threshold, the second threshold, and the third threshold, respectively; If one or more of the following situations occur in combination: the network delay exceeds the first threshold, the packet loss rate exceeds the second threshold, and the bandwidth utilization exceeds the third threshold, the heartbeat network failure is detected; otherwise, the heartbeat network failure is not detected.

4. The Oracle RAC heartbeat network detection switching method according to claim 1, characterized in that: If the heartbeat network failure is detected, switching the network card includes: Use ifenslave to bind multiple network cards to the computing node; Deploy the heartbeat network on the network card; If a heartbeat network failure is detected, the network card of the computing node corresponding to the failed heartbeat network is determined as a failed network card; The computing node is switched from the failed network card to a standby network card.

5. The Oracle RAC heartbeat network detection switching method according to claim 1, characterized in that: If the heartbeat network failure is detected, switching the virtual IP includes: Deploying the heartbeat network on the computing node; If a heartbeat network failure is detected, the computing node corresponding to the failed heartbeat network is determined as a failed computing node; Expelling the faulty computing node from the Oracle RAC cluster; Migrating the virtual IP of the failed computing node to a healthy computing node; Send ARP broadcasts through Oracle Clusterware to update the MAC address table of the network device to ensure that traffic is routed to the healthy computing node; The client redirects the connection to the virtual IP on the healthy computing node.

6. An Oracle RAC heartbeat network detection switching method according to claim 5, characterized in that: After the client redirects the connection to the virtual IP on the healthy computing node, if the heartbeat network failure is detected, switching the virtual IP further includes: Through TAF, the client session is migrated to the healthy computing node.

7. An Oracle RAC heartbeat network detection switching method according to claim 1, characterized in that: Also includes: A dashboard of the network latency, the packet loss rate, and the bandwidth utilization is created using Grafana to display the network latency, the packet loss rate, and the bandwidth utilization.

8. An Oracle RAC heartbeat network detection switching device, used to execute an Oracle RAC heartbeat network detection switching method according to any one of claims 1 to 7, characterized in that: The Oracle RAC heartbeat network detection switching device comprises: A data acquisition module, used for acquiring the number of the sent and received heartbeat packets, the number of bytes received and sent, and the round-trip time of the heartbeat packets in the network card hardware counter through the network card driver or the API; A data analysis module, used for calculating the network delay, the packet loss rate and the bandwidth utilization rate; the data acquisition module is connected to the data analysis module; A fault detection module, used for judging whether the heartbeat network fault is detected according to the network delay, the packet loss rate and the bandwidth utilization rate; the data analysis module is connected to the fault detection module; The automatic recovery module is used to switch the network card or the virtual IP; the fault detection module is connected to the automatic recovery module.

9. An Oracle RAC heartbeat network detection switching device according to claim 8, characterized in that: Also includes: A visualization module is used to create a dashboard of the network delay, the packet loss rate and the bandwidth utilization using Grafana to display the network delay, the packet loss rate and the bandwidth utilization; the automatic recovery module is connected to the visualization module.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, an Oracle RAC heartbeat network detection switching method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A network data protection apparatus and method

    CN109039825A

  • Main and standby network card switching method and device for network card binding mode and storage medium

    CN111726246A

  • Network anomaly monitoring and processing system

    CN119383056A

  • Method for detecting quality of wireless network based on airplay projection screen

    CN119967241A

  • Entangled links, transactions and trees for distributed computing systems

    US20240223360A1

Cited By

  • Bandwidth determination method and device of host channel adapter

    CN120567691A

  • Network switching method and device, electronic equipment, storage medium and program

    CN120658569A