Cluster identification method and device and medium
By obtaining host process information and IP addresses, and using graph database to display component cluster relationships, the problem of automatically identifying large-scale multi-component clusters is solved, and cluster management is achieved without manual intervention.
Patent Information
- Application Number
- CN202410016879.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology cannot automatically identify large-scale multi-component clusters, resulting in network managers needing to manually collect cluster information, which makes management inefficient.
By obtaining the process information of each host to be analyzed, identifying the local and remote IP addresses of each component, and using the graph database to display the component cluster relationship, realizing automatic cluster identification.
You can actively discover clusters without knowing cluster information in advance, which improves management efficiency and provides intuitive cluster relationship display.
Smart Images

Figure CN120263656A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates at least to the field of information technology, and in particular to a cluster identification method, a cluster identification device, and a computer-readable storage medium. Background Art
[0002] In a modern IT (Information Technology) environment, building and maintaining large-scale clusters is a common task. These clusters may contain various components and services, such as databases, message queues, container orchestration, etc. To effectively manage these clusters, it is crucial to identify the relationships and component distributions within the clusters.
[0003] Currently, the relationships between the components of each cluster are in the hands of the personnel who deployed the components. That is, currently, only those who built the cluster know about it. When non-building personnel need to count the cluster, they need to find the personnel who built the cluster and collect cluster information from them in order to finally identify the cluster.
[0004] That is to say, there is currently no method for "discovering" clusters that can automatically identify clusters. Especially for the management of large-scale multi-component clusters, it is impossible to directly identify and discover clusters, which brings inconvenience to network administrators. Summary of the Invention
[0005] The technical problem to be solved by this disclosure is to provide a cluster identification method, a cluster identification device, and a computer-readable storage medium to solve the problem of how to automatically identify clusters.
[0006] In a first aspect, this disclosure provides a cluster identification method, the method including:
[0007] Obtain the process information of each host to be analyzed;
[0008] Obtain each component running on each host to be analyzed, the local Internet Protocol (IP) address and the remote IP address connected by each component according to the process information;
[0009] Identify each component cluster according to the connection relationship between the local IP address and the remote IP address.
[0010] In a second aspect, this disclosure provides a cluster identification device, the device including:
[0011] A collection module, configured to obtain the process information of each host to be analyzed;
[0012] An analysis module, connected to the collection module, configured to obtain each component running on each host to be analyzed, the local Internet Protocol (IP) address and the remote IP address connected by each component according to the process information;
[0013] An identification module, connected to the analysis module, is used to identify each component cluster according to the connection relationship between the local IP address and the remote IP address.
[0014] In a third aspect, the present disclosure provides a computer-readable storage medium in which a computer program is stored. When the computer program is run by a processor, the above-mentioned cluster identification method is implemented.
[0015] The present disclosure provides a cluster identification method, a cluster identification device, and a computer-readable storage medium. By analyzing the process information of each host to be analyzed, the components running on each host to be analyzed, the local IP address and the remote IP address connected to each component are analyzed. According to the connection relationship between the local IP address and the remote IP address, each component cluster is identified. It is possible to actively discover the cluster and automatically identify the cluster without the need to know the cluster information in advance, which can provide a convenient management means for the managers of each host to be analyzed. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of a cluster identification method according to an embodiment of the present disclosure;
[0017] Figure 2 is a flowchart of another cluster identification method according to an embodiment of the present disclosure;
[0018] Figure 3 is a schematic diagram of a cluster identification result according to an embodiment of the present disclosure;
[0019] Figure 4 is a schematic structural diagram of a cluster identification device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the following will further describe the embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0021] It can be understood that the specific embodiments and drawings described herein are only for explaining the present disclosure and not for limiting the present disclosure.
[0022] It can be understood that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0023] It can be understood that, for the convenience of description, only the parts related to the present disclosure are shown in the drawings of the present disclosure, and the parts unrelated to the present disclosure are not shown in the drawings.
[0024] It can be understood that each unit and module involved in the embodiments of the present disclosure may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units and modules may also be integrated into one entity structure.
[0025] It is understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present disclosure may occur in an order different from that marked in the accompanying drawings.
[0026] It is understood that in the flowcharts and block diagrams of the present disclosure, the possible architectures, functions, and operations of systems, devices, equipment, and methods according to various embodiments of the present disclosure are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart may be implemented by a hardware-based system for implementing the specified function, or by a combination of hardware and computer instructions.
[0027] It is understood that the units and modules involved in the embodiments of the present disclosure may be implemented in software or in hardware. For example, the units and modules may be located in the processor.
[0028] Embodiment 1:
[0029] Currently, the common practice of cluster management is to manage cluster information through the means of operation and maintenance personnel management and information entry, which brings many problems: there are risks of loss and incorrect entry in manual entry, and as the number of machines increases, the accuracy rate of managing the cluster will become lower and lower; the workload of manually sorting out the component clusters every week is large; there is a lack of a cluster information display page that can view everything at a glance; the display of the cluster relationship on the traditional page is not friendly enough, and there is a lack of display of the relationship diagram.
[0030] The present disclosure automatically collects host information, and uses the advantages of the graph database to view all component clusters and cluster relationships at a glance. It can automatically collect host information, provide a cluster relationship diagram, display at a glance, and at the same time support different component query functions. The specific solution is as follows:
[0031] As Figure 1 shown, the present disclosure provides a cluster identification method, and the method includes:
[0032] S11. Obtain the process information of each host to be analyzed;
[0033] S12. Obtain each component running on each host to be analyzed, the local Internet Protocol IP address and the remote IP address connected by each component according to the process information;
[0034] S13. Identify each component cluster according to the connection relationship between the local IP address and the remote IP address.
[0035] In this embodiment, first, means for obtaining process information are used to obtain process-related information. Then, based on the goal of automatically identifying a component cluster, data analysis is performed on the process-related information. After the analysis, the cluster obtained according to the analysis results is displayed. The relationship of the displayed cluster can be shown using graph data. This analysis process can utilize the relevant information of the CPU (Central Processing Unit) to analyze the mutual dependency relationship between the servers (hosts) within the cluster. The dependency relationship means that for different IPs (Internet Protocol, and in this article, IP also refers to the IP address) of different servers, they will be connected to each other. The points of mutual connection form a set of clusters. The graph database can vividly display this dependency relationship. Different from the previous cluster monitoring, this embodiment is a new way of "identifying clusters".
[0036] As Figure 2 Figure 5 shows a more specific example of the present disclosure. This example identifies component clusters based on graph data, Linux operation commands, and big data processing technologies. Through graph data, users can query all cluster distributions in the system to achieve an intuitive cluster display effect. The specific process includes: S1. Collect all process-related information of all servers to generate a raw file. Use a shell script to collect the process-related information of the servers (one host is called a server. Usually, there are thousands of servers in a computer room, and each server runs many processes) and generate a process information file; S2. Perform line-by-line parsing and flattening processing on the raw file, that is, parse the process information file line by line according to the rules to generate structured data. Flattening processing is to change one line of data into multiple lines of data. One line of data in the raw file shows multiple open ports in a process, and there is also a field showing the IPs of other machines associated with these ports. According to the rules, one line of data is changed into multiple lines of data to obtain various possible connection relationships between the components running on the local machine and other machines; S3. Perform component analysis on the processed data, extract the component name, local IP, and remote IP, and generate a file, that is, analyze the structured data, classify the raw data according to different component keywords, and extract the local IP and remote IP to generate a file; Finally, to obtain the cluster, perform S4. Import the file into the graph database and set the relationship between the local IP and the remote IP, and S5. Use the graph data interface to display the cluster distribution map, that is, import the generated file using graph data, set the relationship according to the data in the file, and use the graph database query page to display the component cluster distribution; The above process can be continuously performed for cluster identification by setting S6. Timed scheduling.
[0037] In one embodiment, obtaining the process information of each host to be analyzed specifically includes:
[0038] Obtain the process ID (pid) of each process on each host to be analyzed, and obtain the process information in the process information folder corresponding to each pid on each host to be analyzed.
[0039] In one embodiment, obtaining the process information in the process information folder corresponding to each pid on each host to be analyzed specifically includes:
[0040] Obtain the component keywords corresponding to each pid according to the command line cmdline file in the process information folder corresponding to each pid;
[0041] Traverse the process information folder corresponding to each pid to obtain all file descriptors related to the socket for each pid;
[0042] Collect the ports being listened on and the established Transmission Control Protocol (TCP) connections corresponding to each pid according to all file descriptors related to the socket.
[0043] In this embodiment, relevant process information is obtained by using a simple means of obtaining process information. Specifically, it can be: developing a shell script to obtain the host process ID, and obtaining process-related information in the / proc / $pid folder according to the process ID, including the process command line cmd (cmdline file), local IP, multiple local listening ports, and remote connection IP information ( / proc / $pid / net / tcp file).
[0044] More specifically, the main task of step S1 is to complete information collection. The information to be collected includes the following key fields: host_ip (host IP), cmd, listen_port (listening port), tcp_connections (TCP connection). Specifically, for example:
[0045] The following command is used to obtain host_ip:
[0046] host_ip=$(cat / etc / sysconfig / network-scripts / ifcfg-bond* | grep IPADDR | awk -F'=' '{print $2}' | tr '\n' '@')
[0047] The function of this line of command is to extract the IP addresses from all ifcfg-bond network configuration files, concatenate them with the @ symbol, and then assign this string to the variable host_ip, that is, obtain the IP address of the host where the process is located; the sample field data is like: 10.161.54.92@;
[0048] The following command is used to obtain cmd:
[0049] cmd=$(tr -d '\0' < / proc / $pid / cmdline) / / Read data from the / proc / $pid / cmdline file to obtain and display the command line information of a specific process
[0050] The function of this line of command is to read the command line arguments of the process with the process ID $pid, convert the null characters in the arguments to spaces, and then store the processed command line in the variable cmd for obtaining and displaying the command line information of a specific process; sample field data is as follows:
[0051] / usr / jdk64 / jdk1.8.0_77 / / bin / java -Dzookeeper.log.dir= / data / v01 / app / zookeeperCH / bin / ..org.apache.zookeeper.server.quorum.QuorumPeerMain
[0052] ..(The sample is too long, and some unimportant characters are omitted using..);
[0053] listen_port is used to store the TCP port number that the process is listening on. The following are the steps to generate listen_port: Traverse the process's file descriptors: The script first traverses the / proc / $pid / fd directory to find all file descriptors related to sockets. The detailed information of the file descriptors can be obtained through the ls -l command, and the inode information representing the socket can be extracted through awk and grep; Analyze the TCP connection status: For each socket inode, the script reads the / proc / $pid / net / tcp and / proc / $pid / net / tcp6 files to find the TCP connection information related to this inode. These files contain the TCP connection status of the process; Extract the listening port: When the TCP connection status is 0A (hexadecimal, corresponding to the TCP LISTEN state, the script extracts the local port number. It can first split the hexadecimal address (including IP and port), and then only convert the port number to decimal format; Collect the listening ports: All the listening port numbers are collected and separated by spaces, and concatenated into a line and stored in the listen_port variable; sample field data is as follows: 4706218127 18129 (the port numbers opened by the current process, separated by spaces);
[0054] tcp_connections is used to store the TCP connection information established by the process. The following are the steps to generate tcp_connections: Traverse the file descriptors of the process: This step is the same as the first step of listen_port. Traverse the / proc / $pid / fd directory to find all socket-related file descriptors; Analyze the TCP connection status: This is the same as the second step of listen_port. Use the / proc / $pid / net / tcp and / proc / $pid / net / tcp6 files to find the TCP connection information related to each inode; Extract the TCP connection information: When the status of the TCP connection is 01 (in hexadecimal, corresponding to the ESTABLISHED state of TCP), the script extracts the local and remote IP addresses and port numbers, and converts the hexadecimal addresses to decimal format IP addresses and port numbers; Collect the TCP connection information: All connection information in the ESTABLISHED state (including local and remote IP addresses and port numbers) is collected and separated by spaces, and concatenated into a single line and stored in the tcp_connections variable; Field data example: 10.161.54.92:18129-10.161.54.93:42931 10.161.54.92:46275-10.161.54.94:18128
[0055] 10.161.54.92:18129-10.161.54.94:47517
[0056] 10.161.54.92:18127-10.161.54.94:37450 (TCP connection information related to the current process, separated by spaces. It can be seen that the ip:port numbers connected by "-" are a pair of connections. Before "-" is the local ip and port, and after "-" is the remote ip and port);
[0057] Through the above steps, a complete sample data can be obtained, such as:
[0058] 10.161.54.92@^ / usr / jdk64 / jdk1.8.0_77 / / bin / java -Dzookeeper.log.dir= / data / v01 / app / zookeeperCH / bin / .. org.apache.zookeeper.server.quorum.QuorumPeerMain
[0059] ..^47062 18127 18129^10.161.54.92:18129-10.161.54.93:42931 10.161.54.92:46275-10.161.54.94:18128
[0060] 10.161.54.92:18129-10.161.54.94:47517
[0061] 10.161.54.92:18127-10.161.54.94:37450。
[0062] In one embodiment, each component running on each host to be analyzed, the local IP address and the remote IP address connected by each component are obtained according to each process information, specifically including:
[0063] Each component and the classification to which each component belongs are obtained according to the component keyword;
[0064] The local IP address, local port number, remote IP address, and remote port number connected by each component are obtained according to the TCP connections established corresponding to each pid;
[0065] Whether the local port number corresponding to the local IP address is within the port range being listened on by each pid is used to determine whether each component is a client or a server locally.
[0066] In this embodiment, business logic analysis can be performed according to each process information. When the port of the local IP of tcp_connections is not within the port range of listen_port, it can be determined that the local is a client (CLIENT) connecting to the remote service. When the port of the local IP of tcp_connections is within the port range of listen_port, it can be determined that the local is a server (SERVER), that is, the IP and port after "-" are connected locally.
[0067] The business logic analysis results that can be obtained for the foregoing sample data are shown in Table 1:
[0068] Table 1 Process Information Analysis Results
[0069]
[0070] For example, for the first piece of data in Table 1:
[0071] 10.161.54.92:18129-10.161.54.93:42931, where 4 fields are extracted from the tcp field as:
[0072] local_ip: 10.161.54.92, local_port: 18129, remote_ip: 10.161.54.93, remote_port: 42931。
[0073] In one embodiment, the classifications to which each component belongs include:
[0074] Distributed application coordination service zookeeper, relational database management system mysql, remote dictionary service redis, distributed file storage-based database MongoDB, message queue, and / or container orchestration.
[0075] In this embodiment, as can be seen from the foregoing sample data, the component classification is zookeeper. For the cluster discovery of the cluster components for zookeeper, the component keyword is: org.apache.zookeeper.server.quorum.QuorumPeerMain. If you want to identify and discover other component clusters, you can identify other component keywords by yourself. For example, you can also analyze components such as mysql, redis, MongoDB, and other services.
[0076] In one embodiment, after obtaining the local IP address, local port number, remote IP address, and remote port number of each component connection according to the TCP connections established for each pid, the method further includes:
[0077] In response to the local IP address used by each component obtained according to the TCP connections established for each pid being different from the host IP address of each host to be analyzed, connect the host IP address before the local IP address to obtain a new local IP address.
[0078] In this embodiment, by checking the data, it is found that the local local_ip is inconsistent with the host host_ip. It is judged that it may be a container deployment situation. To make the IP accurate, connect the host_ip and local_ip with the "_" symbol, and do the same for the remote_ip.
[0079] In one embodiment, the method specifically includes:
[0080] Obtain the process ID pid of each host to be analyzed based on a shell script, and obtain each process information in the process information folder corresponding to each pid of each host to be analyzed;
[0081] Process each process information in the process information folder corresponding to each pid by using big data processing technologies spark and / or flink, and generate a structured data file;
[0082] Use big data processing technologies such as Spark, Flink, and / or Hive to analyze the structured data file to obtain each component running on each host to be analyzed, the local IP addresses and remote IP addresses connected by each component.
[0083] In the specific example as Figure 2 shown, step S1 may specifically be: by developing a shell script, obtain the host process ID, and obtain process-related information in the path / proc / $pid folder according to the process ID, including the process command line cmd (cmdline file), local IP, multiple local listening ports, and remote connection IP information ( / proc / $pid / net / tcp file); step S2 may specifically be: use big data processing technologies such as Spark and Flink to parse and clean the process information file line by line according to rules to generate structured data; step S3 is specifically: use big data processing technologies such as Spark, Flink, and Hive to analyze the structured data, classify the original data by components according to different component keywords, and extract local IP and remote IP data pairs to generate a file; Hive is suitable for writing SQL for analysis, and Spark and Flink can conveniently clean the data and convert one line into multiple lines, and finally generate a new data file.
[0084] In one embodiment, identifying each component cluster according to the connection relationship between the local IP address and the remote IP address specifically includes:
[0085] Load each host to be analyzed, each component running on each host to be analyzed, and the local IP addresses and remote IP addresses connected by each component into the graph database;
[0086] On the query page of the graph database, display each component and the connection relationship of each component based on the local IP address and the remote IP address, where each set of the same type of components with a connection relationship is a cluster.
[0087] In one embodiment, on the query page of the graph database, displaying each component and the connection relationship of each component based on the local IP address and the remote IP address specifically includes:
[0088] On the query page of the graph database, display each component with the local IP address corresponding to each component. In response to a certain component being a client locally, another component being a server, and there being a connection relationship between the IP addresses of the two, use the certain component as the starting point of the arrow and the other component as the end point of the arrow to connect the two.
[0089] In one embodiment, on the query page of the graph database, displaying each component and the connection relationship of each component based on the local IP address and the remote IP address specifically includes:
[0090] On the query page of the graph database, each component is displayed paged by the category to which the component belongs. In response to the local IP address of a certain component including the host IP address and the local IP address used by the certain component, on the query page of the graph database, the corresponding host to be analyzed is displayed by the host IP address, and the certain component running on the certain host to be analyzed is displayed by the local IP address used by the certain component.
[0091] In this embodiment, the finally displayed cluster example is Figure 3 as shown. The circles represent the components deployed on the hosts. The components with connection relationships among the IPs on multiple hosts form a set of clusters. Each component can be marked and displayed with its own local IP. The direction of the connection arrow can be used to distinguish whether the component is a client or a server. The components can be the same type of components that have been filtered out by sql (database language, Structured Query Language). All the components shown in the same graph are of the same type.
[0092] For example, the following sql can be executed to extract the relationship between the IPs:
[0093] select distinct CONCAT(s.host_ip,'_',s.local_ip),CONCAT(c.host_ip,'_',c.local_ip)
[0094] from psinfo s,psinfo c
[0095] where s.cmd like'%org.apache.zookeeper.server.quorum.QuorumPeerMain%' and s.direction in('SERVER')
[0096] and c.cmd like'%org.apache.zookeeper.server.quorum.QuorumPeerMain%' and c.direction in('CLIENT')
[0097] and s.local_ip=c.remote_ip and s.remote_ip=c.local_ip and s.local_port=c.remote_port and s.remote_port=c.local_port.
[0098] To further illustrate the present disclosure, another sample data analysis process is shown below:
[0099] The sample of the raw data obtained through the shell script is as follows:
[0100] 10.161.54.92@rizhi 1383 1 50Jun21?13-07:42:27 / usr / jdk64 / jdk1.8.0_77 / / bin / java -Dzookeeper.log.dir= / data / v01 / app / zookeeperCH / bin / ..org.apache.zookeeper.server.quorum.QuorumPeer Main
[0101] ..47062 18127 18129 10.161.54.92:18129 - 10.161.54.93:42931 10.161.54.92:46275 - 10.161.54.94:18128
[0102] 10.161.54.92:18129 - 10.161.54.94:47517
[0103] 10.161.54.92:18127 - 10.161.54.94:37450
[0104] Write a Spark program through the Spark offline analysis technology, store the file in the distributed storage HDFS, submit the Spark program to YARN to execute the computing task, and finally generate a new file. The data stored in the new file is a data file in which one piece of information judged according to the connection information is divided into multiple pieces of information. Specifically:
[0105] The fields of the analysis result table include: host_ip, pid, listen_port, local_ip, local_port, remote_ip, remote_port, direction;
[0106] The field analysis includes: host_ip (directly take the value); pid (directly take the value); listen_port (directly take the second-to-last field value); local_ip, (calculate); local_port, (calculate); remote_ip, (calculate); remote_port, (calculate); direction (calculate);
[0107] The logical rules are as follows:
[0108] 1. Range and logic of "direction":
[0109] CLIENT (tcp_connections is separated by spaces and "-", and listen_port is not equal to the first part separated by "-")[0]]
[0110] SERVER (tcp_connections is separated by spaces and "-", and listen_port is equal to the first part separated by "-")[0]]
[0111] NOListen (listen_port is empty and tcp_connections is not empty)
[0112] NOconnect (listen_port is not empty and tcp_connections is empty)
[0113] 2. Data filtering rules are as follows:
[0114] 2.1 Filter data if both listen_port and tcp_connections fields are empty
[0115] 2.2 Filter if the size of the column split by "^" is not equal to 13
[0116] 2.3 Filter if the host_ip field is empty
[0117] 3. Field logic:
[0118] tcp_connections is separated by "-",
[0119] If the 0th port is the same as the remote_port, then the value of "direction" is "server"
[0120] If the 0th port is different from the remote_port, then the value of "direction" is "CLIENT". If listen_port is empty and tcp_connections is not empty, then the value of "direction" is "NOListen"
[0121] If listen_port is not empty and tcp_connections is empty, then the value of "direction" is "NOconnect"
[0122] 4. Logic for storing field rules in the database:
[0123] tcp_connections is separated by spaces, and each data is stored as one record in the database;
[0124] Through the above operations, structured data was obtained, that is, a Hive table was created, and the corresponding data was stored in it. The table and data are as follows:
[0125] create EXTERNAL table psinfo(
[0126] host_ip string,
[0127] pid int,
[0128] cmd string,
[0129] listen_port string,
[0130] local_ip string,
[0131] local_port string,
[0132] remote_ip string,
[0133] remote_port string,
[0134] direction string)
[0135] 10.161.54.92@1383 / usr / jdk64 / jdk1.8.0_77 / / bin / java -Dzookeeper.log.dir= / data / v01 / app / zookeeperCH / bin / .. / logogger=....jmx.log4j.disable=trueorg.apache.zookeeper.server.quorum.QuorumPeerMain / data / v01 / app / zookeeperCH / bin / .. / conf / zoo.cfg 47062 18127 18129 10.161.54.921812710.161.54.94 37450SERVER
[0136] Business logic analysis can generate the table: psinfo
[0137] create EXTERNAL table psinfo(
[0138] host_ip string,
[0139] pid int,
[0140] cmd string,
[0141] listen_port string, -- Listening port list
[0142] local_ip string, -- Local IP
[0143] local_port string,
[0144] remote_ip string, -- Remote IP
[0145] remote_port string,
[0146] direction string -- Value range: SERVER, CLIENT, NOListen, NOconnect (NOListen means there is no port on the local machine but there is a connection, NOconnect means there is a port but no connection) )
[0148] Extract data pairs according to the following rules:
[0149] After checking the data, it is found that there is a situation where local_ip is inconsistent with host_ip. It is judged that it may be the container deployment situation. To make the IP accurate, connect host_ip and local_ip with the "_" symbol, and do the same for remote_ip (obtained according to the association table);
[0150] Exclude the interfering data of connecting to zk through the client, and only keep the zk cluster data;
[0151] The operations performed on the graph data are as follows:
[0152] 1. Create a table
[0153] LOAD CSV FROM "file: / / / 1.csv" AS row
[0154] MERGE(local:IP{address:row[0]})
[0155] MERGE(remote:IP{address:row[1]})
[0156] 2. Establish a relationship: Remote IP connects to local IP
[0157] LOAD CSV FROM "file: / / / 1.csv" AS row
[0158] MATCH(local:IP{address:row[0]})
[0159] MATCH(remote:IP{address:row[1]})
[0160] MERGE(remote)-[:CONNECTS_TO]->(local)
[0161] 3. Relationship display
[0162] MATCH(remote:IP)-[:CONNECTS_TO]->(local:IP)
[0163] WHERE local<>remote
[0164] RETURN local,remote
[0165] In this Embodiment 1, based on the file path of the server process number, all process-related information of the server can be extracted, and key information such as the IP and port for mutual calls between servers can be generated. Based on the graph data, the relationships between servers can be displayed, and further the relationships of the component clusters can be shown, which can facilitate sorting out the information related to the component clusters of all services.
[0166] Embodiment 2:
[0167] As Figure 4 shown, the present disclosure provides a cluster recognition device, and the device includes:
[0168] A collection module 11, configured to obtain the process information of each host to be analyzed;
[0169] An analysis module 12, connected to the collection module 11, and configured to obtain each component running on each host to be analyzed, the local Internet protocol IP address and the remote IP address connected by each component according to the process information;
[0170] An identification module 13, connected to the analysis module 12, and configured to identify each component cluster according to the connection relationship between the local IP address and the remote IP address.
[0171] In an implementation manner, the collection module 11 specifically includes:
[0172] A first collection unit, configured to obtain the process ID pid of each host to be analyzed,
[0173] A second collection unit, connected to the first collection unit, and configured to obtain the process information in the process information folder corresponding to each pid of each host to be analyzed.
[0174] In one embodiment, the second collection unit specifically includes:
[0175] A keyword collection unit, configured to obtain component keywords corresponding to each pid according to the command line cmdline file in the process information folder corresponding to each pid;
[0176] A descriptor collection unit, configured to traverse the process information folder corresponding to each pid to obtain all file descriptors related to the socket corresponding to each pid;
[0177] A listening port collection unit, connected to the descriptor collection unit, configured to collect the ports being listened on corresponding to each pid according to all file descriptors related to the socket;
[0178] A TCP connection collection unit, connected to the descriptor collection unit, configured to collect the Transmission Control Protocol (TCP) connections established corresponding to each pid according to all file descriptors related to the socket.
[0179] In one embodiment, the analysis module 12 specifically includes:
[0180] A component analysis unit, configured to obtain each component and the classification to which each component belongs according to the component keywords;
[0181] A connection analysis unit, configured to obtain the local IP address, local port number, remote IP address, and remote port number to which each component is connected according to the TCP connections established corresponding to each pid;
[0182] A judgment analysis unit, connected to the connection analysis unit, configured to judge whether each component is a client or a server locally according to whether the local port number corresponding to the local IP address is within the range of ports being listened on corresponding to each pid.
[0183] In one embodiment, the classification to which each component belongs includes:
[0184] Distributed application coordination service zookeeper, relational database management system mysql, remote dictionary service redis, distributed file storage-based database MongoDB, message queue, and / or container orchestration.
[0185] In one embodiment, the analysis module 12 further includes:
[0186] A host IP analysis unit, connected to the connection analysis unit, configured to, in response to the local IP address used by each component obtained according to the TCP connections established corresponding to each pid being different from the host IP address of each host to be analyzed, connect the host IP address before the local IP address to obtain a new local IP address.
[0187] In one embodiment:
[0188] The collection module 11 is specifically configured to obtain the process IDs pid of each host to be analyzed based on a shell script, and obtain each process information in the process information folder corresponding to each pid of each host to be analyzed;
[0189] The analysis module 12 includes:
[0190] The processing unit is configured to process each process information in the process information folder corresponding to each pid by using big data processing technologies spark and / or flink, and generate a structured data file;
[0191] The analysis unit is connected to the processing unit and is configured to analyze the structured data file by using big data processing technologies spark, flink and / or hive to obtain each component running on each host to be analyzed, the local IP address and the remote IP address connected by each component.
[0192] In one embodiment, the identification module 13 specifically includes:
[0193] The loading unit is configured to load each host to be analyzed, each component running on each host to be analyzed, and the local IP address and the remote IP address connected by each component into the graph database;
[0194] The display unit is connected to the loading unit and is configured to display each component and the connection relationship of each component based on the local IP address and the remote IP address on the query page of the graph database, wherein each set of like components having a connection relationship is a cluster.
[0195] In one embodiment, the display unit is specifically configured to:
[0196] Display each component on the query page of the graph database with the local IP address corresponding to each component, and in response to that one component is a client locally, another component is a server, and there is a connection relationship between the IP addresses of the two, connect the two with the one component as the starting point of the arrow and the other component as the end point of the arrow.
[0197] In one embodiment, the display unit is specifically configured to:
[0198] Display each component on the query page of the graph database by paging according to the classification to which each component belongs, and in response to that the local IP address of one component includes the host IP address and the local IP address used by the one component, display the corresponding host to be analyzed with the host IP address and display the one component running on the one host to be analyzed with the local IP address used by the one component on the query page of the graph database.
[0199] Example 5:
[0200] Embodiment 5 of the present disclosure provides a computer-readable storage medium storing a computer program, which, when run by a processor, implements the cluster recognition method as described in Embodiment 1 or implements the cluster recognition device as described in Embodiment 2.
[0201] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, computer program modules, or other data. The computer-readable storage medium includes, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.
[0202] In addition, the present disclosure may also provide a computer device including a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the cluster recognition method as described in Embodiment 1, and the computer device may be the cluster recognition device as described in Embodiment 2.
[0203] Among them, the memory is connected to the processor. The memory may adopt flash memory, read-only memory or other memories, and the processor may adopt a central processing unit or a single-chip microcomputer.
[0204] Embodiments 1-3 of the present disclosure provide a cluster recognition method, a cluster recognition device, and a computer-readable storage medium. By analyzing the process information of each host to be analyzed, the components running on each host to be analyzed, the local IP addresses and remote IP addresses connected by each component are analyzed. According to the connection relationship between the local IP address and the remote IP address, each component cluster is recognized, which can actively discover the cluster and automatically identify the cluster without the need to know the cluster information in advance, and can provide a convenient management means for the managers of each host to be analyzed.
[0205] It is understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present disclosure, but the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also regarded as the protection scope of the present disclosure.
Claims
1. A cluster recognition method, characterized in that, The method includes: Obtaining the process information of each host to be analyzed; Obtaining each component running on each host to be analyzed, the local Internet Protocol (IP) address and the remote IP address connected by each component according to the process information; Identifying each component cluster according to the connection relationship between the local IP address and the remote IP address.
2. The method according to claim 1, wherein Obtaining the process information of each host to be analyzed specifically includes: Obtaining the process ID (pid) of each process of each host to be analyzed, and obtaining the process information in the process information folder corresponding to each pid of each host to be analyzed.
3. The method according to claim 2, wherein Obtaining the process information in the process information folder corresponding to each pid of each host to be analyzed specifically includes: Obtaining the component keyword corresponding to each pid according to the command line (cmdline) file in the process information folder corresponding to each pid; Traversing the process information folder corresponding to each pid to obtain all file descriptors related to the socket corresponding to each pid; Collecting the ports being listened on and the established Transmission Control Protocol (TCP) connections corresponding to each pid according to all file descriptors related to the socket.
4. The method according to claim 3, wherein Obtaining each component running on each host to be analyzed, the local IP address and the remote IP address connected by each component according to the process information specifically includes: Obtaining each component and the category to which each component belongs according to the component keyword; Obtaining the local IP address, local port number, remote IP address and remote port number connected by each component according to the TCP connections established corresponding to each pid; Judging whether each component is a client or a server locally according to whether the local port number corresponding to the local IP address is within the range of the ports being listened on corresponding to each pid.
5. The method according to claim 4, wherein The categories to which each component belongs include: Distributed application coordination service zookeeper, relational database management system mysql, remote dictionary service redis, database MongoDB based on distributed file storage, message queue, and / or container orchestration.
6. The method according to claim 4, characterized in that, After obtaining the local IP address, local port number, remote IP address and remote port number connected by each component according to the TCP connections established corresponding to each pid, the method further includes: In response to the local IP address used by each component obtained according to the TCP connections established corresponding to each pid being different from the host IP address of each host to be analyzed, connecting the host IP address before the local IP address to obtain a new local IP address.
7. The method according to any one of claims 2-6, wherein: Obtaining the process ID (pid) of each process of each host to be analyzed based on a shell script, and obtaining the process information in the process information folder corresponding to each pid of each host to be analyzed; Processing the process information in the process information folder corresponding to each pid by using big data processing technologies spark and / or flink, and generating a structured data file; Analyzing the structured data file by using big data processing technologies spark, flink and / or hive to obtain each component running on each host to be analyzed, the local IP address and the remote IP address connected by each component.
8. The method according to any one of claims 1 to 6, characterized in that Identify each component cluster according to the connection relationship between the local IP address and the remote IP address, specifically including: Load each host to be analyzed, each component running on each host to be analyzed, and the local IP address and remote IP address connected by each component into the graph database; Display each component and the connection relationship of each component based on the local IP address and the remote IP address on the query page of the graph database, where each set of components of the same type with a connection relationship is a cluster.
9. The method according to claim 8, wherein Display each component and the connection relationship of each component based on the local IP address and the remote IP address on the query page of the graph database, specifically including: On the query page of the graph database, display each component with the local IP address corresponding to each component. In response to a certain component being a client locally, another component being a server, and there being a connection relationship between the IP addresses of the two, connect the two with the certain component as the starting point of the arrow and the other component as the end point of the arrow.
10. The method according to claim 8, wherein Display each component and the connection relationship of each component based on the local IP address and the remote IP address on the query page of the graph database, specifically including: On the query page of the graph database, display each component paged by the classification to which each component belongs. In response to the local IP address of a certain component including the host IP address and the local IP address used by the certain component, on the query page of the graph database, display the corresponding host to be analyzed with the host IP address and display the certain component running on the certain host to be analyzed with the local IP address used by the certain component.
11. A cluster recognition device, characterized in that, The device includes: A collection module for obtaining the process information of each host to be analyzed; An analysis module, connected to the collection module, for obtaining each component running on each host to be analyzed, the local Internet protocol IP address and remote IP address connected by each component according to the process information; An identification module, connected to the analysis module, for identifying each component cluster according to the connection relationship between the local IP address and the remote IP address.
12. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is run by a processor, the cluster identification method according to any one of claims 1-10 is implemented.