A method, terminal and system for monitoring an instant messaging server
By analyzing server rack temperature and link status, the stability of communication channels is dynamically assessed, heated channels are identified and replaced, and alarm processing is triggered. This solves the problem of insufficient temperature change trend identification in existing technologies and achieves efficient and stable communication link assurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN GOLDEN MILLIMETER TECHNOLOGY CO LTD
- Filing Date
- 2025-08-21
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies are prone to false alarms or missed alarms when faced with slow temperature rises or intermittent thermal interference. They lack the ability to analyze temperature change trends and make it difficult to identify communication quality degradation trends in a timely manner, leading to the accumulation of potential faults and an increased risk of communication interruption.
By collecting data on the internal temperature of the server rack and the ambient temperature, analyzing the characteristics of heat change, and combining the link status, the communication channels are classified into states, communication temperature impact markers are generated, channels affected by heat are identified, channel stability is evaluated, differences in communication availability are analyzed, a channel performance ranking list is generated, the main path fluctuation status is dynamically evaluated, and channels are replaced when necessary, triggering server node alarm linkage processing.
It improves the response efficiency of communication links under thermal interference, ensures that channel performance is quantifiable and traceable, realizes full-link identification and multi-dimensional early warning of communication link operation status, and enhances communication stability and the timeliness of early warning.
Smart Images

Figure CN120956698B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of server technology, and in particular to a method, terminal and system for monitoring instant messaging servers. Background Technology
[0002] In the field of server technology, this involves real-time measurement, change sensing, and over-temperature response processing of the target environment or internal temperature of equipment. Its core aspects include temperature data acquisition and transmission, threshold setting and judgment, alarm signal generation and linkage output. Temperature is sensed through thermistors, and alarm information is transmitted to the management terminal through a judgment mechanism and communication methods to achieve continuous monitoring of the operating status of critical equipment. Traditional real-time temperature monitoring and alarm systems for electrical cabinets address abnormal temperature rises within power control equipment caused by poor heat dissipation, electrical faults, or environmental anomalies. They acquire real-time temperature values inside the cabinet by using thermistors or resistance temperature detectors (RTDs) as detection units, and determine whether the temperature exceeds a set standard based on fixed threshold logic. Then, an audible and visual alarm is triggered by a relay-driven buzzer or warning light. Some systems also use RS485 communication to upload temperature data to a server for monitoring. RS485 is a physical layer communication standard that uses differential signal transmission and supports long-distance, high-interference-resistant serial communication.
[0003] Existing technologies rely on fixed thermistors to collect ambient or internal equipment temperatures in real time, primarily using static thresholds as the basis for alarm judgment. They lack the ability to analyze temperature change trends, leading to false alarms or missed alarms in situations with slow temperature rises or intermittent thermal interference. Alarm responses are mostly based on single temperature exceedance signals driving audible and visual alarms, lacking correlation assessment between actual communication stability and channel performance. This makes it difficult to discern communication quality degradation trends from temperature changes, limiting the timeliness and accuracy of early warning processing. Uploading data to the server via RS485 only achieves basic data aggregation, lacking in-depth analysis capabilities for channel operating status, interruption characteristics, and historical performance changes, making it difficult to support systematic classification and control of communication link status. For example, when the internal temperature of the equipment rises slowly due to poor local heat dissipation, even before reaching the fixed threshold, it has already caused a decline in communication quality. Traditional systems cannot identify and respond in a timely manner, leading to the accumulation of potential faults, alarm delays, and increased risk of communication interruption. These shortcomings limit the dynamic identification capabilities and link stability assurance levels of existing systems in multi-source interference scenarios. Summary of the Invention
[0004] This application provides a method, terminal, and system for monitoring instant messaging servers, which addresses the shortcomings of existing technologies.
[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, a method for monitoring instant messaging servers is provided, including the following steps: Collect data on the internal temperature of the server rack and the ambient temperature, analyze the characteristics of heat change, and classify the status of the server communication channel in conjunction with the link status to generate communication temperature impact markers. Based on the communication temperature impact marker, identify the communication channels affected by heat, retrieve the corresponding channel's operating parameters and interruption information, evaluate channel stability, analyze communication availability differences, and generate a channel performance ranking list. Based on the channel performance ranking list, analyze the score change trend between adjacent periods, determine the performance stability, and mark whether there is fluctuation in the current main path, generating a main path fluctuation indicator status. Based on the main path fluctuation status, select a stable and high-scoring backup path from the channel performance sorting list, replace the channel number and update the current primary channel identifier, and generate a backup link activation path number. Based on the combined status of the backup link activation path number and the communication temperature influence marker, the alarm linkage processing of the server node is determined and triggered, the channel number, server status and alarm information are output, the communication link operation status is identified, and a server communication channel early warning status identifier is generated.
[0006] In one possible design, the communication temperature impact marker includes temperature change magnitude, link status classification, and channel impact level; the channel performance ranking list includes channel score, availability level, and stability coefficient; the main path fluctuation identifier includes score change trend, main path fluctuation level, and time period label; the backup link activation path number includes replacement channel number, backup path priority, and channel update identifier; and the server communication channel warning status identifier includes channel number, server operating status, and alarm information encapsulation content.
[0007] In one possible design scheme, the specific steps for collecting server rack internal temperature and ambient temperature data, analyzing heat change characteristics, and classifying the server communication channel status in conjunction with link status to generate communication temperature impact markers are as follows: Collect temperature variation data and ambient temperature information of monitoring nodes inside the server rack. After performing time series standardization processing on the temperature data inside the server rack and the ambient temperature data, calculate the difference between the standardized value of the temperature inside the server rack and the standardized value of the ambient temperature at the same time point, and calculate the ratio of the standardized value of the temperature inside the server rack to the standardized value of the ambient temperature to obtain the degree of temperature deviation per unit time and generate the node temperature offset rate. Based on the node temperature offset rate, combined with the changes in link connection rate and delay time at the corresponding time, the communication link nodes that are consistent with the temperature offset direction are identified, and the corresponding relationship is compared to generate the temperature link correspondence coefficient. Based on the temperature link correspondence coefficient, combined with the connection stability and temperature offset rate of the link node, it is determined whether the node belongs to the temperature-affected communication channel, a status classification index is established, and a communication temperature-affected label is generated.
[0008] In one possible design scheme, the specific steps of identifying heat-affected communication channels based on the communication temperature impact marker, retrieving the corresponding channel's operating parameters and interruption information, evaluating channel stability, analyzing communication availability differences, and generating a channel performance ranking list are as follows: Based on the heated communication channel identified by the communication temperature influence marker, the operating current, conductor temperature, contact voltage and bit error rate parameters of the channel are called up, the data are extracted and the variation range, range and fluctuation frequency are calculated, the channel stability difference is summarized, and the channel state offset is generated. Based on the channel status offset, combined with the interruption time, recovery delay and number of interruptions, the parameter offset degree and interruption situation are judged accordingly to identify the unstable interval and generate the communication channel availability difference rate. The communication channel availability difference rate is called, and the bit error rate, operating current and recovery delay parameters are combined and sorted according to the average value after channel grouping to generate communication channel performance ranking values.
[0009] In one possible design scheme, the specific steps of analyzing the score change trend between adjacent periods based on the channel performance ranking list, determining performance stability, marking whether there is fluctuation in the current main path, and generating the main path fluctuation indicator status are as follows: Based on the primary path data in the channel performance ranking list, obtain the corresponding score values for adjacent periods, calculate the main path score changes, and aggregate them by path number to generate score fluctuation amplitude values. Based on the score fluctuation range value and the current period score value, determine whether the path score change exceeds the set benchmark, mark the fluctuation situation, and obtain the main path fluctuation status label; The main path fluctuation status label is invoked, and combined with the ranking of the main path in the current period, the top-ranked main paths are filtered out. The fluctuation status labels are then classified and aggregated to obtain the main path fluctuation identifier status.
[0010] In one possible design scheme, the specific steps of selecting a stable and high-scoring backup path from the channel performance ranking list based on the main path fluctuation status, replacing the channel number and updating the current primary channel identifier to generate a backup link activation path number are as follows: Based on the identification result that the main path fluctuation status is unstable, the channel number, score value and stable status parameter in the channel performance ranking list are called to filter the channel number with a stable status, and the score values are sorted to obtain the channel number with the highest score value and generate the score priority channel number. Based on the scoring priority channel number, the current primary channel number and mapping information are called to perform a channel number replacement operation, and the primary channel identifier status is updated synchronously to generate a primary channel update number; The primary channel update number is called to obtain the corresponding activation path number in the link channel configuration table. The path number is assigned to the backup link activation path identifier parameter and synchronously written to the current link path number record field to generate the backup link activation path number.
[0011] In one possible design scheme, the specific steps for determining and triggering server node alarm linkage processing based on the combined state of the backup link activation path number and the communication temperature influence marker, outputting the channel number, server status, and alarm information, identifying the communication link operating status, and generating a server communication channel early warning status identifier are as follows: Based on the combined status of the backup link activation path number and the communication temperature impact marker, the channel number, link identifier and temperature data are extracted. After completing the mapping between the channel and the link, the communication temperature impact marker and the link status parameters are combined and classified, and it is determined whether the communication interference threshold is exceeded, and an abnormal communication status judgment value is generated. Based on the communication status anomaly determination value, the server node channel number and communication anomaly marker are matched, and the server status parameters are combined to make a judgment, extract the server number and running status that meet the conditions, and compare them with the alarm response rule table to generate the server node alarm trigger level value. Based on the alarm trigger level value of the server node, the channel number, server operating status and alarm level are summarized, the operating status identifier number is output by referring to the communication link status identifier encoding table, and the encapsulation information is integrated to generate the server communication channel early warning status identifier value.
[0012] In one possible design, the communication interference threshold is a boundary value used to determine whether the communication link is in an abnormal state, and is usually set according to the stability index of the communication link under different loads, temperatures and interference environments. The server status parameters are a set of multi-dimensional parameters that reflect the current operating status of the server node, and are used to simultaneously characterize server load, resource consumption, and link response indicators. The operating status identifier number is a standardized encoding result representing the current operating status of the communication link, used to abstract the communication status into identifiable and transmissible structured information.
[0013] In a second aspect, a terminal for monitoring an instant messaging server is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for monitoring an instant messaging server.
[0014] Thirdly, a system for monitoring an instant messaging server is provided, the system for monitoring the instant messaging server being used to perform the above-described method for monitoring the instant messaging server, including: The link thermal monitoring module is used to acquire temperature data and ambient temperature monitoring values of multiple areas inside the server rack, detect packet loss rate, latency and link status parameters of the communication channel, identify the thermal interference level of the channel based on the correspondence between temperature data and link status, and mark it with the channel number to generate communication temperature impact markers. The channel stability assessment module is used to retrieve the stability-related parameters of the channel during operation, including transmission delay, retransmission count and interruption records, based on the heated channel number identified in the communication temperature influence marker. The module then judges the communication stability differences of the channel in the current state based on the parameters and generates a channel performance ranking list. The path fluctuation identification module is used to analyze the score changes of the channel in a continuous period based on the channel number of the primary path in the channel performance ranking list, determine whether the score change exceeds the stability judgment standard, and generate the primary path fluctuation identification status. The primary link replacement module is used to select channel numbers with high scores and stable status from the channel performance ranking list based on the channel numbers that have been determined to be unstable in the primary path fluctuation identification status, update the identification of the current primary channel, and generate backup link activation path numbers. The channel early warning linkage module is used to collect server node status, response time and alarm status information based on the channel number enabled in the backup link activation path number and the heat interference level of the corresponding link in the communication temperature influence mark, and to jointly determine the operational risk level of the current link based on the data, and generate a server communication channel early warning status identifier.
[0015] In summary, based on the above methods, terminals, and systems, it can be concluded that: By collecting rack and ambient temperatures and combining them with link status, the accuracy of channel heat identification is improved. Communication stability is dynamically evaluated based on channel operating parameters and interruption information, enhancing the ability to analyze availability differences. The fluctuation status of the main path is determined by the trend of score changes, ensuring that channel performance is quantifiable and traceable. Based on performance ranking, backup path switching and channel identifier updates are realized, improving the response efficiency of communication under thermal interference. Temperature effects and backup link status are jointly used to trigger alarms, outputting server status and alarm information, realizing full-link identification and multi-dimensional early warning of communication link operation status. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a method for monitoring an instant messaging server according to an embodiment of this application. Figure 2 This is a schematic diagram illustrating the specific process of generating communication temperature influence markers in the method for monitoring an instant messaging server according to an embodiment of this application. Figure 3 This is a schematic diagram illustrating the specific process of generating a channel performance ranking list in the method for monitoring an instant messaging server according to an embodiment of this application. Figure 4 This is a schematic diagram illustrating the specific process of generating the main path fluctuation identifier status in the method for monitoring an instant messaging server according to an embodiment of this application. Figure 5 This is a schematic diagram illustrating the specific process of generating a backup link activation path number in the method for monitoring an instant messaging server according to an embodiment of this application. Figure 6 This is a schematic diagram illustrating the specific process of generating a server communication channel early warning status identifier in the method for monitoring an instant messaging server according to an embodiment of this application. Figure 7 This is a schematic diagram of the framework of a monitoring instant messaging server system according to an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] Figure 1 This is a flowchart illustrating a method for monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 1 This application provides a technical solution: a method for monitoring an instant messaging server, comprising the following steps: S1: Collect temperature variation data inside the server rack and ambient temperature information, analyze the current heat change characteristics, and combine with link status data to classify the status of the server communication channel and generate communication temperature impact markers.
[0019] S2: Based on the communication channels marked as affected by heat in the communication temperature influence marker, retrieve the corresponding channel's operating parameters and interrupt information, perform stability assessment on the channel, identify the differences in communication availability of the channel in the current state, and generate a channel performance ranking list.
[0020] S3: Based on the primary path data in the channel performance ranking list, analyze the score change trend between adjacent periods, determine the performance stability, mark whether there is fluctuation in the current primary path, and generate the primary path fluctuation status. If the score of the primary communication path changes beyond the set tolerance multiple times within a continuous period, or if the score trend shows a continuous oscillation, the path is judged to be in an unstable state.
[0021] S4: Based on the identification result that the main path fluctuation status is unstable, select the backup path with stable status and priority in score from the channel performance ranking list, perform channel number replacement and update the current primary channel identifier, and generate the backup link activation path number.
[0022] S5: Based on the joint status judgment of the backup link activation path number and the communication temperature influence mark, trigger the alarm linkage processing of the server node, encapsulate and output the channel number, server status and alarm information, and mark the operating status of the corresponding communication link to generate a server communication channel early warning status mark. The communication temperature impact marker includes the temperature change range, link status classification, and channel impact level; the channel performance ranking list includes the channel score, availability level, and stability coefficient; the main path fluctuation identifier includes the score change trend, main path fluctuation level, and time period label; the backup link activation path number includes the replacement channel number, backup path priority, and channel update identifier; and the server communication channel warning status identifier includes the channel number, server operating status, and alarm information encapsulation content.
[0023] Figure 2 This is a schematic diagram illustrating the specific process of generating communication temperature impact markers in the method for monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 2 The specific steps of S1 are as follows: S101: Collect temperature variation data and ambient temperature information of monitoring nodes inside the server rack. After performing time series standardization processing on the temperature data inside the server rack and the ambient temperature data, calculate the difference between the standardized value of the temperature inside the server rack and the standardized value of the ambient temperature at the same time point, and calculate the ratio of the standardized value of the temperature inside the server rack to the standardized value of the ambient temperature to obtain the degree of temperature deviation per unit time and generate the node temperature offset rate. The process involves collecting temperature variation data from various nodes within the server rack and ambient temperature information. This requires deploying multiple sensor nodes at different locations within the rack. Each sensor module records the temperature value within its node and the ambient temperature value over the same time period. All sensors should be set to sample at fixed time intervals, such as every 5 seconds, and upload the data to the data processing unit via a unified interface. The processing unit constructs complete time series for both node and ambient temperatures. Within each sampling period, the data is standardized by subtracting the average value of its respective series from each data set and then dividing by the standard deviation. After standardization, the difference between the node temperature series and the ambient temperature series at corresponding time points is calculated, along with the ratio between the two values. This ratio represents the node temperature at that time point relative to the ambient temperature. The degree of environmental deviation, for example, if the temperature of a node at a certain time is 45.2℃, while the corresponding ambient temperature is 30.5℃, and if the historical average temperature of the node is 40℃ with a standard deviation of 3℃, and the historical average ambient temperature is 28℃ with a standard deviation of 2.5℃, then the standardized value of the node at that time can be calculated to be approximately 1.73, and the standardized value of the ambient temperature is approximately 1.0, with a difference of 0.73. At the same time, the temperature ratio is approximately 1.48. Multiplying these two values, we get the temperature deviation rate at that time point as 1.0804. The entire process of standardization, difference, and ratio calculation is based on the historical data of the sensing node and the environmental sensor, and can be calculated in real time according to the system configuration. The calculation results will be recorded in the node monitoring database for subsequent data analysis.
[0024] S102: Based on the node temperature offset rate, combined with the changes in link connection rate and delay time at the corresponding time, determine the communication link nodes that are consistent with the temperature offset direction, compare them according to the correspondence, and generate temperature link correspondence coefficients. As a specific embodiment of this application, the specific calculation formula for determining the communication link node that is consistent with the temperature offset direction is as follows: ; Calculate the temperature link correspondence coefficients Based on the corresponding relationship, a comparison is made to generate the temperature link correspondence coefficient; in, Indicates the first The node at the th Temperature offset rate at each time point Indicates the first The average temperature offset rate of each node during this period. Indicates the first The standard deviation of the temperature deviation rate of each node within this period Indicates the first The node at the th Link rate change or latency change at each point in time Indicates the first The average link rate or latency of each node during this period. Indicates the first The standard deviation of the link rate or delay of each node within this period This indicates the number of sampling points within the statistical period. Indicates the first The change in link packet loss rate of each node within the current period Indicates the first The average link packet loss rate of each node during this period. This represents a very small positive constant set to avoid a denominator of zero. Indicates the first Temperature link correspondence coefficients of each node in the current cycle; Temperature offset value For a certain server node sensing module in the 1st The relative offset value obtained by normalizing the standard deviation of the mean of the historical ambient temperature at each time point to the temperature at that node. Through sensing sampling every 5 seconds, node j forms 120 valid sampling points within 10 minutes. For example, time points... to The measured temperature values were 47.2, 46.8, 47.5, 47.9, and 48.3 degrees Celsius, respectively. The corresponding historical average temperature was 42.0 degrees Celsius, and the standard deviation of the historical temperature samples was 2.5 degrees Celsius. Therefore, we have: ; The average offset rate is: ; The standard deviation is obtained by taking the square root of the sample variance, as follows: ; Link parameter items Indicates the link rate at the 1st The values are the changes at specific time points, in Mbps, and were sampled from the link performance monitoring device. The obtained values are as follows: 91.0, 92.5, 89.8, 88.2, 90.1.
[0025] The average value is: ; The standard deviation is calculated as follows: ; The standardized offset product is as follows: ; Similarly, calculate the product of the first 5 time points and sum them: ; Packet loss rate change value This data comes from the packet loss statistics interface of the network monitoring system, and the unit is percentage change. The current period's packet loss rate is 1.6%, and the average of the previous period was 1.0%. Therefore: ; ; Finally, substituting into the formula: ; The results indicate that the standardized correlation between temperature changes and link rate changes within the current period of node j is 0.1599. This value, as the temperature-link correspondence coefficient, reflects the degree of directional consistency and is subsequently used in the link channel classification process to determine whether a node is classified as a temperature-affected communication channel. If this value is below a set judgment threshold (e.g., 0.3), the node's temperature offset rate is considered to have no significant impact on link state changes, and it will not be classified as a temperature-affected node in the state classification. If the value exceeds 0.7, it is considered a strongly affected channel and participates in the main path state score adjustment and link replacement logic. This formula measures the statistical directional consistency between temperature offset rate and link rate or delay variation by calculating the normalized sum of their products. Subtraction between parameters is used to determine the deviation of each term from its periodic average, reflecting the strength of single-point changes; then, division by the standard deviation achieves dimensionlessness, allowing comparisons of physical quantities with different dimensions on a unified scale. The normalized product of temperature and link represents the synergy between the two variables, with positive values indicating the same direction and negative values indicating the opposite direction. Summing the products at multiple time points accumulates the synergistic effect at different time points. Dividing the sum by the number of sampling points represents the average correlation strength, ensuring comparability of values across different sampling scales. The square root in the denominator adjusts the weight of packet loss rate changes on the overall result; a higher packet loss rate results in a greater impact on the final value after square root compression, controlling the amplification effect of link instability on temperature-link consistency and preventing abnormal divergence. The absolute value operation outputs the positive correlation strength, removing the influence of sign direction and ensuring that the final coefficient only represents the degree of stable correlation between thermal characteristics and link state changes. The overall structural design emphasizes multidimensional standardization, normalized weighting, and nonlinear adjustment characteristics to suppress the impact of link fluctuations; The temperature-link correlation coefficient quantifies the correlation between temperature changes of a server node and changes in its communication link status within a specific period. Its value reflects whether the temperature offset rate and changes in link rate or latency exhibit a synchronous trend at multiple time points. This coefficient comprehensively considers the amplitude of temperature fluctuations, the stability of the link status, and the disturbance effect of packet loss rate. It eliminates dimensional differences between different physical quantities through standardization, calculates the intensity of coordinated changes at multiple time points through normalization multiplication, and then performs nonlinear adjustment based on the degree of link packet loss changes, ultimately yielding a dimensionless index characterizing the strength of temperature's effect on link status. A higher value indicates that temperature fluctuations are more likely to directly affect link performance, while a lower value indicates a lack of significant dynamic correlation between the two. It is typically used to determine whether the current node belongs to a temperature-sensitive channel, providing a basis for communication path classification and subsequent path switching mechanisms.
[0026] S103: Based on the temperature link correspondence coefficient, combined with the connection stability and temperature offset rate of the link node, determine whether the node belongs to the temperature-affected communication channel, establish a status classification index, and generate a communication temperature-affected marker. Based on the temperature-link correlation coefficient of each node, combined with link stability and node temperature offset rate, it is determined whether the node belongs to a temperature-affected communication channel. Stability is calculated by the ratio of the duration of the communication connection to the number of disconnections. For example, if the connection duration is 10 hours and only one disconnection occurs during that time, the stability of the node is 36,000 seconds per event, which is much higher than the generally set stability standard threshold of 1,000 seconds per event. Combined with the node's temperature offset rate of 1.2 and the correlation coefficient with link performance of 0.997, it can be considered strongly correlated and has stable communication conditions. Therefore, the node is classified into the temperature-affected communication channel range, and the corresponding status flag is set to 2, indicating a strong impact. If a node has a high temperature offset rate but the communication connection is unstable or there is no obvious correlation, it is not set as an affected node. All judgment results form a status classification index in the data platform and are synchronously written into the status flag field for identification.
[0027] Figure 3 This is a schematic diagram illustrating the specific process of generating a channel performance ranking list for the method of monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 3 The specific steps of S2 are as follows: S201: Based on the heated communication channel identified by the communication temperature influence marker, call the channel's operating current, conductor temperature, contact voltage and bit error rate parameters, extract the data and calculate the range of change, range and fluctuation frequency, summarize the channel stability differences, and generate the channel state offset. After identifying communication channels significantly affected by temperature, a sensor system monitors parameters such as operating current, conductor temperature, contact voltage, and bit error rate. Data is retrieved according to device number or channel identifier and analyzed in specified periodic units. For example, within a certain time period, the operating current data is extracted to be between 28A and 32A, the conductor temperature fluctuates between 72℃ and 78℃, the contact voltage is between 0.28V and 0.32V, and the bit error rate is one in a million per bit. After extraction, the range is calculated by the difference between the maximum and minimum values for each parameter. The range of the operating current is 4A, the conductor temperature is 6℃, the voltage is 0.04V, and the bit error rate is 9 times 10 to the power of -7. The method involves counting the frequency of hourly fluctuations exceeding the average value: current fluctuates 3 times per hour, conductor temperature fluctuates 2 times per hour, and bit error rate fluctuates 1 time per hour. This frequency information is incorporated into the fluctuation assessment. Subsequently, the fluctuation degree of various parameters is normalized to form a unified score. Then, the stability change amplitude of each channel is summarized by weighting each parameter. Within a set period, the scores of the same channel at different times are continuously compared to calculate the dispersion of the score sequence, which is the state offset of the channel. For example, if the score difference value of a certain channel is continuously high during the monitoring period, the offset reaches 0.8, while another channel has an offset of only 0.2 in the same period. The channel with the larger offset shows significant inconsistency in data fluctuation.
[0028] S202: Based on the channel status offset, combined with the interruption time, recovery delay and number of interruptions, the parameter offset degree and interruption situation are judged accordingly, the interval of inconsistency in stability is identified, and the communication channel availability difference rate is generated. The state offset of each communication channel is correlated with its historical outage behavior. For example, a channel with an offset of 0.8 has experienced 4 outages in its past records, with each outage lasting an average of 300 seconds and a recovery delay of 80 seconds. Another channel with an offset of 0.4 has experienced 2 outages, each lasting 150 seconds and with a delay of 40 seconds. The state offset of each channel is divided into three levels: above 0.7 is classified as high variability, the middle range is 0.3 to 0.7, and below 0.3 is classified as low variability. Based on this classification, the outage behavior is aligned with the offset range by time and frequency. The segments where outages occur in concentrated periods are marked and defined as stability inconsistency intervals. The total duration of such intervals and the number of outage events they contain are counted and compared with the total duration and total number of outages in the entire cycle. The outage percentage of each channel in the inconsistency interval is calculated. For example, a channel has an inconsistency interval of 1200 seconds in a total duration of 7200 seconds, with a total of 3 outages, accounting for 75% of the total number of outages. Further, the availability difference rate of this channel is found to be 12.5%, and the high percentage indicates that its communication stability fluctuates significantly in different segments.
[0029] S203: Call the communication channel availability difference rate, combine the bit error rate, operating current and recovery delay parameters, sort them according to the average value after channel grouping, and generate communication channel performance ranking values; The availability difference rate of communication channels is used as the evaluation basis, and integrated with the bit error rate, operating current, and recovery delay parameters of each channel. For example, channel A has an availability difference rate of 12.5%, a bit error rate of one millionth of a bit, a current of 30A, and a recovery delay of 80 seconds. The three types of parameters are processed separately and weighted according to the set weights: bit error rate 40%, current 30%, and recovery delay 30%. The combined score of channel A is 33 points. Channel B has a difference rate of 8%, a bit error rate of one in twenty millionths, a current of 25A, and a delay of 50 seconds, with a weighted score of 22.5 points. Channel C has a difference rate of 2%, a bit error rate of one in one hundred millionths, a current of 20A, and a delay of 20 seconds, with a weighted score of 12 points. According to the total weighted score, channel C has the best performance, followed by B, and then A, thus forming a performance ranking list of communication channels.
[0030] Figure 4 This is a schematic diagram illustrating the specific process of generating the main path fluctuation identifier status in the method for monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 4 The specific steps of S3 are as follows: S301: Based on the primary path data in the channel performance ranking list, obtain the corresponding score values for adjacent periods, calculate the main path score changes, and aggregate them by path number to generate score fluctuation amplitude values. Based on the primary path numbers in the channel performance ranking list, the score values of each primary path are extracted over several consecutive periods. The score differences between adjacent periods are calculated pairwise to form a score change sequence for that path. For example, the path numbered M1 has scores of 82.3, 78.5, and 75.6 in three periods, with adjacent score differences being... 3.8 and 2.9 After traversing all main path numbers to obtain such difference data, the score fluctuation range of each path is further calculated. This range can be approximately estimated by the standard deviation of the score difference sequence. For example, the score fluctuation of path M1 is about 0.64. The larger the score difference sequence, the greater the fluctuation range, indicating that the score of the path is more unstable. The score value is usually calculated based on the performance evaluation results of the communication channel. The evaluation parameters include transmission rate, packet loss rate and average delay. Different parameters are assigned different weights. For example, the rate accounts for 40%, and the packet loss rate and delay each account for 30%. If the actual rate in a certain period is 90Mbps, the packet loss rate is 1.5%, and the delay is 12ms, the score value can be calculated to be about 92 points. After periodically scoring each path, a complete score sequence can be obtained. Then, the difference set between adjacent periods can be obtained. Combined with the path number, the score fluctuation range is collected.
[0031] S302: Based on the score fluctuation range value and the current period score value, determine whether the path score change exceeds the set benchmark, mark the fluctuation situation, and obtain the main path fluctuation status label; Based on the rating fluctuation range and the current period's rating value, a rating fluctuation judgment benchmark is set, including a rating deviation value and a fluctuation range threshold. For example, the rating deviation benchmark is 5, and the fluctuation range threshold is 1. The rating deviation value is the absolute value of the difference between the current rating and the historical average. For example, if path M1 currently has a rating of 75.6 and a historical average rating of 82, the rating deviation value is 6.4, which is greater than the benchmark value of 5. At the same time, the fluctuation range is 1.4, which is higher than the set threshold of 1, and it is judged as a strong fluctuation state. If another path currently has a rating of 88.2 and a historical average of 87.6, the deviation value is 0.6, and the fluctuation range is 0.4, both of which are lower than the threshold. If the value is zero, it is considered a stable state. Each path is processed in the same way, and the fluctuation label of each main path is output. The label level can be set as stable, moderate fluctuation, and strong fluctuation, corresponding to the score deviation value and fluctuation range. A deviation value below 2 and a fluctuation range of less than 0.5 are classified as stable. A deviation value between 2 and 5 or a fluctuation range between 0.5 and 1 are classified as moderate fluctuation. Any value exceeding the above indicators is classified as strong fluctuation. These benchmark values can be set by calculating the distribution data after summarizing the score differences of previous historical data. For example, the deviation value is taken as the upper quartile value, and the fluctuation range is taken as the average level of the fluctuation of all paths as the reference boundary.
[0032] S303: Call the main path fluctuation status label, combine the main path's ranking in the current period, filter the top-ranked main paths, classify and aggregate the fluctuation status labels, and obtain the main path fluctuation identifier status. The ranking of the main path within the current period is based on its current score, which is calculated by combining three indicators: transmission rate, packet loss rate, and average latency of the communication channel. These indicators are weighted at 40%, 30%, and 30% respectively, and then weighted and synthesized after normalization. The higher the score, the higher the ranking. Finally, a ranking list of channels within the current period is formed by ranking the scores from highest to lowest. Extract the main paths with labeled fluctuation states, and combine them with the ranking information of the current period. Set a filtering threshold, such as the top 20, and filter all path numbers within the top 20. For example, paths numbered M1, M3, M5, and M8 are currently ranked 5th, 8th, 12th, and 19th respectively, with corresponding fluctuation states of stable, moderate fluctuation, strong fluctuation, and stable. Among these paths, there are 2 stable paths, 1 moderate fluctuation path, and 1 strong fluctuation path. Aggregate the filtered paths according to their fluctuation states and count their number distribution, resulting in a state distribution of 2 stable paths, 1 moderate fluctuation path, and 1 strong fluctuation path. The aggregated state can be further identified as a path identification mark by using numerical codes or classification labels. The path ranking is arranged from high to low according to the current period's score value, and the sequential number is obtained using a conventional sorting method. The state aggregation process can be completed through filtering, classification, and counting in the data list. The aggregation structure is associated with the original path number to form fluctuation state identification information, which serves as a reference basis for subsequent path processing.
[0033] Figure 5 This is a schematic diagram illustrating the specific process of generating a backup link activation path number in the method for monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 5 The specific steps of S4 are as follows: S401: Based on the identification result that the main path fluctuation status is unstable, call the channel number, score value and stable status parameter in the channel performance ranking list, filter the channel number with stable status, sort the score values, obtain the channel number with the highest score value, and generate the score priority channel number. After identifying the main path as unstable, it is necessary to extract all channel numbers, scores, and stable state parameters from the channel performance ranking list. First, iterate through all channel entries, filtering for channel numbers with a stable state parameter of 1. This parameter value of 1 indicates that the channel status has no significant fluctuations during long-term continuous operation, with a packet loss rate below 0.5% and latency fluctuations below 5ms. The stability indicator is calculated based on 30 consecutive sampling cycles; if the number of channel status changes does not exceed 2 in 30 consecutive cycles, it is marked as stable. After filtering, only channel numbers with stable status and their scores are retained, and these channel scores are sorted in descending order. For example, stable channel numbers include CH002, CH005, and CH007. The corresponding scores are 83.7, 81.2, and 78.9. After sorting, CH002 is prioritized, followed by CH005, and then CH007. When the scores are the same, for example, CH003 and CH004 are both 80.0, it is necessary to further compare their number of consecutive stable operation cycles in the past 72 hours. CH003 has been stable for 66 hours, while CH004 has been stable for 60 hours. Therefore, CH003 is selected first. Finally, the channel number with the highest score is generated as the priority channel number. The entire screening and sorting process can be applied to the main channel switching mechanism of multi-channel communication equipment. When the main channel experiences frequent packet loss, response fluctuations, etc., the scoring mechanism can quickly find the most stable replaceable channel and complete the priority setting.
[0034] S402: Based on the scoring priority channel number, retrieve the current primary channel number and mapping information, perform the channel number replacement operation, and synchronously update the primary channel identifier status to generate the primary channel update number; Based on the priority channel number, such as CH002, extract the current primary channel number, such as CH001, and its corresponding logical information from the channel mapping configuration table, including IP address, physical interface number, port binding status, and primary flag field. Before performing the channel replacement operation, the port logically bound to the current primary channel number CH001 must be unbound. Simultaneously, update the forwarding path entries in the routing policy, redirecting all paths pointing to CH001 to CH002. Update the primary flag value of the corresponding channel flag field in the mapping table, setting the CH001 flag field value to non-primary, and setting the CH002 flag field value to non-primary. The flag field value is set to primary. This field is usually a Boolean or a flag value of 0 or 1. The number marked as 1 will be used as the current primary channel. When the system is running in a dual-channel redundant structure, the operation process involves a dynamic link reconfiguration mechanism to complete the replacement of the primary channel without restarting the system. The data flow is uninterrupted. After the replacement is completed, the current primary channel number field value is updated to CH002 through the recording module. The recording position is in the address block of the link operation table. The number field is usually located at a fixed offset address in the configuration structure. Finally, the updated number is submitted to the status synchronization module to complete the configuration synchronization.
[0035] S403: Call the primary channel update number, obtain the corresponding activation path number in the link channel configuration table, assign the path number to the backup link activation path identifier parameter, and synchronously write it to the current link path number record field to generate the backup link activation path number. After updating the primary channel number, the channel number CH002 is called to query its corresponding active path number, such as PATH07, in the link channel configuration index. This lookup operation matches the channel path mapping row in the configuration structure based on the channel number. Each row in the mapping table contains the channel number and the corresponding path number. If the channel number is CH002, the corresponding path number is PATH07. After the lookup, the path number is written to the standby link active path identifier parameter. This parameter is used to mark the standby path active channel binding information. Then, the current link path number record field is synchronously written. This field is used to record the current valid path number value internally. The data length is 16 bits, the format is an unsigned number, and it is used to be compatible with multiple path number configurations. The field is located at the offset position in the link control structure, such as 0x28. After the writing is completed, PATH07 is recorded as the standby link active path number and submitted to the link status control module. In the multi-link backup architecture, the standby link will maintain its corresponding channel status and data forwarding path based on the current active path number. The entire update process ensures that after the primary channel is switched, the standby link path identifier and the primary channel maintain synchronous update logic consistency.
[0036] Figure 6 This is a schematic diagram illustrating the specific process of generating a server communication channel early warning status identifier in the method for monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 6 The specific steps of S5 are as follows: S501: Based on the combined status of the backup link activation path number and the communication temperature impact marker, extract the channel number, link identifier and temperature data, complete the mapping between the channel and the link, combine and classify the communication temperature impact marker and the link status parameters, and determine whether the communication interference threshold is exceeded, and generate a communication status abnormality judgment value. The backup link activation path number is composed of the numbers between each node in the network path. For example, path ABCD can be marked as 001. The communication temperature impact marker is determined by reading link environment data from temperature sensors and combining it with historical communication error rate data to classify the level. For example, when the link temperature is 65°C and the error rate reaches 1.2E-3, it is marked as level 3. The channel number can be directly obtained from the physical port number, such as GE0 / 1 and GE0 / 2. The link identifier is generated according to the interface and network management numbering rules, such as LINK_ID_01. Temperature data is collected by the embedded temperature control module of the link device and uploaded every 5 seconds. The mapping process between channels and links is matched with the port and link table in the device management system, such as GE 0 / 1 corresponds to LINK_ID_01. When the communication temperature impact flag is jointly classified with the link status parameters, the status is first divided according to the basic indicators such as the current bandwidth utilization, packet loss rate, and latency. For example, if the bandwidth utilization exceeds 80% and the packet loss rate exceeds 2%, it is marked as status level 2. Then, the status level is combined with the temperature impact level to generate a status label. For example, T3S2 represents temperature level 3 and status level 2. This label will be compared with the interference threshold. The interference threshold is preset to be a combination of T3S2 and above, which is considered abnormal, such as T4S3, T5S2, etc. It is determined whether the current status exceeds the threshold and the communication status abnormal judgment value is output accordingly. If it exceeds the threshold, it is marked as abnormal with a value of 1; otherwise, it is 0.
[0037] S502: Based on the communication status anomaly judgment value, match the server node channel number with the communication anomaly mark, combine the server status parameters to make a judgment, extract the server number and running status that meet the conditions, and compare with the alarm response rule table to generate the server node alarm trigger level value. When the communication status anomaly judgment value is 1, the server node judgment process is initiated. The channel number is matched with the communication anomaly marker. The channel information is provided by the link allocation table. For example, GE0 / 1 corresponds to server number S001. If GE0 / 1 is already marked as anomaly, the match with the server node is successful. Subsequently, server operating parameters are extracted, such as CPU utilization, memory usage, and disk I / O response latency. These parameters are from the server management system, and the data is recorded at 10-second intervals. If CPU usage exceeds 85%, memory usage exceeds 90%, and disk I / O response exceeds 10ms, the server is marked as being in a high-load state. Server numbers that meet this condition and their current states are extracted. For example, if server S001 is in a high-load state, the alarm level is matched in conjunction with the abnormal communication marker. The alarm response rule table specifies that: communication anomaly plus high load corresponds to level 3, communication anomaly plus medium load corresponds to level 2, and so on. The alarm trigger level corresponding to server S001 is defined as 3, which serves as the basis for subsequent output.
[0038] S503: Based on the alarm trigger level value of the server node, summarize the channel number, server operating status and alarm level, output the operating status identifier number by referring to the communication link status identifier encoding table, and integrate the encapsulation information to generate the server communication channel early warning status identifier value. Based on an alarm trigger level of 3, the corresponding channel number (e.g., GE0 / 1), server number (e.g., S001), server operating status (e.g., high load), and alarm level value are combined to form a structure. The output structure record contains the channel number, server number, operating status, and alarm level. Then, the level value is converted into an operating status identifier number according to the link status identifier encoding rules. For example, level value 3 corresponds to status number STATE_03. This number is combined with the structure information and formatted for encapsulation. Finally, a unified format information packet is constructed, which includes the channel, server, status, and identifier number. It is named according to the encoding rules, such as PREWARN_GE0 / 1_S001_STATE_03, representing the warning level server channel status identification value, which is stored and sent as an identification mark for the alarm system.
[0039] This application also discloses a terminal for monitoring an instant messaging server, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for monitoring an instant messaging server. The method for monitoring the instant messaging server corresponds one-to-one with the above embodiments and will not be repeated here.
[0040] Figure 7 This is a schematic diagram of the system framework for monitoring an instant messaging server according to an embodiment of this application. Please refer to... Figure 7 A system for monitoring instant messaging servers, comprising: The link thermal monitoring module is used to acquire temperature data and ambient temperature monitoring values of multiple areas inside the server rack, detect packet loss rate, latency and link status parameters of the communication channel, identify the thermal interference level of the channel based on the correspondence between temperature data and link status, and mark it with the channel number to generate communication temperature impact markers. The channel stability assessment module is used to retrieve the stability-related parameters of the channel during operation, including transmission delay, retransmission count and interruption records, based on the heated channel number identified in the communication temperature influence marker. Based on the parameters, it judges the communication stability difference of the channel in the current state and generates a channel performance ranking list. The path fluctuation identification module is used to analyze the score changes of the channel in a continuous period based on the channel number of the primary path in the channel performance ranking list, determine whether the score change exceeds the stability judgment standard, and generate the primary path fluctuation identification status. The primary link replacement module is used to select channel numbers with high scores and stable status from the channel performance ranking list based on the channel numbers that have been determined to be unstable in the primary path fluctuation status, update the current primary channel's identifier, and generate backup link activation path numbers. The channel early warning linkage module is used to collect server node status, response time and alarm status information based on the channel number enabled in the backup link activation path number and the heat interference level of the corresponding link in the communication temperature influence mark, and to jointly determine the operational risk level of the current link based on the data and generate a server communication channel early warning status identifier.
[0041] It should be noted that the system for monitoring the instant messaging server corresponds one-to-one with the embodiments of the method for monitoring the instant messaging server described above, and will not be repeated here.
[0042] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for monitoring an instant messaging server, characterized in that, Includes the following steps: Collect data on the internal temperature of the server rack and the ambient temperature, analyze the characteristics of heat change, and classify the status of the server communication channel in conjunction with the link status to generate communication temperature impact markers. Based on the communication temperature impact marker, identify the communication channels affected by heat, retrieve the corresponding channel's operating parameters and interruption information, evaluate channel stability, analyze communication availability differences, and generate a channel performance ranking list. Based on the channel performance ranking list, analyze the score change trend between adjacent periods, determine the performance stability, and mark whether there is fluctuation in the current main path, generating a main path fluctuation indicator status. Based on the main path fluctuation status, select a stable and high-scoring backup path from the channel performance sorting list, replace the channel number and update the current primary channel identifier, and generate a backup link activation path number. Based on the combined status of the backup link activation path number and the communication temperature influence marker, the alarm linkage processing of the server node is determined and triggered, the channel number, server status and alarm information are output, the communication link operation status is identified, and a server communication channel early warning status identifier is generated.
2. The method for monitoring an instant messaging server according to claim 1, characterized in that, The communication temperature impact marker includes temperature change amplitude, link status classification, and channel impact level; the channel performance ranking list includes channel score, availability level, and stability coefficient; the main path fluctuation identifier includes score change trend, main path fluctuation level, and time period label; the backup link activation path number includes replacement channel number, backup path priority, and channel update identifier; the server communication channel warning status identifier includes channel number, server operating status, and alarm information encapsulation content.
3. The method for monitoring an instant messaging server according to claim 1, characterized in that, The specific steps for collecting server rack internal temperature and ambient temperature data, analyzing heat change characteristics, and classifying server communication channels based on link status to generate communication temperature impact markers are as follows: Collect temperature variation data and ambient temperature information of monitoring nodes inside the server rack. After performing time series standardization processing on the temperature data inside the server rack and the ambient temperature data, calculate the difference between the standardized value of the temperature inside the server rack and the standardized value of the ambient temperature at the same time point, and calculate the ratio of the standardized value of the temperature inside the server rack to the standardized value of the ambient temperature to obtain the degree of temperature deviation per unit time and generate the node temperature offset rate. Based on the node temperature offset rate, combined with the changes in link connection rate and delay time at the corresponding time, the communication link nodes that are consistent with the temperature offset direction are identified, and the corresponding relationship is compared to generate the temperature link correspondence coefficient. Based on the temperature link correspondence coefficient, combined with the connection stability and temperature offset rate of the link node, it is determined whether the node belongs to the temperature-affected communication channel, a status classification index is established, and a communication temperature-affected label is generated.
4. The method for monitoring an instant messaging server according to claim 3, characterized in that, The specific steps for identifying heat-affected communication channels based on the communication temperature impact marker, retrieving the corresponding channel's operating parameters and interruption information, evaluating channel stability, analyzing communication availability differences, and generating a channel performance ranking list are as follows: Based on the heated communication channel identified by the communication temperature influence marker, the operating current, conductor temperature, contact voltage and bit error rate parameters of the channel are called up, the data are extracted and the variation range, range and fluctuation frequency are calculated, the channel stability difference is summarized, and the channel state offset is generated. Based on the channel status offset, combined with the interruption time, recovery delay and number of interruptions, the parameter offset degree and interruption situation are judged accordingly to identify the unstable interval and generate the communication channel availability difference rate. The communication channel availability difference rate is called, and the bit error rate, operating current and recovery delay parameters are combined and sorted according to the average value after channel grouping to generate communication channel performance ranking values.
5. The method for monitoring an instant messaging server according to claim 4, characterized in that, The specific steps for analyzing the score change trend between adjacent periods based on the channel performance ranking list, determining performance stability, marking whether there is fluctuation in the current main path, and generating the main path fluctuation indicator status are as follows: Based on the primary path data in the channel performance ranking list, obtain the corresponding score values for adjacent periods, calculate the main path score changes, and aggregate them by path number to generate score fluctuation amplitude values. Based on the score fluctuation range value and the current period score value, determine whether the path score change exceeds the set benchmark, mark the fluctuation situation, and obtain the main path fluctuation status label; The main path fluctuation status label is invoked, and combined with the ranking of the main path in the current period, the top-ranked main paths are filtered out. The fluctuation status labels are then classified and aggregated to obtain the main path fluctuation identifier status.
6. The method for monitoring an instant messaging server according to claim 5, characterized in that, The specific steps for selecting a stable and high-scoring backup path from the channel performance sorting list based on the main path fluctuation status, replacing the channel number and updating the current primary channel identifier, and generating a backup link activation path number are as follows: Based on the identification result that the main path fluctuation status is unstable, the channel number, score value and stable status parameter in the channel performance ranking list are called to filter the channel number with a stable status, and the score values are sorted to obtain the channel number with the highest score value and generate the score priority channel number. Based on the scoring priority channel number, the current primary channel number and mapping information are called to perform a channel number replacement operation, and the primary channel identifier status is updated synchronously to generate a primary channel update number; The primary channel update number is called to obtain the corresponding activation path number in the link channel configuration table. The path number is assigned to the backup link activation path identifier parameter and synchronously written to the current link path number record field to generate the backup link activation path number.
7. The method for monitoring an instant messaging server according to claim 6, characterized in that, The specific steps for determining and triggering server node alarm linkage processing based on the combined status of the backup link activation path number and the communication temperature influence marker, outputting the channel number, server status, and alarm information, identifying the communication link operating status, and generating a server communication channel early warning status identifier are as follows: Based on the combined status of the backup link activation path number and the communication temperature impact marker, the channel number, link identifier and temperature data are extracted. After completing the mapping between the channel and the link, the communication temperature impact marker and the link status parameters are combined and classified, and it is determined whether the communication interference threshold is exceeded, and an abnormal communication status judgment value is generated. Based on the communication status anomaly determination value, the server node channel number and communication anomaly marker are matched, and the server status parameters are combined to make a judgment, extract the server number and running status that meet the conditions, and compare them with the alarm response rule table to generate the server node alarm trigger level value. Based on the alarm trigger level value of the server node, the channel number, server operating status and alarm level are summarized, the operating status identifier number is output by referring to the communication link status identifier encoding table, and the encapsulation information is integrated to generate the server communication channel early warning status identifier value.
8. The method for monitoring an instant messaging server according to claim 7, characterized in that, The communication interference threshold is a boundary value used to determine whether the communication link is in an abnormal state. It is set according to the stability index of the communication link under different loads, temperatures and interference environments. The server status parameters are a set of multi-dimensional parameters that reflect the current operating status of the server node, and are used to simultaneously characterize server load, resource consumption, and link response indicators. The operating status identifier number is a standardized encoding result representing the current operating status of the communication link, used to abstract the communication status into identifiable and transmissible structured information.
9. A terminal for monitoring an instant messaging server, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for monitoring an instant messaging server as described in any one of claims 1 to 8.
10. A system for monitoring an instant messaging server, characterized in that, The system for monitoring the instant messaging server is used to perform the method for monitoring the instant messaging server according to any one of claims 1-8, wherein the system for monitoring the instant messaging server comprises: The link thermal monitoring module is used to acquire temperature data and ambient temperature monitoring values of multiple areas inside the server rack, detect packet loss rate, latency and link status parameters of the communication channel, identify the thermal interference level of the channel based on the correspondence between temperature data and link status, and mark it with the channel number to generate communication temperature impact markers. The channel stability assessment module is used to retrieve the stability-related parameters of the channel during operation, including transmission delay, retransmission count and interruption records, based on the heated channel number identified in the communication temperature influence marker. The module then judges the communication stability differences of the channel in the current state based on the parameters and generates a channel performance ranking list. The path fluctuation identification module is used to analyze the score changes of the channel in a continuous period based on the channel number of the primary path in the channel performance ranking list, determine whether the score change exceeds the stability judgment standard, and generate the primary path fluctuation identification status. The primary link replacement module is used to select channel numbers with high scores and stable status from the channel performance ranking list based on the channel numbers that have been determined to be unstable in the primary path fluctuation identification status, update the identification of the current primary channel, and generate backup link activation path numbers. The channel early warning linkage module is used to collect server node status, response time and alarm status information based on the channel number enabled in the backup link activation path number and the heat interference level of the corresponding link in the communication temperature influence mark, and to jointly determine the operational risk level of the current link based on the data, and generate a server communication channel early warning status identifier.