Data aggregation method and system
By calculating the server's running status score and calling the corresponding running status table to filter out exception or warning parameters, generating prompt information and increasing sampling frequency, the problem of lack of foresight in server operation is solved, and the efficiency of exception discovery and server stability is improved.
Patent Information
- Application Number
- CN202510832793.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the prior art, the server lacks predictability of abnormal states during operation, resulting in low efficiency in abnormal discovery, which can easily lead to server corruption or loss of business data.
By calculating the server's running status score, calling the corresponding running status table to filter out abnormal or early warning running parameters, and generating prompt information, and at the same time increasing the sampling frequency of early warning running parameters to improve monitoring capabilities.
It improves the efficiency and foresight of abnormal discovery, reduces the risks of server corruption and business data loss, and improves the stability and monitoring effect of the server.
Smart Images

Figure CN120353633A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a data aggregation method and system. Background Art
[0002] During the operation of a server, the operating state of the server can be judged by monitoring the operating parameters of the server. When the operating parameters of the server are abnormal, emergency measures can be taken to avoid the server malfunctioning and affecting the normal operation of network services.
[0003] However, simply judging whether the server is normal based on the operating parameters lacks predictability of abnormal states. Only when the server is abnormal can the server be maintained, or the operating state of the server can be maintained by the built-in adjustment system of the server. This method has a low efficiency of server abnormality, and is likely to cause server damage or business data loss. Summary of the Invention
[0004] This application provides a data aggregation method, including: Obtaining the operating state data of the servers located in the computer room; Calculating an operating state score of the server according to the operating state data; the operating state score is calculated based on the real-time values of the respective operating parameters of the server; If the operating state score is within a first preset range, then calling a first set of operating state tables; the first set of operating state tables includes a plurality of first operating state tables; wherein, the first operating state table includes the first operating parameters of the server and first operating parameter thresholds; obtaining a first operating state table from the first set of operating state tables according to the current service of the server; comparing the real-time value of the first operating parameter of the server with the first operating parameter threshold to obtain abnormal operating parameters; the abnormal operating parameters are the parameters whose real-time values are greater than or equal to the first operating parameter threshold; generating a first prompt message based on the abnormal operating parameters; the first prompt message is used to indicate that the abnormal operating parameters need to be debugged; If the running status score is within the second preset range, call the second set of running status tables; the second set of running status tables includes multiple second running status tables; the second running status table includes the second running parameters of the server and the second running parameter thresholds; the first running parameters are the same as the second running parameters, and the first running parameter thresholds are different from the second running parameter thresholds; obtain the second running status table from the second set of running status tables according to the current business of the server; compare the real-time value of the second running parameter of the server with the second running parameter threshold to obtain the warning running parameter; the warning running parameter is a parameter greater than or equal to the second running parameter threshold; and the warning running parameter is less than the first running parameter threshold; generate a second prompt message based on the warning running parameter; the second prompt message is used to indicate that the warning running parameter needs to be continuously tracked and increase the sampling frequency of the warning running parameter.
[0005] In some feasible embodiments, if the number of servers in the computer room is greater than or equal to 2, the increasing the sampling frequency of the warning running parameter specifically includes: Increase the sampling frequency of the running parameters of the warning server; the warning server is a server with the warning running parameter.
[0006] In some feasible embodiments, the first running status table further includes a business type, and the running status thresholds in the first running status tables corresponding to different business types are different; it further includes: In the case where the business type of the server changes, obtain the first running status table corresponding to the changed business type from the first set of running status tables according to the changed business type.
[0007] In some feasible embodiments, the first running status table further includes a time identifier, and the time identifiers corresponding to the servers running in different time periods are different; obtaining the first running status table from the first set of running status tables according to the current business of the server specifically includes: Obtain the first running status table from the first set of running status tables according to the current business of the server and the system time.
[0008] In some feasible embodiments, if the server includes the warning running parameter, mark the server as a warning server; if the warning server changes to an abnormal server within a preset time, replace the corresponding first running parameter threshold in the first running status table with the second running parameter threshold corresponding to the warning running parameter in the second running status table.
[0009] In some feasible embodiments, replacing the corresponding first operating parameter threshold in the first operating state table with the second operating parameter threshold corresponding to the warning operating parameter in the second operating state table specifically includes: Calculating the difference between the second operating parameter threshold and the first operating parameter threshold; If the difference is greater than or equal to the first difference threshold, calculating a first replacement parameter based on the second operating parameter threshold and the first operating parameter threshold; replacing the first operating parameter according to the first replacement parameter; If the difference is less than the first difference threshold, replacing the first operating parameter based on the second operating parameter threshold.
[0010] In some feasible embodiments, it further includes: When the first operating parameter threshold in the first operating state table is replaced, calculating the real-time score of the first operating state table; If the real-time score is outside the first preset range, increasing the replaced first operating parameter threshold to restore the real-time score of the first operating state table within the first preset range.
[0011] In some feasible embodiments, if the operating state score is within the first preset range, it further includes: Marking the server as an abnormal server; Reducing the service response ratio of the abnormal server and increasing the service response ratios of non-abnormal servers and non-warning servers.
[0012] In a second aspect, an embodiment of the present application provides a data aggregation system, including a sampling device and a control device.
[0013] The sampling device is used to obtain the operating state data of the servers located in the computer room; The control device is used to calculate the operating state score of the server according to the operating state data; the operating state score is calculated based on the real-time values of the respective operating parameters of the server; The control device is further configured to call a first set of operating state tables when the operating state score is within a first preset range; the first set of operating state tables includes a plurality of first operating state tables; wherein, the first operating state table includes first operating parameters of the server and first operating parameter thresholds; obtain a first operating state table from the first set of operating state tables according to the current service of the server; compare a real-time value of the first operating parameter of the server with the first operating parameter threshold to obtain abnormal operating parameters; the abnormal operating parameters are parameters whose real-time value is greater than or equal to the first operating parameter threshold; generate a first prompt message based on the abnormal operating parameters; the first prompt message is used to indicate that the abnormal operating parameters need to be debugged. The control device is further configured to call a second set of operating state tables when the operating state score is within a second preset range; the second set of operating state tables includes a plurality of second operating state tables; the second operating state table includes second operating parameters of the server and second operating parameter thresholds; the first operating parameters and the second operating parameters are the same, and the first operating parameter threshold and the second operating parameter threshold are different; obtain a second operating state table from the second set of operating state tables according to the current service of the server; compare a real-time value of the second operating parameter of the server with the second operating parameter threshold to obtain warning operating parameters; the warning operating parameters are parameters that are greater than or equal to the second operating parameter threshold; and the warning operating parameters are less than the first operating parameter threshold; generate a second prompt message based on the warning operating parameters; the second prompt message is used to indicate that the warning operating parameters need to be continuously tracked and the sampling frequency of the warning operating parameters needs to be increased.
[0014] As can be seen from the above technical content, the embodiments of the present application provide a data aggregation method and system. During the operation of the server, the operating parameters of the server can be collected and the corresponding operating state scores can be calculated. When the operating state score is within the first preset range, a first operating state table is called to screen out abnormal operating parameters by comparison and generate a first prompt message to facilitate debugging of the abnormal operating parameters. When the operating state score is within the second preset range, a second operating state table is called to screen out warning operating parameters by comparison and generate a second prompt message, and the sampling frequency of the warning server is increased to improve the monitoring ability of the warning server, thereby improving the abnormal discovery efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0016] Figure 1 Flow chart of server control branch provided by an embodiment of this application; Figure 2 Flow chart of dynamic replacement of the first running state table provided by an embodiment of this application; Figure 3 Flow chart of server service shunt provided by an embodiment of this application; Figure 4 Schematic diagram of server architecture provided by an embodiment of this application. Detailed implementation manners
[0017] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application.
[0018] During the operation of the server, the operation parameters of the server can be monitored, such as the temperature, vibration frequency, etc. of the server, and then it can be judged whether the server is operating normally based on the operation parameters. However, during the monitoring process, the predictability of abnormal situations is lacking, resulting in corresponding measures being taken only when the server is abnormal. Such a monitoring method has low efficiency in detecting abnormalities and is likely to cause server damage or business data loss.
[0019] To solve the above problems, as Figure 1 shown, an embodiment of this application provides a data aggregation method, including: Obtain the operation status data of the server located in the computer room; In some embodiments, the operation status data of the server may include data at the hardware layer, such as disk vibration frequency, processor temperature, etc., and may also include data at the software layer, such as network congestion rate, message request response rate. Whether the server is operating normally can be judged based on these operation status data.
[0020] Calculate the operation status score of the server according to the operation status data; the operation status score is calculated based on the real-time values of the operation parameters of the server; In some embodiments, the operation status score can be calculated based on the operation status data. Whether the server is normal can be judged more accurately through the operation status score, and the server can be classified based on the operation status score, so as to adopt different monitoring strategies for servers with different operation statuses.
[0021] Among them, according to the actual monitoring requirements, weights can be set for the hardware layer data and software layer data of the server, so as to calculate the running state score of the server based on the weights and the evaluation scores corresponding to the running state data.
[0022] It can be understood that the running state data can be converted into different scores according to different values. For example, the running state score corresponding to a processor temperature of 80°C is higher than that corresponding to a processor temperature of 50°C. Among them, a high running state score does not necessarily mean a good running state, and the score conversion rule can be set according to application requirements. For example, in the embodiment of the present application, the higher the processor temperature, the higher the running state score, and the closer the server is to the abnormal state.
[0023] In this way, the servers can be classified according to the running state scores of the servers, and a multi-level monitoring strategy can be adopted for the servers to improve the abnormal response effect and the abnormal discovery effect.
[0024] If the running state score is within the first preset range, the first running state table set is called; the first running state table set includes multiple first running state tables; among them, the first running state table includes the first running parameters of the server and the first running parameter thresholds; the first running state table is obtained from the first running state table set according to the current service of the server; the real-time value of the first running parameter of the server is compared with the first running parameter threshold to obtain abnormal running parameters; the abnormal running parameters are the parameters whose real-time value is greater than or equal to the first running parameter threshold; a first prompt message is generated based on the abnormal running parameters; the first prompt message is used to indicate that the abnormal running parameters need to be debugged; In some embodiments, the first preset range represents the server abnormal range, that is, if the running state score of the server is within the first preset range, it means that the server is abnormal, that is, the server can be determined as an abnormal server.
[0025] The first running state table can include various running parameters of the server, such as processor temperature, disk vibration frequency, network congestion rate. For different running parameters, running parameter thresholds can be set to determine abnormal running parameters by comparing with the running parameters of abnormal servers, and then adopt targeted processing strategies.
[0026] For example, if the running state score of the server is 90 points and the first preset range is greater than or equal to 90 points, then after calculating the running state score of the server, the first running state table set can be called, and the corresponding first running state table can be found from the first running state table set. Among them, the first running state table also has a corresponding real-time score, and the real-time score is within the first preset range.
[0027] Multiple first operating state tables can be set according to the performance of the server under different service types. Therefore, the corresponding first operating state table can be found according to the service type of the server, which facilitates a more accurate comparison of the relationship between the operating parameters of the server and the first operating parameter threshold, and realizes the abnormal monitoring of the server in different scenarios.
[0028] During the comparison process, the current operating parameters of the server are compared with the first operating parameter threshold to determine the abnormal operating parameters. In this way, the first prompt information for prompting the maintenance personnel can be generated according to the abnormal operating parameters, so as to facilitate the maintenance personnel to maintain the equipment. Moreover, the first prompt information can also be used to prompt the controller inside the server. After detecting the first prompt information, the controller can adopt corresponding adjustment measures. Thereby improving the stability of the server and its adaptability in case of anomalies.
[0029] If the operating state score is within the second preset range, call the second set of operating state tables; the second set of operating state tables includes multiple second operating state tables; the second operating state table includes the second operating parameters of the server and the second operating parameter threshold; the first operating parameters are the same as the second operating parameters, and the first operating parameter threshold is different from the second operating parameter threshold; obtain the second operating state table from the second set of operating state tables according to the current service of the server; compare the real-time value of the second operating parameter of the server with the second operating parameter threshold to obtain the warning operating parameter; the warning operating parameter is a parameter greater than or equal to the second operating parameter threshold; and the warning operating parameter is less than the first operating parameter threshold; generate the second prompt information based on the warning operating parameter; the second prompt information is used to indicate that the warning operating parameter needs to be continuously tracked and increase the sampling frequency of the warning operating parameter.
[0030] In some embodiments, the second preset range represents the server warning range, that is, if the operating state score of the server is within the second preset range, it means that the server is warning, that is, the server can be determined to be a warning server.
[0031] The second operating state table contains the same operating parameters as the first operating state table, but the operating parameter thresholds are different. When the server is in the warning state, no anomaly has occurred, but there is a tendency to occur. Therefore, by calling the second operating state table, the warning server or warning operating parameters can be determined and monitored, which is conducive to improving the predictability of server monitoring.
[0032] For example, if the running status score of the server is 75 points and the second preset range is 70 - 90 points, after calculating the running status score of the server, the second running status table set can be called, and the corresponding second running status table can be found from the second running status table set. Among them, the second running status table also has a corresponding real-time score, and this real-time score is within the second preset range.
[0033] In this way, when the running status score of the server is within the second preset range, the server is regarded as a warning server, that is, a server that is prone to anomalies during subsequent operation. And according to the second running parameter threshold in the second running status table, the warning running parameters of the warning server in the current running status can also be determined. For example, the processor temperature threshold in the second running parameter threshold is 75 °C, and the processor temperature of the warning server is 76 °C, then the processor temperature of the warning server is regarded as the warning running parameter.
[0034] Furthermore, after determining the warning running parameters, a second prompt message can be generated based on the warning running parameters. On the one hand, the second prompt message can be used as a prompt for the operation and maintenance personnel, which is beneficial to server maintenance. On the other hand, it can be used as an adjustment prompt message within the server system to enable the server to adjust the warning running parameters.
[0035] In addition, the sampling frequency of the warning running parameters can also be adjusted to improve the real-time performance of anomaly monitoring. For example, if the processor temperature is the warning running parameter, the adjustment frequency of the warning running parameters can be increased, and then the warning running parameters can be obtained more frequently, thereby improving the monitoring efficiency and being beneficial to timely detecting the running anomalies of the server.
[0036] It can be understood that calculating the running status score of the server first can initially determine the current running status of the server, such as an abnormal status, a warning status, or a normal status. This can facilitate subsequent hierarchical maintenance processing based on the status of the server. Then, combined with the query of the running status table, abnormal running parameters can be determined and maintenance measures can be formulated. And by using the business type of the server as the construction dimension of the running status table, the anomaly detection ability of the server in each scenario can be refined, thereby enhancing the running security of the server in each scenario.
[0037] In some embodiments, if the number of servers in the computer room is greater than or equal to 2, the increasing the sampling frequency of the warning running parameters specifically includes: Increasing the sampling frequency of the running parameters of the warning server; the warning server is a server with the warning running parameters.
[0038] It can be understood that, for server security considerations, the warning server is likely to evolve into an abnormal server, and during the operation of the server, the operating parameters related to the hardware layer and the operating parameters related to the software layer are interrelated. For example, when the processor temperature is too high, it will also affect the execution efficiency of the software layer of the server. Therefore, when determining that the server is a warning server, not only can the sampling frequency of the warning operation parameters be increased, but also the sampling frequency of all warning operation parameters can be increased.
[0039] In this way, by increasing the acquisition frequency of the operating parameters of the warning server, the monitoring of abnormal states can be effectively increased, and thus the abnormalities of the server can be detected more timely. Although it is still when the server has an abnormality that the abnormality can be monitored and the processing means for adjusting the abnormality can be executed, based on the classification of the server operating state and the improvement of the sampling frequency, the real-time performance of abnormality detection can be effectively improved, which is conducive to maintaining the normal operation of the server and provides a certain degree of predictability for abnormality handling.
[0040] In some embodiments, the first operating state table further includes a service type, and the operating state thresholds in the first operating state table corresponding to different service types are different.
[0041] In the case where the service type of the server changes, obtain the first operating state table corresponding to the changed service type from the set of first operating state tables according to the changed service type.
[0042] It can be understood that the first operating state table or the second operating state table corresponding to the service type can be searched in the set of first operating state tables or the set of second operating state tables in combination with the service currently processed by the server, so as to provide more precise anomaly monitoring measures through such a subdivision method.
[0043] For example, when the server processes financial services and requires a faster response rate, the operating parameter threshold can be set relatively high. When the server processes services with lower timeliness requirements, the operating parameter threshold can be set relatively low. Thereby improving the adaptability of the detection to the server in various application scenarios and improving the anomaly detection ability.
[0044] In some embodiments, the first operating state table further includes a time identifier, and the time identifiers corresponding to the servers operating in different time periods are different. That is, obtaining the first operating state table from the set of first operating state tables according to the current service of the server specifically includes: Obtain the first operating state table from the set of first operating state tables according to the current service of the server and the system time.
[0045] It is understandable that the business volume undertaken by the server is also different at different time periods. For example, at midnight, the server should be in a relatively idle state, and at this time, the running parameter threshold should be set relatively low. Another example is that during the day on weekdays, the server should be in a relatively busy state, and at this time, the running parameter threshold should be set relatively high.
[0046] In this way, based on the multi-level dimensions of time identification and business type, richer first running state tables and second running state tables can be preset. Furthermore, the monitoring ability of server abnormal, warning and other states can be improved, which is beneficial to maintaining the normal operation of the server.
[0047] In some embodiments, if the server includes the warning running parameter, the server is marked as a warning server; if the warning server changes to an abnormal server within a preset time, the corresponding second running parameter threshold in the second running state table is used to replace the corresponding first running parameter threshold in the first running state table.
[0048] It is understandable that the probability of a warning server evolving into an abnormal server is higher than that of a normally running server. Therefore, in the scenario of multi-server monitoring, the warning server can be marked and tracked and monitored according to the mark. Moreover, the mark is not limited to the warning server, and the corresponding second running state table of the warning server can also be marked to comprehensively master the relevant information of the warning server.
[0049] In this way, in a scenario with high demand for running stability, if it is detected that the warning server changes to an abnormal server within a preset time, the second running parameter threshold in the second running state table corresponding to the warning server can be used to replace the first running parameter threshold in the first running state table.
[0050] For example, use the second running parameter threshold "processor temperature 75°C" to replace the first running parameter threshold "processor temperature 85°C". In this way, in the same scenario subsequently, when it is detected that the processor temperature of the server reaches 75°C, a prompt message can be generated to indicate that the server may be abnormal and needs maintenance or detection.
[0051] It is understandable that the prompt message here may not be the first prompt message mentioned in the foregoing embodiments. It can be a reminder for a single running parameter exception.
[0052] In this way, in some scenarios where a single running parameter is significantly abnormal and other running parameters are relatively normal, by replacing the parameters prone to abnormalities, the discovery probability of the abnormal state of the server can be increased. And in similar scenarios, maintain the stability of the server.
[0053] In addition, the replacement timeliness can also be set. For example, after one week of replacement, the first operation status table is restored to the previous first operation parameter threshold to maintain the stability of the layout form of the operation status table and improve the adaptability of the operation status table to the monitoring scenario.
[0054] As Figure 2 shown, in some embodiments, replacing the corresponding first operation parameter threshold in the first operation status table according to the second operation parameter threshold corresponding to the warning operation parameter in the second operation status table specifically includes: Calculating the difference between the second operation parameter threshold and the first operation parameter threshold; If the difference is greater than or equal to the first difference threshold, calculating a first replacement parameter based on the second operation parameter threshold and the first operation parameter threshold; replacing the first operation parameter according to the first replacement parameter; If the difference is less than the first difference threshold, replacing the first operation parameter based on the second operation parameter threshold.
[0055] It can be understood that if the difference between the second operation parameter threshold and the first operation parameter threshold is too large, random replacement is likely to cause a sudden drop in the score of the first operation status table, that is, it is likely to cause the real-time score corresponding to the first operation status table to drop, resulting in the loss of the monitoring ability of the first operation status table. Therefore, when replacing, the difference between the second operation parameter threshold and the first operation parameter threshold should be considered to optimize the first operation status table according to the performance of the server in various scenarios without changing the monitoring logic.
[0056] For example, when the second operation parameter "processor temperature 60°C" replaces the first operation parameter "processor temperature 70°C", and the difference of 10°C between the two is less than the difference threshold of 15°C, the first operation parameter threshold can be directly replaced with the second operation parameter threshold. A new first operation status table is constructed.
[0057] Another example is that when the second operation parameter "processor temperature 60°C" replaces the first operation parameter "processor temperature 90°C", and the difference of 30°C between the two is greater than the difference threshold of 15°C, the first operation parameter threshold cannot be directly replaced with the second operation parameter threshold. It is necessary to calculate the first replacement threshold by combining the second operation parameter threshold and the first operation parameter threshold, so as to replace the first operation parameter with the first replacement threshold. Among them, the first replacement threshold can be calculated by means of weighted calculation.
[0058] This can increase the rationality of the dynamic replacement threshold to avoid damaging the overall monitoring logic of the server.
[0059] In some embodiments, replacing the first operating parameter threshold with the second operating parameter threshold will damage the real-time score of the first operating status table, causing the first operating status table to deviate from the set of first operating status tables. Therefore, in the replacement scenario, certain constraints need to be implemented based on the real-time score of the first operating status table.
[0060] If the real-time score is outside the first preset range, increase the replaced first operating parameter threshold so that the real-time score of the first operating status table is restored within the first preset range.
[0061] It can be understood that, based on the calculation method of the operating status score, after the first operating parameter threshold is replaced, the real-time score of the first operating status table can be calculated. When the first preset range is more than 90 points, if the real-time score of the first operating status table is 89 points, the replaced first operating parameter threshold needs to be adjusted so that the real-time score of the first operating status table exceeds 90 points. For example, after the first operating parameter threshold "processor temperature 90°C" is adjusted to "processor temperature 75°C", the real-time score of the first operating status table is less than 90 points. At this time, the processor temperature can be increased so that the real-time score of the first operating status table exceeds 90 points.
[0062] In some embodiments, other operating parameter thresholds can also be fine-tuned to cooperate with the calculation of the real-time score, thereby ensuring the dynamic update of the first operating parameter threshold and also ensuring that the monitoring logic and usage logic of the first operating status table do not change.
[0063] It should be noted that in the embodiments of the present application, the method for calculating the operating status score based on the operating parameters of the server is not limited, and the specific method can be set according to actual requirements.
[0064] As Figure 3 shown, in some embodiments, if the operating status score is within the first preset range, it further includes: Mark the server as an abnormal server; Reduce the business response ratio of the abnormal server, and increase the business response ratios of non-abnormal servers and non-warning servers.
[0065] It can be understood that after calculating the operating status score, a multi-level control strategy can be executed according to the classification situation of the servers based on the operating status score.
[0066] For example, when it is monitored that the operating status score of the server is within the first preset range, it indicates that the server is in an abnormal state. At this time, the business response ratio of the server needs to be adjusted, such as transferring the access nodes of the server-related services, so as to reduce the business response ratio of the server, thereby alleviating the operating pressure of the server, preventing the server from being damaged, and maintaining the normal operation of the services.
[0067] Moreover, transfer the relevant services to the access nodes of other servers, so as to ensure the normal operation of the relevant services. The other servers can be the servers that are in the same server cluster as the current server, or can be the edge servers arranged in advance for relieving the operation pressure of the server.
[0068] The embodiment of the present application further provides a data aggregation system, including a sampling device and a control device.
[0069] The sampling device is used to obtain the operation status data of the servers located in the computer room; The control device is used to calculate the operation status score of the server according to the operation status data; the operation status score is calculated based on the real-time values of the operation parameters of the server; The control device is further used to call a first set of operation status tables when the operation status score is within a first preset range; the first set of operation status tables includes a plurality of first operation status tables; wherein, the first operation status table includes the first operation parameters of the server and the first operation parameter thresholds; obtain the first operation status table from the first set of operation status tables according to the current service of the server; compare the real-time value of the first operation parameter of the server with the first operation parameter threshold to obtain abnormal operation parameters; the abnormal operation parameters are the parameters whose real-time values are greater than or equal to the first operation parameter threshold; generate a first prompt message based on the abnormal operation parameters; the first prompt message is used to indicate that the abnormal operation parameters need to be debugged; The control device is further used to call a second set of operation status tables when the operation status score is within a second preset range; the second set of operation status tables includes a plurality of second operation status tables; the second operation status table includes the second operation parameters of the server and the second operation parameter thresholds; the first operation parameter is the same as the second operation parameter, and the first operation parameter threshold is different from the second operation parameter threshold; obtain the second operation status table from the second set of operation status tables according to the current service of the server; compare the real-time value of the second operation parameter of the server with the second operation parameter threshold to obtain warning operation parameters; the warning operation parameters are the parameters that are greater than or equal to the second operation parameter threshold; and the warning operation parameters are less than the first operation parameter threshold; generate a second prompt message based on the warning operation parameters; the second prompt message is used to indicate that the warning operation parameters need to be continuously tracked and the sampling frequency of the warning operation parameters is increased.
[0070] Such as Figure 4As shown in the figure, an embodiment of the present application provides a data aggregation method and system. During the operation of the server, the operation parameters of the server can be collected and the corresponding operation status scores can be calculated. When the operation status score is within the first preset range, the first operation status table is called to screen out abnormal operation parameters by comparison and generate a first prompt message, so as to facilitate the debugging of the abnormal operation parameters. When the operation status score is within the second preset range, the second operation status table is called to screen out warning operation parameters by comparison and generate a second prompt message, and the sampling frequency of the warning server is increased to improve the monitoring ability of the warning server, thereby improving the abnormal discovery efficiency.
[0071] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other implementation manner extended based on the solution of the present application without creative efforts belongs to the protection scope of the present application.
Claims
1. A data aggregation method, characterized in that, Including: Obtain the operation status data of the servers located in the computer room; Calculate the operation status score of the servers according to the operation status data; The operation status score is calculated based on the real-time values of the respective operation parameters of the servers; If the operation status score is within a first preset range, call a first set of operation status tables; The first set of operation status tables includes multiple first operation status tables; wherein, the first operation status table includes the first operation parameters of the servers and the first operation parameter thresholds; obtain the first operation status table from the first set of operation status tables according to the current business of the servers; compare the real-time values of the first operation parameters of the servers with the first operation parameter thresholds to obtain abnormal operation parameters; the abnormal operation parameters are the parameters whose real-time values are greater than or equal to the first operation parameter thresholds; generate a first prompt message based on the abnormal operation parameters; the first prompt message is used to indicate that the abnormal operation parameters need to be debugged; If the operation status score is within a second preset range, call a second set of operation status tables; the second set of operation status tables includes multiple second operation status tables; the second operation status table includes the second operation parameters of the servers and the second operation parameter thresholds; the first operation parameters are the same as the second operation parameters, and the first operation parameter thresholds are different from the second operation parameter thresholds; obtain the second operation status table from the second set of operation status tables according to the current business of the servers; compare the real-time values of the second operation parameters of the servers with the second operation parameter thresholds to obtain warning operation parameters; the warning operation parameters are the parameters that are greater than or equal to the second operation parameter thresholds; and the warning operation parameters are less than the first operation parameter thresholds; generate a second prompt message based on the warning operation parameters; the second prompt message is used to indicate that the warning operation parameters need to be continuously tracked and the sampling frequency of the warning operation parameters needs to be increased.
2. The data aggregation method according to claim 1, wherein If the number of servers in the computer room is greater than or equal to 2, the increasing of the sampling frequency of the warning operation parameters specifically includes: Increase the sampling frequency of the operation parameters of the warning servers; the warning servers are the servers with the warning operation parameters.
3. The data aggregation method according to claim 1, wherein The first operation status table further includes the business type, and the operation status thresholds in the first operation status tables corresponding to different business types are different; further including: In the case of a change in the business type of the servers, obtain the first operation status table corresponding to the changed business type from the first set of operation status tables according to the changed business type.
4. The data aggregation method according to claim 1, wherein The first operation status table further includes a time identifier, and the time identifiers corresponding to the servers operating in different time periods are different; obtaining the first operation status table from the first set of operation status tables according to the current business of the servers specifically includes: Obtain the first operation status table from the first set of operation status tables according to the current business of the servers and the system time.
5. The data aggregation method according to claim 1, wherein If the server includes the warning operation parameters, mark the server as a warning server; if the warning server changes to an abnormal server within a preset time, replace the corresponding first operation parameter threshold in the first operation status table with the second operation parameter threshold corresponding to the warning operation parameters in the second operation status table.
6. The data aggregation method according to claim 5, characterized in that, The replacing the corresponding first operation parameter threshold in the first operation status table with the second operation parameter threshold corresponding to the warning operation parameters in the second operation status table specifically includes: Calculating the difference between the second operation parameter threshold and the first operation parameter threshold; If the difference is greater than or equal to the first difference threshold, calculating a first replacement parameter based on the second operation parameter threshold and the first operation parameter threshold; replacing the first operation parameter according to the first replacement parameter; If the difference is less than the first difference threshold, replacing the first operation parameter based on the second operation parameter threshold.
7. The data aggregation method according to claim 6, wherein It further includes: Calculating the real-time score of the first operation status table when the first operation parameter threshold in the first operation status table is replaced; If the real-time score is outside the first preset range, increasing the replaced first operation parameter threshold so that the real-time score of the first operation status table returns within the first preset range.
8. The data aggregation method according to claim 1, wherein If the operation status score is within the first preset range, it further includes: Marking the server as an abnormal server; Reducing the service response ratio of the abnormal server and increasing the service response ratios of non-abnormal servers and non-warning servers.
9. A data aggregation system, characterized in that, It includes: A sampling device and a control device; The sampling device is used to obtain the operation status data of the servers located in the computer room; The control device is used to calculate the operation status score of the server according to the operation status data; the operation status score is calculated based on the real-time values of the respective operation parameters of the server; The control device is further used to call the first operation status table set when the operation status score is within the first preset range; the first operation status table set includes multiple first operation status tables; wherein, the first operation status table includes the first operation parameters of the server and the first operation parameter thresholds; obtaining the first operation status table from the first operation status table set according to the current service of the server; comparing the real-time value of the first operation parameter of the server with the first operation parameter threshold to obtain abnormal operation parameters; the abnormal operation parameters are parameters whose real-time value is greater than or equal to the first operation parameter threshold; generating a first prompt message based on the abnormal operation parameters; the first prompt message is used to indicate that the abnormal operation parameters need to be debugged; The control device is further configured to call a second set of operating state tables when the operating state score is within a second preset range; the second set of operating state tables includes a plurality of second operating state tables; the second operating state table includes second operating parameters of the server and second operating parameter thresholds; the first operating parameters are the same as the second operating parameters, and the first operating parameter thresholds are different from the second operating parameters; obtain a second operating state table from the second set of operating state tables according to the current service of the server; compare the real-time value of the second operating parameter of the server with the second operating parameter threshold to obtain a warning operating parameter; the warning operating parameter is a parameter greater than or equal to the second operating parameter threshold; and the warning operating parameter is less than the first operating parameter threshold; generate a second prompt message based on the warning operating parameter; the second prompt message is used to indicate that the warning operating parameter needs to be continuously tracked and increase the sampling frequency of the warning operating parameter.
Citation Information
Patent Citations
Motor fault diagnosis method and related equipment thereof
CN112213640A
Power distribution network tower abnormity monitoring method and system
CN119334401A
Server running state monitoring management system based on AI intelligence
CN119668976A
Coal mine safety monitoring data self-diagnosis method
CN119914362A
Power distribution network fault early warning and positioning system based on distributed traveling wave online measurement
CN119959693A