A server sensor-assisted testing method
By dynamically determining the temperature alarm threshold and optimizing the heating distance, the problems of cumbersome operation, inaccurate temperature control and low heating efficiency during server component testing are solved, achieving safe and effective heating of components and accurate test results.
Patent Information
- Application Number
- CN202411278453.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-09-12
AI Technical Summary
The existing server component sensor testing process is cumbersome, with inaccurate temperature control, easy damage to components, and low heating efficiency.
By obtaining sensor identification and component information, the temperature alarm threshold is dynamically determined, and light control signals are used to indicate temperature exceeds the limit. The heating distance is optimized by combining cluster architecture and test records to achieve accurate temperature monitoring and effective heating.
Reduce component damage, improve test result accuracy, avoid component damage, improve heating efficiency, and ensure safe and effective heating of components.
Smart Images

Figure CN119248598B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of server sensor testing, and in particular to a server sensor-assisted testing method. Background Art
[0002] Currently, server component sensor temperature rise and fall testing is primarily performed by logging into the BMC web page and constantly refreshing the page to obtain real-time sensor temperatures. Server sensors are numerous and diverse, so after refreshing the web page, you must click on the page containing a specific sensor to view it. This method is inconvenient and cumbersome for testing server component sensors. To simulate the actual temperature rise experienced by server components during use, a heat gun is required to heat the corresponding components.
[0003] The following technical problems often occur in the existing simulation test process:
[0004] First, improper operation of the hot air gun can easily cause the temperature of the component to rise too quickly and fail to detect the temperature change in time, resulting in excessive temperature of the component and damaging the corresponding component, shortening its service life or causing irreparable damage;
[0005] Second, the temperature alarm threshold is generally fixed and cannot be changed dynamically, resulting in inaccurate test results or exceeding the temperature limit of some components and damaging the components.
[0006] Third, the heating distance is generally determined by the tester himself. If the distance is too far, the heating efficiency will be low and the components cannot be effectively heated. If the distance is too close, the server will be easily damaged. Summary of the Invention
[0007] This summary is intended to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] The present invention proposes a server sensor-assisted testing method to solve one or more of the technical problems mentioned in the above background technology section.
[0009] The present invention provides a server sensor-assisted testing method, the method comprising:
[0010] Obtaining a server identifier, an Internet Protocol address, a sensor identifier of a target sensor, and login information of the target server in the server cluster to be tested, wherein each server in the server cluster to be tested is configured with multiple sensors, and the multiple sensors are used to collect temperature data of different components of the server;
[0011] The sensor ID of the target sensor is used to query the pre-configured sensor information list to obtain the sensor type corresponding to the sensor ID, the component name corresponding to the target sensor, and the component operation time; based on the sensor type, the component name corresponding to the target sensor, and the component operation time, the temperature alarm threshold corresponding to the target sensor is determined;
[0012] Use the login information to establish a connection with the server corresponding to the server identifier and obtain the temperature data of the target sensor; compare the temperature data with the temperature alarm threshold. If the temperature data is greater than the temperature alarm threshold, generate a light control signal for the prompt light and send the light control signal to the target server so that the target server can control the light of the prompt light.
[0013] Optionally, a temperature alarm threshold corresponding to the target sensor is determined based on the sensor type, the component name corresponding to the target sensor, and the component's operating time, including:
[0014] According to the component name corresponding to the target sensor, query the historical temperature alarm threshold corresponding to the component name and the update timestamp corresponding to the historical temperature alarm threshold;
[0015] Determine whether the historical temperature alarm threshold has expired based on the update timestamp. If the historical temperature alarm threshold has expired, obtain the corresponding component aging curve based on the component name, and query the aging rate corresponding to the component operation time from the component aging curve. The component aging curve represents the corresponding relationship between the component operation time and the aging rate;
[0016] According to the component name corresponding to the target sensor, determine the target temperature load level corresponding to the component name from multiple pre-set temperature load levels, determine the temperature load range corresponding to the target temperature load level, and extract the temperature upper limit value of the temperature load range;
[0017] Determine the temperature alarm threshold corresponding to the target sensor based on the aging rate and temperature upper limit value corresponding to the component operation time.
[0018] Optionally, before querying the pre-configured sensor information list using the sensor identifier of the target sensor to obtain the sensor type corresponding to the sensor identifier, the component name corresponding to the target sensor, and the component operating time, the following steps may be further included:
[0019] Determine a cluster architecture diagram corresponding to the server cluster to be tested, where the cluster architecture diagram includes multiple nodes and edges connecting different nodes, where each node corresponds to a server in the server cluster to be tested, and edges represent associations between different servers;
[0020] Grouping multiple nodes to obtain multiple node groups, the multiple node groups including a root node group, an intermediate node group, and a child node group; configuring a corresponding importance factor for each node group; and
[0021] Determine the temperature alarm threshold corresponding to the target sensor based on the sensor type, the component name corresponding to the target sensor, and the component's operating time, including:
[0022] Determine the reference temperature alarm threshold corresponding to the target sensor based on the sensor type, the component name corresponding to the target sensor, and the component operation time;
[0023] Determine the node group to which the target server belongs, and determine the importance coefficient corresponding to the target server based on the importance factor corresponding to the node group to which it belongs;
[0024] The temperature alarm threshold corresponding to the target sensor is generated according to the importance coefficient and the reference temperature alarm threshold.
[0025] Optionally, determining the importance coefficient corresponding to the target server based on the importance factor corresponding to the node group to which it belongs includes:
[0026] If the node group to which the target server belongs is a child node group, the priorities of the multiple nodes in the child node group are determined respectively, and the nodes in the child node group are sorted according to the priorities to obtain a child node sequence;
[0027] The ranking of the target server in the subnode sequence is determined, and the importance coefficient corresponding to the target server is determined according to the ranking of the target server in the subnode sequence and the importance factor of the subnode group.
[0028] Optionally, each component of the target server is configured with a prompt light, and each prompt light is configured with multiple light modes; the light control signal includes the prompt light number and the target light mode; and
[0029] Sending a light control signal to a target server so that the target server controls the light of the prompt light, including:
[0030] The light control signal is sent to the target server, so that the target server controls the light of the prompt light corresponding to the prompt light number, so that the prompt light corresponding to the prompt light number enters the target light mode.
[0031] The present invention has the following beneficial effects:
[0032] 1. By establishing a connection between the client and the server, the system obtains temperature data from the sensor component. This temperature data is compared with the temperature alarm threshold, generating a lighting control signal that triggers a warning light to reduce damage to the component. Specifically, the system obtains various information about the target server, establishes a connection with the target server, and then obtains temperature data from the target sensor. By obtaining this information, the corresponding temperature alarm threshold is determined. The obtained temperature data is compared with the temperature alarm threshold. When the temperature exceeds the temperature alarm threshold, the client generates a lighting control signal and sends it to the server, causing the warning light to sound an alarm. This reduces damage to the component and prevents further damage.
[0033] 2. The temperature alarm threshold corresponding to the sensor is determined by the reference temperature alarm threshold and the importance coefficient corresponding to the server, thereby improving the accuracy of the test results and avoiding damage to components. Specifically, the main reason for inaccurate test results and damage to components is that the temperature alarm threshold is fixed, but the temperature that each component can withstand is different. Based on this, the present invention determines the reference temperature alarm threshold corresponding to the sensor through the sensor type, the component name corresponding to the sensor, and the running time of the component, determines the importance coefficient corresponding to the target server through the importance factor corresponding to the node group to which the target server belongs, and determines the temperature alarm threshold corresponding to the sensor based on the reference temperature alarm threshold and the importance coefficient. In this way, a temperature alarm threshold is set for each component, which improves the accuracy of the test results and avoids damage to the component due to the temperature alarm threshold exceeding the component's tolerance temperature.
[0034] 3. Through test records, determine the key monitoring component group and recommended heating distance, improve heating efficiency, and enable the components to be effectively heated. Specifically, the main reason for the low heating efficiency is that the heating distance is random, resulting in uneven heating of the components and poor heating effect. Based on this, the present invention obtains the target test record group to obtain the test records, sorts the components according to the number of alarms of different components in the test records, and obtains a component sequence. A target number of components are selected to form a key monitoring component group, and the sum of the heating distances corresponding to different components is averaged to obtain the average heating distance, which is the recommended heating distance. This improves the heating efficiency and enables the components to be effectively heated. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the elements are not necessarily drawn to scale.
[0036] Figure 1The present invention is a flow chart of a server sensor-assisted testing method. DETAILED DESCRIPTION
[0037] The present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0038] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features of the embodiments of the present invention may be combined with each other.
[0039] It should be noted that the concepts of "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0040] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0041] The names of the messages or information exchanged between multiple devices of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0042] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0043] like Figure 1 FIG. 1 is a flow chart of a server sensor-assisted testing method of the present invention, which specifically includes the following steps:
[0044] Step 101: Obtain the server identifier, Internet Protocol address, sensor identifier of the target sensor, and login information of the target server in the server cluster to be tested, wherein each server in the server cluster to be tested is configured with multiple sensors, and the multiple sensors are used to collect temperature data of different components of the server.
[0045] In some embodiments, the execution subject of the flowchart of the sensor-assisted testing method for a server of the present invention may be a client, which may be a desktop computer or a laptop computer.
[0046] In practice, the execution entity can locally store various information about each server in the server cluster to be tested. The server cluster to be tested includes each server. Based on this information, the execution entity can obtain the target server's server ID, Internet Protocol address, sensor ID of the target sensor, and login information for the target server. The target server can be any one or more servers. The server ID is a number that identifies each server, with each server ID corresponding to one server. The Internet Protocol address can be an IP address. The login information can be a username and password.
[0047] In practice, each server in the server cluster under test is equipped with multiple sensors. Each sensor is configured with a sensitive element, a conversion element, and a communication serial port, as needed. The communication serial port is connected to the server and outputs electrical signals to the server via the communication serial port. Each sensor senses temperature changes of different components through the sensitive element, then converts this temperature change information into an electrical signal through the conversion element. This electrical signal is then output to the server via the communication serial port, obtaining temperature data for different components. These components include, but are not limited to, disk array cards, central processing units, and disk backplanes.
[0048] In practice, multiple sensors collect temperature data of different components. In order to distinguish the temperature data collected by each sensor, each sensor can be numbered to obtain a sensor identifier for each sensor. Each sensor corresponds to a component.
[0049] In step 102, the sensor identifier of the target sensor is used to query the pre-configured sensor information list to obtain the sensor type corresponding to the sensor identifier, the component name corresponding to the target sensor, and the component operation time; based on the sensor type, the component name corresponding to the target sensor, and the component operation time, the temperature alarm threshold corresponding to the target sensor is determined.
[0050] In some embodiments, the sensor ID of the target sensor is queried in a pre-configured sensor information list to obtain the sensor type corresponding to the sensor ID, the component name corresponding to the target sensor, and the component runtime. For example, the sensor ID of the target sensor may be sensor 3. For sensor 3, the pre-configured sensor information list is queried to obtain the sensor type corresponding to sensor 3, the component name corresponding to sensor 3, and the component runtime. In practice, the sensor information in the pre-configured sensor information list includes: sensor ID, sensor type, component name corresponding to the sensor, and component runtime.
[0051] Optionally, the temperature alarm threshold corresponding to the target sensor is determined according to the sensor type, the component name corresponding to the target sensor, and the component operation time by the following steps:
[0052] Step 1: According to the component name corresponding to the target sensor, query the historical temperature alarm threshold corresponding to the component name and the update timestamp corresponding to the historical temperature alarm threshold.
[0053] In some embodiments, the executing entity may locally store a component information list. The component information list includes the component name, historical temperature alarm thresholds, and update timestamps corresponding to the historical temperature alarm thresholds. For example, the component name corresponding to the target sensor is the central processing unit (CPU), and the component has been operating for six months. The CPU is searched in the component information list to obtain the CPU's corresponding historical temperature alarm thresholds and the update timestamps corresponding to the historical temperature alarm thresholds.
[0054] Step 2: Determine whether the historical temperature alarm threshold has expired based on the update timestamp. If the historical temperature alarm threshold has expired, obtain the corresponding component aging curve based on the component name, and query the aging rate corresponding to the component operating time from the component aging curve. The component aging curve represents the corresponding relationship between the component operating time and the aging rate.
[0055] In some embodiments, as an example, an expiration time threshold is set, and the expiration time threshold may be 3 months. In practice, if the update timestamp of the historical temperature alarm threshold corresponding to the central processor obtained by the execution subject is July 3, and the current timestamp is September 3, the time difference is 2 months obtained by subtracting the update timestamp of the historical temperature alarm threshold corresponding to the central processor from the current timestamp, and the time difference is less than the expiration time threshold, then the historical temperature alarm threshold has not expired. If the update timestamp of the historical temperature alarm threshold corresponding to the central processor obtained by the execution subject is May 3, and the current timestamp is September 3, the time difference is 3 months obtained by subtracting the update timestamp of the historical temperature alarm threshold corresponding to the central processor from the current timestamp, and the time difference is greater than the expiration time threshold, then the historical temperature alarm threshold has expired.
[0056] In some embodiments, the execution entity may locally store a component aging curve atlas. The component aging curve atlas includes a component aging curve graph corresponding to each component. As an example, if the historical temperature alarm threshold has expired, the component aging curve graph corresponding to the central processing unit is obtained from the locally stored component aging curve atlas, and the aging rate corresponding to the component operation time is queried from the component aging curve graph. The component aging curve graph represents the corresponding relationship between the component operation time and the aging rate. As an example, the operation time of the central processing unit is half a year, and the corresponding aging rate is 0.9.
[0057] Step three: according to the component name corresponding to the target sensor, determine the target temperature load level corresponding to the component name from multiple pre-set temperature load levels, determine the temperature load range corresponding to the target temperature load level, and extract the temperature upper limit value of the temperature load range.
[0058] In some embodiments, the execution entity has locally stored multiple pre-set temperature load levels. The temperature load levels in the pre-set multiple temperature load levels include component name, temperature load level, and temperature load range. As an example, the component name corresponding to the target sensor is the central sensor. The target temperature load level corresponding to the central sensor is queried from the pre-set multiple temperature load levels. The target temperature load level can be level 1, and the temperature load range corresponding to level 1 can be 50-60°C. The upper temperature limit of the temperature load range is 60°C.
[0059] Step 4: Determine the temperature alarm threshold corresponding to the target sensor based on the aging rate and the upper temperature limit corresponding to the component operation time.
[0060] In some embodiments, the temperature alarm threshold for the target sensor is determined based on the aging rate and upper temperature limit corresponding to the component's operating time. For example, if the central sensor has been operating for six months, has an aging rate of 0.9, and an upper temperature limit of 60°C, the temperature alarm threshold for the target sensor is determined by multiplying the aging rate and the upper temperature limit. The temperature alarm threshold for the target sensor is 54°C.
[0061] Step 103, using the login information to establish a connection with the target server corresponding to the server identifier, and obtain the temperature data of the target sensor; compare the temperature data with the temperature alarm threshold, if the temperature data is greater than the temperature alarm threshold, generate a light control signal for the prompt light, and send the light control signal to the target server, so that the target server controls the light of the prompt light.
[0062] In some embodiments, a connection is established with a target server corresponding to the server identifier using login information, and the temperature data of the target sensor is obtained. In practice, the execution subject communicates with the server through the IPMI interface protocol. On this basis, the execution subject can obtain the temperature data of the target sensor. In practice, the temperature data corresponding to the target sensor is compared with the temperature alarm threshold to obtain a comparison result. If the temperature data is less than or equal to the temperature alarm threshold, there is no need to generate a light control signal. If the temperature data is greater than the temperature alarm threshold, a light control signal is generated and sent to the target server. The target server issues a light-on instruction to the prompt light according to the light control signal, and controls the light of the prompt light to turn on. 60 seconds after the light control signal is issued, the target server issues a light-off instruction to control the light of the prompt light to turn off.
[0063] In these embodiments, a client establishes a connection with a server, obtains temperature data from a sensor component, compares the temperature data with a temperature alarm threshold, generates a lighting control signal, and issues a warning through lighting. This ultimately implements a timely warning of overtemperature conditions, minimizing damage to the corresponding component. Specifically, various information about a target server is obtained, a connection is established, and temperature data from a target sensor is obtained. By obtaining various information about the target sensor, a temperature alarm threshold corresponding to the target sensor is determined. The obtained temperature data is compared with the temperature alarm threshold. When the temperature data exceeds the temperature alarm threshold, the client generates a lighting control signal and sends it to the server, causing a warning light to sound an alarm, thereby minimizing damage to the corresponding component.
[0064] In some embodiments, in order to further address the second technical problem described in the background technology section, namely, that "the temperature alarm threshold is generally fixed and cannot be changed dynamically, resulting in inaccurate test results or exceeding the temperature limit of some components and damaging the components", in some embodiments of the present invention, the above method further includes the following steps:
[0065] Step 1: Determine the cluster architecture diagram corresponding to the server cluster to be tested. The cluster architecture diagram includes multiple nodes and edges connecting different nodes, where each node corresponds to a server in the server cluster to be tested, and the edges represent the association relationship between different servers.
[0066] In some embodiments, the execution entity has locally stored a cluster architecture diagram corresponding to each server cluster. Based on this, the cluster architecture diagram corresponding to the server cluster to be tested is determined. The cluster architecture diagram includes multiple nodes and edges connecting different nodes, wherein each node corresponds to a server in the server cluster to be tested, and the edges represent the association relationship between different servers.
[0067] Step 2: Group multiple nodes to obtain multiple node groups, including a root node group, an intermediate node group, and a child node group; and configure a corresponding importance factor for each node group.
[0068] On this basis, the temperature alarm threshold corresponding to the target sensor is determined according to the sensor type, the component name corresponding to the target sensor, and the component operation time, which also includes the following sub-steps:
[0069] Sub-step 1: determining a reference temperature alarm threshold corresponding to the target sensor based on the sensor type, the component name corresponding to the target sensor, and the component operation time;
[0070] Sub-step 2: determining the node group to which the target server belongs, and determining the importance coefficient corresponding to the target server based on the importance factor corresponding to the node group to which it belongs;
[0071] Sub-step three: generating a temperature alarm threshold corresponding to the target sensor according to the importance coefficient and the reference temperature alarm threshold.
[0072] In some embodiments, multiple nodes are grouped to obtain multiple node groups, and the multiple node groups include a root node group, an intermediate node group, and a child node group. The root node group is the starting point in the entire cluster architecture. The intermediate node group is located between the root node group and the child node group. The child node group is located below the intermediate node group. In practice, a corresponding importance factor is configured for each node group. The importance factor represents the importance of each node group in the cluster architecture diagram. As an example, the importance factor of the root node group can be 1, the importance factor of the intermediate node group can be 0.5, and the importance factor of the child node group can be 0.1.
[0073] In some embodiments, a base temperature alarm threshold corresponding to the target sensor is determined based on the sensor type, the name of the component corresponding to the target sensor, and the operating time of the component. The temperature alarm threshold determined in the above embodiment is used as the base temperature alarm threshold. As an example, if the sensor type is a temperature sensor, the name corresponding to the target sensor is a central processing unit, and the operating time of the component is six months, then the base temperature alarm threshold corresponding to the target sensor is 54°C.
[0074] In some embodiments, the node group to which the target server belongs is determined, and the importance coefficient corresponding to the target server is determined based on the importance factor corresponding to the node group. As an example, the node group to which the target server belongs may be a child node group.
[0075] Optionally, the importance coefficient corresponding to the target server is determined based on the importance factor corresponding to the node group to which it belongs. The following steps are also included:
[0076] Step 1: If the node group to which the target server belongs is a child node group, the priorities of the multiple nodes in the child node group are determined respectively, and the nodes in the child node group are sorted according to the priorities to obtain a child node sequence.
[0077] In some embodiments, if the node group to which the target server belongs is a subnode group, each subnode is sorted according to the maximum number of connections to each server corresponding to the multiple subnodes to determine the priority of the multiple nodes in the subnode group. As an example, the multiple nodes in the subnode group may be node A, node B, and node C. The maximum number of connections to the server corresponding to node A may be 8, the maximum number of connections to the server corresponding to node B may be 6, and the maximum number of connections to the server corresponding to node C may be 7. The priority of the multiple nodes is node A takes precedence over node C which takes precedence over node B.
[0078] In practice, the nodes in the child node group are sorted according to priority, and the ranking of node A may be 1, the ranking of node B may be 3, and the ranking of node C may be 2. The resulting child node sequence is node A, node C, and node B.
[0079] Step 2: Determine the ranking of the target server in the subnode sequence, and determine the importance coefficient corresponding to the target server based on the ranking of the target server in the subnode sequence and the importance factor of the subnode group.
[0080] In some embodiments, as an example, the rankings in the subnode sequence are scored, with ranking 1 having a score of 10, ranking 2 having a score of 8, and ranking 3 having a score of 6. The target server may be node C, whose ranking in the subnode sequence is 2, and whose ranking score is 8. In practice, the importance factor of the subnode group may be 0.1, and the product of the ranking score of the target server in the subnode sequence and the importance factor of the subnode group is calculated to obtain an importance coefficient of 0.8 corresponding to the target server.
[0081] In some embodiments, a temperature alarm threshold corresponding to the target sensor is generated based on the importance coefficient and the reference temperature alarm threshold. For example, the importance coefficient may be 0.8, the reference temperature alarm threshold may be 54°C, and the reference temperature alarm threshold is divided by the importance coefficient to obtain a temperature alarm threshold corresponding to the target sensor of 67.5°C.
[0082] Optionally, some embodiments of the present invention further include the following steps:
[0083] In step 1, each component of the target server is configured with a prompt light, and each prompt light is configured with multiple light modes; a light control signal includes a prompt light number and a target light mode.
[0084] On this basis, sending the light control signal to the target server so that the target server controls the light of the prompt light also includes the following sub-steps:
[0085] Sub-step one: sending a light control signal to a target server, so that the target server controls the light of the indicator light corresponding to the indicator light number, so that the indicator light corresponding to the indicator light number enters a target light mode.
[0086] In some embodiments, each component of the target server is configured with a notification light, and each notification light is configured with multiple lighting modes. The lighting modes can be a light-on mode or a light-off mode. In practice, each notification light is numbered to obtain a notification light number. The control signal in the lighting control signal includes the notification light number and the target lighting mode.
[0087] In some embodiments, as an example, the indicator light number may be 3, and the target light mode may be the light-on mode. In practice, a light control signal is sent to the target server, and the target server controls the light of the indicator light numbered 3, thereby controlling the indicator light numbered 3 to enter the light-on mode.
[0088] In these embodiments, the temperature alarm threshold corresponding to the sensor is determined by the reference temperature alarm threshold and the importance coefficient corresponding to the server, thereby improving the accuracy of the test results and avoiding damage to components. Specifically, the main reason for inaccurate test results and damage to components is that the temperature alarm threshold is fixed, but the temperature that each component can withstand is different. Based on this, the present invention determines the reference temperature alarm threshold corresponding to the sensor through the sensor type, the component name corresponding to the sensor, and the running time of the component, determines the importance coefficient corresponding to the target server through the importance factor corresponding to the node group to which the target server belongs, and determines the temperature alarm threshold corresponding to the sensor based on the reference temperature alarm threshold and the importance coefficient. In this way, a temperature alarm threshold is set for each component, which improves the accuracy of the test results and avoids damage to the component due to the temperature alarm threshold exceeding the component's tolerance temperature.
[0089] In some embodiments, in order to further address the third technical issue described in the background technology section, namely, "the heating distance is generally determined by the tester himself. If the distance is too far, the heating efficiency will be low and the components will not be effectively heated. If the distance is too close, the server will be easily damaged." In some embodiments of the present invention, the above method further includes the following steps:
[0090] Step 1: Obtain the test record set corresponding to the server cluster to be tested within the target time period from the test record database, and filter out the test records with temperature alarm conditions from the test record set to form a target test record group. The test records in the target test record group include the test time, server ID, the name of the component where the temperature alarm occurs, and the heating distance.
[0091] In some embodiments, the execution entity has a test record database stored locally. On this basis, the target time period can be set to 3 months, and the test record set corresponding to the server cluster to be tested within 3 months is obtained. The test record set includes the test records corresponding to each server cluster to be tested. In practice, the test records with temperature alarm conditions are filtered out from the test record set to form a target test record group. Among them, the test records in the target test record group include the test time, server identification, the name of the component that caused the temperature alarm, and the heating distance. The heating distance can be 1 cm or 2 cm.
[0092] Step 2: According to the target test record group, the number of temperature alarms and average heating distances corresponding to different components of the server are counted, and different components are sorted in descending order according to the number of temperature alarms to obtain a component sequence.
[0093] In some embodiments, the different components of the server may be disk array cards, central processing units (CPUs), and disk backplanes. The number of temperature alarms and average heating distances corresponding to the different components are counted from the target test record group. The average heating distance is obtained by averaging the sum of the heating distances corresponding to the different components. For example, the number of alarms for the disk array card may be 3, the number of alarms for the CPU may be 5, and the number of alarms for the disk backplane may be 2. The different components are sorted in descending order of alarm count to obtain a component sequence, which may be the CPU, disk array card, and disk backplane.
[0094] Step 3: Select a target number of components from the component sequence in descending order of the number of temperature alarms to form a key monitoring component group.
[0095] In some embodiments, as an example, the target number can be set to 2, and 2 components are selected from the component sequence in descending order of the number of temperature alarms to form a key monitoring component group. The components in the key monitoring component group can be a central processing unit and a disk array card.
[0096] Step 4: For each component in the key monitoring component group, generate operation prompt information corresponding to the component, the operation prompt information includes a recommended heating distance, and the recommended heating distance is generated based on the average heating distance of the component.
[0097] In some embodiments, as an example, a component in the key monitoring component group may be a central processing unit. Based on this, operation prompt information corresponding to the central processing unit is generated. The operation prompt information may be a recommended heating distance. The average heating distance obtained in the above embodiment is used as the recommended heating distance.
[0098] In these embodiments, the key monitoring component group and the recommended heating distance are determined through test records, and the heating efficiency is improved, so that the components can be effectively heated. Specifically, the main reason for the low heating efficiency is that the heating distance is random, which leads to uneven heating of the components and poor heating effect. Based on this, the present invention obtains the test records by obtaining the target test record group, and sorts the components according to the number of alarms of different components in the test records to obtain a component sequence. A target number of components are selected to form a key monitoring component group, and the sum of the heating distances corresponding to different components is averaged to obtain the average heating distance, that is, the recommended heating distance. Thereby, the heating efficiency is improved, the components can be effectively heated and damage is avoided.
[0099] The above descriptions are merely some preferred embodiments of the present invention and illustrate the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present invention.
Claims
1. A server sensor-assisted testing method, applied to a client, characterized in that: include: Obtaining a server identifier, an Internet Protocol address, a sensor identifier of a target sensor, and login information of the target server in the server cluster to be tested, wherein each server in the server cluster to be tested is configured with multiple sensors, and the multiple sensors are used to collect temperature data of different components of the server; Querying a pre-configured sensor information list through the sensor identifier of the target sensor to obtain a sensor type corresponding to the sensor identifier, a component name corresponding to the target sensor, and component operation time; determining a temperature alarm threshold corresponding to the target sensor based on the sensor type, component name corresponding to the target sensor, and component operation time; Using the login information, a connection is established with a target server corresponding to the server identifier, and temperature data of a target sensor is obtained; the temperature data is compared with a temperature alarm threshold, and if the temperature data is greater than the temperature alarm threshold, a light control signal for an indicator light is generated, and the light control signal is sent to the target server, so that the target server controls the light of the indicator light; The step of determining the temperature alarm threshold corresponding to the target sensor according to the sensor type, the component name corresponding to the target sensor, and the component operation time includes: According to the sensor type and component name corresponding to the target sensor, query the historical temperature alarm threshold corresponding to the component name and the update timestamp corresponding to the historical temperature alarm threshold; Determining whether the historical temperature alarm threshold has expired according to the update timestamp, and if the historical temperature alarm threshold has expired, obtaining a corresponding component aging curve graph according to the component name, and querying an aging rate corresponding to the component operation time from the component aging curve graph, wherein the component aging curve graph represents a corresponding relationship between the component operation time and the aging rate; According to the component name corresponding to the target sensor, determining a target temperature load level corresponding to the component name from a plurality of pre-set temperature load levels, determining a temperature load range corresponding to the target temperature load range, and extracting a temperature upper limit value of the temperature load range; Determining a temperature alarm threshold corresponding to the target sensor according to an aging rate corresponding to the operating time of the component and the upper temperature limit; Before querying the pre-configured sensor information list using the sensor identifier of the target sensor to obtain the sensor type corresponding to the sensor identifier, the component name corresponding to the target sensor, and the component operation time, the method further includes: Determine a cluster architecture diagram corresponding to the server cluster to be tested, wherein the cluster architecture diagram includes a plurality of nodes and edges connecting different nodes, wherein each node corresponds to a server in the server cluster to be tested, and the edges represent associations between different servers; Grouping the multiple nodes to obtain multiple node groups, the multiple node groups including a root node group, an intermediate node group, and a child node group; configuring a corresponding importance factor for each node group; and The determining, based on the sensor type, the component name corresponding to the target sensor, and the component operation time, of a temperature alarm threshold corresponding to the target sensor includes: Determine a reference temperature alarm threshold corresponding to the target sensor according to the sensor type, the component name corresponding to the target sensor, and the component operation time; Determine the node group to which the target server belongs, and determine the importance coefficient corresponding to the target server based on the importance factor corresponding to the node group to which it belongs; A temperature alarm threshold corresponding to the target sensor is generated according to the importance coefficient and the reference temperature alarm threshold.
2. The server sensor-assisted testing method according to claim 1, wherein: Determining the importance coefficient corresponding to the target server according to the importance factor corresponding to the node group to which the target server belongs includes: If the node group to which the target server belongs is a child node group, the priorities of the multiple nodes in the child node group are determined respectively, and the nodes in the child node group are sorted according to the priorities to obtain a child node sequence; The ranking of the target server in the sub-node sequence is determined, and the importance coefficient corresponding to the target server is determined according to the ranking of the target server in the sub-node sequence and the importance factor of the sub-node group.
3. The server sensor-assisted testing method according to claim 2, wherein: Each component of the target server is configured with a prompt light, and each prompt light is configured with multiple light modes; the light control signal includes a prompt light number and a target light mode; as well as The step of sending the light control signal to the target server so that the target server controls the light of the prompt light includes: The light control signal is sent to the target server, so that the target server controls the light of the prompt light corresponding to the prompt light number, so that the prompt light corresponding to the prompt light number enters a target light mode.
Citation Information
Patent Citations
Method and device for obtaining early warning threshold value and storage medium
CN107861915A
Method, device and equipment for lightening fault lamp of server system and readable medium
CN113448811A