Machine room fault positioning dynamic environment monitoring system
By deploying a current sensor group and a power supply path database in the computer room, combined with a central processing unit and a human-computer interaction interface, the ambiguity and timeliness issues of locating power supply faults in the computer room are resolved, enabling fast and accurate fault tracing and false alarm suppression, and improving the reliability and operation and maintenance efficiency of the computer room power supply system.
Patent Information
- Application Number
- CN202510915981.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-03
AI Technical Summary
The existing computer room monitoring system has positioning ambiguity, timeliness and reliability contradictions when locating power supply faults. It is unable to quickly and accurately locate the source of the fault and is easily disturbed by instantaneous fluctuations when equipment starts and stops, resulting in a high false alarm rate and a long positioning time.
Deploy a current sensor group and a central processing unit, combine it with the power supply path database and human-computer interaction interface, and implement current value monitoring, fault tracing, and false alarm filtering. Through real-time electrical parameter comparison and topology verification, the threshold is dynamically adjusted to support accurate positioning in dual power switching scenarios.
It has achieved accurate cross-level tracing of fault sources, shortened positioning time by 90%, improved fault determination accuracy by 85%, reduced false alarm rate by 92%, and improved power supply system reliability and operation and maintenance efficiency.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer room monitoring, and more particularly to a computer room fault location dynamic environment monitoring system. Background Art
[0002] In the field of computer room monitoring, there are significant technical bottlenecks in accurately tracing power supply system faults. Traditional monitoring solutions usually deploy current detection devices in the cabinet-level power distribution unit, but they can only identify current anomalies at this level and cannot distinguish whether the fault source is located in the cabinet unit itself, the upstream terminal cabinet branch, the UPS output cabinet circuit breaker, or the AC input cabinet switch. This positioning ambiguity stems from the hierarchical structure characteristics of the power supply link: the AC input cabinet supplies power to the UPS equipment through the circuit breaker, the UPS output cabinet is connected to the terminal cabinet through the branch, and finally the terminal cabinet distributes power to multiple cabinet-level units. Due to the lack of digital modeling of the electrical connection relationship of the entire link, when a fault occurs, the operation and maintenance personnel need to manually review the offline drawings and check the status of the upstream equipment at each level in turn, resulting in an average positioning process time of 20 to 60 minutes.
[0003] The core reasons for the difficulty in locating faults stem from two technical flaws: First, incomplete sensor coverage means a lack of real-time current monitoring at key upstream nodes (such as the main cabinet shunts and UPS output cabinet circuit breakers), making it impossible to analyze the fault transmission path through data correlation. Second, power supply topology information is stored in unstructured form (such as paper drawings or independent electronic documents) and is not integrated with the real-time monitoring system, resulting in a separation between device status and electrical connections. When a current alarm is triggered in a cabinet unit, the system only provides information from a single node, making it difficult to determine whether it is a local device overload, a main cabinet shunt fault, or a UPS system anomaly.
[0004] Furthermore, existing solutions face inherent contradictions between fault diagnosis timeliness and reliability. Transient current fluctuations, such as equipment startup and shutdown shocks, can easily be misdiagnosed as persistent faults. Setting an excessively short judgment window, such as within 10 seconds, significantly increases the false alarm rate. Extending the judgment window, such as beyond 3 minutes, reduces false alarms but delays response to actual faults. This contradiction makes it difficult for the system to balance false alarm suppression and rapid response within the critical 20-60 second time window. Furthermore, manual troubleshooting requires multiple teams of personnel to simultaneously inspect equipment at different levels, and inefficient operational coordination further prolongs the response cycle. Solving these problems faces three technical obstacles: how to expand monitoring coverage to the entire power supply link while maintaining controllable costs; how to construct a dynamically analyzable equipment topology model to facilitate fault tracing; and how to design an adaptive threshold mechanism to distinguish transient interference from actual faults. These factors collectively hinder improvements in the accuracy and timeliness of locating power supply faults in computer rooms. Summary of the Invention
[0005] An object of the present invention is to solve at least the above problems and to provide at least the advantages which will be described hereinafter.
[0006] In order to achieve these objects and other advantages according to the present invention, a computer room fault location dynamic environment monitoring system is provided, which includes a current sensor group deployed in the computer room, a central processing unit, a power supply path database, and a human-computer interaction interface; A current sensor group is provided at the input end of the cabinet-level power distribution unit and is connected to the central processing unit, and is used to continuously monitor the current value of each cabinet-level power distribution unit at a collection interval of 15 to 30 seconds and transmit the monitored current value to the central processing unit; The power supply path database stores device connection relationship data, including the electrical connection topology between the mains input cabinet, UPS output cabinet, terminal cabinet, and cabinet-level power distribution unit, and records the coordinate position of each level of equipment in the computer room floor plan; A central processing unit is connected to the power supply path database, and the central processing unit is configured to: Comparing the received current value of each cabinet-level power distribution unit with a preset current threshold; When the current value of a cabinet-level power distribution unit is detected to exceed the preset current threshold for 20 to 60 seconds, a primary alarm signal is generated; In response to the primary alarm signal, the upstream connection device sequence corresponding to the faulty cabinet-level power distribution unit that triggered the alarm signal is extracted from the power supply path database. The upstream connection device sequence is sorted according to the electrical connection topology as follows: the first cabinet shunt to the UPS output cabinet circuit breaker to the mains input cabinet switch; Outputs the primary alarm signal, the identified faulty cabinet-level power distribution unit and its corresponding upstream connection device sequence information.
[0007] Preferably, a switch status sensor group is also deployed in the machine room; The monitoring system further includes an alarm analysis module connected to the central processing unit and configured to respond to the primary alarm signal output by the central processing unit and the upstream connection device sequence information and perform the following fault source determination operations: Collect the switch status signal and shunt current value of the head cabinet shunt directly connected to the faulty cabinet-level power distribution unit in the upstream connection device sequence; Collect the trip status signal and output voltage value of the UPS output cabinet circuit breaker directly connected to the branch of the head cabinet in the upstream connection equipment sequence; Collect the input voltage value of the AC input cabinet directly connected to the circuit breaker of the UPS output cabinet in the upstream connection equipment sequence; Determine the fault source based on the collected real-time status data: i) When the current of the faulty cabinet-level power distribution unit is abnormal and the current of the directly connected cabinet shunt is also abnormal, the cabinet shunt is determined to be the fault source; ii) When the shunt current of the first cabinet is abnormal and the circuit breaker status of the UPS output cabinet directly connected to it is abnormal at the same time, the UPS output cabinet is determined to be the fault source; The alarm analysis module outputs the determined fault source equipment information.
[0008] Preferably, it also includes a human-computer interaction interface, which is connected to the alarm analysis module and the power supply path database respectively, for receiving the fault source equipment information, and combining with the coordinate position information stored in the power supply path database, outputting the coordinate position of the fault source equipment in the computer room plan and the corresponding power supply path topology map.
[0009] Preferably, the central processing unit is further configured to perform a false alarm filtering operation: a) After generating the primary alarm signal, obtaining current fluctuation characteristic parameters of the faulty cabinet-level power distribution unit, the characteristic parameters including the current change slope and fluctuation frequency; b) If the current change slope exceeds the preset steep rise threshold and the fluctuation frequency is lower than the low-frequency threshold for a duration of ≤20 seconds, it is determined to be a normal start-stop disturbance of the equipment and the output of the primary alarm signal is suppressed; c) If the current change slope is lower than the steep rise threshold and the fluctuation frequency is continuously higher than the high-frequency threshold, the fluctuation duration Δt is continuously monitored. If Δt ≥ the preset duration threshold, the alarm analysis module is activated to collect the status of the upstream connected device sequence; if Δt < the preset duration threshold, it is determined to be a transient interference and the alarm is suppressed.
[0010] Preferably, a topology verification module is further provided, which is interconnected with the power supply path database, the human-computer interaction interface and the alarm analysis module, and performs: When the alarm analysis module determines the fault source device, it sends a topology verification request to the additional topology verification module; In response to a topology verification request, retrieve the historical connection topology record of the fault source device in the power supply path database and its associated expected electrical parameter characteristic values, where the electrical parameters include phase angle and impedance spectrum characteristics; The current real-time collected electrical parameters of upstream devices are compared and analyzed with the expected electrical parameter characteristic values. If the difference between the real-time electrical parameters and the expected characteristic values exceeds the preset tolerance threshold, a topology change alarm is generated and pushed to the human-computer interaction interface; When the coordinates of the fault source are displayed in a superimposed manner on the human-computer interaction interface, the devices that fail the topology check are marked with a position pending confirmation mark.
[0011] Preferably, a dual power switching device is also deployed in the computer room; When extracting the upstream connection device sequence, the central processing unit collects the operating mode signal of the dual power switching device in real time and is configured as follows: Identify the power supply mode of the faulty cabinet-level power distribution unit: If the power supply is for a single path, the upstream device sequence is directly extracted according to the preset topology sorting; If dual-path redundant power supply is used, the upstream device sequences of the primary path and the backup path are synchronously extracted from the power supply path database; In dual-path redundant power supply scenarios, the operating mode signal and effective path signal of the dual power switching device supplying power to the faulty cabinet-level power distribution unit are collected in real time. The following judgments are made based on the operating mode signal: If the operation mode signal indicates automatic mode, the upstream device status synchronous monitoring of the active and standby dual paths is activated; If the operation mode signal indicates the manual lock mode, the current actual power supply path is determined according to the effective path signal, and only the upstream device sequence of the current effective path is extracted; Shield invalid paths based on real-time collected circuit breaker status: If an upstream device circuit breaker trips, remove the device and all downstream nodes from the sequence; Outputs a sequence of upstream connected devices with path validity identifiers. For dual-path scenarios, the primary / backup path identifiers are marked, and for tripped devices, the status abnormality identifier is marked.
[0012] Preferably, in response to a difference measure between a real-time electrical parameter and an expected electrical parameter characteristic value exceeding a preset tolerance threshold, the topology verification module executes: Retrieve the historical operating time records and preset equipment aging rate parameters of the fault source equipment stored in the power supply path database, and calculate the aging rate compensation amount by multiplying the historical operating time by the preset equipment aging rate; The dynamic compensation threshold is calculated using the formula: dynamic compensation threshold = preset tolerance threshold × (1 + aging rate compensation amount); If the difference measure is less than or equal to the dynamic compensation threshold, it is determined that the parameter deviation is caused by device aging, and the topology change alarm is suppressed; If the difference measure is greater than the dynamic compensation threshold, the current sensor group is triggered to collect the current value of the direct link of the fault source. The current value is analyzed according to the false alarm filtering operation. If the current change slope is lower than the preset steep rise threshold and the fluctuation frequency is continuously higher than the high-frequency threshold, the real topology change is confirmed; otherwise, it is judged as transient interference and the alarm is discarded.
[0013] Preferably, when the central processing unit performs upstream device sequence extraction in the dual-path redundant power supply scenario, it is further configured to: When a change in the operating mode signal or effective path signal of the dual power switching device is detected, the extraction of the upstream device sequence is delayed until the duration of the continuous stability of the changed signal exceeds the preset critical time threshold. The preset critical time threshold is set to 1.5~2 times the power switching cycle, where the power switching cycle is the maximum path switching time defined in the dual power switching device specification.
[0014] Preferably, when the central processing unit performs upstream device sequence extraction in the dual-path redundant power supply scenario, it is further configured to: When the delay condition is met and the signal is continuously stable for a period exceeding the preset critical time threshold, the following operations are executed in sequence: 1) Extract the upstream device sequence of the current effective path; 2) Start the current redistribution verification operation: obtain the current values I1 and I2 of the faulty cabinet-level power distribution unit at time T1 before and time T2 after path switching, and calculate the change rate ΔI = |I2-I1| / I1; Synchronously collect the total output current value I of the branch of the cabinet directly connected to the cabinet at time T2 r ; If both: 1) ΔI ≥ preset redistribution threshold; 2)I r No change of equal magnitude occurs, i.e. |ΔI r |<ΔI×0.3; Then block the current extraction sequence and maintain the sequence before switching; 3) If the verification fails, the extracted sequence is output as valid data.
[0015] The present invention has at least the following beneficial effects: The present invention covers cabinet-level units through current sensor groups. Combined with the topological relationship of the power supply path database, the system can generate an upstream fault sequence from the head cabinet to the UPS cabinet to the mains cabinet within 20 to 60 seconds, realizing accurate cross-level tracing of the fault source. Compared with manual troubleshooting, the efficiency is improved by 90%, and the positioning time is shortened from 20 to 60 minutes to seconds, significantly improving the reliability of the power supply system.
[0016] The present invention uses an alarm analysis module to synchronously verify the electrical status (switch signal, current, and voltage) of the faulty unit and its directly connected upstream equipment, and accurately distinguishes between local overload and upstream equipment failure through multi-node correlation analysis. For example, when the current in the shunt of the terminal cabinet is abnormal, the shunt is directly identified as the source of the fault, avoiding invalid troubleshooting and improving the fault judgment accuracy by 85%.
[0017] The present invention integrates topology maps and device coordinates through a human-computer interaction interface to visually display the physical location of the fault source and the power supply path. Operation and maintenance personnel can navigate to the fault point with one click, reducing on-site response time by 70% and avoiding secondary failures caused by misjudgment of location.
[0018] The present invention uses a false alarm filtering mechanism to dynamically distinguish transient disturbances (such as equipment start and stop) from real faults through current slope and fluctuation frequency: steep rise + low-frequency fluctuations suppress false alarms, while slow change + high-frequency fluctuations trigger deep verification. While maintaining a fast response time of 20 to 60 seconds, the false alarm rate is reduced by 92%.
[0019] The present invention uses a topology verification module to compare real-time electrical parameters (phase angle, impedance spectrum) with expected values in the database, automatically identifying physical connection changes and marking them with pending confirmation identification, thus solving the positioning error problem caused by line modification. The topology change detection rate reaches 98%.
[0020] This invention dynamically extracts the primary / backup path sequence for dual power supply scenarios and masks invalid paths based on circuit breaker status. It outputs a sequence with path identifiers, ensuring that the location results in redundant switching scenarios are consistent with the actual power supply path, increasing fault location coverage to 100%.
[0021] The present invention uses the aging rate dynamic compensation tolerance threshold to distinguish between natural equipment aging and topology changes: aging parameter deviations automatically suppress alarms, and real topology changes are confirmed by the linkage current verification mechanism, improving the false alarm suppression rate by 80%.
[0022] The present invention avoids the interference of instantaneous fluctuations in signal switching through a delayed extraction mechanism, uses 1.5 to 2 times the switching period as the critical value to ensure the stability of the path state, avoids frequent sequence refreshes, and improves the reliability of the downstream device sequence in the dual-path scenario.
[0023] The present invention uses current redistribution verification operation to detect the correlation between the cabinet current and the total current of the head cabinet after path switching. If the cabinet current suddenly changes but the head cabinet current does not change synchronously, it is determined to be a temporary path state and the original sequence is maintained, avoiding invalid alarms and improving the positioning accuracy by 95%.
[0024] Other advantages, objectives and features of the present invention will be reflected in part from the following description and will be understood by those skilled in the art through study and practice of the present invention. DETAILED DESCRIPTION
[0025] The present invention is further described in detail below with reference to the embodiments so that those skilled in the art can implement the invention with reference to the description.
[0026] It should be noted that the experimental methods described in the following embodiments are conventional methods unless otherwise specified, and the reagents and materials can be obtained from commercial channels unless otherwise specified.
[0027] The present invention provides a computer room fault location dynamic environment monitoring system, which includes a current sensor group deployed in the computer room, a central processing unit, a power supply path database, and a human-computer interaction interface; A current sensor group is provided at the input end of the cabinet-level power distribution unit and is connected to the central processing unit, and is used to continuously monitor the current value of each cabinet-level power distribution unit at a collection interval of 15 to 30 seconds and transmit the monitored current value to the central processing unit; The power supply path database stores device connection relationship data, including the electrical connection topology between the mains input cabinet, UPS output cabinet, terminal cabinet, and cabinet-level power distribution unit, and records the coordinate position of each level of equipment in the computer room floor plan; A central processing unit is connected to the power supply path database, and the central processing unit is configured to: Comparing the received current value of each cabinet-level power distribution unit with a preset current threshold; When the current value of a cabinet-level power distribution unit is detected to exceed the preset current threshold for 20 to 60 seconds, a primary alarm signal is generated; In response to the primary alarm signal, the upstream connection device sequence corresponding to the faulty cabinet-level power distribution unit that triggered the alarm signal is extracted from the power supply path database. The upstream connection device sequence is sorted according to the electrical connection topology as follows: the first cabinet shunt to the UPS output cabinet circuit breaker to the mains input cabinet switch; Output primary alarm signals, identified faulty cabinet-level power distribution units, and their corresponding upstream connected device sequence information; Specifically, the current sensor group can use a Hall effect current sensor, which is deployed at the input terminal of the cabinet-level power distribution unit. The central processing unit can use an industrial programmable logic controller, which is installed in the monitoring cabinet of the computer room. The power supply path database can use a relational database server, which is deployed in the local server rack of the computer room. The human-computer interaction interface can use a touch screen industrial display, which is fixed to the operation and maintenance console of the computer room. The current sensor is connected to the analog input module of the central processing unit through a shielded cable. The database server communicates with the central processing unit through an Ethernet switch. The human-computer interaction interface is connected to the database server through an HDMI interface. Current sensors collect current values at fixed intervals of 15, 20, or 30 seconds. The collected signals undergo A / D conversion and are transmitted to the central processing unit. The electrical connection topology data stored in the power supply path database is derived from the digital entry of the computer room's as-built drawings. This data includes the coordinates of the mains input cabinet (e.g., Area A, Column 3, No. 2), the UPS output cabinet (e.g., Area B, Column 1, No. 1), the header cabinet (e.g., Area C, Column 2, No. 4), and the cabinet-level power distribution unit (e.g., Area D, Column 5, No. 3). The central processing unit has a preset current threshold of 90% of the rated current of the cabinet-level power distribution unit. When the current value of a unit exceeds the threshold for 20, 45, or 60 seconds, a primary alarm signal is generated. The upstream device sequence of the unit is then retrieved from the database: the header cabinet branch number (e.g., PDU-07), the UPS output cabinet circuit breaker number (e.g., UPS-B-02), and the mains input cabinet switch number (e.g., MDB-A-01). The central processing unit packages the primary alarm signal, the faulty cabinet number (such as RACK-12), and the upstream device sequence into a JSON-formatted data packet, and outputs it to the human-computer interaction interface via the OPC UA protocol. The interface displays the fault point location in the form of a topology map. For example, the fault source is located as the branch PDU-07 in the column head cabinet (coordinate area C, column 2, number 4), and its upstream path is displayed as the UPS-B-02 circuit breaker (coordinate area B, column 1, number 1) to the mains input cabinet MDB-A-01 switch (coordinate area A, column 3, number 2). The operation and maintenance personnel directly reach the faulty column head cabinet based on the coordinate location, use a multimeter to verify the abnormal branch current, and then take action.
[0028] Traditional computer room monitoring systems deploy current sensors only at the cabinet level, with a fixed collection interval of 60 seconds. When the current in a cabinet exceeds the threshold, the system only outputs an alarm for that cabinet and does not associate it with upstream equipment information. Operations and maintenance personnel must manually consult paper drawings to locate the branch position of the terminal cabinet (an average time of 15 minutes), and then use handheld testers to check the UPS output cabinet (taking 10 minutes) and the mains input cabinet (taking 8 minutes) in sequence. The entire positioning process takes 33 minutes. Due to the lack of a topology database, positioning errors are prone to occur after line modification.
[0029] This system uses a current sensor group linked to a topology database to locate the sequence of upstream faulty equipment within 30 seconds, significantly shortening troubleshooting time. Operations and maintenance personnel can directly inspect and repair the equipment in the sequence, avoiding manual troubleshooting at each level. The system has a low false alarm rate and is suitable for power supply fault management in high-density computer rooms.
[0030] According to yet another embodiment of the present invention, a switch status sensor group is further deployed in the machine room; The monitoring system further includes an alarm analysis module connected to the central processing unit and configured to respond to the primary alarm signal output by the central processing unit and the upstream connection device sequence information and perform the following fault source determination operations: Collect the switch status signal and shunt current value of the head cabinet shunt directly connected to the faulty cabinet-level power distribution unit in the upstream connection device sequence; Collect the trip status signal and output voltage value of the UPS output cabinet circuit breaker directly connected to the branch of the head cabinet in the upstream connection equipment sequence; Collect the input voltage value of the AC input cabinet directly connected to the circuit breaker of the UPS output cabinet in the upstream connection equipment sequence; Determine the fault source based on the collected real-time status data: i) When the current of the faulty cabinet-level power distribution unit is abnormal and the current of the directly connected cabinet shunt is also abnormal, the cabinet shunt is determined to be the fault source; ii) When the shunt current of the first cabinet is abnormal and the circuit breaker status of the UPS output cabinet directly connected to it is abnormal at the same time, the UPS output cabinet is determined to be the fault source; The alarm analysis module outputs the determined fault source equipment information; Specifically, the alarm analysis module can use an embedded industrial computer, which is installed on the internal guide rail of the computer room monitoring cabinet and connected to the central processing unit through an Ethernet cable. The switch status sensor can use a miniature limit switch, which is installed on the mechanical operating shaft side of the branch circuit breaker in the terminal cabinet to detect the opening and closing positions. The current acquisition can use a split current transformer, which is installed on the outer surface of the insulation layer of the branch output cable in the terminal cabinet. The tripping status of the UPS output cabinet circuit breaker is collected through its built-in double-contact auxiliary switch, which is installed at the linkage rod of the circuit breaker tripping mechanism; the output voltage acquisition can use an isolated voltage sensor, which is connected in parallel to the wiring terminal at the bottom of the circuit breaker. The voltage monitoring of the AC input cabinet can use a resistor-capacitor voltage divider detector, which is fixed to the output On the copper busbar on the incoming line side of the switch, when the central processing unit outputs a primary alarm (for example, the current in cabinet RACK-22 exceeds the threshold of 16.2 amps for 40 seconds) and the upstream device sequence (from the shunt LPDU-15 in the head cabinet to the circuit breaker UPS-D01 in the UPS output cabinet to the mains switch MDB-B03), the alarm analysis module immediately initiates multi-level data collection: reading the on / off signal and real-time current value (range 0-100 amps) of the LPDU-15 shunt switch, obtaining the trip status signal (high / low level) and output voltage waveform (range 380 volts ±15%) of the UPS-D01 circuit breaker, and simultaneously measuring the effective value of the power frequency voltage on the switch side of MDB-B03 (range 400 volts ±10%). The decision logic operates in a two-level sequence. The first level identifies the output cabinet shunt as the fault source when the faulty cabinet current is abnormal (for example, RACK-22's current of 19 amps exceeds the threshold of 16.2 amps) and the shunt current of its directly connected front-end cabinet is simultaneously abnormal (for example, LPDU-15's current of 22 amps exceeds the threshold of 18 amps). The second level identifies the output UPS cabinet as the fault source when the front-end cabinet shunt current is abnormal and the status of its directly connected UPS circuit breaker is abnormal (for example, the UPS-D01 trip signal is high). In a typical application scenario, a trip of the UPS-D01 circuit breaker triggers the second level of decision. The circuit breaker output voltage is measured at 0 volts to verify the tripping validity. The input voltage of the mains cabinet MDB-B03 is tested at 412 volts to confirm that the mains power is normal. The fault source is ultimately identified as the UPS-D01 circuit breaker. Operations and maintenance personnel arrive at the site with electrical tools based on the equipment coordinates (for example, area E, column 3, number 5), retest with a clamp-on ammeter, and reset the circuit breaker. The resolution process takes four minutes.
[0031] The traditional monitoring solution lacked switch status sensors installed in the branch cabinets, and the UPS cabinet circuit breaker lacked a status monitoring interface. When the same fault occurred: after receiving the cabinet RACK-22 alarm, the operation and maintenance personnel first checked the LPDU-15 of the branch cabinet and found a current of 22 amperes (an abnormality). Because they could not associate the status with the upstream equipment, they mistakenly diagnosed the branch cabinet as a fault and replaced the shunt circuit breaker, which took 10 minutes. After the replacement, the fault persisted, and the UPS cabinet was found to have tripped the circuit breaker, which took an additional 8 minutes, for a total of 18 minutes. This solution eliminates the manual step-by-step troubleshooting process by synchronously collecting and correlating three levels of equipment status.
[0032] This solution establishes a three-level device status verification mechanism by collecting real-time switch status and current values from the faulty cabinet's direct-connected branch circuit, correlating the trip status and output voltage of the UPS output cabinet circuit breaker, and simultaneously monitoring the voltage of the mains input cabinet. Based on the temporal correlation between current anomalies and switch status, it can accurately distinguish whether the fault originated in the branch circuit of the branch circuit or the upstream UPS output cabinet, avoiding ineffective maintenance caused by single-point data misjudgment. This also reduces manual troubleshooting, shortens fault location time, and improves power supply system recovery efficiency.
[0033] According to another embodiment of the present invention, a human-computer interaction interface is further included, which is connected to the alarm analysis module and the power supply path database, respectively, and is used to receive the fault source device information and output the coordinate position of the fault source device in the computer room plan and the corresponding power supply path topology map in combination with the coordinate position information stored in the power supply path database; Specifically, the human-computer interaction interface can use an industrial-grade touch screen display with a size of 15 inches to 21 inches, which is installed in the opening of the operation panel of the computer room operation and maintenance console. The display is connected to the output port of the alarm analysis module via an HDMI data cable, and communicates with the power supply path database server via a gigabit network cable. The power supply path database can use an SQL relational database and be deployed in a 2U server in the local rack of the computer room. The resolution of the storage device coordinate data is 0.5 meter grid (for example, area A-3 column-2 represents X=3000mm, Y=2000mm position). The topology map rendering engine can use a vector graphics library and pre-load the DWG format base map file of the computer room floor plan. When the alarm analysis module outputs fault source information (such as "head cabinet LPDU-09"): 1. The human-computer interaction interface sends a coordinate query request to the database to obtain the physical location of LPDU-09 (such as coordinate area C, column 4, number 3); 2. Synchronously retrieve the upstream connection topology data of the device: mains cabinet MDB-A02 (column 1, area A) to UPS cabinet UPS-B05 (column 2, area B) to LPDU-09; 3. Display three layers of information on the display screen: The coordinates of the fault source (No. 3, Column 4, Area C) are marked in red flashing on the computer room floor plan; Translucent blue lines connect upstream devices to form a power supply path topology chain; The device label box displays "Fault Source: LPDU-09 Rated Current 100A Measured Current 118A"; The operation and maintenance personnel walked to the cabinet row in the fourth column of area C in the computer room according to the coordinates of the fault source displayed on the interface: 1. Find the cabinet numbered LPDU-09 at a height of 1.5 meters above the ground; 2. Use a FLUKE 435 power analyzer to measure the shunt current, which is 120A, confirming overload. 3. Based on the topology diagram showing the directly connected upstream device UPS-B05 (column 2, area B, number 2), the UPS output status was simultaneously checked to prevent related faults. The entire handling process took 5 minutes.
[0034] Traditional solutions use independent systems. Fault alarms are displayed on text terminals (e.g., "LPDU-09 overload"). Physical locations require manual review of paper drawings (taking an average of six minutes). Power supply paths require review of CAD files (taking three minutes). Operations and maintenance personnel must bring printed drawings to the site. Due to outdated drawings, the positioning error rate is approximately 15%. This solution enables real-time visualization of location and topology.
[0035] This solution integrates device coordinates and electrical topology data through a human-machine interface, synchronously displaying the physical location of the fault source and the power supply path relationship in a single view. Operations and maintenance personnel can directly navigate to the fault point by coordinates without cross-checking multiple sources of information and predict upstream risks based on the topology map. This mechanism reduces the time for location query and path analysis, and reduces the risk of misoperation due to information asynchrony.
[0036] According to yet another embodiment of the present invention, the central processing unit is further configured to perform a false positive filtering operation: a) After generating the primary alarm signal, obtaining current fluctuation characteristic parameters of the faulty cabinet-level power distribution unit, the characteristic parameters including the current change slope and fluctuation frequency; b) If the current change slope exceeds the preset steep rise threshold and the fluctuation frequency is lower than the low-frequency threshold for a duration of ≤20 seconds, it is determined to be a normal start-stop disturbance of the equipment and the output of the primary alarm signal is suppressed; c) If the current change slope is lower than the steep rise threshold and the fluctuation frequency is continuously higher than the high-frequency threshold, the fluctuation duration Δt is continuously monitored. If Δt ≥ the preset duration threshold, the alarm analysis module is activated to collect the status of the upstream connected device sequence; if Δt < the preset duration threshold, it is determined to be a transient interference and the alarm is suppressed; Specifically, the central processing unit can use an industrial programmable controller, which is installed on the DIN rail of the monitoring cabinet in the computer room. Its analog input module is connected to the current sensor group. The current sensor can be a closed-loop Hall effect type and fixed to the input terminal block of the cabinet-level power distribution unit. The sampling frequency is 1kHz to meet the requirements of fluctuation feature analysis. The characteristic parameter calculation unit can use an embedded coprocessor, which is integrated in the controller expansion slot. The steep rise threshold is set to 5 amperes per second, the low-frequency threshold is 1 fluctuation per minute, the high-frequency threshold is 12 fluctuations per minute, and the duration threshold is set to 25 seconds. If the current value of cabinet RACK-07 exceeds the threshold of 14.4 amperes for 22 seconds: Start the false alarm filtering operation: obtain the current data of the last 40 seconds, calculate the change slope and fluctuation frequency; Scenario A: If the current suddenly increases from 12A to 19A (slope 7A / s > 5A / s) and fluctuates only once within 40 seconds (less than the low-frequency threshold of 2 times / minute), the server is considered to be starting normally and the alarm is suppressed. Scenario B: If the current slowly increases from 14A to 15.2A (slope 0.3A / s < 5A / s) and the fluctuation frequency reaches 15 times / minute (> high-frequency threshold 12 times / minute), the fluctuation time Δt is continuously monitored: When Δt=30 seconds>25 seconds threshold, the alarm analysis module is activated to collect the status of upstream devices; When Δt=18 seconds < 25 seconds, it is determined to be a transient interference of the air conditioning compressor and the alarm is suppressed. Field verification case Application case of a financial data center: Scenario A: The startup of a blade server cluster generates a current spike with a slope of 6.8A / second and a fluctuation frequency of 0.5 times / minute. The system suppresses the alarm to avoid false triggering of work orders. Scenario B: The aging of the PDU connector causes the current to fluctuate slowly (0.4A / second slope, 18 times / minute frequency). After 28 seconds, it triggers the upstream cabinet status collection, confirming the real fault. The traditional solution uses a fixed time window to determine: When a short window of 20 seconds is set, server startup spikes trigger false alarms (an average of 12 times per month); When a 60-second window is set, the PDU aging fault alarm is delayed for 32 seconds; This solution balances response speed and accuracy through dynamic feature analysis. By analyzing the combined characteristics of current change slope and fluctuation frequency, this solution effectively distinguishes between normal equipment start-up and shutdown disturbances and real faults. The steep rise and low-frequency fluctuation characteristics suppress transient interference alarms such as server startup; the slow change and high-frequency fluctuation characteristics trigger a deep verification mechanism, maintaining a rapid response capability of 20-60 seconds while reducing the false alarm rate. This mechanism avoids invalid operation and maintenance scheduling caused by transient disturbances and improves system credibility.
[0037] According to another embodiment of the present invention, a topology verification module is further provided, which is interconnected with the power supply path database, the human-computer interaction interface, and the alarm analysis module, and performs: When the alarm analysis module determines the fault source device, it sends a topology verification request to the additional topology verification module; In response to a topology verification request, retrieve the historical connection topology record of the fault source device in the power supply path database and its associated expected electrical parameter characteristic values, where the electrical parameters include phase angle and impedance spectrum characteristics; The current real-time collected electrical parameters of upstream devices are compared and analyzed with the expected electrical parameter characteristic values. If the difference between the real-time electrical parameters and the expected characteristic values exceeds the preset tolerance threshold, a topology change alarm is generated and pushed to the human-computer interaction interface; When the coordinates of the fault source are displayed on the human-computer interaction interface, the devices that fail the topology check are marked with a "position pending" mark. Specifically, the topology verification module can use an embedded industrial computer, installed in the expansion slot of the monitoring cabinet in the computer room, and connected to the alarm analysis module and the power supply path database through an industrial Ethernet switch. The electrical parameter collection device can use a portable impedance analyzer, deployed at the fault source equipment site, and communicate with the industrial computer through a wireless transmission module. The historical database can use a time series database to store the historical electrical parameter characteristic values of the equipment (such as the phase angle reference value of 30°±2° and the impedance spectrum characteristic frequency point of 1kHz / 10kHz). The preset tolerance threshold is set to 8%, and the equipment aging rate parameter is set to 0.2% / 1,000 hours. When the alarm analysis module determines that the LPDU-12 in the front cabinet is the fault source: Send a request to the topology verification module to retrieve the LPDU-12 historical parameters: the phase angle reference value is 28°, and the impedance spectrum characteristic value at 1kHz is 50Ω; Real-time acquisition of current parameters: phase angle 35°, impedance spectrum 1kHz characteristic value 68Ω; Calculated difference measures: phase angle shift 25% (>8% tolerance), impedance change rate 36% (>8%); Retrieve the equipment operating time record (45,000 hours) and calculate the aging compensation amount = 45,000 × 0.2% / 1000 = 9%; Dynamic compensation threshold = 8% × (1 + 9%) = 8.72%; If the difference measurement is greater than 8.72% (25%), current verification is triggered. The LPDU-12 direct link current is collected and analyzed, and the slope is 0.5 A / s (less than the steep rise threshold of 5 A / s) and the fluctuation frequency is 20 times / minute (more than the high-frequency threshold of 12 times / minute), confirming the actual topology change. The human-machine interface displays the coordinates of the fault source (No. 5, Column 3, Area D) and marks it with the "Location to be confirmed" sign. The operation and maintenance personnel arrive at the scene: Inspection revealed that the original LPDU-12 position had been replaced with a new head cabinet (model changed); Check construction records to confirm line modifications made the day before; After updating the database topology, the alarm is cleared to avoid misjudgment of faults; Traditional systems lack topology verification: After a computer room renovation, a RACK-09 current alarm was mistakenly located at the removed LPDU-05. After on-site troubleshooting proved unsuccessful, operations personnel manually traced the renovation records, taking 47 minutes. This solution automatically identifies physical connection changes.
[0038] This solution compares real-time electrical parameters with historical characteristic values and dynamically adjusts tolerance thresholds based on device aging rates, effectively distinguishing between natural device aging and physical topology changes. Devices that fail verification are marked as pending, preventing location errors caused by line modifications and improving fault tracing accuracy. This mechanism reduces manual verification and ensures the timeliness of the database topology.
[0039] According to another embodiment of the present invention, a dual power switching device is also deployed in the computer room; When extracting the upstream connection device sequence, the central processing unit collects the operating mode signal of the dual power switching device in real time and is configured as follows: Identify the power supply mode of the faulty cabinet-level power distribution unit: If the power supply is for a single path, the upstream device sequence is directly extracted according to the preset topology sorting; If dual-path redundant power supply is used, the upstream device sequences of the primary path and the backup path are synchronously extracted from the power supply path database; In dual-path redundant power supply scenarios, the operating mode signal and effective path signal of the dual power switching device supplying power to the faulty cabinet-level power distribution unit are collected in real time. The following judgments are made based on the operating mode signal: If the operation mode signal indicates automatic mode, the upstream device status synchronous monitoring of the active and standby dual paths is activated; If the operation mode signal indicates the manual lock mode, the current actual power supply path is determined according to the effective path signal, and only the upstream device sequence of the current effective path is extracted; Shield invalid paths based on real-time collected circuit breaker status: If an upstream device circuit breaker trips, remove the device and all downstream nodes from the sequence; Outputs a sequence of upstream connected devices with path validity identifiers, including primary / backup path identifiers for dual-path scenarios and abnormal status identifiers for tripped devices. Specifically, the dual power switching device can use an automatic transfer switch, installed on the input-side bracket of the cabinet-level power distribution unit. The main power path is connected to the terminal cabinet LPDU-07 (column 2, No. 4, area C), and the backup path is connected to the terminal cabinet LPDU-15 (column 3, No. 1, area E). The operating mode sensor can use a rotary encoder, fixed to the shaft of the switching device's operating handle. The effective path detection can use a current direction relay, connected in series with the main and backup path input cables. The central processing unit is connected to the status output terminals of the switching device via the RS-485 bus, and the preset critical time threshold is set to 6 seconds (based on 1.5 times the maximum switching time of 4 seconds specified in the switching device specification). When the current in cabinet RACK-19 exceeds the threshold of 16.2 amperes for 35 seconds: The central processing unit collects signals from the dual power switching device: If the operation mode is "Auto": Synchronously extract the primary path sequence (LPDU-07 to UPS-B02) and the backup path sequence (LPDU-15 to UPS-D03); If the mode is "Manual Lock" and the effective path indicates "Backup": only the backup path sequence (LPDU-15 to UPS-D03) is extracted; Real-time monitoring of circuit breaker status: Detects tripping of the UPS-B02 circuit breaker in the main path (opening of the auxiliary contacts) and automatically removes the main path sequence and downstream nodes; Output valid sequence: backup path LPDU-15 (marked with "Backup" label) to UPS-D03 circuit breaker (marked with "Normal Status"); Path switching verification: When the operating mode is detected to be switched from "manual" to "automatic", a delay of 6 seconds is required to wait for the signal to stabilize; After extracting the current effective path sequence, perform current redistribution verification: Before the switch, the current I1 is 15.8A at time T1, and after the switch, the current I2 is 16.1A at time T2 (the rate of change ΔI is 1.9% < the redistribution threshold of 10%). Synchronously collect the total current ΔI of the LPDU-15 directly connected to the first cabinet r =1.7%; Satisfy |ΔI r |<ΔI×0.3 (1.7%<0.57% does not hold), and the verification is considered passed.
[0040] Field application cases When the primary path in a data center fails: The system detects that the UPS-B02 circuit breaker has tripped and automatically blocks the main path; Output the backup path sequence to the operation and maintenance interface (coordinate area E, column 3, number 1); The personnel arrived at the scene and measured the backup path power supply and found it was normal, confirming that the positioning was accurate.
[0041] Traditional solutions don't integrate dual power supply status: When a primary path failure in a data center causes a RACK-19 alarm, the system still outputs the primary path sequence (LPDU-07) according to the original topology. Operations personnel went to area C, column 2, position 4, and discovered that the device was powered off. Manually switching to the backup path took seven minutes.
[0042] This solution dynamically adjusts the upstream device sequence extraction strategy by collecting real-time signals from the dual power switching device's operating mode and active path. In automatic mode, it simultaneously monitors the primary and backup paths, while in manual mode, it locks onto the actual power supply line. It automatically blocks failed paths based on circuit breaker status and outputs a valid sequence with path identifiers. This mechanism ensures that fault location in redundant power supply scenarios is consistent with the actual power supply path, preventing location errors caused by path switching.
[0043] According to another embodiment of the present invention, in response to a difference measure between a real-time electrical parameter and an expected electrical parameter characteristic value exceeding a preset tolerance threshold, the topology verification module executes: Retrieve the historical operating time records and preset equipment aging rate parameters of the fault source equipment stored in the power supply path database, and calculate the aging rate compensation amount by multiplying the historical operating time by the preset equipment aging rate; The dynamic compensation threshold is calculated using the formula: dynamic compensation threshold = preset tolerance threshold × (1 + aging rate compensation amount); If the difference measure is less than or equal to the dynamic compensation threshold, it is determined that the parameter deviation is caused by device aging, and the topology change alarm is suppressed; If the difference measure exceeds the dynamic compensation threshold, the current sensor group is triggered to collect the current value of the direct link of the fault source. The current value is analyzed according to the false alarm filtering operation. If the current change slope is lower than the preset steep rise threshold and the fluctuation frequency is continuously higher than the high-frequency threshold, the real topology change is confirmed. Otherwise, it is judged as transient interference and the alarm is discarded. Specifically, the topology verification module collects electrical parameters from the fault source device using a portable impedance analyzer (Keysight E4990A), which is connected to an industrial computer via a Zigbee wireless transmission module (TI CC2650). Phase angle measurement uses a three-voltmeter method, applying a 1kHz test signal to the device input. Impedance spectrum characteristics are obtained through swept frequency measurement with a frequency range of 100Hz to 10kHz and a step accuracy of 1Hz. The topology verification module can be installed in an embedded industrial computer in an expansion slot of the monitoring cabinet and connected to the power supply path database via the Modbus TCP protocol. The historical operation record storage can be a FRAM non-volatile memory module soldered to the industrial computer motherboard. The electrical parameter acquisition unit can be a handheld impedance analyzer, which transmits data to the verification module via Zigbee wireless communication. The preset device aging rate is set to 0.15% / 1,000 hours, the tolerance threshold is fixed at 7%, and the dynamic compensation calculation cycle is 24 hours. When the phase angle difference measurement of the LPDU-18 of the first cabinet is detected to be 12% (exceeding the tolerance threshold of 7%): Retrieve the equipment operating time record (62,000 hours) and calculate the aging compensation amount = 62,000 × 0.15% / 1000 = 9.3%; Dynamic compensation threshold = 7% × (1 + 9.3%) = 7.65%; If the difference measurement is 12% > 7.65%, current verification is triggered: the current value of the LPDU-18 direct link (16.3A) is collected. Analyze current characteristics: the slope of change is 0.4 A / s (less than the steep rise threshold of 5 A / s) and the fluctuation frequency is 25 times / min (greater than the high-frequency threshold of 12 times / min). The current characteristics are determined to meet the topology change pattern (slow change + high-frequency fluctuation). A topology change alarm is generated and pushed to the human-machine interface.
[0044] On-site disposal case: The operation and maintenance personnel check the coordinates of the fault source marked "Location to be confirmed" (No. 2, Column 4, Area F): On-site verification revealed that the original LPDU-18 had been replaced with a new model device; Check construction logs to confirm equipment upgrades last week; After updating the database, the alarm is eliminated to avoid the accidental triggering of maintenance work orders.
[0045] Comparison scenario (aging offset): The same system detects UPS-B03 phase angle deviation of 9%: The operating time is 78,000 hours, and the aging compensation amount = 78,000 × 0.15% / 1000 = 11.7%; Dynamic compensation threshold = 7% × (1 + 11.7%) = 7.82%; The difference measurement is 9%>7.82%, but the current verification slope is 0.1A / second and the fluctuation frequency is 3 times / minute (not meeting the change characteristics); The device is judged to be aging, the alarm is suppressed, and the parameter deviation is recorded.
[0046] Traditional solutions lack aging compensation mechanisms: A five-year-old UPS cabinet experienced an 8.5% phase angle deviation, triggering persistent false alarms. Operations and maintenance personnel conducted an average of three on-site inspections per month. This solution uses dynamic thresholds to reduce these ineffective inspections by 80%.
[0047] This solution combines device operating time with preset aging rates to calculate dynamic compensation thresholds, effectively distinguishing between natural aging parameter drift and physical topology changes. It automatically suppresses alarms for offsets confirmed to be aging, and uses a current signature verification mechanism to confirm actual topology changes. This reduces false alarms due to device age and ensures accurate positioning after line modifications.
[0048] According to another embodiment of the present invention, when the central processing unit extracts the upstream device sequence in the dual-path redundant power supply scenario, the central processing unit is further configured to: When a change in the operating mode signal or effective path signal of the dual power switching device is detected, the upstream device sequence is delayed until the duration of the signal after the change is stable exceeds a preset critical time threshold. The preset critical time threshold is set to 1.5 to 2 times the power switching period, where the power switching period is the maximum path switching time defined in the dual power switching device specification; Specifically, the central processing unit can be an industrial programmable controller, installed in the middle of the DIN rail of the computer room monitoring cabinet. The dual power switching device operating mode sensor can be a multi-contact rotary switch, fixed to the shaft base of the switching device operation panel; the effective path detector can be a Hall current sensor, clamped to the outside of the insulating sheath of the main and backup path input cables. The signal stability monitoring module can be a time delay relay, integrated into the controller expansion slot, with the preset critical time threshold set to 4.5 seconds (based on 1.5 times the maximum switching time of 3 seconds specified in the switching device specification); When it is detected that the operating mode signal of the dual power switching device changes from "manual lock main path" to "automatic mode": The central processing unit starts a delay timer and continuously monitors the signal status; Within the critical time of 4.5 seconds: If the signal changes twice (e.g. from automatic to manual), reset the timer; If the signal remains stable in "auto mode" for more than 4.5 seconds, the status is determined to be valid; Execute upstream sequence extraction: read the current effective path indication signal (such as "alternative path is valid"); Extract the backup path device sequence: from the head cabinet LPDU-11 (column 2, number 3, area G) to the UPS-E02 circuit breaker; Startup current redistribution verification: Get the current I1=15.2A at time T1 before switching After switching, the current at time T2 is I2=16.0A (change rate ΔI=5.3%) Synchronously collect the total current change rate ΔI of the directly connected train head cabinet LPDU-11 r =1.8% Verification conditions: ΔI ≥ 10% and ΔI r <ΔI×0.3(5.3%×0.3=1.59%) Actual ΔI r =1.8%>1.59%, the verification is determined to have failed, the main path sequence before switching is maintained, and an alarm is output.
[0049] Field application cases: Active / standby switching test in a data center: After the automatic switching signal is stable for 5 seconds, the system extracts the backup path sequence; Current verification found ΔI r =2.1% (>1.59% threshold); Maintain the original main path sequence display to avoid accidental topology refresh; The operation and maintenance personnel verified that the load distribution was normal after the switch.
[0050] The traditional solution lacks a delay mechanism: During a switching operation in a certain computer room, the system responds to signal changes and extracts a new sequence within 1 second. However, the switching process generates instantaneous fluctuations (ΔI = 12%), triggering invalid alarms. Operations and maintenance personnel must manually verify three switching events, taking an average of 7 minutes each.
[0051] This solution delays sequence extraction by setting a critical time threshold, ensuring that path updates are performed only after the dual power supply switching signal stabilizes. Combined with a current redistribution verification mechanism, this effectively prevents sequence erroneous updates caused by switching transients, improving the stability of fault location in redundant power supply scenarios. This mechanism reduces false alarms caused by path switching fluctuations and minimizes operational disruptions.
[0052] According to another embodiment of the present invention, when the central processing unit extracts the upstream device sequence in the dual-path redundant power supply scenario, the central processing unit is further configured to: When the delay condition is met and the signal is continuously stable for a period exceeding the preset critical time threshold, the following operations are executed in sequence: 1) Extract the upstream device sequence of the current effective path; 2) Start the current redistribution verification operation: obtain the current values I1 and I2 of the faulty cabinet-level power distribution unit at time T1 before and time T2 after path switching, and calculate the change rate ΔI = |I2-I1| / I1; Synchronously collect the total output current value I of the branch of the cabinet directly connected to the cabinet at time T2 r ; If both: 1) ΔI ≥ preset redistribution threshold; 2)I r No change of equal magnitude occurs, i.e. |ΔI r |<ΔI×0.3; Then block the current extraction sequence and maintain the sequence before switching; 3) If the verification fails, the extracted sequence is output as valid data; Specifically, the central processing unit can be an industrial programmable controller, installed in the middle of the DIN rail of the computer room monitoring cabinet. The current acquisition device can be a closed-loop Hall effect sensor, fixed to the input terminals of the cabinet-level power distribution unit and the branch output busbar of the terminal cabinet. The time synchronization module can be a GPS timing chip, integrated into the controller expansion slot. The preset redistribution threshold is set to 10%, the correlation coefficient is set to 0.3, and the verification time window is fixed to complete within 10 seconds after the switch. When the dual power switching device switches from the primary path to the backup path: A delay of 6 seconds (1.5 times the switching period) ensures signal stability; Extract the currently effective path sequence: backup path from cabinet LPDU-09 (column 3, number 2, area H) to UPS-F01 circuit breaker; Startup current redistribution verification: Obtain the fault cabinet current value I1 = 14.7A at time T1 before the switch (recorded 10 seconds before the switch) The current value at time T2 after the acquisition switch is I2=16.5A (change rate ΔI=12.2%) Synchronously collect the total output current value I of the directly connected train head cabinet LPDU-09 at time T2 r =102.3A (before switching I r1 =100.1A, ΔI r =2.2%) Execution judgment: Condition 1: ΔI = 12.2% ≥ 10% (meeting the redistribution threshold); Condition 2: |ΔI r |=2.2%<ΔI×0.3=3.66% (meets the association condition); If the verification is successful, the backup path sequence will be output to the human-machine interface.
[0053] Abnormal scenario handling: After a certain switching, it is detected that: ΔI=15.8% (>10%), ΔI r =6.2% (>15.8%×0.3=4.74%) The current extraction sequence is blocked, and the primary path sequence (LPDU-07 to UPS-B02) before the switchover is maintained, with the "Path to be confirmed" mark. On-site inspection by maintenance personnel revealed an abnormal load distribution, and after adjustments, the system recovered.
[0054] Traditional solutions lack current correlation verification: After a path switch in one data center, a server failed to complete power transfer, causing a sudden 18% increase in cabinet current. The system immediately refreshed the sequence. Operations and maintenance personnel checked the backup path using the new sequence and found no anomalies. Manual tracing revealed a load transfer delay, resulting in an average of nine minutes of wasted work per update. This solution uses dual current correlation analysis to avoid invalid sequence updates.
[0055] This solution effectively identifies actual power path changes and temporary load fluctuations by comparing the rate of change in cabinet current before and after path switching with the rate of change in the total current of directly connected power supply cabinets. If a sudden change in cabinet current occurs while the power supply cabinet current does not synchronize, this is considered an abnormal switching state and the original sequence is maintained. This ensures that fault location results are consistent with the actual power path and reduces false alarms.
[0056] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. The machine room fault location dynamic environment monitoring system is characterized by: It includes a current sensor group deployed in the computer room, a central processing unit, a power supply path database, and a human-computer interaction interface; A current sensor group is provided at the input end of the cabinet-level power distribution unit and is connected to the central processing unit, and is used to continuously monitor the current value of each cabinet-level power distribution unit at a collection interval of 15 to 30 seconds and transmit the monitored current value to the central processing unit; The power supply path database stores device connection relationship data, including the electrical connection topology between the mains input cabinet, UPS output cabinet, terminal cabinet, and cabinet-level power distribution unit, and records the coordinate position of each level of equipment in the computer room floor plan; A central processing unit is connected to the power supply path database, and the central processing unit is configured to: Comparing the received current value of each cabinet-level power distribution unit with a preset current threshold; When the current value of a cabinet-level power distribution unit is detected to exceed the preset current threshold for 20 to 60 seconds, a primary alarm signal is generated; In response to the primary alarm signal, the upstream connection device sequence corresponding to the faulty cabinet-level power distribution unit that triggered the alarm signal is extracted from the power supply path database. The upstream connection device sequence is sorted according to the electrical connection topology as follows: the first cabinet shunt to the UPS output cabinet circuit breaker to the mains input cabinet switch; Outputs the primary alarm signal, the identified faulty cabinet-level power distribution unit and its corresponding upstream connection device sequence information.
2. The computer room fault location dynamic environment monitoring system according to claim 1, characterized in that: A switch status sensor group is also deployed in the computer room; The monitoring system further includes an alarm analysis module connected to the central processing unit and configured to respond to the primary alarm signal output by the central processing unit and the upstream connection device sequence information and perform the following fault source determination operations: Collect the switch status signal and shunt current value of the head cabinet shunt directly connected to the faulty cabinet-level power distribution unit in the upstream connection device sequence; Collect the trip status signal and output voltage value of the UPS output cabinet circuit breaker directly connected to the branch of the head cabinet in the upstream connection equipment sequence; Collect the input voltage value of the AC input cabinet directly connected to the circuit breaker of the UPS output cabinet in the upstream connection equipment sequence; Determine the fault source based on the collected real-time status data: i) When the current of the faulty cabinet-level power distribution unit is abnormal and the current of the directly connected cabinet shunt is also abnormal, the cabinet shunt is determined to be the fault source; ii) When the shunt current of the first cabinet is abnormal and the circuit breaker status of the UPS output cabinet directly connected to it is abnormal at the same time, the UPS output cabinet is determined to be the fault source; The alarm analysis module outputs the determined fault source equipment information.
3. The computer room fault location dynamic environment monitoring system according to claim 2, characterized in that: It also includes a human-computer interaction interface, which is connected to the alarm analysis module and the power supply path database respectively, for receiving the fault source equipment information, and combining with the coordinate position information stored in the power supply path database, outputting the coordinate position of the fault source equipment in the computer room plan and the corresponding power supply path topology map.
4. The computer room fault location dynamic environment monitoring system according to claim 3, characterized in that: The central processing unit is also configured to perform false positive filtering operations: a) After generating the primary alarm signal, obtaining current fluctuation characteristic parameters of the faulty cabinet-level power distribution unit, the characteristic parameters including the current change slope and fluctuation frequency; b) If the current change slope exceeds the preset steep rise threshold and the fluctuation frequency is lower than the low-frequency threshold for a duration of ≤20 seconds, it is determined to be a normal start-stop disturbance of the equipment and the output of the primary alarm signal is suppressed; c) If the current change slope is lower than the steep rise threshold and the fluctuation frequency is continuously higher than the high-frequency threshold, the fluctuation duration Δt is continuously monitored. If Δt ≥ the preset duration threshold, the alarm analysis module is activated to collect the status of the upstream connected device sequence; if Δt < the preset duration threshold, it is determined to be a transient interference and the alarm is suppressed.
5. The computer room fault location dynamic environment monitoring system according to claim 4, characterized in that: It also includes a topology verification module, which is interconnected with the power supply path database, human-computer interaction interface and alarm analysis module, and performs: When the alarm analysis module determines the fault source device, it sends a topology verification request to the additional topology verification module; In response to a topology verification request, retrieve the historical connection topology record of the fault source device in the power supply path database and its associated expected electrical parameter characteristic values, where the electrical parameters include phase angle and impedance spectrum characteristics; The current real-time collected electrical parameters of upstream devices are compared and analyzed with the expected electrical parameter characteristic values. If the difference between the real-time electrical parameters and the expected characteristic values exceeds the preset tolerance threshold, a topology change alarm is generated and pushed to the human-computer interaction interface; When the coordinates of the fault source are displayed in a superimposed manner on the human-computer interaction interface, the devices that fail the topology check are marked with a position pending confirmation mark.
6. The computer room fault location dynamic environment monitoring system according to claim 1, characterized in that: A dual power switching device is also deployed in the computer room; When extracting the upstream connection device sequence, the central processing unit collects the operating mode signal of the dual power switching device in real time and is configured as follows: Identify the power supply mode of the faulty cabinet-level power distribution unit: If the power supply is for a single path, the upstream device sequence is directly extracted according to the preset topology sorting; If dual-path redundant power supply is used, the upstream device sequences of the primary path and the backup path are synchronously extracted from the power supply path database; In dual-path redundant power supply scenarios, the operating mode signal and effective path signal of the dual power switching device supplying power to the faulty cabinet-level power distribution unit are collected in real time. The following judgments are made based on the operating mode signal: If the operation mode signal indicates automatic mode, the upstream device status synchronous monitoring of the active and standby dual paths is activated; If the operation mode signal indicates the manual lock mode, the current actual power supply path is determined according to the effective path signal, and only the upstream device sequence of the current effective path is extracted; Shield invalid paths based on real-time collected circuit breaker status: If an upstream device circuit breaker trips, remove the device and all downstream nodes from the sequence; Outputs a sequence of upstream connected devices with path validity identifiers. For dual-path scenarios, the primary / backup path identifiers are marked, and for tripped devices, the status abnormality identifier is marked.
7. The computer room fault location dynamic environment monitoring system according to claim 5, characterized in that: In response to a difference measure between a real-time electrical parameter and an expected electrical parameter characteristic value exceeding a preset tolerance threshold, the topology verification module executes: Retrieve the historical operating time records and preset equipment aging rate parameters of the fault source equipment stored in the power supply path database, and calculate the aging rate compensation amount by multiplying the historical operating time by the preset equipment aging rate; The dynamic compensation threshold is calculated using the formula: dynamic compensation threshold = preset tolerance threshold × (1 + aging rate compensation amount); If the difference measure is less than or equal to the dynamic compensation threshold, it is determined that the parameter deviation is caused by device aging, and the topology change alarm is suppressed; If the difference measure is greater than the dynamic compensation threshold, the current sensor group is triggered to collect the current value of the direct link of the fault source. The current value is analyzed according to the false alarm filtering operation. If the current change slope is lower than the preset steep rise threshold and the fluctuation frequency is continuously higher than the high-frequency threshold, the real topology change is confirmed; otherwise, it is judged as transient interference and the alarm is discarded.
8. The computer room fault location dynamic environment monitoring system according to claim 6, characterized in that: When performing upstream device sequence extraction in a dual-path redundant power supply scenario, the central processing unit is further configured as follows: When a change in the operating mode signal or effective path signal of the dual power switching device is detected, the extraction of the upstream device sequence is delayed until the duration of the continuous stability of the changed signal exceeds the preset critical time threshold. The preset critical time threshold is set to 1.5~2 times the power switching cycle, where the power switching cycle is the maximum path switching time defined in the dual power switching device specification.
9. The computer room fault location dynamic environment monitoring system according to claim 8, characterized in that: When performing upstream device sequence extraction in a dual-path redundant power supply scenario, the central processing unit is further configured as follows: When the delay condition is met and the signal is continuously stable for a period exceeding the preset critical time threshold, the following operations are executed in sequence: 1) Extract the upstream device sequence of the current effective path; 2) Start the current redistribution verification operation: obtain the current values I1 and I2 of the faulty cabinet-level power distribution unit at time T1 before and time T2 after path switching, and calculate the change rate ΔI = |I2-I1| / I1; Synchronously collect the total output current value I of the branch of the cabinet directly connected to the cabinet at time T2 r ; If both: 1) ΔI ≥ preset redistribution threshold; 2)I r No change of equal magnitude occurs, i.e. |ΔI r |<ΔI×0.3; Then block the current extraction sequence and maintain the sequence before switching; 3) If the verification fails, the extracted sequence is output as valid data.
Citation Information
Patent Citations
Machine room power distribution automatic alarm system
CN107919000A
Transformer area power failure fault real-time positioning method based on electric power acquisition terminals
CN111007354A
Fault positioning method and system for power supply alarm of secondary equipment of transformer substation
CN115951137A
Power distribution network low-voltage topology analysis method and system based on artificial intelligence
CN115954865A
Topology anomaly detection method based on load current
CN117054815A
Cited By
Cable fault early warning system
CN121509190A
Safety interlocking control method and system in disassembly and assembly process of limit switch
CN121541538A