Electric power inspection robot inspection method and system based on multi-sensor fusion
By employing an environment-adaptive sensor weight adjustment and multimodal feature fusion method, the problems of unstable positioning and diagnostic accuracy of power inspection robots in complex environments were solved, achieving high-precision navigation and accurate fault identification, thereby improving the inspection efficiency and reliability of power equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-03-10
AI Technical Summary
Existing power inspection robots suffer from unstable navigation accuracy in complex industrial environments, insufficient accuracy of multi-sensor data fusion algorithms, and weak fault identification capabilities of intelligent diagnostic algorithms, resulting in problems such as large positioning errors, inconsistent data, low diagnostic accuracy, and high early fault missed rate.
We employ an environment-adaptive sensor weight dynamic adjustment, multi-sensor spatiotemporal synchronization and spatial registration, and an attention-based multimodal feature fusion method. By querying a weight mapping table, we dynamically adjust the sensor fusion weights to achieve weighted fusion of multimodal positioning data, perform spatiotemporal alignment and feature stitching, and use the attention mechanism to calculate the fault type.
It improves the positioning accuracy of power inspection robots in complex environments, realizes collaborative analysis of multi-sensor data, enhances the ability to identify fault characteristics, and achieves accurate fault identification and early warning.
Smart Images

Figure CN121640232A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a power inspection robot inspection method and system based on multi-sensor fusion. Background Technology
[0002] With the need for intelligent development in the power energy industry, power equipment inspection technology is gradually evolving from traditional manual inspection to intelligent and automated processes. Currently, some power inspection robots are being used in substations and power plants both domestically and internationally. These robots are typically equipped with sensors such as visible light cameras and infrared thermal imagers, automatically moving along preset paths and collecting equipment operation data. Existing power inspection robots mainly use lidar or visual navigation technology for positioning, collecting equipment status information using single sensors or simple combinations of multiple sensors, and determining whether there are any abnormalities through threshold judgments or simple pattern recognition algorithms. For example, inspection robots from GE in the United States and Kawasaki Heavy Industries in Japan began to be used for substation equipment inspection at the end of the 20th century. In China, institutions such as the Hefei Institutes of Physical Science of the Chinese Academy of Sciences and Hefei University of Technology have also made significant progress in the field of intelligent inspection robots, developing inspection robot systems with autonomous navigation and multi-sensor data acquisition capabilities.
[0003] However, existing power inspection robot technology still has many shortcomings. First, its navigation accuracy and stability are insufficient in complex industrial environments. Power plant turbine areas experience extreme environments such as high temperatures, high pressures, and strong magnetic fields. Traditional navigation systems are easily interfered with, leading to significant positioning errors. In particular, distortion of magnetometer data in strong magnetic fields can reduce positioning accuracy by 30% to 50%, affecting the accuracy and reliability of inspections. Second, the accuracy and real-time performance of multi-sensor data fusion algorithms need improvement. Existing technologies face difficulties in time synchronization, spatial registration, and data consistency of heterogeneous data collected by different sensors. They often employ simple data overlay or independent analysis methods, making it difficult to achieve deep fusion and collaborative analysis, resulting in insufficient comprehensiveness and accuracy in equipment status monitoring. Third, intelligent diagnostic algorithms have weak ability to identify early signs of equipment failure. Existing algorithms mainly rely on threshold judgment and simple pattern recognition, lacking the ability to perform deep learning and trend analysis on massive historical data. This makes it difficult to achieve early warning and accurate prediction of failures, easily leading to missed or false alarms.
[0004] Further analysis revealed a progressive causal relationship among the aforementioned problems. Because the navigation system cannot dynamically adjust sensor fusion weights based on environmental interference, the robot's positioning accuracy decreases in strong magnetic fields or high-temperature environments, preventing it from accurately reaching the preset inspection position and consequently affecting the spatial accuracy of multi-sensor data acquisition. When sensor data acquisition positions deviate significantly, even with time synchronization, data from different sensors actually correspond to different parts of the equipment. This spatial inconsistency renders data fusion meaningless, and the locations of sound sources, heat sources, and vibration sources cannot be accurately associated with the same equipment component. When spatiotemporally inconsistent data is input into the diagnostic algorithm, the lack of inherent correlation between acoustic features, thermal imaging features, and vibration features means that simple feature stitching cannot reveal the coupled manifestations of faults across multiple physical domains. This leads to the diagnostic algorithm's inability to accurately identify fault types, especially for complex faults exhibiting abnormal characteristics simultaneously in the acoustic, thermal, and vibration physical domains. Single-modal features or fixed-weight feature fusion struggles to distinguish different fault types, ultimately resulting in low fault diagnosis accuracy and a high rate of missed early fault detections. Therefore, the core issues of power inspection robot technology are how to achieve high-precision autonomous navigation in complex power environments, how to accurately align heterogeneous data from multiple sensors in time and space, how to build a deep fusion mechanism of multi-physical domain features, and how to achieve accurate fault identification and early warning based on fused features. Summary of the Invention
[0005] This application provides a power inspection robot inspection method and system based on multi-sensor fusion, which enables accurate identification and early warning of power equipment faults through dynamic adjustment of sensor weights with environmental adaptation, precise spatiotemporal synchronization and spatial registration, and multimodal feature fusion based on attention mechanism.
[0006] Firstly, this application provides a power line inspection robot inspection method based on multi-sensor fusion, the method comprising: Step S1: Query the weight mapping table based on the ambient magnetic field strength and ambient temperature to obtain the sensor fusion weights; Step S2: Perform weighted fusion of the multimodal localization data according to the sensor fusion weights to obtain the robot pose; Step S3: Based on the robot's pose, acquire sound signals, thermal images, and vibration signals, and perform spatiotemporal alignment on the sound signals, thermal images, and vibration signals to obtain synchronization data; Step S4: Extract the voiceprint features, thermal image features, and vibration features from the synchronization data respectively, and stitch them together to obtain a three-dimensional feature vector; Step S5: Calculate the fusion weights of each modality in the three-dimensional feature vector through an attention mechanism, and perform weighted fusion to obtain the fault type.
[0007] Secondly, this application provides a power inspection robot system based on multi-sensor fusion, the power inspection robot system based on multi-sensor fusion includes: The query module is used to query the weight mapping table based on the ambient magnetic field strength and ambient temperature to obtain the sensor fusion weights; The weighting module is used to perform weighted fusion of multimodal localization data according to the sensor fusion weights to obtain the robot pose; The alignment module is used to acquire sound signals, thermal images, and vibration signals based on the robot's pose, and to perform spatiotemporal alignment of the sound signals, thermal images, and vibration signals to obtain synchronization data; The stitching module is used to extract the acoustic features, thermal imaging features and vibration features from the synchronous data respectively, and stitch them together to obtain a three-dimensional feature vector. The calculation module is used to calculate the fusion weights of each modality in the three-dimensional feature vector through an attention mechanism, and perform weighted fusion to obtain the fault type.
[0008] Thirdly, a power inspection robot based on multi-sensor fusion is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the power inspection robot based on multi-sensor fusion to execute the aforementioned power inspection robot based on multi-sensor fusion method.
[0009] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, cause the computer to execute the above-described power inspection robot inspection method based on multi-sensor fusion.
[0010] The technical solution provided in this application solves the problem of unstable positioning accuracy of power inspection robots in complex power environments such as strong magnetic fields and high temperatures by using an environment-adaptive sensor fusion weight dynamic adjustment mechanism. Based on real-time measured environmental magnetic field strength and temperature, a pre-stored weight mapping table is queried to dynamically obtain the fusion weight coefficients of LiDAR, inertial measurement unit, visual odometry, and magnetometer. The multimodal positioning data is then weighted and fused according to these weight coefficients to obtain the robot's pose. In strong magnetic field environments, the weight of the magnetometer, which is susceptible to electromagnetic interference, is automatically reduced while the weight of the LiDAR is increased. In high-temperature environments, the weight of the visual odometry, which is susceptible to thermal distortion, is automatically reduced while the weight of the inertial measurement unit is increased. This allows the navigation system to adaptively select the most reliable sensor combination based on the degree of environmental interference, avoiding the negative impact of interfered sensor data on overall positioning accuracy in traditional fixed-weight fusion methods, and ensuring the robot's high-precision positioning capability in various complex environments. The multi-sensor spatiotemporal alignment method employed in this invention solves the technical problem of the difficulty in collaborative analysis of heterogeneous sensor data. Based on the robot's pose reaching the inspection position, it synchronously collects sound signals, thermal images, and vibration signals. Using the sound signal acquisition timestamp as the reference time axis, linear interpolation is used to map the timestamps of thermal images and vibration signals with different sampling rates to the reference time axis to achieve time synchronization. Furthermore, the spatial coordinates of each sensor are calculated through sound source localization algorithms, thermal image analysis, and vibration sensor installation positions. These coordinates are then mapped to a unified equipment coordinate system to achieve spatial registration, resulting in synchronized data that is strictly aligned in both time and space dimensions. This allows sound signature features, thermal image features, and vibration features to be accurately associated with the operating state of the same part of the equipment at the same time, laying a data foundation for subsequent multimodal feature fusion and fault diagnosis, and overcoming the feature association errors caused by spatiotemporal inconsistencies in traditional methods.
[0011] The three-dimensional feature vector construction method adopted in this invention solves the technical problem of insufficient feature representation capability of a single sensor. It performs short-time Fourier transform on the sound signal in the synchronous data and divides it into multiple sub-frequency bands to extract acoustic features. It segments the high-temperature region of the thermal image sequence through a region growing algorithm and calculates the temperature time sequence features to extract thermal image features. It performs fast Fourier transform on the three-axis components of the vibration signal to extract the vibration frequency peak features. These three types of features are spliced together in a preset order to obtain a three-dimensional feature vector, which enables the equipment operating status to be comprehensively represented from the three physical domains of acoustics, thermals, and mechanical vibration. The coupled performance of a single fault type in multiple physical domains can be completely captured. Compared with the traditional single sensor method, which can only reflect a single aspect of the fault, the three-dimensional feature vector integrates information from multiple physical domains and enhances the ability to distinguish fault features. This invention employs a multimodal feature fusion method based on an attention mechanism, which solves the technical problem that fixed-weight fusion cannot adapt to the feature distribution of different fault types. The three-dimensional feature vector is decomposed into three feature components: acoustic signature, thermal image, and vibration. These are then converted into encoded vectors by an encoding network and concatenated to form a multimodal encoding matrix. A query matrix, key matrix, and value matrix are obtained through linear transformation. The product of the query matrix and the transpose of the key matrix is calculated, normalized, and then subjected to a softmax function to obtain the attention score matrix. The attention score matrix is multiplied by the value matrix to obtain the fused feature vector. The attention mechanism can adaptively learn the importance weights of each modal feature under different fault types. The acoustic and vibration modes have higher weights for bearing wear faults, the thermal imaging mode has higher weights for local overheating faults, and the vibration mode has higher weights for mechanical loosening faults. This achieves adaptive matching between fault type and feature weights. Compared with the uniform processing of fixed weight fusion methods, the attention mechanism dynamically allocates modal importance according to specific fault modes, making the fused features more accurately reflect the essential characteristics of the fault. Finally, the fused feature vector is mapped to the probability distribution of each fault type through the fault classification network, and the fault type with the highest probability is selected as the diagnosis result, realizing accurate identification of power equipment faults and solving the technical problems of low diagnostic accuracy and weak early fault identification capability of traditional methods. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of an embodiment of the power inspection robot inspection method based on multi-sensor fusion in this application. Figure 2This is a schematic diagram of an embodiment of the power inspection robot inspection system based on multi-sensor fusion in this application. Figure 3 This is a schematic block diagram of the structure of a power inspection robot based on multi-sensor fusion in an embodiment of the present invention. Detailed Implementation
[0014] This application provides a method and system for power line inspection using a multi-sensor fusion-based robot. The terms "first," "second," "third," "fourth," etc. (if present)," in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0015] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the power inspection robot inspection method based on multi-sensor fusion in this application includes: Step S1: Query the weight mapping table based on the ambient magnetic field strength and ambient temperature to obtain the sensor fusion weights; Step S2: Weight the multimodal localization data according to the sensor fusion weights to obtain the robot pose; Step S3: Acquire sound signals, thermal images, and vibration signals based on the robot's pose, and perform spatiotemporal alignment on the sound signals, thermal images, and vibration signals to obtain synchronization data; Step S4: Extract the acoustic features, thermal imaging features, and vibration features from the synchronous data respectively, and stitch them together to obtain a three-dimensional feature vector; Step S5: Calculate the fusion weights of each modality in the three-dimensional feature vector through the attention mechanism, and perform weighted fusion to obtain the fault type.
[0016] It is understood that the executing entity of this application can be a power inspection robot system based on multi-sensor fusion, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiments use a server as an example for illustration.
[0017] Specifically, in step S1, the magnetic field strength sensor and the temperature sensor measure the magnetic field strength and ambient temperature of the inspection environment, respectively. When the magnetic field strength is 1500 microtesla and the ambient temperature is 60 degrees Celsius, the magnetic field strength is compared with the preset magnetic field threshold of 1000 microtesla and the ambient temperature is compared with the preset temperature threshold of 50 degrees Celsius. The current environmental interference level is determined to be level 3. The weight mapping table is consulted to obtain the weight coefficients of the lidar (0.45), the inertial measurement unit (IMU) (0.30), the visual odometry (VAU) (0.20), and the magnetometer (0.05). Among these, the magnetometer weight is reduced by 66.7% compared to the normal environment, while the lidar weight is increased by 28.6%. In step S2, the LiDAR scans the environment to generate 3D point cloud data containing 30,000 points, each point containing 3D coordinates and reflection intensity. The inertial measurement unit (IMU) collects three-axis acceleration and three-axis angular velocity data at a frequency of 100 Hz. The visual odometry extracts 500 ORB feature points from the binocular camera images and calculates pose change data through feature matching. The magnetometer measures three-axis magnetic field data. These raw data are converted into pose parameter formats respectively. The LiDAR pose is calculated to obtain position coordinates and attitude quaternions using the iterative nearest point algorithm. The IMU pose is obtained through integration. The visual odometry pose is obtained through essential matrix decomposition. The magnetometer pose is calculated through magnetic field direction. Then, the LiDAR pose is multiplied by 0.45, the IMU pose by 0.30, the visual odometry pose by 0.20, and the magnetometer pose by 0.05 to obtain four weighted pose components. The position coordinates of these four weighted pose components are summed, and the attitude quaternions are normalized according to the weights to obtain the robot pose. In step S3, the robot moves to an inspection position 2 meters in front of the turbine equipment based on its pose. The acoustic sensor collects 50,000 sampling points for 2 seconds of sound signal at a sampling rate of 25 kHz. The infrared thermal imager collects 50 frames of 256x192 pixel thermal images for 2 seconds at a frame rate of 25 Hz. The vibration sensor collects 2,000 triaxial vibration data points for 2 seconds of vibration signal at a sampling rate of 1 kHz. The timestamp of the sound signal collection is extracted as the reference time axis. The timestamp of the 25th frame of the thermal image is 1 second. Linear interpolation is used to map it to the corresponding moment on the reference time axis. The timestamps of 000 vibration sampling points are 1 second and are also mapped to the reference time axis. The time synchronization error is controlled within 1 millisecond. The sound source localization algorithm estimates the sound source location coordinates by calculating the time difference of the sound received by the microphone array. The spatial location is 1.8 meters away from the sensor, with a horizontal deflection angle of 15 degrees and a vertical deflection angle of 10 degrees. The thermal image analysis calculates the heat source location coordinates by finding the centroid of the highest temperature area. The vibration sensor is installed on the device base and estimates the vibration source location coordinates through the vibration propagation model. These three location coordinates are transformed from their respective sensor coordinate systems to a unified device coordinate system to obtain synchronized data.In step S4, a short-time Fourier transform is performed on the audio signal, with a window length of 1024 sampling points corresponding to 40.96 milliseconds and an overlap rate of 50%, resulting in a 512-row, 97-column spectrum matrix. The spectrum is divided into 16 sub-bands, including 30-100 Hz, 100-200 Hz, and up to 8000-12000 Hz. Five feature parameters are calculated for each sub-band: average amplitude, standard deviation, peak amplitude, peak frequency, and spectral entropy. An 80-dimensional voiceprint feature vector is generated for each of the 16 sub-bands. The highest temperature, lowest temperature, average temperature, and temperature standard deviation are calculated for each of the 50 frames of the thermal image sequence. Pixels with temperatures higher than the average temperature plus twice the standard deviation are clustered using a region growing algorithm. The high-temperature region is classified into three categories. The area, perimeter, circularity, highest temperature, and average temperature of each high-temperature region are extracted. The temperature rise rate is obtained by calculating the temperature change rate of the high-temperature region between adjacent frames. The temperature fluctuation range within 50 frames is obtained by calculating the temperature fluctuation amplitude. These parameters are combined to form a 64-dimensional thermal image feature vector. The X-axis, Y-axis, and Z-axis components of the vibration signal are subjected to Fast Fourier Transform to obtain the vibration spectrum. In the spectrum of each axis, the top 5 frequency peaks and their corresponding amplitudes are extracted by sorting them from largest to smallest. A total of 30 parameters are extracted from the three axes to form a vibration feature vector. The 80-dimensional, 64-dimensional, and 30-dimensional feature vectors are concatenated in the order of acoustic text, thermal image, and vibration to obtain a 174-dimensional three-dimensional feature vector.In step S5, the 174-dimensional three-dimensional feature vector is decomposed into 80-dimensional voiceprint feature components, 64-dimensional thermal image feature components, and 30-dimensional vibration feature components. The voiceprint feature components are input into a three-layer fully connected voiceprint coding network containing 128 neurons, 64 neurons, and 32 neurons, respectively. Each layer uses the ReLU activation function and outputs a 32-dimensional voiceprint coding vector. The thermal image feature components and vibration feature components are also input into their respective coding networks to obtain 32-dimensional thermal image coding vectors and 32-dimensional vibration coding vectors. The three 32-dimensional coding vectors are concatenated row-wise to form a 3x32 multimodal coding matrix. The query matrix, key matrix, and value matrix are obtained by multiplying the multimodal coding matrix with three 32x32 weight matrices, respectively. The query matrix and key matrix are transposed and multiplied to obtain a 3x3 attention score matrix. Each row of the score matrix is normalized by dividing by the square root of 32, and then the softmax function is applied to each row to convert the score into a probability value. The first row of the resulting attention score matrix represents the voiceprint modality's response to the three modalities. The attention levels are 0.65, 0.20, and 0.15, respectively. The second row represents the attention level of the thermal imaging mode, and the third row represents the attention level of the vibration mode. After multiplying the attention score matrix with the value matrix, global average pooling is performed to obtain a 32-dimensional fused feature vector. This vector is input into the first fully connected layer containing 64 neurons for linear transformation and ReLU activation to obtain an intermediate feature vector. The intermediate feature vector is input into the second fully connected layer containing 30 neurons for linear transformation to obtain a 30-dimensional power equipment fault score vector. The 30 dimensions correspond to 30 preset power equipment fault types, such as bearing wear, equipment imbalance, misalignment, mechanical loosening, and oil film eddy. The softmax function is applied to the fault score vector to convert the scores of each dimension into probability values between 0 and 1, and the sum of all probability values is 1 to obtain the power equipment fault probability distribution vector. The probability distribution vector is traversed to find the dimension index with the largest probability value. Based on the mapping relationship between this index and the preset fault type, the fault type is determined to be bearing wear, and the diagnostic confidence is a probability value of 0.92.
[0018] In one specific embodiment, step S1 includes: The magnetic field strength is measured at the current location using a magnetic field strength sensor, and the ambient temperature is measured at the current location using a temperature sensor. The environmental interference level is obtained by comparing the magnetic field strength value with a preset magnetic field threshold and the ambient temperature value with a preset temperature threshold. Based on the environmental interference level, the corresponding LiDAR weight coefficient, inertial measurement unit weight coefficient, visual odometry weight coefficient, and magnetometer weight coefficient are queried in the pre-stored weight mapping table. The weighting coefficients of lidar, inertial measurement unit, visual odometry, and magnetometer are used as the sensor fusion weights.
[0019] Specifically, the magnetic field strength sensor and temperature sensor measure the physical environmental parameters of the inspection robot's current location. The magnetic field strength sensor uses a triaxial magnetometer chip to continuously monitor the surrounding magnetic field strength, outputting the magnetic field strength value in microtesla (µT). The temperature sensor uses a digital temperature and humidity chip to measure the ambient temperature value, outputting the temperature value in degrees Celsius. In the turbine area of the power plant, a strong magnetic field is generated by the operation of large generators and transformers, and the heat generated by the equipment operation causes the ambient temperature to rise. After the inspection robot enters this area, the magnetic field strength sensor measures a magnetic field strength of 1800 µT, and the temperature sensor measures an ambient temperature of 65 degrees Celsius. The magnetic field strength value is compared with a preset magnetic field threshold to determine the degree of magnetic field interference. The preset magnetic field threshold includes five levels with threshold boundaries of 500 microtesla, 1000 microtesla, 2000 microtesla, and 5000 microtesla. The ambient temperature value is compared with a preset temperature threshold to determine the degree of temperature interference. The preset temperature threshold includes five levels with threshold boundaries of 30 degrees Celsius, 50 degrees Celsius, 70 degrees Celsius, and 90 degrees Celsius. A magnetic field strength value of 1800 microtesla, falling between 1000 and 2000 microtesla, is classified as level 3 magnetic field interference. An ambient temperature value of 65 degrees Celsius, falling between 50 and 70 degrees Celsius, is classified as level 3 temperature interference. The higher of the magnetic field interference level and the temperature interference level is taken as the environmental interference level. When the two are equal, that level value is directly adopted. Therefore, the environmental interference level is determined to be level 3. The corresponding sensor weight coefficients are queried in the pre-stored weight mapping table based on the environmental interference level. The weight mapping table is stored in the non-volatile memory of the robot controller. The table structure is five rows and four columns. The row index is the environmental interference level from 1 to 5, and the column index is the weight coefficient of LiDAR, the weight coefficient of inertial measurement unit, the weight coefficient of visual odometry, and the weight coefficient of magnetometer. Based on the environmental interference level of 3, the third row of data is queried, and the LiDAR weight coefficient is read as 0.45, the inertial measurement unit weight coefficient as 0.30, the visual odometry weight coefficient as 0.20, and the magnetometer weight coefficient as 0.05.These four weighting coefficients are used as the sensor fusion weights. The LiDAR weighting coefficient of 0.45 indicates that LiDAR positioning data accounts for 45% of the fusion calculation; the Inertial Measurement Unit (INS) weighting coefficient of 0.30 indicates that INS positioning data accounts for 30%; the Visual Odometry (VIO) weighting coefficient of 0.20 indicates that VIO positioning data accounts for 20%; and the Magnetometer weighting coefficient of 0.05 indicates that Magnetometer positioning data accounts for only 5%. Compared to environmental interference level 1, the Magnetometer weighting coefficient is reduced from 0.15 to 0.05, while the LiDAR weighting coefficient is increased from 0.35 to 0.45. The Magnetometer is easily affected by electromagnetic interference in strong magnetic field environments, which can lead to measurement data distortion. Reducing its weight avoids distorted data affecting the fusion positioning accuracy. LiDAR, based on laser ranging, is not affected by magnetic fields, so increasing its weight enhances the anti-interference capability of fusion positioning. The sum of the four weighting coefficients always equals 1 to ensure the mathematical consistency of the fusion calculation.
[0020] When the inspection robot moved from a normal environment to a strong magnetic field area near the generator, the magnetic field strength sensor detected a sudden change in the magnetic field strength from 600 microtesla to 2500 microtesla, and the temperature sensor detected an increase in the ambient temperature from 35 degrees Celsius to 75 degrees Celsius. The magnetic field strength of 2500 microtesla exceeds the threshold of 2000 microtesla but is below the threshold of 5000 microtesla, classifying it as level 4 magnetic field interference. The ambient temperature of 75 degrees Celsius exceeds the threshold of 70 degrees Celsius but is below the threshold of 90 degrees Celsius, classifying it as level 4 temperature interference. Taking the higher of the two values, the environmental interference level is determined to be level 4. Querying the fourth row of the weight mapping table, the weight coefficients for the LiDAR and inertial measurement units are found to be 0.50 and 0.35 respectively, indicating visual odometry. The weighting coefficients are 0.13 for the fusion sensor and 0.02 for the magnetometer. The magnetometer weighting coefficient is further reduced to 0.02, retaining only a minimal reference role. The weighting coefficient for the lidar is increased to 0.50, becoming the dominant positioning data source. The weighting coefficient for the inertial measurement unit is increased to 0.35, serving as an important auxiliary positioning data source. The weighting coefficient for the visual odometry is reduced to 0.13 because camera lenses are prone to thermal distortion in high-temperature environments, affecting image quality. This dynamic weighting adjustment mechanism adjusts the fusion weights of each sensor in real time according to the degree of environmental interference. In strong magnetic field environments, the weight of the magnetometer is reduced while the weights of the lidar and inertial measurement unit are increased. In high-temperature environments, the weight of the visual odometry is reduced to ensure the accuracy and stability of fusion positioning in complex power environments.
[0021] In one specific embodiment, step S2 includes: The surrounding environment is scanned by lidar to generate three-dimensional point cloud data, three-axis acceleration data and three-axis angular velocity data are collected by inertial measurement unit, image feature points are extracted by visual odometry to calculate pose change data, and three-axis magnetic field data are measured by magnetometer. The three-dimensional point cloud data, three-axis acceleration data, three-axis angular velocity data, pose change data, and three-axis magnetic field data are converted into corresponding pose parameters to obtain the pose of the lidar, the pose of the inertial measurement unit, the pose of the visual odometry, and the pose of the magnetometer. The weighted pose components are obtained by multiplying the LiDAR pose with the LiDAR weight coefficient, the inertial measurement unit pose with the inertial measurement unit weight coefficient, the visual odometry pose with the visual odometry weight coefficient, and the magnetometer pose with the magnetometer weight coefficient. The robot pose is obtained by summing the weighted pose components.
[0022] Specifically, the lidar rotates and scans the surrounding environment at a frequency of 10 Hz, with each scan angle being 270 degrees and an angular resolution of 0.25 degrees. Each scan generates 1080 ranging points, each containing distance and angle values. After 3D coordinate transformation, 3D point cloud data containing X, Y, and Z coordinates is generated. Each point in the point cloud data also includes reflection intensity information, representing the energy intensity of the laser reflected back. The inertial measurement unit continuously collects motion data at a frequency of 100 Hz. The three-axis accelerometer measures the linear acceleration in the X, Y, and Z axes, respectively, and the three-axis gyroscope measures the angular velocity around the X, Y, and Z axes, respectively. The unit of acceleration data is meters per second squared, and the unit of angular velocity data is degrees per second. Visual odometry acquires images from both left and right perspectives using binocular cameras. After grayscale processing, the images are analyzed using the ORB feature extraction algorithm to detect keypoints. The ORB algorithm first uses FAST corner detection to identify locations with dramatic grayscale changes as keypoints. Then, it calculates the binary descriptors around each keypoint. Keypoints in adjacent frames are matched using Hamming distance. Successfully matched keypoint pairs are then triangulated to calculate their 3D spatial positions. Changes in keypoint positions between consecutive frames reflect changes in the robot's pose. The pose change data includes both positional and orientation changes. A magnetometer measures the three-axis magnetic field strength, outputting X-axis, Y-axis, and Z-axis magnetic field strengths in microtesla (µT). The relationship between the geomagnetic field direction and the robot's orientation is used to calculate the pose angle.
[0023] 3D point cloud data is converted into LiDAR pose using an iterative nearest-point algorithm. This algorithm registers the current frame point cloud with the previous frame point cloud. First, for each point in the current frame point cloud, the nearest point in the previous frame is searched as its corresponding point. Then, a rotation matrix and translation vector that minimizes the sum of squared distances between all corresponding point pairs are calculated. These rotation and translation vectors describe the pose change of the current frame relative to the previous frame. Accumulating the pose changes across all frames yields the LiDAR pose, which includes position coordinates and attitude quaternions. Position coordinates are X, Y, and Z coordinates, and the attitude quaternion consists of four values representing rotation in 3D space. Three-axis acceleration and angular velocity data are converted into inertial measurement unit (IMU) pose through integration. Integrating the angular velocity over time yields the change in attitude angle. Accumulating these changes gives the current attitude angle, which is then converted into an attitude quaternion. Subtracting gravitational acceleration from acceleration yields the robot's motion acceleration. Integrating the motion acceleration over time yields the velocity, and integrating the velocity over time yields the position coordinates. The IMU pose includes both position coordinates and attitude quaternions. Pose change data directly represents the increment of visual odometry pose. Accumulating the increments yields the visual odometry pose, which also includes position coordinates and attitude quaternions. Three-axis magnetic field data calculates the robot's heading angle using the angle between the magnetic field direction and the geomagnetic field direction. The heading angle, combined with the tilt angle, is converted into attitude quaternions. Since the magnetometer does not provide position information, the magnetometer pose only contains attitude quaternions, while the position coordinates are set to zero vectors.
[0024] The lidar pose coordinates are multiplied by the lidar weighting coefficient, and the attitude quaternion is multiplied by the lidar weighting coefficient to obtain the lidar weighted pose component. The inertial measurement unit (IMU) pose coordinates are multiplied by the IMU weighting coefficient, and the attitude quaternion is multiplied by the IMU weighting coefficient to obtain the IMU weighted pose component. The visual odometry (VAU) pose coordinates are multiplied by the VVA weighting coefficient, and the attitude quaternion is multiplied by the VVA weighting coefficient to obtain the VVA weighted pose component. The magnetometer pose coordinates zero vector is multiplied by the magnetometer weighting coefficient, and the result is still a zero vector. The attitude quaternion is multiplied by the magnetometer weighting coefficient to obtain the magnetometer weighted pose component. The position coordinates of the four weighted pose components are summed separately. The X coordinates of the LiDAR weighted position coordinates, the inertial measurement unit weighted position coordinates, the visual odometry weighted position coordinates, and the magnetometer weighted position coordinates are added to obtain the fused X coordinate. The Y and Z coordinates are also added separately to obtain the position coordinates of the robot pose. The attitude quaternions of the four weighted pose components are normalized according to their weights and then summed. First, the four attitude quaternions are transformed to the same reference direction to avoid the problem of double coverage of quaternions. Then, the four components of the four attitude quaternions are weighted and summed separately. The summation result is normalized to make the quaternion magnitude 1, so as to obtain the attitude quaternion of the robot pose.
[0025] As the inspection robot moves alongside the steam turbine equipment, the lidar scans the turbine casing and pipes, generating point cloud data. Using an iterative nearest-point algorithm, continuous frames of point cloud data are registered to calculate the robot's positional changes relative to the equipment, which is then converted into a lidar pose. The inertial measurement unit (IMU) collects the robot's acceleration and angular velocity during acceleration and turning, and calculates the robot's trajectory and attitude changes through integration, converting this into an IMU pose. The visual odometry extracts texture feature points from the equipment surface in the camera images, matches these feature points across consecutive frames to calculate the robot's pose changes, and converts this into a visual odometry pose. The magnetometer measures the surrounding magnetic field strength, calculates the robot's heading angle based on the magnetic field direction, and converts this into a magnetometer pose. The lidar pose is multiplied by a weighting coefficient of 0.45, and the IMU pose is multiplied by a weighting coefficient of 0.30 to obtain the visual odometry pose. The pose is multiplied by a weighting coefficient of 0.20, and the magnetometer pose is multiplied by a weighting coefficient of 0.05 to obtain four weighted pose components. The position coordinates are summed separately to obtain the fused position coordinates. The pose quaternions are weighted, normalized, and then summed to obtain the fused pose quaternions. The robot pose integrates the positioning data of four sensors. LiDAR has high positioning accuracy but high computational cost, inertial measurement unit has fast response but has cumulative error, visual odometry has rich features but is affected by lighting, and magnetometer provides absolute direction but is susceptible to magnetic field interference. Through a weighted fusion mechanism, the advantages of each sensor are complemented. The weighting coefficients are dynamically adjusted according to the level of environmental interference. In strong magnetic field environments, the weight of magnetometer is reduced and the weight of LiDAR is increased. In high temperature environments, the weight of visual odometry is reduced and the weight of inertial measurement unit is increased. The fused robot pose has a positioning error controlled within 5 mm in complex power environments.
[0026] In one specific embodiment, step S3 includes: Based on the robot's pose, it arrives at the preset inspection position, collects the sound signal of the equipment operation through the voiceprint sensor, collects the thermal image sequence of the equipment surface through the infrared thermal imager, and collects the vibration signal of the equipment through the vibration sensor. The acquisition timestamps of the sound signal, the thermal images of each frame in the thermal image sequence, and the vibration signal are extracted. Using the acquisition timestamp of the sound signal as the reference time axis, the timestamps of the thermal image sequence and the vibration signal are mapped to the reference time axis through linear interpolation to obtain the time-synchronized sound signal, thermal image sequence, and vibration signal. The sound source location coordinates corresponding to the sound signal are calculated by the sound source localization algorithm, the heat source location coordinates corresponding to the heat image sequence are calculated by thermal image analysis, and the vibration source location coordinates corresponding to the vibration signal are calculated based on the installation location of the vibration sensor. The coordinates of the sound source, heat source, and vibration source are mapped to a unified equipment coordinate system to obtain synchronization data.
[0027] Specifically, the robot moves to a pre-planned inspection position based on its pose information. This pre-set inspection position is stored in the inspection task database and includes information such as device number, device type, inspection position coordinates, and sensor orientation angle. After reaching the inspection position, the robot adjusts its gimbal to align the sensor with the device to be inspected. The acoustic signature sensor uses a microphone array with four microphones arranged in a square, synchronously collecting the device's operating sound at a 25 kHz sampling rate for 2 seconds, generating a sound signal containing 50,000 sampling points. Each sampling point records the sound amplitude value and the timestamp of the acquisition time, with a timestamp accuracy of 0.04 milliseconds. An infrared thermal imager continuously captures images of the device surface at a 25 Hz frame rate for 2 seconds, generating a thermal image sequence containing 50 frames. Each frame has a resolution of 256 x 192 pixels, and each pixel records the temperature value at the corresponding location and the acquisition timestamp of that frame, with a timestamp accuracy of 40 milliseconds. The vibration sensor uses a triaxial accelerometer mounted on the equipment base to collect equipment vibration at a sampling rate of 1 kHz for 2 seconds, generating a vibration signal containing 2000 sampling points. Each sampling point records the X-axis vibration acceleration, Y-axis vibration acceleration, Z-axis vibration acceleration, and the timestamp of the acquisition time, with a timestamp accuracy of 1 millisecond.
[0028] The timestamp of the first sampling point in the audio signal is extracted as the starting point of the reference time axis. Timestamps from the first frame, second frame, up to the 50th frame are extracted from the thermal image sequence. Similarly, timestamps from the first sampling point, second frame, up to the 2000th sampling point are extracted from the vibration signal. A 25 kHz sampling rate for the audio signal means 25,000 sampling points are collected per second, with a time interval of 0.04 milliseconds between adjacent sampling points. A 25 Hz frame rate for the thermal image means 25 frames are collected per second, with a time interval of 40 milliseconds between adjacent frames. A 1 kHz sampling rate for the vibration signal means 1,000 sampling points are collected per second, with a time interval of 1 millisecond between adjacent sampling points. The different sampling rates of the three sensors result in the timestamps not being on the same time axis. Using the acquisition timestamps of the sound signals as the reference time axis, ranging from 0 milliseconds to 2000 milliseconds, the timestamps of the thermal image sequence are mapped to this reference time axis. The first frame of the thermal image has a timestamp of 0 milliseconds, directly corresponding to 0 milliseconds on the reference time axis; the second frame has a timestamp of 40 milliseconds, directly corresponding to 40 milliseconds on the reference time axis; the 25th frame has a timestamp of 960 milliseconds, directly corresponding to 960 milliseconds on the reference time axis; and the 50th frame has a timestamp of 1960 milliseconds, directly corresponding to 1960 milliseconds on the reference time axis. Similarly, the timestamps of the vibration signals are mapped to the reference time axis. The first vibration sampling point has a timestamp of 0 milliseconds, corresponding to 0 milliseconds on the reference time axis; the second vibration sampling point has a timestamp of 1 millisecond, corresponding to 1 millisecond on the reference time axis; the 1000th vibration sampling point has a timestamp of 999 milliseconds, corresponding to 999 milliseconds on the reference time axis; and the 2000th vibration sampling point has a timestamp of 1999 milliseconds, corresponding to 1999 milliseconds on the reference time axis. By using linear interpolation to process timestamps that are not integer multiples, the sound signal corresponds to the 2000th sampling point at 80 milliseconds on the reference time axis. The thermal image sequence is located between the 2nd and 3rd frames at 80 milliseconds. The thermal image at 80 milliseconds is calculated by linear interpolation as a weighted average of the thermal images of the 2nd and 3rd frames according to the time distance. The vibration signal corresponds to the 80th sampling point at 80 milliseconds.
[0029] The sound source localization algorithm calculates the sound source location by the time difference of the sound signals received by the microphone array. Since the four microphones receive the same sound signal at different times, and the sound travels different distances from the sound source to different microphones, the arrival times differ. The algorithm calculates the time difference between any two microphones, multiplies this time difference by the speed of sound, and obtains the distance difference between the two microphones and the sound source. The four microphones form six pairs, generating six distance difference equations. Solving these equations calculates the three-dimensional spatial coordinates of the sound source, including the X-axis, Y-axis, and Z-axis distances relative to the center of the microphone array. Thermal image analysis identifies the heat source location by finding the region with the highest temperature in the thermal image. It iterates through all pixels in the thermal image, finding the pixel with the highest temperature value as the center of the heat source. The row and column numbers of the center pixel are converted to coordinates in the thermal image coordinate system. Combining the thermal imager's field of view and the distance from the robot to the device, triangulation is used to calculate the three-dimensional spatial coordinates of the heat source, including the X-axis, Y-axis, and Z-axis distances relative to the thermal imager. The vibration sensor is installed at a fixed position on the equipment base. The vibration signal collected by the vibration sensor comes from the vibration of the internal moving parts of the equipment, which is transmitted to the base through the equipment structure. The vibration source position coordinates are directly set as the coordinates of the vibration sensor installation position. The vibration source position coordinates include the X-axis distance, Y-axis distance, and Z-axis distance relative to the reference point of the equipment base.
[0030] The coordinate system for the sound source location coordinates is the microphone array coordinate system, the coordinate system for the heat source location coordinates is the thermal imager coordinate system, and the coordinate system for the vibration source location coordinates is the device coordinate system. The origins and coordinate axis directions of these three coordinate systems are different. These three coordinate systems are unified into the device coordinate system. The device coordinate system has its origin at the center of the device base, with the X-axis pointing forward, the Y-axis pointing to the left, and the Z-axis pointing upward. The position offset and rotation angle of the microphone array coordinate system relative to the device coordinate system are known. A coordinate transformation matrix is used to transform the sound source location coordinates from the microphone array coordinate system to the device coordinate system. This transformation includes rotation followed by translation. The rotation matrix is constructed based on the installation angle of the microphone array, and the translation vector is the position offset of the microphone array center relative to the center of the device base. Similarly, the position offset and rotation angle of the thermal imager coordinate system relative to the device coordinate system are also known. A similar coordinate transformation matrix is used to transform the heat source location coordinates to the device coordinate system. The vibration source location coordinates are already in the device coordinate system and therefore do not require transformation. After the three location coordinates are unified into the device coordinate system, the location coordinates of the sound source, heat source, and vibration source form synchronous data. The synchronous data is not only aligned to the reference time axis in the time dimension, but also aligned to the unified device coordinate system in the spatial dimension. The sound signal, thermal image sequence, and vibration signal are associated with their respective spatial location coordinates. Subsequent feature extraction and fault diagnosis are based on the spatiotemporally aligned synchronous data.
[0031] When the inspection robot checks the steam turbine equipment, it moves to an inspection position two meters in front of the turbine. A microphone array collects the turbine's operating sounds, an infrared thermal imager captures thermal images of the turbine bearings, and vibration sensors collect vibrations from the turbine base. The sound signals collect the periodic sounds generated by the rotor's rotation inside the turbine and the friction sounds from the bearings. The thermal image sequence displays the temperature distribution of the bearing housing, and the vibration signals record the base vibrations caused by the rotor's rotation. The sound signal timestamp starts at 0 milliseconds; the 10th frame of the thermal image has a timestamp of 360 milliseconds, mapped to a reference time axis of 360 milliseconds; and the 360th sampling point of the vibration signal has a timestamp of 359 milliseconds, mapped to a reference time axis of 359 milliseconds. At 9 milliseconds, the data points of the three sensors were synchronized at a reference time axis of around 360 milliseconds. The sound source localization algorithm calculated that the sound mainly came from the middle and slightly right of the turbine. The thermal image analysis calculated that the heat source was located at the right bearing of the turbine, and the vibration source was located at the center of the base. After the coordinates of the three positions were transformed into the equipment coordinate system, the sound source position coordinates and the heat source position coordinates were close on the X and Y axes, indicating that the sound and heat came from the same bearing. The vibration source position coordinates were lower on the Z axis than the sound source and heat source positions, indicating that the vibration was transmitted from the bearing to the base. The time-space aligned synchronization data revealed that the bearing simultaneously generated abnormal sound, high temperature and vibration. The consistency of the three physical phenomena in time and space pointed to the bearing wear failure.
[0032] In one specific embodiment, step S4 includes: A short-time Fourier transform is performed on the audio signal in the synchronization data to convert the time-domain audio signal into a spectrum matrix. The spectrum matrix is divided into multiple sub-bands, and the average amplitude, standard deviation, peak amplitude, peak frequency and spectral entropy of each sub-band are calculated to obtain the voiceprint features. The global maximum temperature, minimum temperature, average temperature and temperature standard deviation are calculated for the thermal image sequence in the synchronous data. The high temperature region is segmented by the region growing algorithm and the area, perimeter, circularity, maximum temperature and average temperature of each high temperature region are extracted. The temperature rise rate and temperature fluctuation amplitude of the high temperature region are calculated to obtain the thermal image features. Fast Fourier transform is performed on the three-axis components of the vibration signal in the synchronous data to extract the top five frequency peaks with the largest amplitude and their corresponding amplitudes in the vibration spectrum of each axis, thus obtaining the vibration characteristics. By splicing voiceprint features, thermal imaging features, and vibration features in a preset order, a three-dimensional feature vector is obtained.
[0033] Specifically, a short-time Fourier transform (SFT) is performed on the audio signal containing 50,000 sampling points in the synchronization data. The SFT divides the time-domain signal into frames, with each frame selecting 1,024 sampling points corresponding to a duration of 40.96 milliseconds. Adjacent frames overlap by 512 sampling points, i.e., an overlap rate of 50%. A Hamming window function is added to each frame of audio signal to reduce spectral leakage. A 1,024-point fast Fourier transform is then performed on the windowed signal to obtain the spectrum of that frame. The spectrum contains 512 frequency components covering a frequency range of 0 to 12.5 kHz, with a frequency resolution of 24.4 Hz. The 50,000 sampling points are divided into 97 frames, resulting in a 512-row, 97-column spectrum matrix. Each element of the spectrum matrix represents the amplitude of a specific frequency component at a specific time. The spectrum matrix is divided into 16 sub-bands according to the frequency dimension. The first sub-band covers 30 to 100 Hz and corresponds to rows 1 to 3 of the spectrum matrix. The second sub-band covers 100 to 200 Hz and corresponds to rows 4 to 8. The third sub-band covers 200 to 500 Hz and corresponds to rows 9 to 20. Sub-bands are divided sequentially up to the 16th sub-band, covering 16,000 to 20,000 Hz. For each sub-band, the average amplitude is calculated, which is the average value of all frequency components in the sub-band at all times. The standard deviation is calculated, which is the degree of dispersion of the amplitude relative to the average value. The peak amplitude is calculated, which is the largest amplitude in the sub-band. The frequency corresponding to the peak amplitude is recorded as the peak frequency. The spectral entropy is calculated to represent the degree of disorder in the frequency distribution of the sub-band. The spectral entropy is calculated by normalizing the amplitude of each frequency component to a probability distribution and then calculating the Shannon entropy. Five feature parameters are calculated for each of the 16 sub-bands to obtain an 80-dimensional voiceprint feature vector.
[0034] Temperature statistics are calculated for a sequence of 50 thermal images included in the synchronized data. The highest temperature pixel is identified as the global maximum temperature, and the lowest temperature pixel is identified as the global minimum temperature. The arithmetic mean of all pixel temperatures is calculated as the average temperature, and the standard deviation of all pixel temperatures relative to the average temperature is calculated as the temperature standard deviation. A region growing algorithm is used to segment high-temperature regions. The algorithm starts with a pixel whose temperature is higher than the average temperature plus twice the standard deviation as a seed point. It checks the four adjacent pixels (up, down, left, and right) of the seed point, merging adjacent pixels with temperatures above a threshold into the high-temperature region and using them as new seed points to continue expanding. This expansion process is repeated until no adjacent pixels meet the conditions. A connected set of high-temperature pixels constitutes a high-temperature region. Multiple high-temperature regions are obtained by traversing all pixels that meet the seed point conditions. Shape and temperature features are extracted for each high-temperature region. The number of pixels contained in the high-temperature region is calculated as its area, and the number of pixels at the boundary of the high-temperature region is calculated as its perimeter. The circularity is calculated as 4 times pi multiplied by the area divided by the square of the perimeter. A circularity close to 1 indicates that the region's shape is close to a circle. The highest pixel value within the high-temperature region is calculated as the region's highest temperature, and the average temperature of all pixels within the high-temperature region is calculated as the region's average temperature. The temperature rise rate of the high-temperature region is calculated, and the positional change of the high-temperature region between consecutive frames is tracked. The change in the average temperature of the high-temperature region between corresponding frames is divided by the time interval to obtain the temperature rise rate. The temperature fluctuation amplitude of the high-temperature region is calculated, and the maximum and minimum values of the average temperature of the high-temperature region are statistically analyzed over 50 frames. The difference between the maximum and minimum values is used to obtain the temperature fluctuation amplitude. By combining global temperature statistical features, high-temperature region shape features, high-temperature region temperature features, and temperature temporal features, a 64-dimensional thermal image feature vector is obtained.
[0035] The vibration signal containing 2000 triaxial sampling points in the synchronous data was processed separately for the X-axis, Y-axis, and Z-axis components. A 2048-point Fast Fourier Transform (FFT) was performed on the 2000 sampling points of the X-axis vibration signal. The FFT converted the time-domain vibration signal into a frequency-domain vibration spectrum. The vibration spectrum contained 1024 frequency components covering the frequency range of 0 to 500 Hz, with a frequency resolution of 0.488 Hz. The frequency component with the largest amplitude was found by traversing the 1024 frequency components of the vibration spectrum, and its frequency and amplitude were recorded. After setting the frequency component to zero, the frequency component with the second largest amplitude was searched. This process was repeated 5 times to obtain the top 5 frequency peaks with the largest amplitude in the X-axis vibration spectrum and their corresponding amplitudes. The 5 frequency peaks correspond to the main frequency components of the equipment vibration. The same processing method was used to extract 5 frequency peaks and amplitudes from the Y-axis and Z-axis vibration signals, respectively. A total of 30 parameters were extracted from the three axes. The 15 frequency peaks and 15 corresponding amplitudes formed a 30-dimensional vibration feature vector.
[0036] The three types of feature vectors are concatenated according to a preset order of acoustic features, thermal imaging features, and vibration features. The first to 80th elements of the 80-dimensional acoustic feature vector are arranged sequentially. The first to 64th elements of the 64-dimensional thermal imaging feature vector are appended to the acoustic feature vector as the 81st to 144th elements. The first to 30th elements of the 30-dimensional vibration feature vector are appended to the thermal imaging feature vector as the 145th to 174th elements, resulting in a 174-dimensional three-dimensional feature vector. The three-dimensional feature vector contains the operating state features of the equipment in the three physical domains of acoustics, thermals, and mechanical vibration. The acoustic features characterize the frequency components and energy distribution of the sound of the equipment, the thermal imaging features characterize the temperature distribution and heat accumulation on the surface of the equipment, and the vibration features characterize the vibration frequency and intensity of the internal operating parts of the equipment. The combination of the three types of features comprehensively describes the multi-physics operating state of the equipment.
[0037] In one specific embodiment, step S5 includes: The three-dimensional feature vector is decomposed into voiceprint feature components, thermal image feature components, and vibration feature components, which are then input into the voiceprint coding network, thermal image coding network, and vibration coding network, respectively, for encoding processing to obtain the voiceprint coding vector, thermal image coding vector, and vibration coding vector. The voiceprint encoding vector, thermal image encoding vector, and vibration encoding vector are concatenated to form a multimodal encoding matrix. The multimodal encoding matrix is then mapped to a query matrix, a key matrix, and a value matrix through linear transformation. The query matrix and the transpose of the key matrix are multiplied together. The result is normalized and then the attention score matrix is calculated using the softmax function. The attention score matrix is then multiplied with the value matrix to obtain the fused feature vector. The fused feature vectors are input into the fault classification network for classification processing, and the probability distribution of each fault type is output. The fault type with the highest probability is selected as the fault type.
[0038] Specifically, the 174-dimensional feature vector is decomposed according to feature type: the first 80 elements are voiceprint feature components, the 81st to 144th elements are thermal image feature components, and the 145th to 174th elements are vibration feature components. The voiceprint feature components are input into the voiceprint coding network, which contains three fully connected layers. The first layer contains 128 neurons, the second layer contains 64 neurons, and the third layer contains 32 neurons. Each fully connected layer undergoes a linear transformation using a weight matrix and a bias vector, followed by the ReLU activation function. The ReLU activation function sets negative values to zero and retains positive values. The 80-dimensional voiceprint feature component is multiplied by the 128x80 weight matrix of the first layer and then added to the 128-dimensional bias vector to obtain a 128-dimensional vector. After applying ReLU activation, this 128-dimensional vector is input into the second layer. The 128-dimensional vector is then multiplied by the 64x128 weight matrix of the second layer and then added to the 64-dimensional bias vector to obtain a 64-dimensional vector. After applying ReLU activation, this 64-dimensional vector is input into the third layer. The 64-dimensional vector is then multiplied by the 32x64 weight matrix of the third layer and then added to the 32-dimensional bias vector to obtain a 32-dimensional voiceprint coding vector. Thermal image feature components and vibration feature components are input into thermal image coding network and vibration coding network with the same structure, respectively. The 64-dimensional thermal image feature components are processed through a three-layer fully connected network to obtain a 32-dimensional thermal image coding vector, and the 30-dimensional vibration feature components are processed through a three-layer fully connected network to obtain a 32-dimensional vibration coding vector. The three coding networks encode the feature components of different dimensions into a unified 32-dimensional vector.
[0039] The 32-dimensional voiceprint encoding vector, 32-dimensional thermal image encoding vector, and 32-dimensional vibration encoding vector are concatenated row-wise, with the voiceprint encoding vector as the first row, the thermal image encoding vector as the second row, and the vibration encoding vector as the third row, forming a 3x32 multimodal encoding matrix. The multimodal encoding matrix is then mapped to a query matrix, a key matrix, and a value matrix using three linear transformation matrices. The query matrix is obtained by multiplying the multimodal encoding matrix by a 32x32 query weight matrix. Each row of the multimodal encoding matrix is multiplied by the query weight matrix to obtain the corresponding row of the query matrix, resulting in a 3x32 query matrix. The key matrix is obtained by multiplying the multimodal encoding matrix by a 32x32 key weight matrix, also resulting in a 3x32 key matrix. The value matrix is obtained by multiplying the multimodal encoding matrix by a 32x32 value weight matrix, also resulting in a 3x32 value matrix. The parameters of the three weight matrices are learned through training.
[0040] Matrix multiplication is performed between the query matrix and the transpose of the key matrix. The transpose of the key matrix is 32x3. Multiplying the 3x32 query matrix by the 3x32 transpose of the key matrix yields a 3x3 original attention score matrix. The first element of the first row and first column of the original attention score matrix is the dot product of the first row of the query matrix and the first column of the transpose of the key matrix, representing the degree of attention of the voiceprint modality to itself. The second element of the first row and second column is the dot product of the first row of the query matrix and the second column of the transpose of the key matrix, representing the degree of attention of the voiceprint modality to the thermal imaging modality. This process is repeated to obtain nine elements. The original attention score matrix is then normalized by dividing each element of the original matrix by the square root of 32, which is 5.657. Normalization prevents the matrix multiplication result from being too large, which could cause the gradient of the softmax function to vanish. The attention score matrix is calculated by applying the softmax function to the normalized matrix. The softmax function is applied row-wise. For each of the three elements in the first row of the attention score matrix, an exponential function is calculated. The exponent of the first element in the first row and first column is divided by the sum of the exponents of the three elements in the first row to obtain the first element in the first row and first column after softmax. The sum of the three elements in the first row after softmax is 1. The second and third rows are processed in the same way, resulting in a 3x3 attention score matrix. Each row of the attention score matrix represents the attention weight distribution of the corresponding modality to the three modalities. The attention score matrix is then multiplied by the value matrix. Multiplying the 3x3 attention score matrix by the 3x32 value matrix yields a 3x32 fusion coding matrix. The first row of the fusion coding matrix is a weighted combination of the first row of the attention score matrix and the three rows of the value matrix, representing the encoded vector after fusing the voiceprint modality with other modal information. Global average pooling is then applied to the three rows of the fusion coding matrix, and the average of the corresponding columns in the three rows is calculated to obtain a 32-dimensional fusion feature vector.
[0041] A 32-dimensional fused feature vector is input into the fault classification network, which contains two fully connected layers. The first layer contains 64 neurons. The fused feature vector is multiplied by a 64x32 weight matrix, and after adding a 64-dimensional bias vector, ReLU activation is applied to obtain a 64-dimensional intermediate feature vector. The second layer contains 30 neurons corresponding to 30 types of power equipment faults. The intermediate feature vector is multiplied by a 30x64 weight matrix, and after adding a 30-dimensional bias vector, a 30-dimensional power equipment fault score vector is obtained. Each element of the fault score vector represents the score for the corresponding fault type. The softmax function is applied to the fault score vector to calculate the probability distribution of each fault type. The first element of the fault score vector corresponds to the bearing wear fault score. The exponential function value of this element is calculated, and this exponential value is divided by the sum of the exponential values of the 30 elements of the fault score vector to obtain the probability of the bearing wear fault. The probabilities of the 30 fault types are calculated sequentially to obtain a 30-dimensional power equipment fault probability distribution vector. The sum of all elements in the probability distribution vector is 1. Traverse the 30 elements of the probability distribution vector to find the element with the highest probability value, record the index position of the element, and determine the fault type according to the mapping relationship between the index position and the preset power equipment fault type. Index 0 corresponds to bearing wear, index 1 corresponds to equipment imbalance, index 2 corresponds to misalignment, index 3 corresponds to mechanical looseness, and index 4 corresponds to oil film whirl. When the index with the highest probability is 0, the fault type is determined to be bearing wear.
[0042] In one specific embodiment, the fused feature vector is input into a fault classification network for classification processing, outputting the probability distribution of each fault type, and selecting the fault type with the highest probability as the fault type, including: The fused feature vector is input into the first fully connected layer for linear transformation and activation processing to obtain the intermediate feature vector. The intermediate feature vector is input into the second fully connected layer for linear transformation to obtain the power equipment fault score vector. The number of dimensions of the power equipment fault score vector is the same as the number of preset power equipment fault types. The power equipment fault score vector is subjected to softmax normalization to convert the score values of each dimension into probability values, thus obtaining the power equipment fault probability distribution vector. Traverse the probability values in the power equipment failure probability distribution vector, select the dimension index with the largest probability value, and obtain the failure type according to the mapping relationship between the dimension index and the preset power equipment failure type. The preset power equipment failure type includes at least one of bearing wear, equipment imbalance, misalignment, mechanical loosening or oil film eddy.
[0043] Specifically, a 32-dimensional fused feature vector is input into the first fully connected layer. The first fully connected layer contains 64 neurons and corresponding weight matrices and bias vectors. The weight matrix has a dimension of 64 rows and 32 columns, and the bias vector has a dimension of 64. The 32 elements of the fused feature vector are multiplied one by one with the 32 elements of the first row of the weight matrix and then summed. The first element of the bias vector is added to obtain the first element of the output of the first fully connected layer. The fused feature vector is calculated with the second row of the weight matrix to obtain the second element. This process is repeated to obtain a 64-dimensional linear transformation result. Each element of the linear transformation result is processed by the ReLU activation function. The ReLU activation function determines whether the element value is greater than zero. If it is greater than zero, the original value is kept. If it is less than or equal to zero, it is set to zero. After activation, a 64-dimensional intermediate feature vector is obtained. The intermediate feature vector retains the positive value information in the linear transformation result and suppresses the negative value information.
[0044] A 64-dimensional intermediate feature vector is input into the second fully connected layer. The second fully connected layer contains 30 neurons corresponding to 30 preset power equipment fault types. The weight matrix of the second fully connected layer has a dimension of 30 rows and 64 columns, and the bias vector has a dimension of 30. The 64 elements of the intermediate feature vector are multiplied one by one with the 64 elements of the first row of the weight matrix and then summed. The first element of the bias vector is added to obtain the first element of the power equipment fault score vector. The first element corresponds to the score of the first fault type, namely bearing wear. The intermediate feature vector is calculated with the second row of the weight matrix to obtain the score of the second element, corresponding to the equipment imbalance. This process is repeated to obtain the 30-dimensional power equipment fault score vector. The 30 elements of the score vector correspond to the scores of bearing wear, equipment imbalance, misalignment, mechanical looseness, oil film eddy, and 25 other power equipment fault types. The score values are real numbers without a fixed range.
[0045] The power equipment fault score vector is normalized using softmax. The softmax function converts the score values into probability values. First, the exponential function value of each element in the score vector is calculated. The bearing wear score is the first element, and the natural exponent of the element is calculated to obtain the first exponent value. The equipment imbalance score is the second element, and the natural exponent of the element is calculated to obtain the second exponent value. The exponent values of all 30 elements are calculated in turn. Then, the sum of the 30 exponent values is used as the normalization denominator. The exponent value of bearing wear is divided by the sum of the exponent values to obtain the probability value of bearing wear. The exponent value of equipment imbalance is divided by the sum of the exponent values to obtain the probability value of equipment imbalance. This process yields a 30-dimensional power equipment fault probability distribution vector. Each element of the probability distribution vector represents the probability of the corresponding fault type occurring. All 30 probability values are between 0 and 1, and their sum equals 1. The softmax function converts the relative magnitudes of the score vectors into a probability distribution. Fault types with higher scores have higher probabilities, and fault types with lower scores have lower probabilities.
[0046] Iterate through the 30 probability values of the power equipment failure probability distribution vector, record the first probability value and its index 0, compare the second probability value with the current maximum probability value. If the second probability value is larger, update the maximum probability value and its index to 1. Compare the 30 probability values in turn to find the element with the largest probability value and record the dimension index of that element. Determine the failure type based on the mapping relationship between the dimension index and the preset power equipment failure type. The mapping relationship is defined as follows: index 0 corresponds to bearing wear, index 1 corresponds to equipment imbalance, index 2 corresponds to misalignment, index 3 corresponds to mechanical looseness, index 4 corresponds to oil film eddy, and indices 5 to 29 correspond to the other 25 power equipment failure types. When the dimension index with the largest probability value is 0, the failure type is determined to be bearing wear. When the dimension index is 3, the failure type is determined to be mechanical looseness. The preset power equipment failure types include, but are not limited to, bearing wear, equipment imbalance, misalignment, mechanical looseness, and oil film eddy, as well as common power equipment failures such as bearing cracks, rotor bending, foundation looseness, coupling failure, and gear meshing failure.
[0047] The above describes the power inspection robot inspection method based on multi-sensor fusion in the embodiments of this application. The following describes the power inspection robot inspection system based on multi-sensor fusion in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the power inspection robot inspection system based on multi-sensor fusion in this application includes: The query module is used to query the weight mapping table based on the ambient magnetic field strength and ambient temperature to obtain the sensor fusion weights; The weighting module is used to perform weighted fusion of multimodal localization data according to the sensor fusion weights to obtain the robot pose; The alignment module is used to acquire sound signals, thermal images, and vibration signals based on the robot's pose, and to perform spatiotemporal alignment of the sound signals, thermal images, and vibration signals to obtain synchronization data; The stitching module is used to extract the acoustic features, thermal imaging features and vibration features from the synchronous data respectively, and stitch them together to obtain a three-dimensional feature vector. The calculation module is used to calculate the fusion weights of each modality in the three-dimensional feature vector through an attention mechanism, and perform weighted fusion to obtain the fault type.
[0048] above Figure 2 The multi-sensor fusion-based power inspection robot system in this embodiment of the invention is described in detail from the perspective of modular functional entities. The multi-sensor fusion-based power inspection robot equipment in this embodiment of the invention is described in detail from the perspective of hardware processing.
[0049] Reference Figure 3This invention also provides a power inspection robot based on multi-sensor fusion, which can be a server, and its internal structure can be as follows: Figure 3 As shown, the power inspection robot based on multi-sensor fusion includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed as a computer, provides computing and control capabilities. The memory of the power inspection robot includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the power inspection robot stores the data corresponding to this embodiment. The network interface of the power inspection robot is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0050] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the power inspection robot inspection equipment based on multi-sensor fusion applied thereto.
[0051] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the power inspection robot inspection method based on multi-sensor fusion.
[0052] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0053] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a power inspection robot based on multi-sensor fusion (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0054] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A power inspection robot inspection method based on multi-sensor fusion, characterized in that, The method comprises: Step S1: querying a weight mapping table according to an environmental magnetic field intensity and an environmental temperature to obtain a sensor fusion weight; Step S2: performing weighted fusion on multi-modal positioning data according to the sensor fusion weight to obtain a robot pose; Step S3: based on the robot pose, collecting a sound signal, a thermal image, and a vibration signal, performing space-time alignment on the sound signal, the thermal image, and the vibration signal to obtain synchronous data; Step S4: respectively extracting a voiceprint feature, a thermal image feature, and a vibration feature in the synchronous data to splice a three-dimensional feature vector; Step S5: calculating a fusion weight of each modality in the three-dimensional feature vector through an attention mechanism to perform weighted fusion to obtain a fault type. 2.The multi-sensor fusion based power inspection robot inspection method according to claim 1, wherein, The step S1 comprises: measuring a magnetic field intensity value of a current position through a magnetic field intensity sensor and measuring an environmental temperature value of the current position through a temperature sensor; comparing the magnetic field intensity value with a preset magnetic field threshold value and comparing the environmental temperature value with a preset temperature threshold value to obtain an environmental interference level; querying corresponding laser radar weight coefficients, inertial measurement unit weight coefficients, visual odometry weight coefficients, and magnetometer weight coefficients in a pre-stored weight mapping table according to the environmental interference level; taking the laser radar weight coefficients, the inertial measurement unit weight coefficients, the visual odometry weight coefficients, and the magnetometer weight coefficients as the sensor fusion weight. 3.The multi-sensor fusion based power inspection robot inspection method according to claim 2, wherein, The step S2 comprises: generating three-dimensional point cloud data by scanning a surrounding environment through a laser radar, collecting three-axis acceleration data and three-axis angular velocity data through an inertial measurement unit, extracting image feature points to calculate pose change data through a visual odometry, and measuring three-axis magnetic field data through a magnetometer; respectively converting the three-dimensional point cloud data, the three-axis acceleration data and three-axis angular velocity data, the pose change data, and the three-axis magnetic field data into corresponding pose parameters to obtain a laser radar pose, an inertial measurement unit pose, a visual odometry pose, and a magnetometer pose; multiplying the laser radar pose by the laser radar weight coefficient, multiplying the inertial measurement unit pose by the inertial measurement unit weight coefficient, multiplying the visual odometry pose by the visual odometry weight coefficient, and multiplying the magnetometer pose by the magnetometer weight coefficient to obtain weighted pose components; performing summation operation on the weighted pose components to obtain the robot pose. 4.The multi-sensor fusion based power inspection robot inspection method according to claim 1, wherein, The step S3 comprises: arriving at a preset inspection position according to the robot pose, collecting a sound signal of equipment operation through a voiceprint sensor, collecting a thermal image sequence of an equipment surface through an infrared thermal imager, and collecting a vibration signal of the equipment through a vibration sensor; extracting a collection timestamp of the sound signal, a collection timestamp of each frame of thermal image in the thermal image sequence, and a collection timestamp of the vibration signal, taking the collection timestamp of the sound signal as a reference time axis, mapping the timestamps of the thermal image sequence and the vibration signal to the reference time axis through linear interpolation to obtain a time-synchronized sound signal, thermal image sequence, and vibration signal; The sound source position coordinates, the heat source position coordinates and the vibration source position coordinates are mapped into a unified device coordinate system to obtain the synchronization data. The step S4 comprises: 5.The multi-sensor fusion based power inspection robot inspection method according to claim 1, wherein, The sound signals in the synchronization data are subjected to short-time Fourier transform to convert the time-domain sound signals into a frequency spectrum matrix, the frequency spectrum matrix is divided into a plurality of sub-frequency bands and the average amplitude, standard deviation, peak amplitude, peak frequency and spectrum entropy of each sub-frequency band are calculated to obtain the voiceprint features; The thermal image sequence in the synchronization data is calculated to obtain the thermal image features, including the global maximum temperature, the minimum temperature, the average temperature and the temperature standard deviation, the high-temperature regions are segmented by a region growing algorithm and the area, perimeter, circularity, regional maximum temperature and regional average temperature of each high-temperature region are extracted, and the temperature rise rate and temperature fluctuation amplitude of the high-temperature regions are calculated; The three-axis components of the vibration signals in the synchronization data are subjected to fast Fourier transform respectively, the first five frequency peaks with the maximum amplitude and the corresponding amplitudes in the amplitude spectrum of each axis are extracted to obtain the vibration features; The voiceprint features, the thermal image features and the vibration features are spliced in a preset order to obtain the three-dimensional feature vector. The step S5 comprises: 6.The multi-sensor fusion based power inspection robot inspection method according to claim 1, wherein, The three-dimensional feature vector is decomposed into a voiceprint feature component, a thermal image feature component and a vibration feature component, and is input into a voiceprint encoding network, a thermal image encoding network and a vibration encoding network respectively for encoding processing to obtain a voiceprint encoding vector, a thermal image encoding vector and a vibration encoding vector; The voiceprint encoding vector, the thermal image encoding vector and the vibration encoding vector are spliced to form a multi-modal encoding matrix, and the multi-modal encoding matrix is mapped into a query matrix, a key matrix and a value matrix respectively through linear transformation; The query matrix is subjected to matrix multiplication operation with the transpose of the key matrix, the attention score matrix is calculated by the softmax function after normalization processing of the operation result, and the fusion feature vector is obtained by matrix multiplication operation of the attention score matrix and the value matrix; The fusion feature vector is input into a fault classification network for classification processing, and the probability distribution of each fault type is output, and the fault type with the maximum probability is selected as the fault type. The fusion feature vector is input into a fault classification network for classification processing, and the probability distribution of each fault type is output, and the fault type with the maximum probability is selected as the fault type, comprising:
7. The multi-sensor fusion based power inspection robot inspection method according to claim 6, wherein, The fusion feature vector is input into a first full connection layer for linear transformation and activation processing to obtain an intermediate feature vector; The intermediate feature vector is input into a second full connection layer for linear transformation to obtain a power equipment fault score vector, and the dimension number of the power equipment fault score vector is the same as the number of preset power equipment fault types. The power equipment failure score vector is subjected to softmax normalization, and score values in each dimension are converted into probability values to obtain a power equipment failure probability distribution vector; Each probability value in the power equipment failure probability distribution vector is traversed, and a dimension index with the largest probability value is selected. According to a mapping relationship between the dimension index and a preset power equipment failure type, the failure type is obtained. The preset power equipment failure type includes at least one of bearing wear, equipment imbalance, poor centering, mechanical looseness, or oil film whirling.
8. A multi-sensor fusion based power inspection robot inspection system, characterized in that, The power inspection robot inspection method based on multi-sensor fusion according to any one of claims 1-7 is implemented, and the power inspection robot inspection system based on multi-sensor fusion comprises: The query module is configured to query a weight mapping table according to the environmental magnetic field intensity and the environmental temperature to obtain a sensor fusion weight. The weighting module is configured to perform weighted fusion on the multi-modal positioning data according to the sensor fusion weight to obtain a robot pose. The alignment module is configured to collect sound signals, thermal images, and vibration signals based on the robot pose, perform time-space alignment on the sound signals, the thermal images, and the vibration signals to obtain synchronized data. The splicing module is configured to extract a voiceprint feature, a thermal image feature, and a vibration feature in the synchronized data respectively, and splice to obtain a three-dimensional feature vector. The computing module is configured to calculate a fusion weight of each modality in the three-dimensional feature vector through an attention mechanism, perform weighted fusion, and obtain a failure type.
9. A multi-sensor fusion based power inspection robot inspection device, characterized in that, The computer program is run on the processor to implement the power inspection robot inspection method based on multi-sensor fusion according to any one of claims 1-7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is run on the processor to implement the power inspection robot inspection method based on multi-sensor fusion according to any one of claims 1-7.
Citation Information
Cited By
Intelligent inspection system for electrical power distribution cabinet in constructional engineering
CN121899553A