A power inspection control method and device, a terminal device and a storage medium
By preprocessing, feature alignment, and fusion of multimodal data in power line inspection, and combining this with a deep learning model to generate decision instructions, the problem of the inability to integrate multimodal data has been solved, thus improving the accuracy and stability of power line inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ELECTRIC POWER RES INST OF GUANGDONG POWER GRID CO LTD
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-21
Smart Images

Figure CN122437237A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of inspection technology, and in particular to a power inspection control method, device, terminal equipment and storage medium. Background Technology
[0002] The stable operation of power systems is crucial for the normal functioning of the social economy and the protection of people's lives. Power system inspection, as a key link in ensuring the safe and reliable operation of the power system, primarily involves regular or real-time inspections of power equipment and lines to promptly detect equipment faults, defects, and potential safety hazards, such as overheating, loosening, and damage to equipment, broken strands in lines, and foreign object attachments. This allows for timely repairs and handling, preventing the escalation of faults and power outages, and ensuring the continuity and stability of power supply. With the continuous development of robotics technology, power inspection robots are increasingly being applied in the field of power inspection, overcoming the limitations of manual inspection to some extent. Power inspection robots can be equipped with various sensors, such as cameras, infrared thermal imagers, and lidar, enabling them to automatically collect multimodal data such as images, temperature, and point clouds from power equipment, achieving automated inspection of power facilities.
[0003] Existing power inspection and control methods typically process the collected multimodal data independently, which fails to fully integrate and utilize the information between the multimodal data, making it difficult to comprehensively reflect the operating status of power equipment and resulting in low accuracy of power inspection and control. Summary of the Invention
[0004] This invention provides a power inspection and control method, device, terminal equipment, and storage medium, which can solve the technical problem in the prior art that the collected multimodal data is processed independently, the information between the multimodal data cannot be fully integrated and utilized, and it is difficult to comprehensively reflect the operating status of power equipment, resulting in low accuracy of power inspection and control.
[0005] This invention provides a power inspection and control method, comprising: Acquire multimodal raw data in power inspection scenarios, preprocess the multimodal raw data to obtain a standardized multimodal dataset; Extract the single-modal features corresponding to each modality data in the standardized multimodal dataset, align the single-modal features based on the attention mechanism to obtain aligned single-modal features, calculate the association weight between each single-modal feature and the inspection task, and use a cross-modal attention fusion algorithm to fuse multiple single-modal features according to the association weight to generate a unified multimodal feature sequence. The multimodal unified feature sequence is input into a deep learning model, and the cross-modal attention mechanism and hierarchical encoding module in the deep learning model are used to generate decision instructions for the inspection task based on the multimodal unified feature sequence. Based on the decision instructions, the power inspection robot is controlled to perform corresponding power inspection actions.
[0006] Furthermore, the multimodal raw data includes image data, point cloud data, sound data, and pose data. The preprocessing of the multimodal raw data to obtain a standardized multimodal dataset includes: The image data is denoised using adaptive median filtering, and the denoised image data is then normalized and its resolution is unified to obtain single-modal image data. The point cloud data is filtered and statistically analyzed to remove outliers. The point cloud data with outliers removed is then subjected to coordinate system transformation to obtain single-modal point cloud data. Wavelet thresholding is used to denoise the sound data, and the denoised sound data is converted into Mel spectrum to obtain single-mode sound data; The attitude data is corrected by Kalman filtering, and the corrected attitude data is then normalized to obtain single-mode attitude data. The single-modal image data, single-modal point cloud data, single-modal sound data, and single-modal pose data are integrated and processed to obtain a standardized multimodal dataset.
[0007] Furthermore, the step of generating decision instructions for the inspection task based on the multimodal unified feature sequence using the cross-modal attention mechanism and hierarchical encoding module in the deep learning model includes: The intermodal attention weights of the multimodal unified feature sequence are calculated based on the cross-modal attention mechanism, and the cross-modal interaction features of the multimodal unified feature sequence are calculated based on the intermodal attention weights. A hierarchical coding module is used to extract multi-scale coding features based on the cross-modal interaction features; The multi-scale encoded features are mapped to the decision dimension through a fully connected layer, and the decision feature representation is output. The decision instructions for the inspection task are generated based on the decision characteristics.
[0008] Furthermore, the step of generating decision instructions for inspection tasks based on the decision feature representation includes: The decision feature representation is subjected to task adaptation processing to obtain motion control features; Perform forward and inverse kinematics operations on the motion control features to obtain the target values of the joint angles; The joint control command is calculated based on the target value of the joint angle. When the rationality verification of the joint control command is passed, a decision command is generated based on the joint control command and the drive parameters of the actuator.
[0009] Furthermore, based on the decision instructions, the power inspection robot is controlled to perform corresponding power inspection actions, including: The joint control instructions and actuator drive parameters in the decision instructions are converted into a hardware-compatible format to obtain a compatible format decision instructions. The hardware format decision instructions are parsed through the hardware control interface to obtain the corresponding joint control instructions and drive parameters; The joint control actuator drives the joint motor to rotate according to the joint control, and drives the actuator to run according to the drive parameters.
[0010] Furthermore, the calculation of the association weight between each single-modal feature and the inspection task includes: The feature selection method in machine learning is used to calculate the mutual information value between each single modal feature and the inspection task; The association weight between each single-modal feature and the inspection task is determined based on the mutual information value.
[0011] The present invention also provides a power inspection and control device, comprising: The preprocessing module is used to acquire multimodal raw data in the power inspection scenario, and preprocess the multimodal raw data to obtain a standardized multimodal dataset. The fusion processing module is used to extract the single-modal features corresponding to each modality data in the standardized multimodal dataset, perform alignment processing on the single-modal features based on the attention mechanism to obtain aligned single-modal features, calculate the association weight between each single-modal feature and the inspection task, and use a cross-modal attention fusion algorithm to fuse multiple single-modal features according to the association weight to generate a unified multimodal feature sequence. The decision instruction generation module is used to input the multimodal unified feature sequence into the deep learning model, and use the cross-modal attention mechanism and hierarchical encoding module in the deep learning model to generate decision instructions for the inspection task based on the multimodal unified feature sequence. The inspection control module is used to control the power inspection robot to perform corresponding power inspection actions based on the decision instructions.
[0012] The present invention also provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the power inspection control method as described above.
[0013] The present invention also provides a computer-readable storage medium, comprising: a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the power inspection control method as described above.
[0014] The following benefits can be obtained by implementing the present invention: This invention integrates multimodal data into a unified feature sequence through feature alignment and feature fusion processing. This enables information from different modalities to complement each other, allowing for the full integration and utilization of information between multimodal data. This comprehensively reflects the operating status of power equipment and significantly improves the accuracy of power inspection and control.
[0015] Furthermore, this invention utilizes the cross-modal attention mechanism and hierarchical coding module in deep learning models to perform in-depth analysis on the fused multimodal unified feature sequences, accurately generating decision instructions for inspection tasks, thereby effectively improving the stability and reliability of inspection control. Attached Figure Description
[0016] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a power inspection and control method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a power inspection and control device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0020] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0021] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0022] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0023] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0024] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0025] See Figure 1 To address the technical problem in existing technologies where multimodal data is processed independently, resulting in insufficient integration and utilization of information between multimodal data and a lack of comprehensive reflection of the operating status of power equipment, thus leading to low accuracy in power inspection and control, an embodiment of the present invention provides a power inspection and control method, comprising: S1. Obtain multimodal raw data in the power inspection scenario, preprocess the multimodal raw data to obtain a standardized multimodal dataset; In this embodiment of the invention, the power inspection scenario refers to the work scenario of inspecting and maintaining various equipment and lines in the power system. These work scenarios involve areas such as substations, transmission lines, and distribution equipment. Multimodal raw data includes different types of data, such as image data, point cloud data, sound data, and attitude data. This embodiment of the invention can preprocess the multimodal raw data to improve data quality, unify data formats, and reduce data noise, making it more suitable for subsequent analysis and modeling.
[0026] S2. Extract the single-modal features corresponding to each modality data in the standardized multimodal dataset, align the single-modal features based on the attention mechanism to obtain aligned single-modal features, calculate the association weight between each single-modal feature and the inspection task, and use a cross-modal attention fusion algorithm to fuse multiple single-modal features according to the association weight to generate a multimodal unified feature sequence. In this embodiment of the invention, the purpose of feature alignment is to match and correspond data features of different modalities semantically or spatially, so that they can be compared and fused in the same semantic space; feature fusion is to combine the aligned features of different modalities to generate a unified feature representation containing multiple modal information; the multimodal unified feature sequence is a sequence formed by arranging the fused multimodal features in a certain order.
[0027] In this embodiment of the invention, aligning single-modal features based on an attention mechanism may include: First, the input unimodal features undergo a linear transformation, mapping them to three different feature spaces: a query vector, a key vector, and a value vector. The query vector represents the target feature representation to be aligned, the key vector measures the similarity or matching degree with the query vector, and the value vector contains the original feature information to be aggregated. By calculating the dot product or scaled dot product between the query vector and the key vector, an attention score matrix is obtained, reflecting the correlation strength between different positions or features. Subsequently, the attention scores in the attention matrix are normalized using the Softmax function to obtain the attention weight distribution. These weights represent the importance of each source feature when constructing the aligned features. Finally, the attention weights and value vectors are weighted and summed to generate the aligned feature representation. This process allows the model to adaptively focus on the features most relevant to the current task while suppressing interference from irrelevant or noisy features. In this embodiment, features with low weights can be fine-tuned to correct intermodal feature biases and ensure semantic matching of features.
[0028] In this embodiment of the invention, a cross-modal attention fusion algorithm is used to fuse multiple single-modal features according to the association weights to generate a multimodal unified feature sequence, which may include: The association weights corresponding to each single modal feature are used as attention weighting coefficients to weight and enhance different single modal features to obtain weighted single modal features. The weighted single modal features are then aggregated and fused to obtain a multimodal unified feature sequence.
[0029] This invention extracts single-modal features to ensure that information from each modality is fully preserved. By calculating association weights, it can highlight features related to the inspection task. It uses a cross-modal attention fusion algorithm to generate a unified feature sequence, which can integrate the advantages of multiple modalities, making the features more representative and comprehensive. This allows the model to better understand complex inspection scenarios, improves its decision-making ability for inspection tasks, and enhances the system's adaptability to different inspection situations.
[0030] S3. Input the multimodal unified feature sequence into the deep learning model, and use the cross-modal attention mechanism and hierarchical encoding module in the deep learning model to generate the decision instructions for the inspection task based on the multimodal unified feature sequence. In this embodiment of the invention, a decision feature representation can be generated based on a multimodal unified feature sequence, and then a decision instruction can be further generated based on the decision feature representation. The deep learning model can be an end-to-end VLA model.
[0031] In this embodiment of the invention, the cross-modal attention mechanism is a technique that enables the model to automatically focus on important parts of the input data; the hierarchical encoding module is a module that performs hierarchical processing and encoding of the input multimodal unified feature sequence; the decision feature representation is a feature vector output by the deep learning model after processing by the cross-modal attention mechanism and the hierarchical encoding module, which is related to the robot's inspection task decision. This feature vector contains key information about how to perform the inspection task, such as where the robot should go, which equipment to inspect, and what inspection method to use.
[0032] In this embodiment of the invention, the robot motion control mapping module can generate joint control commands and actuator drive parameters based on decision feature representation.
[0033] The robot motion control mapping module converts decision feature representations into joint control commands and actuator drive parameters that the robot can understand and execute, thereby establishing a mapping relationship between decision features and actual robot actions. Based on the information in the decision features, it calculates the motion angles, speeds, and actuator operating parameters of each robot joint. Among the joint control commands and actuator drive parameters, the joint control commands are used to control the movement of the robot joints, typically including information such as the joint's motion angle, speed, and time. The actuator drive parameters are used to drive the robot's actuators, such as the opening and closing force of the robotic arm's gripper and the focal length adjustment parameters of the camera lens.
[0034] S4. Based on decision-making instructions, control the power inspection robot to perform corresponding power inspection actions.
[0035] The embodiments of the present invention integrate multimodal data into a unified feature sequence through feature alignment and feature fusion processing, which enables the information of different modal data to complement each other, and allows the information between multimodal data to be fully integrated and utilized, thereby comprehensively reflecting the operating status of power equipment and significantly improving the accuracy of power inspection and control.
[0036] Furthermore, this embodiment of the invention utilizes the cross-modal attention mechanism and hierarchical coding module in the deep learning model to perform in-depth analysis on the fused multimodal unified feature sequence, accurately generate decision instructions for the inspection task, thereby effectively improving the stability and reliability of the inspection control.
[0037] In this embodiment of the invention, the construction of the deep learning model in step S3 includes: S301. Obtain the historical multimodal unified feature sequence and determine the historical decision instructions of the historical multimodal unified feature sequence, and use the historical multimodal unified feature sequence as training samples. S302. Input the multimodal unified feature sequence of the training samples into the cross-modal attention mechanism module, calculate the intermodal attention weights, and generate cross-modal interaction features based on the intermodal attention weights; S303. A hierarchical coding module is used to extract multi-scale coding features based on cross-modal interaction features; S304. The multi-scale encoded features are mapped to the decision dimension through a fully connected layer, and the decision feature representation is output. S305. Generate a predictive decision instruction for the inspection task based on the decision feature representation.
[0038] S306. Calculate the loss function value between the prediction decision instruction and the label, and based on the loss function value, update the network parameters of the cross-modal attention mechanism, the hierarchical coding module and the fully connected layer through the backpropagation algorithm.
[0039] S307. Iterate through steps S302 to S306 until the loss function converges or the preset number of training rounds is reached, and a trained deep learning model is obtained.
[0040] In one embodiment, step S2, calculating the association weight between each single-modal feature and the inspection task, may further include: S231. Using the feature selection method in machine learning, calculate the mutual information value between each single modal feature and the inspection task; In this embodiment of the invention, the correlation between each single-modal feature and the inspection task can be calculated based on the mutual information method or the chi-square test method. For example, the mutual information value between the single-modal feature and the task label can be calculated using the mutual information method. The larger the value, the stronger the correlation between the modal feature and the inspection task, and its weight should be correspondingly higher. Simultaneously, the experience of experts in the field of power inspection can be combined to evaluate and score the degree of correlation between each modal feature and the inspection task. The expert scoring results are then combined with data-driven methods to determine the final correlation weight. Mutual information is used to measure the amount of shared information between two variables.
[0041] In this embodiment of the invention, a single unimodal feature can be used as the analysis unit. The target output of the inspection task is used as the category label, and a sample set corresponding to the feature values and task labels is constructed. First, the overall distribution of the inspection task labels is statistically analyzed. Based on the statistical results, the inherent uncertainty entropy of the inspection task is calculated. The entropy of the inspection task quantifies the degree of uncertainty inherent in the label itself. This degree reflects the unpredictability of the task result without relying on any feature. Under the condition of fixing the current unimodal feature value, the distribution of the task label is re-statistically analyzed, that is, the conditional entropy after knowing the unimodal feature is calculated. This conditional entropy is used to represent the remaining uncertainty of the task result after knowing the feature information. The mutual information value is obtained by subtracting the remaining conditional entropy after knowing the unimodal feature from the inherent uncertainty entropy of the task label. This value represents the amount of information provided by the unimodal feature to reduce the uncertainty of the inspection task. The larger the value, the stronger the correlation between the feature and the task and the higher the discriminative power. The above steps are repeated for each unimodal feature to obtain its corresponding mutual information value, which is used for subsequent feature importance ranking and feature selection.
[0042] S232. Determine the association weight between each single modal feature and the inspection task based on the mutual information value.
[0043] In this embodiment of the invention, all mutual information values can be normalized to obtain a normalized result. This normalized result is the association weight between the single-modal feature and the inspection task. The normalized result satisfies the condition that it is between 0 and 1, and the sum of all normalized data is 1. For example, if the mutual information value calculated for single-modal feature A is 2, the mutual information value calculated for single-modal feature B is 3, and the mutual information value calculated for single-modal feature C is 5, the weights corresponding to single-modal features A, B, and C after normalization are 0.2, 0.3, and 0.5, respectively.
[0044] In one embodiment, the multimodal raw data includes image data, point cloud data, sound data, and pose data. Step S1 involves preprocessing the multimodal raw data to obtain a standardized multimodal dataset, including: S11. Adaptive median filtering is used to denoise the image data. The denoised image data is then normalized and its resolution is unified to obtain single-modal image data. In this embodiment of the invention, single-modal image data is obtained through adaptive median filtering, normalization, and unified resolution. Adaptive median filtering can remove speckle noise, grain noise, etc. from the image. For example, a photo of equipment taken by an inspection robot (1920×1080) → noise removal → normalization → adjustment to single-modal image data of size 224×224.
[0045] S12. Filter and statistically analyze the point cloud data and remove outliers. Then, perform coordinate system transformation on the point cloud data with outliers removed to obtain single-modal point cloud data. In this embodiment of the invention, invalid points and sparse noise points in the point cloud can be filtered out first, and then statistical methods can be used to remove obviously abnormal outliers (equipment errors, distant noise points). Finally, all point clouds are unified under a standard coordinate system. The labeled coordinate system can be the robot coordinate system. For example, the pipeline point cloud scanned by LiDAR → remove drift noise points → transform from the radar coordinate system to the robot coordinate system.
[0046] S13. Denoise the sound data using wavelet thresholding and convert the denoised sound data into Mel spectrum to obtain single-mode sound data. In this embodiment of the invention, wavelet thresholding can be used to remove ambient noise and electromagnetic interference from the sound, and the denoised time-domain sound can be converted into a Mel spectrogram (which is more suitable for sound features recognized by the model). For example, standby recording → removing ambient noise → generating a 64×64 Mel spectrogram.
[0047] S14. Kalman filtering is used to correct the attitude data, and the corrected attitude data is normalized to obtain single-mode attitude data. In this embodiment of the invention, Kalman filtering can be used to correct the drift, jitter, and random noise of sensors such as IMUs, making the attitude angle more stable. Then, normalization is performed to unify the numerical range of different attitude parameters, resulting in smooth, accurate, and numerically consistent single-modal attitude data. For example, the robot tilt angle (including jitter noise) collected by the IMU → Kalman smoothing → normalized to [-1, 1].
[0048] S15. Integrate and process the single-modal image data, single-modal point cloud data, single-modal sound data, and single-modal pose data to obtain a standardized multimodal dataset.
[0049] In this embodiment of the invention, the four sets of data—image, point cloud, sound, and pose—can be aligned in time and space, their format, dimensions, and numerical range can be unified, and invalid samples can be removed to obtain a complete, standard, and directly usable multimodal dataset for model training.
[0050] This invention preprocesses various types of multi-source raw data, enabling the adoption of adaptive denoising methods for different data characteristics, effectively improving data quality. Furthermore, it can further verify data quality and remove anomalies to ensure data reliability, providing a reliable data foundation, and thus improving the reliability and accuracy of power inspection and control.
[0051] In a specific embodiment, the data corresponding to point clouds can be a set of three-dimensional spatial coordinates; the data corresponding to images can be a two-dimensional light intensity / color matrix; the data corresponding to posture is a one-dimensional temporal vector, neither in image nor text form; and the data corresponding to sound is sound-related data. Image data, point cloud data, sound data, and posture data are data collected based on different sensor types, such as image data collected by cameras, point cloud data collected by LiDAR, sound data collected by microphones, and posture data collected by inertial measurement units, which respectively correspond to modal data of four independent perception channels: vision, geometry, hearing, and motion.
[0052] In one embodiment, step S3, utilizing the cross-modal attention mechanism and hierarchical encoding module in the deep learning model, generates decision instructions for the inspection task based on the multimodal unified feature sequence, including: S31. Calculate the intermodal attention weights of the multimodal unified feature sequence based on the cross-modal attention mechanism, and calculate the cross-modal interaction features of the multimodal unified feature sequence based on the intermodal attention weights. In this embodiment of the invention, the multimodal unified feature sequence can be dimensionally adapted, and its dimensions and length can be converted and regularized according to the input requirements of the deep learning model. Positional encoding can be added to obtain the model input features.
[0053] In this embodiment of the invention, input features can be passed to a cross-modal attention module to calculate inter-modal attention weights, strengthen the association of inspection-related features, suppress redundancy, and obtain cross-modal interaction features; One approach is to use a self-attention mechanism similar to that in Transformers to calculate inter-modal attention weights. That is, assuming there are... Each modality has a feature matrix, and the features of each modality are represented as a vector. ,in It is the feature dimension; Define query matrix , bond Sum matrix The query vector is obtained through linear transformation. Key vector Sum value vector ; Calculate attention score Then, the attention scores are normalized into attention weights using the softmax function. , , Indicates the first The modality pair of the first Attention weights for each modality.
[0054] Based on the calculated attention weights Sum value vector Calculate modal interaction features : , This allows the features of each modality to incorporate information from other modalities, resulting in cross-modal interaction features that take into account the relationships between modalities.
[0055] S32. A hierarchical coding module is used to extract multi-scale coding features based on cross-modal interaction features; In this embodiment of the invention, a multi-level coding module is used to extract and integrate features from low and high dimensions through a multi-level Transformer encoder to obtain multi-scale coded features.
[0056] In this embodiment of the invention, low-dimensional feature extraction involves inputting cross-modal interaction features into the bottom layer of the hierarchical encoding module. The bottom layer encoder can employ a convolutional neural network (CNN) or a simple fully connected layer to perform preliminary feature extraction and dimensionality reduction on the input features, resulting in a low-dimensional feature representation. Multi-layer encoder processing involves sequentially inputting the low-dimensional features into multiple layer encoders. Each layer encoder can employ a Transformer encoder structure, using a self-attention mechanism and a feedforward neural network to further encode and abstract the features. In each layer, the features gradually transition from local information to global information, and the dimensionality may be adjusted according to design requirements. High-dimensional feature integration involves concatenating the feature representations from different layers after multi-layer encoder processing using a concatenation operation to obtain a feature vector containing multi-scale information.
[0057] S33. Map the multi-scale encoded features to the decision dimension through a fully connected layer to output the decision feature representation; In this embodiment of the invention, multi-scale encoded features can be mapped to the decision dimension through a fully connected layer, and fine-tuned by combining a loss function to output a decision feature representation that reflects key information about the defect location.
[0058] In this embodiment of the invention, the specific process of outputting the decision feature representation can be as follows: The multi-scale encoded features are input into the fully connected layer. During training, the cross-entropy loss function is used, and the mean squared error loss function is used for the regression task. The parameters of the fully connected layer are fine-tuned according to the gradient of the loss function through the backpropagation algorithm, so that the output decision feature representation can better reflect key information such as defect location. After mapping and fine-tuning by the fully connected layer, the final output feature vector is the decision feature representation that can reflect the key information of defect location.
[0059] S34. Generate decision instructions for inspection tasks based on decision characteristics.
[0060] This invention adapts features to meet model input requirements through dimensionality adaptation, adds positional encoding to preserve sequence information, strengthens the association of inspection-related features through a cross-modal attention module, suppresses redundant information, and uses a hierarchical encoding module to form multi-scale encoded features through multi-layer encoding and fine-tuning through fully connected layer mapping combined with a loss function. The output is a decision feature representation that reflects key information about the defect location, which helps the robot accurately understand the inspection task and provides a reliable basis for generating precise control commands, thereby effectively improving the accuracy and efficiency of power inspection.
[0061] In one embodiment, step S34, generating decision instructions for the inspection task based on the decision feature representation, includes: S341. Perform task adaptation processing on the decision feature representation to obtain motion control features; In this embodiment of the invention, the target task to be completed in this inspection can be extracted from the decision feature representation first, such as: going to a certain point, inspecting a certain component, avoiding obstacles, inspecting along a specified path, etc. According to the actual conditions of this inspection task, such as the spatial structure of the inspection scene, the distribution of obstacles and the limitations of the equipment's movement range, etc., an adaptive motion model is selected from the preset motion modes. For example, a linear motion model is selected when going to the target point, and a circular motion model is selected when performing surround detection. Based on the selected motion model, the decision feature representation is transformed into motion control features such as motion trajectory, motion speed, and attitude angle.
[0062] In this embodiment of the invention, a suitable motion model, such as linear motion, circular motion, and joint motion, can be selected according to the task requirements of the inspection task, such as the task environment and task constraints. Based on the model, the decision command is calculated to generate motion control features.
[0063] S342. Perform forward and inverse kinematics operations on the motion control features to obtain the target values of the joint angles; In this embodiment of the invention, the target value can also be optimized by combining the robot joint degree of freedom constraints to avoid exceeding the range of motion, thereby improving the reliability of inspection control.
[0064] S343. Calculate the joint control command based on the target value of the joint angle. When the rationality verification of the joint control command passes, generate the decision command based on the joint control command and the drive parameters of the actuator.
[0065] In this embodiment of the invention, the joint control command includes the joint rotation direction and the joint rotation speed. This embodiment of the invention can also simultaneously generate drive parameters based on actuator performance parameters, including drive current sampling and drive voltage parameters.
[0066] In this embodiment of the invention, the joint control commands and drive parameters can be validated for reasonableness, values exceeding the safety threshold can be eliminated, and finally compliant joint control execution and actuator drive parameters can be output to obtain the final decision command.
[0067] This invention provides a solution for obtaining motion control features by performing task adaptation processing on the decision feature representation, performing forward and inverse kinematics operations on the motion control features to obtain joint angle target values, and calculating joint control commands based on the joint angle target values. This enables the robot to accurately complete inspection actions based on the joint control commands and drive parameters, effectively improving the reliability and stability of inspection.
[0068] In one embodiment, step S4, based on decision instructions, controls the power inspection robot to perform corresponding power inspection actions, including: S41. Convert the joint control instructions and actuator drive parameters in the decision instructions into a hardware-compatible format to obtain a compatible format decision instructions; S42. Parse the hardware format decision instructions through the hardware control interface to obtain the corresponding joint control instructions and drive parameters; S43. Drive the joint motor to rotate according to the joint control, and drive the actuator to run according to the drive parameters.
[0069] In this embodiment of the invention, the joint pivot rotation and actuator operation form an inspection action, which may include robotic arm gripping and body movement, etc.
[0070] This invention parses hardware format decision instructions through a hardware control interface to obtain corresponding joint control instructions and drive parameters. It drives the joint motor to rotate according to the joint control instructions and drives the actuator to run according to the drive parameters. This ensures seamless integration between software control and hardware execution, enabling the robot to accurately complete inspection tasks according to the plan. This effectively improves the automation level and execution efficiency of power inspection, reduces human intervention and operational errors, and thus effectively improves the reliability and stability of inspection control.
[0071] In one embodiment, after controlling the power inspection robot to perform corresponding power inspection actions based on decision instructions in step S4, the control parameters of the power inspection robot can be adjusted, and the adjusted control parameters can be used to optimize the accuracy of subsequent inspection actions, including: The robot's posture sensors and force sensors collect real-time inspection action feedback data, which includes: actual joint angles, actuator output force, and robot position deviation.
[0072] Record the correspondence between feedback data and instructions; Compare the feedback data with the target value of the instruction, calculate the deviation, and adjust the control parameters using the PID algorithm; The control parameters will be adjusted to optimize the accuracy of subsequent inspection actions.
[0073] This invention uses attitude and force sensors to collect feedback data in real time, recording the correspondence with commands to gain a comprehensive understanding of the robot's actual operation. By comparing the feedback data with the command target value to calculate the deviation, the control parameters are adjusted using a PID algorithm, enabling rapid response and correction of deviations. The adjusted parameters are then used for subsequent inspections, continuously optimizing motion accuracy. The feedback data covers key information such as the actual joint angles and actuator output forces, providing a comprehensive basis for parameter optimization. This allows the robot to adapt to complex and ever-changing inspection environments, improving the stability and accuracy of robot inspections.
[0074] Implementing the embodiments of the present invention has the following beneficial effects: This invention integrates multimodal data into a unified feature sequence through feature alignment and feature fusion processing. This enables information from different modalities to complement each other, allowing for the full integration and utilization of information between multimodal data. This comprehensively reflects the operating status of power equipment and significantly improves the accuracy of power inspection and control.
[0075] Furthermore, this invention utilizes the cross-modal attention mechanism and hierarchical coding module in deep learning models to perform in-depth analysis on the fused multimodal unified feature sequences, accurately generating decision instructions for inspection tasks, thereby effectively improving the stability and reliability of inspection control.
[0076] like Figure 2 As shown, based on the above method embodiments, corresponding apparatus embodiments are provided; An embodiment of the present invention provides a power inspection and control device, comprising: Preprocessing module 10 is used to acquire multimodal raw data in power inspection scenarios, preprocess the multimodal raw data, and obtain a standardized multimodal dataset. In this embodiment of the invention, the power inspection scenario refers to the work scenario of inspecting and maintaining various equipment and lines in the power system. These work scenarios involve areas such as substations, transmission lines, and distribution equipment. Multimodal raw data includes different types of data, such as image data, point cloud data, sound data, and attitude data. This embodiment of the invention can preprocess the multimodal raw data to improve data quality, unify data formats, and reduce data noise, making it more suitable for subsequent analysis and modeling.
[0077] The fusion processing module 20 is used to extract the single-modal features corresponding to each modality data in the standardized multimodal dataset, perform alignment processing on the single-modal features based on the attention mechanism to obtain aligned single-modal features, calculate the association weight between each single-modal feature and the inspection task, and use a cross-modal attention fusion algorithm to fuse multiple single-modal features according to the association weight to generate a multimodal unified feature sequence. In this embodiment of the invention, the purpose of feature alignment is to match and correspond data features of different modalities semantically or spatially, so that they can be compared and fused in the same semantic space; feature fusion is to combine the aligned features of different modalities to generate a unified feature representation containing multiple modal information; the multimodal unified feature sequence is a sequence formed by arranging the fused multimodal features in a certain order.
[0078] In this embodiment of the invention, aligning single-modal features based on an attention mechanism may include: First, the input unimodal features undergo a linear transformation, mapping them to three different feature spaces: a query vector, a key vector, and a value vector. The query vector represents the target feature representation to be aligned, the key vector measures the similarity or matching degree with the query vector, and the value vector contains the original feature information to be aggregated. By calculating the dot product or scaled dot product between the query vector and the key vector, an attention score matrix is obtained, reflecting the correlation strength between different positions or features. Subsequently, the attention scores in the attention matrix are normalized using the Softmax function to obtain the attention weight distribution. These weights represent the importance of each source feature when constructing the aligned features. Finally, the attention weights and value vectors are weighted and summed to generate the aligned feature representation. This process allows the model to adaptively focus on the features most relevant to the current task while suppressing interference from irrelevant or noisy features. In this embodiment, features with low weights can be fine-tuned to correct intermodal feature biases and ensure semantic matching of features.
[0079] In this embodiment of the invention, a cross-modal attention fusion algorithm is used to fuse multiple single-modal features according to the association weights to generate a multimodal unified feature sequence, which may include: The association weights corresponding to each single modal feature are used as attention weighting coefficients to weight and enhance different single modal features to obtain weighted single modal features. The weighted single modal features are then aggregated and fused to obtain a multimodal unified feature sequence.
[0080] This invention extracts single-modal features to ensure that information from each modality is fully preserved. By calculating association weights, it can highlight features related to the inspection task. It uses a cross-modal attention fusion algorithm to generate a unified feature sequence, which can integrate the advantages of multiple modalities, making the features more representative and comprehensive. This allows the model to better understand complex inspection scenarios, improves its decision-making ability for inspection tasks, and enhances the system's adaptability to different inspection situations.
[0081] The decision instruction generation module 30 is used to input the multimodal unified feature sequence into the deep learning model, and use the cross-modal attention mechanism and hierarchical encoding module in the deep learning model to generate decision instructions for the inspection task based on the multimodal unified feature sequence. In this embodiment of the invention, a decision feature representation can be generated based on a multimodal unified feature sequence, and then a decision instruction can be further generated based on the decision feature representation. The deep learning model can be an end-to-end VLA model.
[0082] In this embodiment of the invention, the cross-modal attention mechanism is a technique that enables the model to automatically focus on important parts of the input data; the hierarchical encoding module is a module that performs hierarchical processing and encoding of the input multimodal unified feature sequence; the decision feature representation is a feature vector output by the deep learning model after processing by the cross-modal attention mechanism and the hierarchical encoding module, which is related to the robot's inspection task decision. This feature vector contains key information about how to perform the inspection task, such as where the robot should go, which equipment to inspect, and what inspection method to use.
[0083] In this embodiment of the invention, the robot motion control mapping module can generate joint control commands and actuator drive parameters based on decision feature representation.
[0084] The robot motion control mapping module converts decision feature representations into joint control commands and actuator drive parameters that the robot can understand and execute, thereby establishing a mapping relationship between decision features and actual robot actions. Based on the information in the decision features, it calculates the motion angles, speeds, and actuator operating parameters of each robot joint. Among the joint control commands and actuator drive parameters, the joint control commands are used to control the movement of the robot joints, typically including information such as the joint's motion angle, speed, and time. The actuator drive parameters are used to drive the robot's actuators, such as the opening and closing force of the robotic arm's gripper and the focal length adjustment parameters of the camera lens.
[0085] The inspection control module 40 is used to control the power inspection robot to perform corresponding power inspection actions based on decision commands.
[0086] The embodiments of the present invention integrate multimodal data into a unified feature sequence through feature alignment and feature fusion processing, which enables the information of different modal data to complement each other, and allows the information between multimodal data to be fully integrated and utilized, thereby comprehensively reflecting the operating status of power equipment and significantly improving the accuracy of power inspection and control.
[0087] Furthermore, this embodiment of the invention utilizes the cross-modal attention mechanism and hierarchical coding module in the deep learning model to perform in-depth analysis on the fused multimodal unified feature sequence, accurately generate decision instructions for the inspection task, thereby effectively improving the stability and reliability of the inspection control.
[0088] In one embodiment, the raw multimodal data includes image data, point cloud data, sound data, and pose data. The raw multimodal data is preprocessed to obtain a standardized multimodal dataset, including: Adaptive median filtering is used to denoise the image data. The denoised image data is then normalized and its resolution is unified to obtain single-modal image data. The point cloud data is filtered and statistically analyzed to remove outliers. The point cloud data with outliers removed is then subjected to coordinate system transformation to obtain single-modal point cloud data. Wavelet thresholding is used to denoise the sound data, and the denoised sound data is converted into Mel spectrum to obtain single-mode sound data; Kalman filtering is used to correct the attitude data, and the corrected attitude data is then normalized to obtain single-mode attitude data. The single-modal image data, single-modal point cloud data, single-modal sound data, and single-modal pose data are integrated and processed to obtain a standardized multimodal dataset.
[0089] In one embodiment, the decision instructions for the inspection task are generated based on the multimodal unified feature sequence using the cross-modal attention mechanism and hierarchical encoding module in a deep learning model, including: The intermodal attention weights of the multimodal unified feature sequence are calculated based on the cross-modal attention mechanism, and the cross-modal interaction features of the multimodal unified feature sequence are calculated based on the intermodal attention weights. A hierarchical coding module is used to extract multi-scale coding features based on cross-modal interaction features; The multi-scale encoded features are mapped to the decision dimension through a fully connected layer, and the decision feature representation is output. Decision instructions for inspection tasks are generated based on decision characteristics.
[0090] In one embodiment, generating decision instructions for the inspection task based on decision feature representation includes: Task adaptation processing is performed on the decision feature representation to obtain motion control features; Perform forward and inverse kinematics operations on the motion control characteristics to obtain the target values of the joint angles; The joint control command is calculated based on the target value of the joint angle. When the rationality verification of the joint control command is passed, a decision command is generated based on the joint control command and the drive parameters of the actuator.
[0091] In one embodiment, based on decision instructions, the power inspection robot is controlled to perform corresponding power inspection actions, including: The joint control instructions and actuator drive parameters in the decision instructions are converted into a hardware-compatible format to obtain a hardware-compatible decision instruction. The hardware format decision instructions are parsed through the hardware control interface to obtain the corresponding joint control instructions and drive parameters; The joint control actuator drives the joint motor to rotate and drives the actuator to run according to the drive parameters.
[0092] In one embodiment, the association weight between each single-modal feature and the inspection task is calculated, including: The feature selection method in machine learning is used to calculate the mutual information value between each single modal feature and the inspection task; The association weight between each single modal feature and the inspection task is determined based on the mutual information value.
[0093] Implementing the embodiments of the present invention has the following beneficial effects: This invention integrates multimodal data into a unified feature sequence through feature alignment and feature fusion processing. This enables information from different modalities to complement each other, allowing for the full integration and utilization of information between multimodal data. This comprehensively reflects the operating status of power equipment and significantly improves the accuracy of power inspection and control.
[0094] Furthermore, this invention utilizes the cross-modal attention mechanism and hierarchical coding module in deep learning models to perform in-depth analysis on the fused multimodal unified feature sequences, accurately generating decision instructions for inspection tasks, thereby effectively improving the stability and reliability of inspection control.
[0095] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the power inspection and control method provided by any of the above-described method embodiments of the present invention.
[0096] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0097] Based on the above embodiments of the power inspection and control method, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the power inspection and control method of any embodiment of the present invention.
[0098] For example, in this embodiment, the computer program can be divided into one or more modules, one or more modules are stored in memory and executed by a processor to complete the present invention. One or more module elements can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a terminal device.
[0099] Terminal devices can be computing devices such as desktop computers, laptops, handheld computers, and cloud servers. Terminal devices may include, but are not limited to, processors and memory.
[0100] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device through various interfaces and lines.
[0101] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the power inspection control method of any of the above-described method embodiments of the present invention.
[0102] The modules / units integrated into the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0103] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A power line inspection and control method, characterized in that, include: Acquire multimodal raw data in power inspection scenarios, preprocess the multimodal raw data to obtain a standardized multimodal dataset; Extract the single-modal features corresponding to each modality data in the standardized multimodal dataset, align the single-modal features based on the attention mechanism to obtain aligned single-modal features, calculate the association weight between each single-modal feature and the inspection task, and use a cross-modal attention fusion algorithm to fuse multiple single-modal features according to the association weight to generate a unified multimodal feature sequence. The multimodal unified feature sequence is input into a deep learning model, and the cross-modal attention mechanism and hierarchical encoding module in the deep learning model are used to generate decision instructions for the inspection task based on the multimodal unified feature sequence. Based on the decision instructions, the power inspection robot is controlled to perform corresponding power inspection actions.
2. The power inspection and control method as described in claim 1, characterized in that, The multimodal raw data includes image data, point cloud data, sound data, and pose data. The preprocessing of the multimodal raw data to obtain a standardized multimodal dataset includes: The image data is denoised using adaptive median filtering, and the denoised image data is then normalized and its resolution is unified to obtain single-modal image data. The point cloud data is filtered and statistically analyzed to remove outliers. The point cloud data with outliers removed is then subjected to coordinate system transformation to obtain single-modal point cloud data. Wavelet thresholding is used to denoise the sound data, and the denoised sound data is converted into Mel spectrum to obtain single-mode sound data; The attitude data is corrected by Kalman filtering, and the corrected attitude data is then normalized to obtain single-mode attitude data. The single-modal image data, single-modal point cloud data, single-modal sound data, and single-modal pose data are integrated and processed to obtain a standardized multimodal dataset.
3. The power inspection and control method as described in claim 1, characterized in that, The step of generating decision instructions for the inspection task based on the multimodal unified feature sequence, utilizing the cross-modal attention mechanism and hierarchical encoding module in the deep learning model, includes: The intermodal attention weights of the multimodal unified feature sequence are calculated based on the cross-modal attention mechanism, and the cross-modal interaction features of the multimodal unified feature sequence are calculated based on the intermodal attention weights. A hierarchical coding module is used to extract multi-scale coding features based on the cross-modal interaction features; The multi-scale encoded features are mapped to the decision dimension through a fully connected layer, and the decision feature representation is output. The decision instructions for the inspection task are generated based on the decision characteristics.
4. The power inspection and control method as described in claim 3, characterized in that, The step of generating the decision instruction for the inspection task based on the decision feature representation includes: The decision feature representation is subjected to task adaptation processing to obtain motion control features; Perform forward and inverse kinematics operations on the motion control features to obtain the target values of the joint angles; The joint control command is calculated based on the target value of the joint angle. When the rationality verification of the joint control command is passed, a decision command is generated based on the joint control command and the drive parameters of the actuator.
5. The power inspection and control method as described in claim 1, characterized in that, Based on the decision instructions, the power inspection robot is controlled to perform corresponding power inspection actions, including: The joint control instructions and actuator drive parameters in the decision instructions are converted into a hardware-compatible format to obtain a compatible format decision instructions. The hardware format decision instructions are parsed through the hardware control interface to obtain the corresponding joint control instructions and drive parameters; The joint control actuator drives the joint motor to rotate according to the joint control, and drives the actuator to run according to the drive parameters.
6. The power inspection and control method as described in claim 1, characterized in that, The calculation of the association weight between each single-modal feature and the inspection task includes: The feature selection method in machine learning is used to calculate the mutual information value between each single modal feature and the inspection task; The association weight between each single-modal feature and the inspection task is determined based on the mutual information value.
7. A power line inspection and control device, characterized in that, include: The preprocessing module is used to acquire multimodal raw data in the power inspection scenario, and preprocess the multimodal raw data to obtain a standardized multimodal dataset. The fusion processing module is used to extract the single-modal features corresponding to each modality data in the standardized multimodal dataset, perform alignment processing on the single-modal features based on the attention mechanism to obtain aligned single-modal features, calculate the association weight between each single-modal feature and the inspection task, and use a cross-modal attention fusion algorithm to fuse multiple single-modal features according to the association weight to generate a unified multimodal feature sequence. The decision instruction generation module is used to input the multimodal unified feature sequence into the deep learning model, and use the cross-modal attention mechanism and hierarchical encoding module in the deep learning model to generate decision instructions for the inspection task based on the multimodal unified feature sequence. The inspection control module is used to control the power inspection robot to perform corresponding power inspection actions based on the decision instructions.
8. The power inspection and control device as described in claim 7, characterized in that, The multimodal raw data includes image data, point cloud data, sound data, and pose data. The preprocessing of the multimodal raw data to obtain a standardized multimodal dataset includes: The image data is denoised using adaptive median filtering, and the denoised image data is then normalized and its resolution is unified to obtain single-modal image data. The point cloud data is filtered and statistically analyzed to remove outliers. The point cloud data with outliers removed is then subjected to coordinate system transformation to obtain single-modal point cloud data. Wavelet thresholding is used to denoise the sound data, and the denoised sound data is converted into Mel spectrum to obtain single-mode sound data; The attitude data is corrected by Kalman filtering, and the corrected attitude data is then normalized to obtain single-mode attitude data. The single-modal image data, single-modal point cloud data, single-modal sound data, and single-modal pose data are integrated and processed to obtain a standardized multimodal dataset.
9. A terminal device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the power inspection control method as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, include: A stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the power inspection control method as described in any one of claims 1-6.