Power grid monitoring bionic robot with multi-mode sensing and voice interaction functions

Through the bionic robot with multimodal perception and voice interaction functions, the existing power grid monitoring equipment has solved the problem of single perception capabilities and poor voice interaction, and achieved high-precision abnormality recognition and efficient human-machine collaboration, which has improved the intelligence and adaptability of power grid monitoring.

CN120412583APending Publication Date: 2025-08-01XIANGYANG POWER SUPPLY COMPANY OF STATE GRID HUBEI ELECTRIC POWER
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510764595.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing power grid monitoring equipment and robots have a single perception capability, making it difficult to integrate multi-source information for high-precision identification, lack of intelligent judgment and risk prediction in a multi-factor environment, and lack of a highly robust voice interaction mechanism, making it difficult to effectively communicate human-computer in a high-noise and high-interference environment.

Method used

A bionic robot with multimodal perception and voice interaction functions is designed. Through the acquisition unit, the sound, image, temperature and micro-vibration information are synchronized, and the space-time calibration and preprocessing are performed. The recognition unit performs cross-judgment, identify abnormal areas, and the construction unit judges the risk of multi-factor coupling, generates task description information. The analysis unit performs speech enhancement and tone change extraction in a high-noise environment, and the control unit generates motion control instructions and voice feedback.

Benefits of technology

It significantly improves the accuracy of identification of electrical, structural and environmental abnormalities, enhances the robustness of monitoring, realizes efficient human-machine collaboration, has the ability to adjust independently, and improves the intelligence and flexibility of task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412583A_ABST
    Figure CN120412583A_ABST
Patent Text Reader

Abstract

The invention relates to a power grid monitoring bionic robot with multi-mode sensing and voice interaction functions. The acquisition unit acquires sound, images, temperature and power grid equipment surface micro-vibration information, completes data preprocessing and time and space calibration, and generates a sensing data packet. And the identification unit performs cross judgment based on the abrupt change characteristics and correlation of different types of data, identifies electrical, structural or environmental anomalies, and generates abnormal region description information. The construction unit analyzes whether a multi-factor coupling risk exists in combination with a power grid environment type and a historical rule, and generates task description information including processing suggestions and development path prediction. The analysis unit supports speech enhancement and tone extraction in a high-interference environment, and extracts a user interaction intention in combination with task description information. And the control unit jointly judges a behavior correction strategy according to the task description information and the interaction instruction information, generates a corresponding motion control instruction and voice feedback, and realizes autonomous monitoring and interaction response in a complex power grid environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bionic robots, and particularly to a bionic robot for power grid monitoring with multi-modal perception and voice interaction functions. Background Art

[0002] In the prior art, the monitoring of the power grid operation state mainly relies on fixed cameras, infrared imaging devices, local sensing terminals or simplified inspection robots. These systems mostly focus on the single-modal perception ability, such as image recognition or temperature detection, and make a preliminary judgment through the background analysis system. In addition, some intelligent inspection robots already have the ability to execute tasks mobilely, but their perception means and interaction ability are still relatively limited, especially the understanding of the environment and the adaptability to tasks in complex power scenarios are insufficient.

[0003] The existing power grid monitoring devices or robots generally have the following problems: First, the perception ability is single, and it is difficult to fuse multi-source information such as images, sounds, temperatures and micro-vibrations to achieve high-precision recognition; Second, in the event of sudden abnormal events, it is impossible to make intelligent judgments and risk predictions based on multi-factor environments; Third, there is a lack of a highly robust voice interaction mechanism, and it is difficult to adapt to effective human-machine communication in high-noise and high-interference environments; Fourth, the behavior control system is weakly associated with the perception feedback, and it is impossible to achieve task chain correction and active adaptation.

[0004] Based on the above problems, the present invention proposes a bionic robot with multi-modal perception and voice interaction functions suitable for complex power grid monitoring scenarios. Summary of the Invention

[0005] The present application provides a bionic robot for power grid monitoring with multi-modal perception and voice interaction functions to improve the power grid monitoring accuracy and the human-machine cooperation efficiency.

[0006] The present application provides a bionic robot for power grid monitoring with multi-modal perception and voice interaction functions, including:

[0007] An acquisition unit, configured to synchronously acquire sound, image, temperature and micro-vibration information on the surface of power grid equipment, preprocess the acquired data and perform time-space calibration, and generate a perception data packet in a unified format;

[0008] An identification unit, configured to receive the perception data packet, make a cross-judgment based on the mutation characteristics and correlations between different data types, identify electrical anomalies, structural anomalies or environmental anomaly regions, and obtain anomaly region description information including the anomaly type, location and time;

[0009] A building unit, which is configured to receive the abnormal area description information, combine the environmental type and historical grid rules, determine whether the abnormality has a multi-factor coupling risk, and generate task description information containing processing suggestions and development path predictions.

[0010] An analysis unit, which is configured to receive user voice input, perform voice enhancement and tone change extraction in an interference environment including high noise, strong wind or low voltage, and combine the task description information to identify the user's instruction intention and obtain interaction instruction information.

[0011] A control unit, which is configured to receive the interaction instruction information and the task description information, jointly judge the behavior correction strategy, and generate motion control instructions and voice feedback content.

[0012] The beneficial effects of this application mainly include: (1) By synchronously collecting sound, image, temperature and micro-vibration information on the device surface, and performing unified spatio-temporal calibration and preprocessing, the comprehensive recognition ability of electrical abnormalities, structural abnormalities and environmental abnormalities can be significantly improved, and the accuracy and robustness of monitoring can be enhanced. (2) Combining historical grid rules and environmental types, a multi-factor coupling risk analysis mechanism is constructed, and development path prediction information is generated, enabling the robot to have the ability to predict abnormal trends in the future and improve the fault prevention and control level. (3) Through voice enhancement and tone recognition functions in interference environments such as high noise, strong wind or low air pressure, it is ensured that the robot can accurately understand user instructions in harsh power sites and achieve efficient human-robot collaboration. (4) By jointly inputting the user's interaction intention and task description information into the control unit to form a dynamically corrected behavior strategy, the robot has the ability to autonomously adjust during tasks such as inspection and assistance, significantly improving the intelligence and flexibility of task completion. Description of the Drawings

[0013] Figure 1 is a schematic diagram of a grid monitoring bionic robot with multi-modal perception and voice interaction functions provided by the first embodiment of this application. Detailed Embodiments

[0014] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this application. Therefore, this application is not limited by the specific embodiments disclosed below.

[0015] The first embodiment of this application provides a grid monitoring bionic robot with multi-modal perception and voice interaction functions. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of this application. The following combines Figure 1 to describe a grid monitoring bionic robot with multi-modal perception and voice interaction functions provided by the first embodiment of this application.

[0016] The system includes a collection unit 101, an identification unit 102, a construction unit 103, an analysis unit 104, and a control unit 105.

[0017] The collection unit 101 is used to synchronously collect sound, image, temperature, and micro-vibration information on the surface of power grid equipment, preprocess the collected data, and perform time and space calibration to generate perception data packets in a unified format.

[0018] The collection unit 101 is designed to synchronously collect and uniformly process various key physical information at the power grid operation site, aiming to provide a high-quality, spatio-temporally aligned data basis for subsequent anomaly identification and task decision-making. The collection unit 101 mainly includes four types of sensor components: an audio collection component, an image collection component, a temperature collection component, and a micro-vibration collection component, corresponding to the acquisition tasks of sound information, image information, temperature information, and micro-vibration information on the surface of power grid equipment. The above components are structurally integrated at multiple layout positions of the bionic robot body and are uniformly scheduled and time-sequenced controlled by a main control collection control circuit board to ensure a high degree of synchronization in the collection process.

[0019] In terms of sound information collection, the audio collection component consists of an omnidirectional high-sensitivity digital microphone array, which is arranged at the upper end of the robot body and is equipped with a wind noise shielding structure. The collected original sound signal is first filtered by a band-pass filter to remove frequency bands below 50 Hz and above 12 kHz, and then subjected to a short-time Fourier transform to extract its spectral envelope, pitch change, and band energy distribution characteristics for subsequent judgment of acoustic patterns related to electrical anomalies such as partial discharge, arc sound, and breakdown sound.

[0020] In terms of image information collection, the image collection component includes two types of visual devices: a high-dynamic-range visible light camera and a low-temperature-drift thermal imager, which are respectively arranged at the bionic head position and the side arm position of the robot. After the image information is collected, it is processed by an image enhancement module to improve the edge sharpness and contrast, and histogram equalization is used to correct the exposure deviation caused by voltage fluctuations. Subsequently, through template matching and feature point detection algorithms, significant markings are made on the states such as deformation on the surface of power grid equipment, pollution on the surface of insulators, and sag of conductors.

[0021] In terms of temperature information collection, a non-contact infrared temperature measurement unit is used, which is arranged above the rotation pan-tilt of the robot head. This unit scans the temperature distribution map of power grid equipment in real time and extracts multi-dimensional parameters such as temperature rise rate, hot spot stability, and temperature gradient through time-continuous window analysis, so as to reflect whether there are overheating, short-circuit, or overloading phenomena in electrical equipment.

[0022] In terms of micro-vibration information acquisition, the acquisition unit 101 arranges piezoelectric vibration sensors at the flexible probe parts where the robot body contacts the grid equipment, and monitors the subtle mechanical responses of the equipment housing and its connecting components in real time. The micro-vibration signals are processed by fast Fourier transform and envelope demodulation algorithms to obtain the mode characteristics of mechanical looseness, structural fatigue or low-frequency resonance.

[0023] After the above four types of perception data are acquired, the acquisition unit 101 calibrates all the data according to a unified timestamp, and constructs a three-dimensional time-space index model in combination with the position coordinate information acquired by the robot pose sensor. This model aligns multi-source data with millisecond-level time accuracy and uniformly encapsulates it into a structured perception data packet, including: the acquisition time, spatial position, original data, feature results, synchronization marker identification, etc. of each modality.

[0024] The acquisition unit 101 is equipped with a local cache mechanism and data integrity verification logic to ensure that data packets in a complex electromagnetic environment will not have time drift or loss phenomena. At the same time, this unit and the identification unit 102 perform data interaction through a high-speed bus interface or a wireless high-speed communication module to ensure that the perception data packet can be transmitted to the identification unit within the millisecond-level time delay range, providing a time efficiency guarantee for subsequent anomaly detection and task response.

[0025] In summary, the acquisition unit 101 can not only cover the main abnormal manifestation forms in the grid operation environment, but also effectively convert the original perception information into highly available data input through processing processes such as structured preprocessing, time-space alignment, and feature extraction, supporting the intelligent perception and response capabilities of the whole machine in a high-voltage power environment. The design of this unit ensures the stability, accuracy and execution efficiency of the robot when facing multi-source perception tasks.

[0026] The identification unit 102 is used to receive the perception data packet, and based on the mutation characteristics and correlations between different data types, make cross-judgments to identify electrical anomaly, structural anomaly or environmental anomaly areas, and obtain anomaly area description information including anomaly type, location and time.

[0027] The identification unit 102 is used to receive the perception data packet generated by the acquisition unit 101, and comprehensively analyze the image information, sound information, temperature information and micro-vibration information contained therein to identify electrical anomaly, structural anomaly or environmental anomaly areas that may exist at the grid site, and generate corresponding anomaly area description information. The core lies in using the correlation and difference between different perception modalities in time and space to make joint judgments, thereby improving the accuracy and scene adaptability of anomaly identification.

[0028] Inside the recognition unit 102, the perception data packet is first unpacked and data classified. The image information, sound information, temperature information, and micro-vibration information are aligned according to the timestamp and spatial coordinates to construct a multi-modal data window based on the current moment. The time span of this window can be set from 1 to 5 seconds, and the spatial range is dynamically adjusted according to the current pose information of the robot to ensure that the recognition process always covers the multi-source states within the same physical area.

[0029] In the processing of image information, the recognition unit 102 calls the edge enhancement result and region feature extraction algorithm to obtain visual signs of cracks, deformations, inclinations, or detachments on the device surface, and calculates the inter-frame difference with the historical image frames to identify the mutation region. If obvious pollution patches on the insulator or abnormal wire suspension height are found, the system will record the corresponding image frame, device position information, and mutation boundary map.

[0030] For sound information, the recognition unit 102 inputs the spectral envelope and short-time energy change sequence into the acoustic discrimination model to identify whether there are arc sounds, partial discharge sounds, or other typical electrical discharge sounds, and combines the background noise baseline to determine whether the sound source exceeds the safety threshold. If the sound change occurs at the same position as the image region, the system will enhance the abnormal determination level of this region.

[0031] The temperature information consists of infrared temperature measurement data. The recognition unit establishes a temperature rise trend curve for each detection point, calculates the temperature mutation rate, maximum temperature difference, and local hot spot distribution. If the temperature of a certain part of the device continues to rise within a short period of time and coincides with the region with structural defects in the image, the recognition unit will mark this region as high risk.

[0032] The micro-vibration information is processed through the frequency domain envelope feature, low-frequency resonance point, and impact pulse recognition algorithm to identify mechanical looseness, base fatigue, or abnormal vibration at the connection part. When such vibration characteristics appear together with the pulse characteristics of the image edge fracture and sound, the recognition unit combines them into the manifestation of structural abnormality.

[0033] The key of the recognition unit 102 lies in fusing the above modalities, constructing an abnormal mutation relationship graph through cross criteria, and classifying each abnormal node in the graph. If a node simultaneously meets the multi-modal mutation characteristics, such as image crack + abnormal sound source + obvious temperature rise + micro-vibration deviating from the reference value, the system classifies this node as a "high-confidence abnormal point" and further expands it into an abnormal region. The description information of the abnormal region includes fields such as abnormal type (electrical / structural / environmental), central coordinate position (calculated by multi-source fusion), first discovery time, mutation degree score, etc.

[0034] For example, during the inspection of the power grid, the image information received by the recognition unit 102 shows an obvious structural edge break between the image frames before and after time t1 on the outer shell of a certain device. The system calculates the length of the broken edge and the pixel gradient change rate, and finds that the change amplitude is greater than 30% of the set image mutation threshold, and the area of the broken area exceeds 5% of the reference standard device surface baseline value, thus determining that there is a deformation anomaly at the image level.

[0035] Meanwhile, the recognition unit 102 synchronously receives sound information. After short-time Fourier transform processing, in the audio segment corresponding to before and after t1, a group of high-frequency pulse signals are detected, with the frequency concentrated in the range of 1.8 to 2.2 kHz, the peak amplitude much higher than the average value of the surrounding background, and the signal having irregular sharp fluctuations. The system compares this frequency band with the on-site background noise model and finds that it has the characteristic spectrum of typical arc breakdown sound, so it is marked as a potential discharge anomaly at the sound level.

[0036] At this time, within the interval from t1 to t1+Δt of the infrared thermal imager, the temperature of the abnormal part in the corresponding image rises from 38°C to 59°C, and the temperature rise rate reaches 10°C per minute, and the hot spot area is limited within 50 pixels around the broken area identified in the image. The system determines that there is an obvious abnormal temperature rise phenomenon in this local area, and combines with the normal operating condition model of the device to rule out the possibility of normal thermal changes caused by load fluctuations.

[0037] Furthermore, the micro-vibration information shows that the vibration amplitude of the device outer shell increases in the range of 10 Hz–40 Hz during the same period. Among them, the first-order main vibration peak drifts from the original 14 Hz to 17 Hz, and the deviation degree exceeds 10%. At the same time, the envelope curve shows periodic asymmetric jitter. Such signal characteristics correspond to the low-frequency resonance characteristics after equipment loosening or shell structure fracture.

[0038] The recognition unit 102 aligns the above image information, sound information, temperature information, and micro-vibration information in time, and confirms that their changes all occur within ±2 seconds of time t1, and spatially they all focus on the right insulation arm area of device number X101. Based on this, the system confirms that there are significant mutation characteristics of cross-modal consistency at this position, and the abnormal signs confirm each other.

[0039] Finally, the recognition unit 102 calibrates this area as "electrical anomaly", records the abnormal type as "electrical discharge + structural damage", the abnormal position as "right insulation arm of device X101", the abnormal occurrence time as "t1", the image crack mutation rate as 32%, the sound peak deviation as +18 dB, the temperature rise rate as 10°C per minute, the vibration frequency deviation as +3 Hz, and assigns a "high confidence" flag to this abnormal record for subsequent task planning and user voice interaction processing.

[0040] In addition, the recognition unit 102 also has a dynamic tracking mechanism, which continuously updates the previously discovered abnormal areas on the robot's patrol path. For situations such as the expansion of the abnormal range, the enhancement of signals, or the evolution of the abnormal type, they will be recorded in the description information and pushed to the construction unit 103 for task analysis.

[0041] The entire recognition process runs in real time on the edge computing chip or local embedded platform, and the response time is controlled within 500 milliseconds to ensure that the robot can instantly identify, instantly annotate, and instantly respond during the patrol process. The system also supports data synchronization with the dispatching center for manual verification and access to the operation and maintenance system, further enhancing reliability.

[0042] Through the above mechanism, the recognition unit 102 can not only independently complete abnormal recognition, but also achieve multi-modal cross-enhanced judgment, greatly improving the autonomous recognition ability and perception intelligence level of the power grid monitoring bionic robot in complex power fields.

[0043] Furthermore, the recognition unit is specifically used for:

[0044] Perform synchronous time registration on the image information, temperature information, and micro-vibration information contained in the perception data packets generated by the acquisition unit, extract the mutation time points of each perception modality, and construct a modality mutation alignment sequence for marking potential high-risk moments when multiple modalities mutate simultaneously within the same time window;

[0045] Based on the modality mutation alignment sequence, perform spatial overlap comparison on the pixel color difference mutation regions of the image information, the local temperature rise slope mutation regions of the temperature information, and the main frequency offset mutation regions of the micro-vibration information, calculate the mutation intersection degree of each modality in the same region, and generate an abnormal overlap score value for characterizing the consistency degree of multi-modal mutations in this region;

[0046] Fuse the abnormal overlap score value with the mutation confidence levels of each modality to construct a heterogeneous modality confidence voting structure. The confidence voting structure is used to trigger the judgment that the spatial region is an abnormal region to be annotated when the confidence levels of any two of the three modalities exceed the preset threshold;

[0047] Combined with a preset modality conflict counterexample set, perform reverse contradiction mode comparison on the abnormal region to be annotated. If its modality combination characteristics highly match the false abnormal types recorded in history, cancel the annotation of this abnormal region; otherwise, officially mark it as an abnormal region and generate abnormal region description information including the abnormal type, location, and time for use by the construction unit.

[0048] In the grid monitoring bionic robot with multimodal perception and speech interaction functions according to the present invention, the recognition unit not only undertakes the basic function of receiving and processing the perception data packets generated by the acquisition unit, but also further realizes the high-confidence recognition of abnormal areas through cross-mutation collaborative analysis between modalities.

[0049] During the robot's patrol task execution, the acquisition unit continuously collects image information, temperature information, and micro-vibration information at a certain frequency, and uniformly packages the above data into structured perception data packets. The timestamp, three-dimensional position information, modality source identifier, and original data feature vector contained in the perception data packet constitute the basic input for spatio-temporal multimodal data processing. After receiving the perception data packet, the first operation performed by the recognition unit is the synchronous time registration operation of the modality data. This registration is based on the task time axis shared by all modalities, and uses the mutation points in each type of modality data as anchor points to find the closest change moments in the image, temperature, and micro-vibration data. For example, when there is an edge color difference mutation in the image frame at time t0, the system will search in the temperature curve and vibration spectrogram to see if there are corresponding temperature rise inflection points or frequency jump points within a 50ms time window before and after. Once two or more modality mutation events fall within this time window, the system will mark this time period as the "potential high-risk synchronous moment" in the modality mutation alignment sequence.

[0050] After the modality mutation alignment sequence is determined, the recognition unit further performs a joint analysis on the spatial area where the mutation occurs. For the image information, the system uses the regional color difference change rate and the image gradient change direction as the spatial markers for cracks or structural abnormalities; for the temperature information, the system calculates the local temperature curve slope and marks the abnormal temperature rise boundary in the image coordinate system; for the micro-vibration information, the system marks the vibration source position through the frequency domain peak shift and the low-frequency resonance bandwidth expansion. These spatial areas from different modalities will be used to calculate the modality spatial overlap degree after coordinate unified projection and boundary standardization processing. The recognition unit uses the ratio of the spatial intersection area (i.e., the area of the region where at least two regions of the three modalities overlap divided by the union area) as a quantification index to generate an abnormal overlap score value. This score value is used to characterize whether there is an enhanced consistency between the mutations in the modalities, that is, whether multiple modalities jointly point to an abnormal area in terms of spatial position.

[0051] For example, during the daily patrol task execution of the grid monitoring bionic robot, the robot patrols to the switchgear group of an outdoor substation. The acquisition unit real-time collects image information, temperature information, and micro-vibration information, and packages them into a unified format of perception data packet. The recognition unit starts to process this data packet, recognizes that there is a suspected abnormal state at the current moment, and then enters the modality mutation collaborative recognition process.

[0052] First, the recognition unit detects a crack boundary on the surface of an insulating arm from the image information, and the edge gradient mutation is clear. After processing, the system marks the upper left corner of the mutated area of the image as the point (120, 340) in the image coordinate system, and the lower right corner as the point (170, 385). The unit of the image coordinate system is pixels. Therefore, the length of the mutated area of the image is 170 - 120 = 50 pixels, the width is 385 - 340 = 45 pixels, and the area is 50 × 45 = 2250 pixels².

[0053] At the same time, the thermal imaging module collects the temperature information of this area and detects a non-uniform hot spot area within the same spatial range. The system uses the temperature slope mutation detection algorithm to extract the boundary of the local temperature rise area. The upper left corner of this hot spot area is (125, 345), and the lower right corner is (175, 388). It is calculated that the length of the temperature mutation area is 175 - 125 = 50 pixels, the width is 388 - 345 = 43 pixels, so the area is 50 × 43 = 2150 pixels².

[0054] Furthermore, when the robot approaches the switchgear cabinet, the micro-vibration sensor detects an abnormal frequency drift. The system projects the abnormal vibration response onto the image plane coordinates by aligning with the robot's posture and the environmental depth map, forming a micro-vibration mutation area. The upper left corner of this area is (130, 350), and the lower right corner is (172, 392). The corresponding length is 172 - 130 = 42 pixels, and the width is 392 - 350 = 42 pixels. Therefore, the area is 42 × 42 = 1764 pixels².

[0055] To determine whether these three types of mutations are spatially consistent, the recognition unit further calculates the overlapping area of these three regions. The recognition unit uses the standard two-dimensional rectangle intersection calculation method to find the common part of the three rectangle intersection regions on the image plane, that is:

[0056] The horizontal overlapping interval is: max(120, 125, 130) = 130 to min(170, 175, 172) = 170, and the length is 170 - 130 = 40 pixels;

[0057] The vertical overlapping interval is: max(340, 345, 350) = 350 to min(385, 388, 392) = 385, and the height is 385 - 350 = 35 pixels;

[0058] Therefore, the spatial intersection of these three mutated regions is a rectangular region with a length of 40 pixels and a width of 35 pixels, and the area is 40 × 35 = 1400 pixels², which is the spatial overlapping area of the three-modal mutation region.

[0059] Next, the recognition unit calculates the abnormal overlap score value according to the definition. This value is used to quantify the degree of spatial consistency of the three types of modal mutation regions. The calculation formula is as follows:

[0060] Abnormal overlap score value = Area of the overlapping region ÷ Average area of the modal mutation regions

[0061] Among them, the area of the overlapping region is 1400 pixels²;

[0062] Average area of the modal mutation regions = (2250 + 2150 + 1764) / 3 = 6164 / 3 ≈ 2054.67 pixels²;

[0063] Substitute into the formula:

[0064] Abnormal overlap score value = 1400 ÷ 2054.67 ≈ 0.6814

[0065] This score value is approximately 0.68, which is significantly higher than the significant overlap threshold set in this system (for example, the default threshold is 0.6). Therefore, this region will be determined as a significantly overlapping mutation region, with strong multi-modal consistency, and enter the next confidence voting mechanism for preliminary judgment of the abnormal region.

[0066] Based on this, the recognition unit will weight and fuse this abnormal overlap score value with the confidence levels of each modal mutation, construct a heterogeneous modal confidence voting structure, and make a trigger judgment in combination with the set threshold. If the confidence levels of any two types of modalities are higher than their respective preset thresholds and the abnormal overlap score value exceeds the set threshold, it will enter the formal "abnormal region to be labeled" process; otherwise, it will be recorded as a low-confidence candidate in the buffer.

[0067] Through the above calculation process, it can be clearly seen that the abnormal overlap score value is not only based on the precise alignment of the spatial positions of each type of modality and the evaluation of the rectangular area, but also comprehensively considers the degree of spatial consistency among the three. Its core logic is based on the spatially coincident relationship and area ratio quantification with clear physical meanings, has high engineering implementation feasibility, and can quickly and reliably complete the determination of multi-modal collaborative mutation consistency without relying on complex deep learning structures.

[0068] However, in the actual power operation scenario, there are still some modal combinations that exhibit high mutation overlap under non-abnormal conditions. For example, the wind blowing the wire may cause false abnormal conditions such as deformation, temperature increase, or micro-vibration disturbance in the image. To solve this problem, the recognition unit further introduces a "modal conflict counterexample set" mechanism for reverse judgment. This counterexample set is pre-constructed by the system during actual operation or simulation training, and includes several sets of modal combination templates under known non-abnormal states. These templates include, but are not limited to, false abnormal feature structures such as "image blur + temperature jitter + micro-vibration stability", "image edge vibration + hot spot movement + no main peak in the spectrum", etc. When the recognition unit identifies an abnormal area to be labeled, it will compare the modal combination features of this area with the templates in the modal conflict counterexample set. The comparison method uses the weighted cosine similarity between modal feature vectors and the time synchronization error tolerance, and a matching threshold is set. If the comparison result shows that the modal combination of this area is highly consistent with a certain counterexample template (such as the similarity exceeds 0.85 and the time drift is less than 30 ms), the system determines that this area may be a "non-genuine abnormality" and cancels its abnormal label to reduce false alarms.

[0069] On the contrary, if the matching degree of the modal combination of the labeled area does not reach the exclusion standard with any of the conflict counterexample templates, the recognition unit will officially identify it as an abnormal area and generate standardized abnormal area description information. This information includes fields such as abnormal type (electrical abnormality, structural abnormality, or environmental abnormality), abnormal spatial location (three-dimensional coordinates and image pixel area), first recognition time (high-confidence mutation alignment moment), mutation modal composition and confidence structure, and spatial overlap score. This abnormal area description information will be directly output for use by the construction unit to further determine whether this abnormality has a multi-factor coupling risk and decide the task behavior strategy.

[0070] Throughout the process, various calculation methods and parameter configurations used by the recognition unit can be dynamically adjusted according to the task configuration file. The modal mutation synchronization window width, overlap score threshold, mutation confidence threshold, and counterexample template matching rule are all adjustable variables, allowing the system to be adaptively deployed according to the actual power equipment type, environmental characteristics, or operation and maintenance strategy. In addition, the recognition unit is also designed with an abnormal evolution cache mechanism. For the identified abnormal areas to be labeled, it will continuously monitor whether the modal mutation expands, strengthens, or transfers in subsequent multiple sampling periods to form a dynamic time series graph of the abnormal area, providing the original input for the subsequent trend prediction of the construction unit.

[0071] In summary, the recognition unit of the present invention has the ability to jointly recognize multimodal perception mutations in the power grid operation scenario. Through the four-level linkage mechanism of "modal mutation time series alignment", "spatial overlap quantitative analysis", "heterogeneous modal voting mechanism" and "false alarm suppression by counterexample matching", it not only improves the accuracy, anti-interference ability and scenario adaptability of anomaly recognition, but also establishes a semantic input interface for the construction unit through structured output, significantly optimizing the intelligent judgment link of the entire bionic robot system.

[0072] The construction unit 103 is configured to receive the abnormal area description information, and in combination with the environmental type and historical power grid rules where it is located, judge whether the anomaly has a multi-factor coupling risk, and generate task description information containing processing suggestions and development path predictions.

[0073] The construction unit 103 is used to receive the abnormal area description information output by the recognition unit 102, and on this basis, in combination with the operating environment characteristics of the power grid equipment and historical operation and maintenance rules, comprehensively judge whether the current anomaly involves a multi-factor coupling risk, and further generate task description information with processing suggestions and development path prediction content for use in subsequent voice parsing and behavior control links.

[0074] In a specific implementation, the construction unit 103 first parses the abnormal area description information, extracts elements such as the abnormal type, abnormal occurrence location, initial trigger time, mutation degree score, and associated modal quantity. The system retrieves basic information such as the standard type, operating level, and load level of the device in the structure configuration table according to the device number corresponding to the abnormal location, and combines the geographical tags provided by the GIS system to obtain the environmental type of the area where the device is located, including background parameters such as altitude, climate zone, humidity level, salt fog level, and electromagnetic interference level.

[0075] Based on the above input, the building unit 103 further analyzes whether it constitutes a "multi-factor coupling risk". The so-called multi-factor coupling risk refers to the situation where two or more abnormal factors that overlap in time, are associated in space, and may form a synergistic effect in mechanism act on the same device or local area simultaneously, resulting in a rapid deterioration of the device state or inducing a systemic failure. This type of risk does not depend on the severity of a single-mode abnormality, but rather on the premise that mild mutations occur simultaneously in multiple sensing modes, and an amplification effect is generated by their interactive influence. The building unit 103 determines this risk mainly based on three aspects: First, the mutation characteristics of two or more modes among the image information, temperature information, sound information, or micro-vibration information simultaneously exist in the report of the identification unit 102, the spatial coincidence degree of the mutation positions is not less than 80%, and the timestamp difference does not exceed 1 second; Second, the target device is in a known high-interference and high-complexity operating background in the power grid environment, such as conditions like heavy load, humidity, high temperature difference, high corrosion, and proximity to a vibration source; Third, there are fault cases similar to this combined condition in the historical rule library, and there is a relatively high probability of accident induction in past cases.

[0076] For example, if the identification unit 102 detects a slight edge depression in the image of a certain switch cabinet, the temperature data shows a rapid rise of more than 10°C within 3 minutes at this position, and the vibration sensor records a new low-frequency fluctuation signal below 10 Hz, the system will determine that the device is in a "structure + heat + vibration" multi-mode interference state; at the same time, this device is located in a high-salt fog area along the coast, and the operation and maintenance records show that bracket cracking and insulation breakdown have occurred under similar conditions in the past, then the building unit 103 will mark this position as a "medium-strong multi-factor coupling risk", and record the risk level as level three and the handling suggestion as "manual on-site review is required and planned power-off disassembly inspection" in the task description information.

[0077] The building unit 103 also calls an empirical model constructed based on the similar fault evolution path to make a preliminary prediction on the time development trend of the current abnormal state. For example, in the above case, if the historical data shows that similar cracks develop at an average expansion rate of 0.5 millimeters per hour under the background of salt spray corrosion, and the temperature rise curve conforms to the typical characteristics in the early stage of thermal breakdown, the building unit will generate an evolution prompt of "it is expected that the crack will expand to the warning boundary within 24 hours", and suggest that the dispatching center handle it within a limited time.

[0078] The task description information finally forms a structured output, and the fields include: abnormal number, device number, abnormal type, confidence level, multi-mode coincidence score, list of environmental risk factors, handling suggestion, expected evolution trend, and recommended handling time window, etc. If the abnormal confidence level is insufficient or the modal sources are less than three categories, the building unit will mark the task as "low confidence level, keep observing", and submit it for background retention.

[0079] Through the above processing flow, the construction unit 103 realizes the transformation from underlying multi-modal perception data to high-level risk semantics, enhances the robot's ability to judge potential coupling faults and assign task priorities in complex power grid scenarios, provides clear, accurate, and executable task semantic support for the parsing unit 104 and the control unit 105, and significantly improves the intelligence, adaptability, and forward-looking response ability of the system.

[0080] The following is an example to illustrate the complete working process of the construction unit 103 in the present invention, covering the whole process from receiving the information output by the recognition unit, matching the power grid rules, judging the multi-factor coupling risk, to generating task description information.

[0081] During the daily inspection task of a 220 kV substation in the mountainous outdoors, when the bionic power grid monitoring robot patrolled to the side of a reactor numbered TB-04, the acquisition unit 101 obtained information in multiple modalities: the image information showed that there was an obvious crack about 12 cm long and 2 cm wide on the surface of the steel frame of the equipment base, and the edge of the crack showed an irregular expansion trend; in the temperature information, the infrared imaging data in the same area showed a hot spot with an area of about 20 cm 2 ², and its temperature rose from 45 °C to 63 °C in a short time, and the temperature rise rate exceeded 6 °C per minute; the micro-vibration information recorded a continuous non-resonant periodic vibration in the frequency band of 7–15 Hz, and the main frequency point drift amplitude reached 12%; no obvious abnormality was found in the sound information.

[0082] The recognition unit 102 performs fusion processing on the above mutation data and outputs the abnormal area description information, including the abnormal type (structural abnormality + thermal abnormality), spatial position (right rear side of the base of the TB-04 reactor), mutation time (t0 = 10:23:41), image crack mutation score of 0.78, hot spot temperature rise score of 0.82, micro-vibration stability score of 0.75, and indicates that there is a multi-modal mutation in this area.

[0083] After receiving the abnormal area description information, the construction unit 103 first matches the equipment file corresponding to the abnormal position, and confirms that TB-04 belongs to an outdoor large grounding reactor, which has been continuously operating for 4.5 years, and the last maintenance was 8 months ago. By reading the reactor structure layout diagram, the system identifies that the abnormal area corresponds to the connection point of the load-bearing base, which is a typical high-stress concentration area. Subsequently, the construction unit calls the environmental data of this area from the GIS and meteorological databases, and identifies that the equipment is currently in a high-altitude area with a large temperature difference between day and night, the temperature difference between day and night exceeds 20 °C, and the air humidity is higher than 85% all year round.

[0084] The construction unit further searched for matching records in the local fault case database and discovered that over the past five years, multiple devices of the same model had experienced incidents such as cracking and shedding, grounding anomalies, or porcelain bushing explosions within 72 hours, all caused by the simultaneous presence of three factors: structural cracks, hot spots, and vibration frequency drift. Based on the current equipment's actual condition, modal catastrophe scores, and environmental factors, the construction unit determined that the anomaly was a "high-confidence multi-factor coupling risk," with structural fatigue as the dominant factor and high temperature differences and foundation looseness as the inducing factors.

[0085] Based on this, the construction unit 103 generates structured task description information as follows:

[0086] Anomaly Number: E-20240519-037

[0087] Related equipment number: TB-04

[0088] Abnormality type: structural abnormality + thermal abnormality + vibration abnormality

[0089] Abnormal location: right rear side of the device base

[0090] Risk level: Level 3 (high)

[0091] Modal overlap: image / temperature / micro-vibration, spatial overlap 92%, temporal overlap 1.5 seconds

[0092] Environmental trigger factors: high temperature difference, high humidity, long continuous operation cycle

[0093] Evolution prediction: The crack is expected to grow to more than 20 cm within 72 hours. The hot spot is unstable and there is a risk of steel frame fracture.

[0094] Recommended treatment plan: Prioritize manual review and use thermal imaging and acoustic emission methods for local inspection; if further crack extension is found, it is recommended to arrange power outage and disassembly inspection within 48 hours

[0095] Suggested feedback priority: high, needs to be pushed to the dispatch center immediately

[0096] Finally, the task description information is sent to the parsing unit 104 to provide a basis for querying task details, confirming or canceling instructions in voice interaction, and serves as a behavioral decision input for the control unit 105 to determine whether it is necessary to execute path fallback, prompt broadcast or stay on standby.

[0097] Through the above process, the construction unit 103 realizes the structural semantic enhancement of the recognition results, multi-factor cross-judgment and future-oriented task prediction output, so that the robot has the ability to actively judge the abnormal risk level and action priority, and meet the high standards of intelligent inspection in complex power grid environments.

[0098] Furthermore, the building unit is specifically configured to:

[0099] Extract the edge change contour related to the image information, the temperature rise rate data related to the temperature information, and the main frequency offset amplitude related to the micro-vibration information in the abnormal area description information, and construct a corresponding multi-factor abnormal feature sequence in chronological order to depict the evolution trend of the abnormal area in different modal dimensions;

[0100] Based on the multi-factor abnormal feature sequence, construct an evolution path state diagram, map the image crack growth trend, the temperature slope rise trend, and the vibration spectrum perturbation trend to the feature nodes in the diagram state respectively, and establish a modal joint path model through the state transition within a continuous time window;

[0101] Perform fuzzy linguistic variable mapping on the feature nodes in the evolution path state diagram, divide the change stage of the abnormal state into fuzzy levels including slight, enhanced, and unstable, and assign a modal membership function to each node to construct a fuzzy trend sequence for predicting the development direction of the abnormal area;

[0102] Combine the fuzzy trend sequence with the spatial position and risk level in the abnormal area description information to generate task description information including processing suggestions, task priorities, and expected evolution time windows.

[0103] In the grid monitoring bionic robot with multi-modal perception and voice interaction functions described in the present invention, the building unit not only undertakes the conventional processing responsibilities for the abnormal area description information, but also is further used for dynamically modeling the abnormal state, predicting the development trend, and outputting task semantic information.

[0104]

[0105]

[0105] In image information, the extraction of the edge change contour can adopt edge contrast analysis between consecutive frames, or mark the boundary changes of structural variation regions such as cracks, fractures, and spalling based on the contour tracking algorithm. This boundary change can be expressed by numerical indicators such as crack length, edge pixel density, or boundary curvature change rate. In temperature information, the construction unit obtains the derivative of the temperature change over time in the local area and performs smoothing processing to extract the temperature rise rate. A rapid change in the temperature rise rate often indicates problems such as overheating of internal loads, loose electrical connections, or thermal dissipation failure. In micro-vibration information, by performing time-frequency transformation on the signal (such as short-time Fourier transform), the main vibration frequency of each time slice is extracted, and whether there are significant drifts, harmonic enhancements, or bandwidth expansions is monitored, thereby obtaining dynamic characteristics reflecting mechanical looseness, resonance effects, or local structural imbalance.

[0106] Arranging the above three types of modal features in chronological order forms a multi-factor anomaly feature sequence. At each moment in this feature sequence, it includes an image feature quantity (such as crack length), a temperature rise rate value, a main vibration frequency offset value and its time stamp, spatial index, and acquisition confidence. This sequence forms the basis for subsequent evolutionary modeling.

[0107] After constructing the anomaly feature sequence, the construction unit will establish an evolutionary path state diagram in the form of a graph structure. In this state diagram, the image features, temperature features, and micro-vibration features in each time slice are respectively embedded in the graph as nodes. Each layer in the graph structure corresponds to a time window. The three modal nodes within the same time window are horizontally connected by virtual edges to record the multi-modal collaboration state at the same moment. Between different time windows, the front and rear nodes of the same modality are vertically connected by directed edges, indicating the time evolution path of the anomaly features of this modality. For example, if the crack length is 1.2 mm at t1, 1.8 mm at t2, and 2.6 mm at t3, an image path from t1 to t3 is formed in the graph; if the temperature rise rate increases from 3 °C per minute to 6 °C per minute, a temperature rise path is similarly established. The same is true for the main vibration frequency drift path of micro-vibration.

[0108] The state transition in the graph structure can be appended with transition weights, such as information on transition amplitude, transition rate change, etc., to depict the speed and direction of abnormal changes. This structure not only retains the horizontal collaboration relationship between modalities but also accurately describes the time evolution process of each type of anomaly factor, forming a multi-dimensional time series feature network.

[0109] Next, the construction unit performs semantic modeling on each node in the evolution path state diagram so as to be further converted into trend level labels that can be understood by the task planning and behavior control unit. The present invention uses a fuzzy logic method to divide numerical abnormal states. For the image crack growth trend, fuzzy linguistic variables such as "no crack", "slight crack", "significant crack", and "rapid crack" can be defined, and a corresponding membership function is established for each variable, such as a triangular membership function based on crack length; for the temperature rise rate, linguistic levels such as "slow temperature rise", "enhanced temperature rise", and "rapid temperature rise" can be established; for the change in the main vibration frequency, corresponding fuzzy levels such as "stable frequency", "frequency drift", and "abnormal frequency" can be established.

[0110] The construction unit inputs the eigenvalue of each node into its corresponding membership function model and calculates the membership degrees on each fuzzy linguistic variable. For example, if the temperature rise rate is 5.2 °C / min, it can simultaneously have a membership degree of 0.7 for "enhanced temperature rise" and a membership degree of 0.3 for "rapid temperature rise". In this way, the entire graph structure is mapped into a fuzzy trend network, forming a fuzzy trend sequence in chronological order, such as "slight crack → significant crack → rapid crack", or "slow temperature rise → enhanced temperature rise → rapid temperature rise".

[0111] The fuzzy trend sequence can be used to infer the abnormal development direction, identify whether it tends to be stable, whether there are signs of getting out of control, and whether there is an evolution mode with the aggravation of hidden risks. Compared with the fixed threshold judgment, the present invention effectively avoids the problems of fuzzy boundaries or premature misjudgment through fuzzy logic processing, and can identify the potential path deviating from the normal working condition at the early stage when the trend just forms.

[0112] Finally, the construction unit fuses the fuzzy trend sequence with the spatial position, start time, equipment number, and risk level in the abnormal area description information to generate complete task description information. The task description information includes, but is not limited to: whether the abnormal area needs to be processed preferentially, how long it is expected to enter the severe state, whether it is recommended that the dispatching center arrange for manual confirmation or stop the patrol and continue to observe. The task description information can further include the development rate of the abnormal area (for example, the crack propagation speed is 0.6 mm / hour), the modal dominant risk type (for example, structure abnormality dominates), and the recommended processing window (such as recommending re-inspection within 24 hours).

[0113] The building unit outputs the above information in the form of structured data and synchronously transmits it to the parsing unit and the control unit. The parsing unit can adjust the voice interaction strategy according to the semantic content of the task. For example, when the user issues an instruction of "Is it dangerous?", the system can answer that "There is a trend of crack intensification in this area, and it is expected to reach the critical risk state within 36 hours. It is recommended to intervene and check."; The control unit can then adjust the robot path, select to stay and observe or interrupt the current patrol inspection and give priority to scheduling this area.

[0114] Through the above processing flow, the building unit realizes a complete intelligent chain from anomaly detection to dynamic trend modeling and then to semantic task planning. Different from the traditional anomaly response system that only triggers processing actions based on static indicators, the building unit of the present invention has cross-modal, multi-dimensional, and time-evolution perception capabilities, and can judge the development state of anomalies through a fuzzy inference mechanism, significantly improving the robot system's response ability and predictive control ability for complex, implicit, and gradual anomalies.

[0115] The parsing unit 104 is used to receive the user's voice input, perform voice enhancement and tone change extraction in an interference environment including high noise, strong wind or low pressure, and combine the task description information to identify the user's instruction intention and obtain the interactive instruction information.

[0116] The parsing unit 104 is used to receive, process and understand the user's voice input, and finally convert it into interactive instruction information that can be called by the control unit 105. The core function of this unit is to enable the bionic robot for power grid monitoring to still accurately recognize the human operation intention in a high-interference environment and dynamically adjust the semantic interpretation result according to the current task background, so as to achieve natural and efficient voice interaction.

[0117] In actual operation, the bionic robot is in places with complex noise environments and frequent air pressure fluctuations such as power grid substations, line inspection channels or switchgear rooms. To improve the robustness of speech recognition, the parsing unit 104 first captures the user's voice signal through a group of digital array microphones and performs beamforming processing in the built-in local voice enhancement module to enhance the directivity of the sound source and suppress the background noise in non-target directions. Subsequently, the system calls a multi-channel adaptive noise suppression algorithm to isolate and filter the interference components such as arc sound, electric fan sound, and distant equipment sound in the continuous speech stream, so as to obtain a clear voice signal for subsequent semantic processing.

[0118] After completing the preprocessing of the voice signal, the parsing unit 104 inputs the obtained clean voice signal into the speech recognition engine. The recognition engine adopts an end-to-end speech recognition method based on short-time spectral features and transcribes it into a continuous text instruction. The system retains auxiliary information such as intermediate time tags, pause lengths, and intonation levels during the recognition process to construct a tone change model. This model can recognize emotional intensification, urgency, or negative tones in the statement. For example, if the user raises the volume and repeats the keyword "stop" in the statement, the system will interpret it as a high-priority instruction to request interrupting the task; if the user's intonation drops and is accompanied by a slow speaking speed, it may be interpreted as a review or slow-down behavior instruction.

[0119] To avoid misjudgment of a single voice content, the parsing unit 104 integrates the task description information from the construction unit 103 as the context during semantic understanding. The role of this mechanism is to limit the user's semantic scope through the task background and improve the accuracy of semantic parsing. For example, when the task description information shows that the current abnormal device is a high-risk electrical component and the user's voice contains instructions such as "continue" or "execute", the system will attach an intention confidence judgment to this instruction. Only when it clearly recognizes a non-negative tone and matches the condition that the task status allows continuation, will it generate a valid interaction instruction; otherwise, it will request the user to confirm the voice.

[0120] The parsing unit 104 is also designed with a multi-round instruction fault tolerance mechanism. When the user's voice is ambiguous or the recognition confidence is lower than the threshold, the system will request the user to repeat or rephrase it in the form of voice feedback, such as "Please confirm whether to continue patrolling this area", to ensure that the interaction does not accidentally trigger key control actions. This mechanism reconstructs the intention reasoning process by combining the previous round of voice, the current task status, and the current voice, thereby improving the actual interaction efficiency.

[0121] The output result of the interaction instruction information is in a structured format, including fields such as the user instruction type (continue, stop, review, confirm, alarm, etc.), confidence score, associated task number, voice trigger timestamp, and intonation auxiliary label. This information will be synchronously transmitted to the control unit 105 together with the task description information provided by the construction unit 103 for generating specific behavior execution plans and voice response contents.

[0122] In summary, through multi-level anti-interference voice enhancement, intonation change modeling, context-constrained semantic recognition, and multi-round confirmation interaction mechanisms, the parsing unit 104 realizes the stable voice interaction ability of the present invention in a high-interference power grid scenario, enabling the bionic robot to have practical functions such as flexibly adjusting the inspection behavior under human voice guidance, actively responding to instructions, and assisting in early warning processing.

[0123] Furthermore, the parsing unit is specifically used for:

[0124] Extract the peak sound pressure, the range of speech rate variation, and the frequency energy distribution information of the user's voice input. Combine with the task risk level in the task description information to generate an intonation intensity score, and construct a linkage identifier between the intonation and the task risk level to mark the potential operation risk level of the current voice command;

[0125] Based on the linkage identifier, determine whether the user's voice input has a coupled state of high intonation intensity and high task risk level. If the preset linkage threshold is met, the parsing unit generates an interactive instruction information with a semantic confirmation request field and outputs a confirmation prompt content to guide the user to perform a clear semantic restatement;

[0126] After generating the interactive instruction information, the parsing unit performs fine-tuning processing on the spectrum range of the original user voice input, including compressing the resonance section of the input signal, enhancing the band-pass of the transition frequency band, and suppressing the background interference spectrum section, and outputs a registered voice enhancement signal;

[0127] Jointly input the voice enhancement signal and the semantic confirmation request field into the semantic recognition process, update the semantic confidence score in the interactive instruction information, and combine with the response time requirement and the task type field in the task description information to generate a high-confidence interactive instruction information finally used for the behavior judgment of the control unit.

[0128] In a grid monitoring bionic robot with multi-modal perception and voice interaction functions, to achieve the robustness and safety of voice control in complex environments, the parsing unit in the system undertakes the key tasks of extracting key information from the original voice signal, judging potential risks, improving recognition accuracy, and finally forming reliable instructions for the control unit to use.

[0129] First, during the process of the user interacting with the bionic robot through voice, the parsing unit captures the voice input in real time and extracts three key features from the original signal: peak sound pressure, range of speech rate variation, and frequency energy distribution. The peak sound pressure is used to reflect the intonation intensity of the user in the voice input, and this indicator is usually obtained through short-time energy calculation. After the parsing unit frames the input voice, it calculates the maximum energy value of each frame and takes the global peak of the entire sentence as the representative sound pressure intensity indicator for this input. The range of speech rate variation is obtained by detecting the time interval between phonemes. When the system performs preliminary speech recognition, it marks the boundaries of language units such as vowels and consonants and calculates their change rates on the time axis. If the speech rate shows significant unevenness, that is, suddenly speeds up or slows down in the sentence, it may mean that the user is expressing an emergency or a hesitant command. The frequency energy distribution information analyzes the spectral structure of the voice through fast Fourier transform, statistically calculates the energy concentration degree of each frequency band, and is used to judge whether there are non-verbal interferences such as wind noise, mechanical noise, or electromagnetic harmonic interference.

[0130] The above three indicators are combined into a parameter vector for analyzing the characteristics of speech intonation. The parsing unit fuses this parameter vector with the task risk level in the task description information. The task description information usually contains the risk level of this abnormal area, such as low, medium or high, indicating whether there are problems such as electrical overload, structural fatigue or environmental anomalies in this area. The system sets a group of intonation intensity scoring rules and calculates the intonation intensity score value. For example, when the peak sound pressure is higher than 90 dB, the speech rate mutation rate is greater than 20%, and the frequency energy is concentrated above 2 kHz, a higher intonation score value can be assigned to this input.

[0131] The parsing unit establishes a linkage mapping relationship between this intonation score value and the task risk level, and constructs an intonation-task risk linkage identifier. This identifier is a semantic label used to indicate whether the current user speech input has significant characteristics in both attributes of "strong intonation" and "high task risk". This linkage identifier can be judged by Boolean logic (such as the score is greater than 0.8 and the risk level is high), or fuzzy logic can be used to assign a certain linkage confidence value.

[0132] When it is recognized that the input has a strong intonation and the target task corresponds to a high risk level, the system believes that this instruction has a high degree of uncertainty or safety sensitivity. To prevent incorrect control actions caused by misunderstandings, the parsing unit will automatically generate an interactive instruction information with a semantic confirmation request field. This field clearly indicates that the system requires the user to further confirm their instruction and outputs a voice prompt message, such as "Your current operation involves a high-risk task. Please confirm whether to continue the execution". This confirmation mechanism provides necessary guarantee for safe interaction in the power grid scenario, especially when the user's intonation suddenly increases or the environment changes, it can effectively avoid incorrect triggering of commands leading to misoperation of high-voltage equipment or interruption of the inspection process.

[0133] After generating the confirmation request, the parsing unit enters the voice signal quality enhancement phase. In power operation environments, voice signals are often affected by background interference such as wind noise, electrical current howling, and equipment vibration, which can easily lead to recognition errors or reduced confidence. To this end, the parsing unit performs spectral fine-tuning on the original voice signal. First, the system compresses the resonance segments. The system identifies the overlapping resonance areas between the user's vocal tract and the background equipment, narrows the bandwidth, and flattens the energy to suppress "howling"-like misidentification features. Secondly, bandpass enhancement is performed within the transition region of the voice band (typically 1kHz to 2kHz) to improve clarity and penetration in the human voice band. Finally, the system performs spectral suppression on the frequency band containing background noise. For example, when continuous white noise or low-frequency vibration signals are detected, a frequency masking algorithm is used to suppress the signal in this frequency band below the recognition threshold. After these processes, the system outputs a registered voice enhancement signal, which significantly improves robustness against background interference while preserving the user's semantic features.

[0134] The enhanced speech signal does not enter the semantic recognition module independently, but is instead inputted jointly with the semantic confirmation request field generated in the previous step. During the semantic recognition process, the system considers both the enhanced speech content and contextual information, including whether the user has confirmed the request. The semantic recognition result includes not only the text content but also a semantic confidence score. This score is updated to a dynamic confidence value that integrates speech quality, confirmation status, and historical recognition accuracy. The parsing unit integrates the final confidence score of the interactive instruction information with the response time requirement and task type field in the task description information. For example, if the task type is "structural crack risk warning" and the response time requirement is "must respond within 3 minutes," the system will appropriately lower the confidence threshold to increase response speed. If the task type is "routine inspection supplementary instructions," the semantic confidence threshold can be increased to ensure semantic accuracy.

[0135] Ultimately, the parsing unit outputs high-confidence interaction command information, which includes semantic content that has been processed and confirmed multiple times, and also integrates the task background, response strategy, and interaction safety mechanism, ensuring that the information the control unit relies on when making behavioral decisions is stable, reliable, and consistent with the actual semantics of the scene.

[0136] The entire process not only improves the practicality of the voice interaction system in complex environments, but also combines the risk level system and time response mechanism of power grid equipment operation to realize a closed-loop processing chain of "intonation understanding - task matching - interaction safety - command enhancement".

[0137] The control unit 105 is configured to receive the interaction instruction information and the task description information, jointly determine the behavior correction strategy, and generate motion control instructions and voice feedback content.

[0138] The control unit 105 serves as the decision-making and execution core in the power grid monitoring bionic robot. Its function is to comprehensively judge whether the current behavior needs to be corrected or replanned based on the interaction instruction information and task description information, and accordingly generate corresponding motion control instructions and voice feedback content, so as to drive the robot to complete specific actions such as patrol, stay, bypass, report or assist. Its design goal is to achieve stable, predictable, human instruction-sensitive and self-correcting multi-task behavior control in a complex power operation environment.

[0139] In implementation, the control unit 105 first receives the interaction instruction information from the parsing unit 104. After being subjected to speech recognition and semantic parsing, this information has been structured into clear instruction types, such as "continue to move forward", "return to the abnormal point", "pause the patrol", "execute the manual review suggestion", etc., and also includes additional elements such as intonation intensity, confidence score, trigger moment, etc. At the same time, the control unit also synchronously receives the task description information from the construction unit 103, which contains key contents such as the type, location, risk level, recommended response method and expected development path of the abnormal area.

[0140] The behavior decision-making process of the control unit 105 is carried out based on the state-instruction fusion model. The system matches the current task state (such as robot pose, task progress, abnormal point distribution) with the user's voice interaction intention to judge whether it is necessary to interrupt the original path, change the target area, adjust the execution priority or modify the motion strategy. For example, when the task description information shows that the front area is a high-risk thermal fault area, but the user's instruction intention is "bypass", the system will calculate the conflict degree between this instruction and the task goal. If the voice confidence of this instruction exceeds the set threshold, or there are safety hazards in the current path, the bypass strategy will be preferentially executed; otherwise, the user will be prompted to confirm whether to forcibly interrupt the task.

[0141] After determining the behavior strategy, the control unit translates this strategy into a low-level motion instruction set, including control signals such as forward, backward, turn, attitude adjustment, speed regulation, stay and wait, target point locking, etc. These signals will be decomposed into synchronous drive signals of multi-degree-of-freedom motors according to the characteristics of the robot body structure, and real-time obstacle avoidance and repositioning will be carried out to ensure smooth movement in complex terrains, avoid high-voltage equipment, and maintain a stable walking posture.

[0142] At the same time, the control unit will also generate supporting voice feedback content to inform the on-site operation and maintenance personnel or the remote monitoring end of the current behavior response situation in natural language. For example, when executing the "return to the abnormal point" instruction, the system will broadcast through voice "Returning to the right area of the switch cabinet where a crack was detected last time. Please pay attention to the surrounding environment", ensuring the transparency and interpretability of human-machine collaboration. The voice feedback content will dynamically adjust the speech rate, intonation and content details according to the risk level in the task description information and the urgency of the user's instruction.

[0143] If the robot encounters a sudden change in the environment while executing control commands, such as a tilted ground surface, frontal obstructions, or increased electromagnetic interference, the control unit immediately calls backtracking node information (if available) to determine whether to abort the current action and switch to risk avoidance mode, or notify the user for further instructions. This mechanism ensures the robot has basic environmental adaptability in unpredictable power operation sites.

[0144] In summary, control unit 105 integrates task intent with user interaction information to formulate a clear behavioral strategy and issue specific, executable motion control commands. Simultaneously, it establishes a real-time, closed-loop human-machine information flow through voice feedback, ensuring the robot performs precise, safe, and adaptable inspections and auxiliary tasks in complex power grid scenarios. Its system architecture fully supports flexible task scheduling, exception handling responses, and personalized voice interaction, ultimately guaranteeing the implementation of the multimodal perception and semantic interaction capabilities of the present invention.

[0145] For example, during high-temperature, high-load operation in a primary main transformer area, the bionic robot identified a suspected overheating hazard during its inspection process through recognition unit 102 and construction unit 103. Its task description information indicated that the abnormal area had a rapid temperature rise and significant temperature difference, and historical data showed cases of insulation failures caused by thermal aging in similar equipment. The system recommended prioritizing manual review. Simultaneously, the on-site operation and maintenance personnel issued a voice command to "pause the current inspection and return to the hot spot to recheck." The parsing unit 104 recognized the voice content as a high-confidence action correction intention and clarified the instruction type as "task path rollback and key review."

[0146] After receiving the interactive command information and the original task description, the control unit 105 first determines whether the user command conflicts with the current robot behavior path and confirms that the command has task-level priority. Next, the system reconstructs the path based on the three-dimensional coordinate position information of the previous anomaly point and the robot's current position. It then calls the motion planning module to generate a reverse navigation path, returning to the location of the abnormal heat source along the original path. During the motion control command issuance process, the control unit dynamically adjusts based on the known obstacle locations in the path, ground slope information, and current electromagnetic environment parameters to ensure smooth and safe movement.

[0147] When the robot approaches the target area, the control unit immediately calls the acquisition unit 101 to re-collect multimodal data on the abnormal area, including image, temperature, and micro-vibration data, and notifies the recognition unit to perform high-density resampling analysis. Simultaneously, the control unit generates voice feedback, announcing through the loudspeaker: "We have returned to the abnormal area as instructed. We are re-collecting temperature rise and structural status data. Please wait," allowing on-site personnel to understand the execution status in real time.

[0148] During this process, the tone intensity in the interactive instruction information and the risk level in the task description information jointly affect the behavior scheduling strategy of the control unit. When the user's tone is imperative and the task risk is "high", the system will execute directly without waiting for reconfirmation; if the tone is advisory and the task level is low, the control unit may first confirm by voice "Are you sure you want to abort the original inspection path and execute a return re-inspection?" to ensure safety and interactive rationality.

[0149] Through the above process, the control unit 105 realizes the joint judgment of the interactive semantics and the task context, dynamically generates the motion behavior control and the voice interaction feedback content, enabling the robot to have the ability to adapt to sudden instructions, dynamically correct behaviors and cooperate with people interactively.

[0150] Furthermore, the control unit is also used for:

[0151] During the process of the power grid monitoring bionic robot executing the inspection task or the assistance task, based on the discriminative spatio-temporal change characteristics in the continuously collected image information, temperature information and micro-vibration information by the acquisition unit, automatically mark multiple backtracking nodes for sensing state restoration in the task path, and construct a motion backtracking mark sequence according to the marking order;

[0152] When the position of the abnormal area description information output by the recognition unit changes, the risk level rises, or there is a behavior conflict with the interactive instruction information output by the parsing unit, the control unit calls the historical node closest to the current task position in the motion backtracking mark sequence as the new behavior recovery anchor point;

[0153] After the control unit backtracks to the behavior recovery anchor point, it restores the sensing state data corresponding to the behavior recovery anchor point, re-performs multi-modal information acquisition and analysis with higher priority on the target area, and updates the behavior decision-making strategy according to the difference in the recognition results before and after backtracking, so as to improve the response ability and abnormal recognition accuracy of the power grid monitoring bionic robot in complex scenarios such as high-voltage tower base areas, cable trench intersection areas or blind corner areas of occlusion.

[0154] In the power grid monitoring bionic robot with multi-modal perception and voice interaction functions proposed by the present invention, in order to further enhance the task adaptability and behavior recoverability of the robot in complex environments, the control unit is designed to have the ability to restore the sensing state at the path level, and can automatically mark key backtracking nodes based on the real-time collected multi-modal data, and dynamically adjust the behavior path based on these nodes when encountering emergencies or task conflicts.

[0155] In the actual application in the power grid field, robots often need to perform long-term inspection and auxiliary operation tasks in areas with complex structures, cramped spaces, and frequent interference signals, such as high-voltage tower bases, corners in distribution rooms, cable trench intersections, or underground ventilation channels. These areas not only have physical access obstacles but also are accompanied by high-frequency electromagnetic interference, low-light conditions, equipment vibration sources, or wind noise interference, posing multiple challenges to robot path planning, status judgment, and information collection. In such an environment, if the robot encounters situations such as instruction conflicts, enhanced abnormal perception, or path failure during the task execution, and cannot effectively backtrack to a stable state point and re-establish the perception chain, it will lead to task execution failure, missed inspections, or misjudgments. Therefore, the "motion backtracking marker sequence" mechanism is introduced in the control unit of the present invention.

[0156] The first step of this mechanism is the marking and recording of backtracking nodes. During the process of the robot performing inspection tasks or assisting tasks, the acquisition unit continuously obtains image information, temperature information, and micro-vibration information. The control unit does not simply use this information for immediate judgment but performs spatial position binding and temporal structuring processing on these modal data to monitor the spatio-temporal change characteristics in real time. When the control unit detects that the structural edge change between image frames remains stable, the temperature difference gradient in the thermal imaging map is lower than the threshold within a continuous period, the micro-vibration spectrogram has no obvious disturbance, and the robot's traveling speed is in a low-speed or stationary state, the system will consider that this point has state stability and perception recoverability and thus automatically mark this point as a "backtracking node".

[0157] Each backtracking node not only records the image, temperature, and vibration information at that moment but also binds the three-dimensional coordinates, sampling timestamp, the current pose of the robot, the relative navigation direction, and the standard deviation score of multi-modal data at this position. The control unit will save the above node information in an ordered data structure, which is the "motion backtracking marker sequence". The marking order is usually dynamically updated with the forward direction of the robot's task path as the main axis. To avoid recording redundancy and data storage burden, the system introduces a time window mechanism and a redundancy judgment mechanism, and new nodes will only be added when the modal change rate between adjacent nodes is higher than the set difference or the spatial distance exceeds the set interval.

[0158] After constructing the motion backtracking marker sequence, the control unit uses this sequence as a backup reference mechanism for path scheduling and status restoration. When the robot continues to execute the task and the description information of the abnormal area output by the recognition unit changes, such as the abnormal position deviating from the previous recognition result by more than the set distance, or the abnormal risk level changing from "medium" to "high", or when there is a logical conflict between the interactive instruction information output by the parsing unit and the current behavior state (such as the task path is to move forward and the user instruction is "return to the previous fault point"), the control unit will immediately trigger the motion backtracking decision-making process.

[0159] This process first calculates the distance between the current position of the robot and the historical nodes in the motion backtracking marker sequence. Combining the current direction, path smoothness, and environmental conditions, it selects the nearest node with a high-confidence perception state record as the "behavior recovery anchor point". The selection of this anchor point gives priority to modal stability indicators, such as parameters like image noise density, thermal map gradient amplitude, and vibration spectrum uniformity. If multiple nodes meet the distance and confidence conditions, the system will select forward according to the node number order to ensure the minimum action interference of the robot's backtracking path.

[0160] Once the recovery anchor point is determined, the control unit will backtrack the robot's path to the position of this anchor point, reactivate the perception state data corresponding to this node, and synchronously restore the robot's attitude, visual angle, perception configuration, etc. to the state at that moment. On this basis, the acquisition unit re-performs high-density multi-modal information acquisition on the target area, and the coverage angle, distance, focal length, or frame rate can all be appropriately adjusted based on this anchor point to increase the coverage range of abnormal details. The control unit calculates the spatial consistency, modal trend consistency, and change in mutation degree of the recognition result by comparing the newly acquired data with the original data of this node and the current abnormal area description information, so as to judge whether the evolution of the abnormal event exceeds the original expectation, or whether it is a false alarm, occlusion, transient disturbance, etc.

[0161] For example, inside a cable trench intersection, the robot records an abnormal point where image distortion, temperature rise, and micro-vibration exist simultaneously at time t1 and writes it into the sequence as a backtracking node. However, after the robot moves forward to the next position at time t2, the recognition unit finds that the position of this abnormal area has shifted and the risk level has increased. The parsing unit simultaneously receives a voice command from the operation and maintenance personnel "confirm whether the abnormal position has expanded". At this time, the control unit will make a joint judgment based on the above information and consider that the backtracking mechanism needs to be triggered. The system locates the backtracking node at time t1 and makes the robot return to the position of this node again, restores the sensing configuration at time t1, and performs image and thermal imaging overlay sampling on this area again. By comparing the image contour growth rate and the change rate of the hot spot area, it is confirmed that the abnormality in this area has expanded to twice the original size. The system adjusts the task priority accordingly and pushes a new task description to the interaction module.

[0162] Finally, the control unit reconstructs the behavior strategy based on the updated difference in the recognition result, and judges whether to continue the patrol, switch to fixed-point monitoring, wait for manual intervention, or perform a secondary confirmation scan. The new behavior strategy will be converted into corresponding motion control instructions and sent to each actuator to complete the action. At the same time, the control unit generates voice feedback content, such as "The abnormal area has been confirmed to have expanded, and is switching to the stationary state and reporting the abnormal details", to ensure that on-site or remote operation and maintenance personnel can grasp the robot's response logic and execution status in real time.

[0163] This control flow enables the robot to quickly retreat to the most reliable position and state at the lowest cost when encountering path interference, data deviation, false alarm risks, or user command conflicts, re-establish a stable perception chain, and make precise decisions by combining the data changes before and after. Compared with the traditional one-way control flow of "forward - rescan - error reporting - waiting", the behavior backtracking and decision update mechanism of the present invention can form a three-way coupling loop at the path level, data level, and instruction level, significantly improving the task adaptability, response agility, and fault identification accuracy of the robot in high-risk power scenarios.

[0164] In addition, during the process of implementing the above functions, the control unit also reserves interfaces with the motion control model, path planning algorithm, and perception quality scoring system, facilitating future in-depth learning optimization of the selection of backtracking nodes, strategy priority ranking, and abnormal evolution model by combining artificial intelligence algorithms, providing support for the stability and intelligent upgrade of the long-term operation of the system.

[0165] Furthermore, when constructing the motion backtracking marker sequence, the control unit is also used to calculate a confidence value for each backtracking node in the motion backtracking marker sequence. The calculation of the confidence value includes the following steps:

[0166] The control unit calculates the modal time drift tolerance index based on the temporal stability of the image information, temperature information, and micro-vibration information collected by the acquisition unit at the position of the first backtracking node, as the drift consistency score representing the degree of multi-modal perception synchronization offset;

[0167] The control unit extracts the interference sensitivity level corresponding to the first backtracking node according to the drift consistency score, in combination with the occlusion interference map pre-established in the power grid environment where the first backtracking node is located, and generates a perception suppression factor, which is used to reflect the environmental vulnerability of the image information, temperature information, and micro-vibration information at this node;

[0168] The control unit generates a confidence benchmark value based on the perception suppression factor and the drift consistency score, by jointly counting the number of successful multi-modal state discriminations generated by the first backtracking node during the execution of the robot task;

[0169] The control unit fuses and weights the confidence benchmark value with the spatial priority level of the backtracking node in the motion path, outputs the confidence value, and preferentially selects the backtracking node with a higher confidence value during the behavior recovery process, ensuring the stable recovery of the task and the reconstruction of the perception state consistency under multi-abnormal modal interference.

[0170] In the grid monitoring bionic robot of the present invention, in order to achieve more stable and highly reliable path backtracking and state recovery capabilities in a complex grid operation environment, the control unit is further configured to have the ability to evaluate the confidence of each node in the motion backtracking marker sequence. Specifically, this function enables the robot to select the most appropriate backtracking node based on scientific and quantitative methods when there is a large uncertainty in multimodal data, strong interference points in the path, or deviation in recognition results, avoiding misjudgment, repeated actions, or system behavior oscillation caused by backtracking to nodes that are not recoverable or have contaminated signals.

[0171] During the process of the bionic robot performing inspection or auxiliary tasks, the control unit receives the multimodal perception data output by the acquisition unit in real time and continuously performs structured analysis on this data. When the control unit determines that the current state is stable, the perception distribution is uniform, and the data fluctuation is below the set threshold based on the image information, temperature information, and micro-vibration information, it will mark this position as a backtracking node and record it in the motion backtracking marker sequence. To ensure the accuracy of the backtracking strategy, when generating each backtracking node, the control unit also needs to quantitatively evaluate the reliability, reproducibility, and representativeness of this node. This quantitative result is the "confidence value". This confidence value will be an important basis for node screening when triggering backtracking in the future.

[0172] The evaluation of the confidence value consists of multiple stages. First is the temporal consistency analysis of the modal data. Due to problems such as sampling frequency differences, timestamp errors, and processing delays in various sensors in the perception unit, even if multiple modalities are logically synchronized for acquisition, there may be temporal misalignment in the actual data. Therefore, the control unit needs to calculate the temporal drift difference of each modal data at this node based on the sampling time window of the image information, temperature information, and micro-vibration information. For example, if the image frame time is t0, the temperature sampling time is t0 + 40ms, and the micro-vibration peak appears at t0 + 45ms, then the maximum temporal drift is 45ms. The system sets a standard modal synchronization tolerance (such as 50ms). If the multimodal drift values at the node are all less than this threshold, the modal temporal drift tolerance score of this node will be relatively high, indicating that this node has good synchronous consistency in time. The control unit outputs a drift consistency score based on such statistics, usually a floating-point value between 0 and 1. The higher the value, the more ideal the modal synchronization and the stronger the collaborative interpretation ability between the data.

[0173] After obtaining the drift consistency score, the system further corrects the confidence level by combining environmental interference factors. In the on-site operation of the power grid, robots are often in adverse environments such as the presence of wind barriers, obstructions, uneven terrain, or strong electromagnetic interference. Therefore, the present invention establishes an occlusion interference map for a specific deployment area. Based on long-term historical monitoring data or map models, this map records the possible perception occlusion factors in each monitoring area, such as image acquisition blind spots in areas close to walls, infrared temperature measurement attenuation positions in high-humidity environments, and electromagnetic vibration interference points caused by proximity to large transformers. Each area is assigned an interference sensitivity level, usually represented by levels 0 to 5, ranging from completely open and interference-free to strong interference and multiple occlusions. The control unit queries the occlusion interference map based on the position information of the current backtracking node to obtain the interference sensitivity level of this node. Then, in combination with the aforementioned modal drift score, the control unit derives a "perception suppression factor", whose physical meaning is the possible degree to which the perception data of this node is restricted by the environment, and is used to penalize the node's confidence level. For example, if a node has high modal time consistency itself but is in a strongly occluded area or a high magnetic field perturbation area, the perception suppression factor of this node will be high, and the system will lower its confidence level.

[0174] After completing the evaluation of time consistency and environmental interference, the control unit also statistically analyzes the historical performance of each backtracking node to measure whether its perception state is representative. During the robot's task, the data recorded by different backtracking nodes may be used multiple times to determine whether the current environmental state is abnormal. If a node is repeatedly confirmed as "stable state", "consistent judgment result", and "high confidence in anomaly recognition", the number of successful state discriminations of this node will be cumulatively increased, indicating that its credibility as a "perception state anchor point" is stronger. The control unit statistically analyzes the number of successful multi-modal state discriminations of this node within a certain time window and normalizes it with the number of modalities (usually three) to generate a "confidence level reference value". This reference value reflects the historical stability of this node as a "reference state" during task execution.

[0175] Finally, the control unit combines the node spatial position and navigation structure in the task path information to fuse and weight the confidence level reference value with the "spatial priority level". The spatial priority level refers to whether the position of this node in the entire task path is easy to access, safe to retreat, conducive to perspective arrangement, etc. For example, a node located in the middle of a straight path with an unobstructed view has a high spatial priority level; while a node close to an obstacle, located on a slope, or at a relatively long distance has a lower spatial priority level. The system comprehensively calculates the confidence level reference value and the spatial priority level according to the preset weight factor and outputs the final confidence level value. The control unit stores all the generated confidence level values together with the corresponding nodes in the motion backtracking marker sequence structure and continuously updates them during task execution.

[0176] When the robot needs to trigger behavior backtracking during a task, the control unit first screens the nodes in the motion backtracking marker sequence that meet the basic conditions such as close position, complete modalities, and complete states, and then selects the one with the highest confidence value as the behavior recovery anchor point, and retreats to the spatial position and perception state corresponding to this node. On this basis, multi-modal acquisition, anomaly recognition, and strategy adjustment are carried out again to ensure that the system can still recover to the most reliable state point and reconstruct the task logic chain from this point in complex situations such as path perturbation, perception drift, or instruction conflict.

[0177] For example, when the robot travels to a cable support area and enters a pending confirmation state due to detecting signals of abnormal temperature and mild structural cracks. After the subsequent recognition results deviate and the voice instruction prompts "Please return to the previous stable point for retesting", the control unit triggers the confidence ranking mechanism. Three historical backtracking nodes that meet the conditions are found in the sequence. The first node has a modality drift score of 0.91, a low perception inhibition factor (in an unobstructed area), the state discrimination has been successful 4 times, and the spatial priority level is medium; the second node has slightly better perception data but is located at the corner of the channel; the third node has a large modality time difference. The system finally selects the first node as the anchor point for backtracking and successfully completes state reconstruction and anomaly comparison.

[0178] Through the above mechanism, the control unit not only has the structural logic of path backtracking, but also has the quantitative judgment ability for the quality of multi-modal perception signals. Its confidence scoring model integrates multiple dimensions such as temporal consistency, environmental interference, historical stability, and spatial accessibility, enabling the robot to still select the best path recovery node in a multi-factor interference environment, avoiding mis-touching, misjudging, and executing repetitive tasks, and significantly improving the stability and response accuracy of the system in high-risk power scenarios.

[0179] Although this application is disclosed above with preferred embodiments, it is not used to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application should be subject to the scope defined by the claims of this application.

Claims

1. A bionic robot for power grid monitoring with multi-modal perception and voice interaction functions, characterized in that, Including: A collection unit, configured to synchronously collect sound, image, temperature, and micro-vibration information on the surface of power grid equipment, preprocess the collected data and perform time-space calibration, and generate perception data packets in a unified format; An identification unit, configured to receive the perception data packets, perform cross-judgment based on the mutation characteristics and correlations between different data types, identify electrical anomalies, structural anomalies, or environmental anomaly regions, and obtain anomaly region description information including anomaly types, locations, and times; A construction unit, configured to receive the anomaly region description information, combine the environmental type and historical power grid rules, determine whether the anomaly has a multi-factor coupling risk, and generate task description information containing processing suggestions and development path predictions; An analysis unit, configured to receive user voice input, perform voice enhancement and tone change extraction in an interference environment including high noise, strong wind, or low voltage, and combine the task description information to identify the user's instruction intention and obtain interaction instruction information; A control unit, configured to receive the interaction instruction information and the task description information, jointly determine a behavior correction strategy, and generate motion control instructions and voice feedback content.

2. The bionic robot for power grid monitoring with multi-modal perception and voice interaction functions according to claim 1, characterized in that The control unit is further configured to: During the process of the power grid monitoring bionic robot performing a patrol task or an assistance task, based on the discriminative spatio-temporal change characteristics in the continuously collected image information, temperature information, and micro-vibration information by the collection unit, automatically mark multiple backtracking nodes for perception state restoration in the task path, and construct a motion backtracking mark sequence according to the marking order; When the location of the anomaly region description information output by the identification unit changes, the risk level rises, or there is a behavior conflict with the interaction instruction information output by the analysis unit, the control unit calls the historical node closest to the current task location in the motion backtracking mark sequence as a new behavior recovery anchor point; After the control unit backtracks to the behavior recovery anchor point, it restores the perception state data corresponding to the behavior recovery anchor point, re-performs multi-modal information collection and analysis with higher priority on the target region, and updates the behavior decision-making strategy according to the difference in the recognition results before and after backtracking, so as to improve the response ability and anomaly recognition accuracy of the power grid monitoring bionic robot in complex scenarios including high-voltage tower base regions, cable trench intersection regions, or blind corner regions of occlusions.

3. The bionic robot for power grid monitoring with multi-modal perception and voice interaction functions according to claim 2, characterized in that, When constructing the motion backtracking mark sequence, the control unit is further configured to calculate a confidence value for each backtracking node in the motion backtracking mark sequence, and the calculation of the confidence value includes the following steps: The control unit calculates a modal time drift tolerance index based on the temporal stability of the image information, temperature information, and micro-vibration information collected by the collection unit at the first backtracking node location, as a drift consistency score representing the degree of multi-modal perception synchronization offset; The control unit extracts the interference sensitivity level corresponding to the first backtracking node according to the drift consistency score, in combination with the pre-established occlusion interference map in the power grid environment where the first backtracking node is located, and generates a perception suppression factor to reflect the environmental vulnerability of the image information, temperature information, and micro-vibration information at this node; The control unit generates a confidence benchmark value by jointly counting the number of successful multi-modal state discriminations of the first backtracking node during the execution of the robot task based on the perception inhibition factor and the drift consistency score; The control unit fuses and weights the confidence benchmark value with the spatial priority level of the backtracking node in the motion path, outputs a confidence value, and preferentially selects the backtracking node with a higher confidence value during the behavior recovery process to ensure the stable recovery of the task and the reconstruction of the perception state consistency under multi-abnormal mode interference.

4. The bionic robot for power grid monitoring with multi-modal perception and voice interaction functions according to claim 1, characterized in that, The recognition unit is specifically used for: Performing synchronous time registration on the image information, temperature information, and micro-vibration information included in the perception data packet generated by the acquisition unit, extracting the mutation time points of each perception mode, and constructing a modal mutation alignment sequence for marking potential high-risk moments when multiple modes simultaneously mutate within the same time window; Based on the modal mutation alignment sequence, performing spatial overlap comparison on the pixel color difference mutation region of the image information, the local temperature rise slope mutation region of the temperature information, and the main frequency offset mutation region of the micro-vibration information, calculating the mutation intersection degree of each mode in the same region, and generating an abnormal overlap score value for characterizing the consistency degree of multi-modal mutations in this region; Weightedly fusing the abnormal overlap score value with the confidence of each modal mutation to construct a heterogeneous modal confidence voting structure, and the confidence voting structure is used to trigger the judgment that this spatial region is an abnormal region to be marked when the confidence of any two of the three modes exceeds a preset threshold; Combined with a preset modal conflict counterexample set, performing a reverse contradiction mode comparison on the abnormal region to be marked. If its modal combination characteristics highly match the false abnormal types recorded in history, cancel the annotation of this abnormal region. Otherwise, officially mark it as an abnormal region and generate abnormal region description information including the abnormal type, location, and time for use by the construction unit.

5. The bionic robot for power grid monitoring with multimodal perception and voice interaction functions according to claim 1, wherein The construction unit is specifically used for: Extracting the edge change contour related to the image information, the temperature rise rate data related to the temperature information, and the main frequency offset amplitude related to the micro-vibration information in the abnormal region description information, and constructing a corresponding multi-factor abnormal feature sequence in chronological order for depicting the evolution trend of the abnormal region in different modal dimensions; Based on the multi-factor abnormal feature sequence, constructing an evolution path state diagram, mapping the image crack growth trend, the temperature slope rise trend, and the vibration spectrum perturbation trend to the feature nodes in the diagram state respectively, and establishing a modal joint path model through state transitions within consecutive time windows; Performing fuzzy linguistic variable mapping on the feature nodes in the evolution path state diagram, dividing the change stage of the abnormal state into fuzzy levels including slight, enhanced, and unstable, and assigning a modal membership function to each node to construct a fuzzy trend sequence for predicting the development direction of the abnormal region; Combining the fuzzy trend sequence with the spatial position and risk level in the abnormal region description information to generate task description information including processing suggestions, task priorities, and expected evolution time windows.

6. The bionic robot for power grid monitoring with multimodal perception and voice interaction functions according to claim 1, characterized in that, The parsing unit is specifically used for: Extract the peak sound pressure, speed change amplitude, and frequency energy distribution information of the user's voice input, combine with the task risk level in the task description information, generate an intonation intensity score, and construct a linkage identifier between the intonation and the task risk level to mark the potential operation risk level of the current voice command; Based on the linkage identifier, determine whether the user's voice input has a coupling state of high intonation intensity and high task risk level. If the preset linkage threshold is met, the parsing unit generates interaction instruction information with a semantic confirmation request field and outputs a confirmation prompt content to guide the user to perform a clear semantic restatement; After generating the interaction instruction information, the parsing unit performs spectrum range fine-tuning processing on the original user voice input, including compressing the resonance section of the input signal, enhancing the bandpass of the transition frequency band, and suppressing the background interference spectrum section, and outputs a registered voice enhancement signal; Jointly input the voice enhancement signal and the semantic confirmation request field into the semantic recognition process, update the semantic confidence score in the interaction instruction information, and combine with the response time requirement and task type field in the task description information to generate high-confidence interaction instruction information finally used for the behavior judgment of the control unit.

Citation Information

Cited By

  • Multi-modal sensing fusion robot humanoid operation method and system

    CN120588247A

  • Operation and maintenance technology service remote guidance interaction method and system combined with agent assistance

    CN120670563A

  • Data center operation and maintenance robot control method based on large model

    CN120928698A

  • A large model-based data center operation and maintenance robot control method

    CN120928698B

  • End-to-end data dynamic collaboration system and method based on multi-modal data chain

    CN120931345A