Patrol method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610734308.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-04
AI Technical Summary
然而相关巡检装置存在无法识别变电站内设备故障的问题,具有不能及时消除故障的安全隐患
[0014]As described above, the inspection method provided in this application, on the one hand, adjusts the beamforming direction and position of the audio input module by detecting the sound source and position information of dynamically moving personnel, thereby achieving dynamic voice enhancement to follow the personnel and ensuring the continuity of voice command recognition during operations. On the other hand, based on the adjustment of the audio input module to follow the movement of the personnel, the equipment topology, equipment sound source information, and target equipment can be dynamically updated. By comprehensively analyzing information under different states, the target equipment can be located more accurately, thus more accurately determining the fault type. The combination of these two methods enables personnel to perform maintenance work more accurately and efficiently, improving the safety of the operating environment.
Smart Images

Figure CN122695698A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic digital data processing technology, and in particular to an inspection method, apparatus, equipment and storage medium. Background Technology
[0002] With the intelligent transformation of the power industry, inspection devices are being gradually introduced into substations. However, these devices often fail to identify equipment faults within the substation, posing a safety hazard due to the inability to promptly eliminate these faults. Furthermore, the enhancement direction of the microphone array module in these devices is tied to keyword-based positioning information. Users must repeatedly trigger keywords to make the microphone array module adaptively adjust its enhancement direction, impacting operational efficiency and further increasing safety risks. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide an inspection method, apparatus, equipment and storage medium.
[0004] To achieve the above objectives, this application provides an inspection method applied to an inspection device. The inspection device includes an audio input module, the beamforming direction and position information of which are dynamically adjusted based on the personnel voice source information and position information of the target personnel. The method includes: responding to the target voiceprint information received by the audio input module as device sound, extracting feature information from the target voiceprint information to obtain target feature information; obtaining device sound source information based on the position information of the audio input module and the target voiceprint information; obtaining a target device based on the device topology, device position information, and the device sound source information; the device topology is constructed from candidate devices determined by the target feature information; obtaining a target fault type based on the target feature information and fault feature information of multiple fault types of the target device; and responding to the target voiceprint information received by the audio input module as instruction information, performing maintenance on the target device according to the target fault type.
[0005] Optionally, extracting feature information from the target voiceprint information to obtain target feature information includes: converting the target voiceprint information into spectral data; using a first feature extraction module to extract features from the spectral data within a local time-frequency range to obtain the spectral features; using a second feature extraction module to extract temporal dependencies from the spectral features to obtain the temporal features; using an attention mechanism to calculate the correlation degree of each frame of feature data in the temporal features relative to the target device; obtaining weight information based on the correlation degree; and performing weighted summation processing on each frame of feature data in the temporal features based on the weight information to obtain the target feature information.
[0006] Optionally, obtaining device sound source information based on the location information of the audio input module and the target voiceprint information includes: the audio input module comprising multiple audio input units arranged in a ring; determining the location information of each audio input unit; calculating the physical distance and time difference based on the location information of different audio input units and the time of receiving the target voiceprint information; obtaining the distance difference between the sound source and different audio input units based on the time difference and the speed of sound; calculating the horizontal azimuth and vertical elevation angle based on the distance difference and the physical distance; and obtaining the device sound source information based on the physical distance, the horizontal azimuth, and the vertical elevation angle.
[0007] Optionally, obtaining the target device based on the device topology, device location information, and device sound source information includes: determining multiple candidate devices from multiple devices based on the target feature information; adjusting the positioning algorithm parameters based on the device topology formed by the multiple candidate devices; and obtaining the target device based on the positioning algorithm parameters, the device location information, and the device sound source information.
[0008] Optionally, the method further includes: obtaining the person's sound source information based on the location information of the target person and the location information of the audio input module; the person's sound source information includes the sound source direction angle, actual distance, and horizontal direction; if the actual distance does not exceed the pickup distance, updating the beamforming direction of the audio input module in response to the sound source direction angle; if the actual distance exceeds the pickup distance, obtaining the movement distance based on the actual distance and the pickup distance; and controlling the audio input module to move towards the target user based on the movement distance and the horizontal direction, thereby updating the location information of the audio input module.
[0009] Optionally, if the actual distance does not exceed the pickup distance, updating the beamforming direction of the audio input module in response to the sound source direction angle includes: the audio input module comprising multiple audio input units; determining the guide vector of the audio input module in response to the sound source direction angle; calculating the channel weight of each audio input unit based on the guide vector; and updating the beamforming direction of the audio input module based on the channel weight of each audio input unit.
[0010] Optionally, the method further includes: determining a fault priority based on the importance of the target device and the target fault type; generating a fault diagnosis report based on the type of the target device, the target fault type, the sound source information of the device, the fault priority, and the fault duration; obtaining historical fault data and corresponding fault handling strategies based on the fault diagnosis report; and correspondingly, in response to an instruction from a person whose voiceprint type is that of the target personnel, repairing the target device based on the target fault type, including: in response to an instruction from a person whose voiceprint type is that of the target personnel, repairing the target device based on the fault diagnosis report, the historical fault data, and the fault handling strategy.
[0011] This application provides an inspection device, comprising: a camera for tracking a target person and determining the location information of the target person; an audio input module for receiving target voiceprint information; a processing module for adjusting the beamforming direction and position information of the audio input module according to the voice source information and location information of the target person; in response to the target voiceprint information received by the audio input module being device sound, extracting feature information of the target voiceprint information to obtain target feature information; obtaining device sound source information according to the position information of the audio input module and the target voiceprint information; obtaining a target device according to the device topology, device location information, and device sound source information; the device topology is constructed from candidate devices determined by the target feature information; obtaining a target fault type according to the target feature information and fault feature information of multiple fault types of the target device; and in response to the target voiceprint information received by the audio input module being instruction information of the target person, performing maintenance on the target device according to the target fault type.
[0012] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement an inspection method.
[0013] This application provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to execute an inspection method.
[0014] As described above, the inspection method provided in this application, on the one hand, adjusts the beamforming direction and position of the audio input module by detecting the sound source and position information of dynamically moving personnel, thereby achieving dynamic voice enhancement to follow the personnel and ensuring the continuity of voice command recognition during operations. On the other hand, based on the adjustment of the audio input module to follow the movement of the personnel, the equipment topology, equipment sound source information, and target equipment can be dynamically updated. By comprehensively analyzing information under different states, the target equipment can be located more accurately, thus more accurately determining the fault type. The combination of these two methods enables personnel to perform maintenance work more accurately and efficiently, improving the safety of the operating environment. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram illustrating an application scenario of an inspection method according to an embodiment of this application;
[0017] Figure 2 This is a flowchart illustrating an inspection method according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] With the intelligent transformation of the power industry, inspection equipment is being gradually introduced into substations. Sound source localization and voice enhancement technologies based on multi-microphone arrays are becoming key to improving operational efficiency. Among related technologies, the microphone array module mainly has two core functions: first, sound source localization, which identifies the direction and angle of the sound source by triggering keywords; second, voice enhancement, which forms beamforming in the localized direction to improve the quality of the target voice and suppress noise from other directions.
[0021] However, related technologies in substation scenarios have significant drawbacks: substations have dense equipment, vast work areas, and a large range of personnel movement. In voice enhancement solutions, the enhancement direction of the microphone array module is strongly bound to keyword-based location information, requiring repeated keyword triggering as the user moves, severely impacting work efficiency and posing a safety hazard due to the inability to promptly eliminate faults. Furthermore, substations contain numerous high-voltage electrical devices (such as transformers, circuit breakers, and disconnectors), and faults in these devices can lead to serious safety accidents. Related technologies only focus on distinguishing between human voices and environmental noise, failing to identify abnormal noises from equipment faults, let alone accurately determine the type and specific fault mode of the faulty equipment. This makes it difficult for maintenance personnel to quickly locate problems and take countermeasures, posing a significant safety risk.
[0022] Related technologies can propose an improved solution for tracking users using two-dimensional imaging equipment, but this solution can only acquire two-dimensional angle information, which is not accurate enough and cannot monitor distance. It also fails to recognize users when they are outside the sound pickup range.
[0023] In view of this, this application proposes an inspection method. On the one hand, by detecting the sound source information and location information of dynamically moving target personnel, the beamforming direction and location information of the audio input module are adjusted to achieve dynamic voice enhancement following the target personnel, ensuring the continuity of voice command recognition during operation. On the other hand, based on the adjustment of the audio input module to follow the movement of the target personnel, the equipment topology, equipment sound source information, and target equipment can be dynamically updated. By comprehensively analyzing information under different states, the target equipment can be located more accurately, thereby more accurately determining the fault type. The combination of these two methods enables target personnel to perform maintenance work more accurately and efficiently, improving the safety of the operating environment.
[0024] refer to Figure 1 This is a schematic diagram illustrating an application scenario of the inspection method provided in this application embodiment. The application scenario includes a camera 101, an audio input module 102, and a processing module 103. The camera 101, audio input module 102, and processing module 103 can all be connected via wired or wireless communication networks to achieve real-time communication.
[0025] The imaging device 101 can be a two-dimensional or three-dimensional imaging device. Two-dimensional imaging devices may include, but are not limited to, monocular cameras, industrial cameras, etc., while three-dimensional imaging devices may include, but are not limited to, depth cameras, binocular stereo vision cameras, LiDAR, etc. For example, the imaging device 101 may be an industrial-grade RGB-D type, supporting skeletal recognition within 20 meters, with a frame rate exceeding 30fps, and possessing dustproof and electromagnetic interference protection functions, adapting to the harsh environment of substations. The imaging device 101 can be used to track target personnel and obtain their location information.
[0026] The audio input module 102 may include at least one audio input unit, which may include, but is not limited to, devices such as microphones, pickup arrays, audio receiver interfaces, or communication modules. The audio input module 102 can be used to collect target voiceprint information, which may come from the device or from the target person.
[0027] In this embodiment, the number and configuration of audio input units are not limited. In one exemplary embodiment, multiple audio input units can be arranged in a ring configuration. For example, multiple audio input units can be distributed on multiple parallel rings with the same radius or different heights, or multiple audio input units can be arranged on the same plane but on concentric circles with different radii. In a scenario using a ring configuration of 8 audio input units, the sampling rate can be set to exceed 48kHz, the time delay measurement accuracy can be set to not exceed 1μs, the pickup distance can be set to 6 meters, and it can have a design to resist high-voltage electromagnetic interference. It should be noted that this is only an example and is not limited to these setting parameters.
[0028] This application embodiment does not limit the positional relationship between the audio input module 102 and the shooting device 101. In one exemplary embodiment, the audio input module 102 can be arranged adjacent to the shooting device 101. For example, the distance between the shooting device 101 and the audio input module 102 can be less than or equal to 10cm, and the installation height can be 1.5-2.0 meters above the ground to ensure the coverage of user tracking and sound source acquisition. The shooting device 101 can be set with a spatial reference point (0, 0, 0), and its central axis can be parallel to the ground (X=0, Y=0); the audio input module 102 can be placed horizontally, and its central coordinates (X2, Y2, Z2) are determined by hardware design.
[0029] The processing module 10 can be used to execute an inspection method based on the location information of the target personnel acquired by the shooting device 101 and the target voiceprint information acquired by the audio input module 102. The inspection method may include: in response to the target voiceprint information received by the audio input module being device sound, obtaining target feature information based on the spectral and temporal characteristics of the target voiceprint information; obtaining device sound source information based on the location information of the audio input module and the target voiceprint information; obtaining the target device based on the device topology, device location information, and device sound source information; the device topology is constructed from candidate devices determined based on the target feature information; obtaining the target fault type based on the target feature information and the fault feature information of multiple fault types of the target device; and in response to the target voiceprint information received by the audio input module being instruction information of the target personnel, performing maintenance on the target device according to the target fault type.
[0030] In this embodiment, the processing module 103 can also dynamically adjust the beamforming direction and position information of the audio input module 102 based on the target person's voice source information and location information. Based on the adjusted beamforming direction and position information of the audio input module 102, the aforementioned inspection method is executed.
[0031] In one alternative embodiment, the processing module 103 may integrate an artificial intelligence (AI) model to handle computational tasks in the inspection method. In a preferred exemplary embodiment, the processing module 103 may include a neural network processing unit (NPU) that can support real-time voiceprint feature extraction and hybrid model computation with high performance.
[0032] In an optional embodiment, the processing module 103 may further include a dual-layer voiceprint feature database, which is communicatively connected to the audio input module 102 to receive target voiceprint information. The dual-layer voiceprint feature database may include a first-layer voiceprint feature database and a second-layer voiceprint feature database. The first-layer voiceprint feature database may include a device type feature database and a human voice semantic feature database, which can be used to determine the type of target voiceprint information acquired by the audio input unit. Based on the device type feature database, it can be determined whether the target voiceprint information belongs to device sound; based on the human voice semantic feature database, it can be determined whether the target voiceprint information belongs to instruction information issued by personnel. In an exemplary embodiment, the device type feature database may cover core equipment of the substation, such as transformers, circuit breakers, disconnect switches, GIS equipment, instrument transformers, etc., with each type of equipment having more than 8,000 samples, containing operating voiceprint information under different working conditions. The second-layer voiceprint feature database may include a fault type feature database, which can be used to determine the fault type of the equipment. The fault type feature database can store multiple fault feature information for multiple fault types of multiple devices. Each device can correspond to multiple fault types, and each fault type can correspond to multiple fault feature information. In an exemplary embodiment, each type of device in the fault type feature database can correspond to 5-8 typical faults. For example, Geographic Information System (GIS) devices can include faults such as partial discharge, gas leakage, and poor conductor contact. Each type of fault sample can exceed 5000, covering voiceprint information of different fault degrees. In this embodiment, the storage space of the processing module 103 is not limited; exemplaryly, the storage space can exceed 32GB.
[0033] In an optional embodiment, the application scenario may further include a storage unit, which is communicatively connected to the processing module 103 to interact with it. For example, the processing module 103 can retrieve the required data from the storage unit or store the processed results within it. The storage unit can store parameter information for all power equipment within the substation, including but not limited to equipment location information, equipment model, rated operating parameters, historical fault data, and fault handling strategies. This application does not limit the implementation of the storage unit. In an exemplary embodiment, the storage unit may use an industrial-grade Flash memory chip with a capacity exceeding 16GB, supporting batch import and update of equipment parameters and encrypted storage of historical data.
[0034] In an optional embodiment, the processing module 103 can generate a fault handling log based on the operation of the target personnel in handling the fault, and update the device parameter information of the dual-layer voiceprint feature database and / or storage unit based on the fault handling log.
[0035] In this embodiment of the application, the shooting device 101, the audio input module 102 and the processing module 103 may be integrated into the inspection device, but are not limited thereto.
[0036] In an optional embodiment, the application scenario may also include a terminal device, which can be communicatively connected to the processing module 103 and can obtain data detected by the audio input module and the shooting device, intermediate data, and final processing results from the processing module 103. For example, the terminal device can obtain parameter information of the target device or the target fault type for the target personnel to retrieve.
[0037] In an optional embodiment, the target user can send command information to the processing module 103 via the audio input module 102, in which case the type of command information is voice command information. Alternatively, the target user can send command information to the processing module 103 via a terminal device, in which case the type of command information is operation command information. For example, after troubleshooting, the processing module 103 can, in response to the operation command information received by the terminal device or the voice command information received by the audio input module 102, reset the state of the target device from a fault state to a warning state.
[0038] The following is combined with Figure 1 The application scenarios described above illustrate the inspection method according to exemplary embodiments of this application. It should be noted that the above application scenarios are merely shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application can be applied to any applicable scenario.
[0039] like Figure 2 As shown in the embodiment of this application, an inspection method is provided, applied to an inspection device. The inspection device includes an audio input module. The beamforming direction and position information of the audio input module can be dynamically adjusted according to the sound source information and position information of the target personnel. The inspection method may include: S1. In response to the target voiceprint information received by the audio input module being the device sound, extract the feature information of the target voiceprint information to obtain the target feature information.
[0040] S2. Obtain the device sound source information based on the location information of the audio input module and the target voiceprint information.
[0041] S3. Based on the device topology, device location information, and device sound source information, the target device is obtained; the device topology is constructed from the candidate devices determined based on the target feature information.
[0042] S4. Based on the target feature information and the fault feature information of multiple fault types of the target equipment, the target fault type is obtained.
[0043] S5. Responding to the target voiceprint information received by the audio input module as instruction information, the target device is inspected and repaired according to the target fault type.
[0044] In this embodiment, the target personnel can refer to objects located within a preset geographical space of the target substation. The target substation can refer to any substation, and the preset geographical space can be customized according to the actual physical boundaries of different substations. In an optional embodiment, to achieve refined management and control, an intelligent identification mechanism can be introduced: when an object enters the preset geographical space, a facial recognition algorithm is triggered to verify its identity and determine whether it is an authorized worker. For personnel who pass the verification, their movement trajectory can be continuously tracked, thereby ensuring the safe and efficient conduct of maintenance work.
[0045] Target voiceprint information can refer to the sound signals emitted by equipment or target personnel collected by the audio input module within the target substation. Target feature information can refer to the feature vector extracted from the target voiceprint information using a neural network to determine the target fault type. In an exemplary embodiment, the audio input module can collect omnidirectional voiceprint information within the substation in real time, which may include operator voices, equipment operating sounds, environmental noise, etc., and can be synchronously transmitted to the processing module. The processing module can extract the feature information of the target voiceprint information and compare it with the first-layer voiceprint feature database to determine the category of the target voiceprint information, or it can use the database itself. If it matches the human voice semantic feature database, it can be determined as the instruction information of the target personnel. The voice enhancement effect can be strengthened by combining location information to ensure the accuracy of instruction recognition in noisy environments. If it matches the equipment type feature database, it can determine the type of equipment that emitted the target voiceprint information, and then further identification can be performed based on the second-layer voiceprint feature database. In the embodiments of this application, the extracted feature information can be applied differently, and the neural network used can also be different. Therefore, the feature information extracted for the purpose of determining the type of target voiceprint information is not necessarily the same as the target feature information.
[0046] In this embodiment, the beamforming direction and position information of the audio input module can be dynamically adjusted according to the target person's voice source information and position information, that is, it dynamically changes as the target person moves. The beamforming direction refers to the directional receiving beam formed by signal superposition in the spatial sound field through adjusting the phase delay and signal gain of each audio input unit in the audio input module. This beamforming direction can be used to enhance the target voiceprint information, suppress noise interference, and ensure that the command information issued by the target user during movement can be clearly identified. The voice source information refers to the calculated spatial position of the target person emitting the target voiceprint information.
[0047] Based on the location information of the audio input module and the target voiceprint information, the device sound source information can be obtained. Device sound source information refers to the calculated spatial location of the device emitting the target voiceprint information. The target device refers to the device emitting the target voiceprint information. Due to noise and other reasons, the target device may not be directly determined based on the device sound source information, requiring further correction. Therefore, this embodiment obtains the target device based on the device topology, device location information, and device sound source information. The device topology refers to the topology constructed by multiple candidate devices in the target substation. Candidate devices refer to devices whose fault feature information matches the target feature information. Fault feature information refers to the feature information corresponding to a possible fault in the device. It should be noted that matching the target feature information refers to fault feature information whose similarity to the target feature information is greater than a preset similarity threshold, which can be set according to specific circumstances.
[0048] Because the beamforming direction and position information of the audio input module can be dynamically adjusted to follow the target person, this dynamic adjustment may cause differences in the received voiceprint information caused by the same fault in the same device. Consequently, the constructed device topology and device sound source information may also differ, leading to variations in the obtained target device. By comprehensively analyzing multiple device topologies, multiple device sound source information, and multiple target devices obtained by the audio input module following the target person's movement, the final target device is obtained. This application does not limit the implementation method of the comprehensive analysis. For example, based on the distance between the audio input module and the device sound source information, different weights can be assigned to target devices obtained from different device topologies and different device sound source information, and then the weights can be combined to obtain the final target device. Alternatively, the number of times the same candidate device is used as a target device can be counted, and the target device exceeding a preset number of times can be used as the final target device.
[0049] After identifying the target device, the target fault type can be obtained based on the target feature information and the fault feature information of multiple fault types of the target device. The target fault type can refer to the fault type corresponding to the fault feature information in the target device that matches the target feature information. In an exemplary embodiment, based on the device type determined by the first-layer voiceprint feature database, the fault type feature library corresponding to the device can be called, such as the fault feature sub-library for transformers (core grounding fault, winding short circuit, abnormal oil level), and the fault feature sub-library for circuit breakers (arc chamber fault, operating mechanism jamming), etc., for secondary precise comparison: if it matches the normal operation feature sub-library, it can be determined as harmless background noise, and the current configuration can be maintained; if it matches a certain fault type feature sub-library, for example, the matching degree exceeds 75%, it can be determined as that type of fault; if it does not match any feature sub-library, but the difference from normal features exceeds 70%, it can be determined as an unknown fault.
[0050] After obtaining the target fault type, the target device can be repaired based on the target voiceprint information received by the audio input module, in response to the target fault type.
[0051] In this embodiment, on the one hand, by detecting the sound source and location information of a dynamically moving target person, the beamforming direction and location information of the audio input module are adjusted to achieve dynamic voice enhancement following the target person, ensuring the continuity of voice command recognition during operations. On the other hand, based on the adjustment of the audio input module to follow the target person's movement, the equipment topology, equipment sound source information, and target equipment can be dynamically updated. By comprehensively analyzing information under different states, the target equipment can be located more accurately, thereby more accurately determining the fault type. The combination of these two aspects enables the target person to perform maintenance work more accurately and efficiently, improving the safety of the operating environment.
[0052] In an optional embodiment, extracting feature information from the target voiceprint information to obtain target feature information may include: converting the target voiceprint information into spectrum data; using a first feature extraction module to extract features from the spectrum data within a local time-frequency range to obtain spectrum features; using a second feature extraction module to extract temporal dependencies from the spectrum features to obtain temporal features; using an attention mechanism to calculate the correlation degree of each frame of feature data in the temporal features relative to the target device; obtaining weight information based on the correlation degree; and performing weighted summation processing on each frame of feature data in the temporal features based on the weight information to obtain target feature information.
[0053] Spectral data refers to the distribution of sound energy along the time and frequency axes. The first feature extraction module, based on a deep learning unit using a Convolutional Neural Network (CNN), extracts fine-grained features of the spectral data within a local time-frequency range through sliding scans of the convolutional kernels, yielding spectral features such as the pulse peaks of partial discharge sounds or high-frequency fluctuations of mechanical friction. The second feature extraction module employs a Long Short-Term Memory (LSTM) network to capture the dynamic dependencies and evolutionary patterns of spectral features over long time, obtaining temporal features, such as the complete process of a device fault gradually evolving from a faint abnormal sound to a continuous loud noise. Based on this, an attention mechanism is introduced to dynamically weight the extracted temporal features, focusing on and strengthening the key time nodes and feature channels with the highest fault-identifying value, while suppressing interference from redundant information such as environmental background noise. By deeply integrating the local perception capability of the first feature extraction module, the temporal modeling capability of the second feature extraction module, and the feature optimization capability of the attention mechanism, the accuracy of power equipment fault classification can be improved, achieving an accuracy rate exceeding 95%.
[0054] In one optional embodiment, obtaining device sound source information based on the location information of the audio input module and the target voiceprint information may include: the audio input module comprising multiple audio input units arranged in a ring; determining the location information of each audio input unit; calculating the physical distance and time difference based on the location information of different audio input units and the time of receiving the target voiceprint information; obtaining the distance difference between the sound source and different audio input units based on the time difference and the speed of sound; calculating the horizontal azimuth and vertical elevation angle based on the distance difference and the physical distance; and obtaining device sound source information based on the physical distance, the horizontal azimuth, and the vertical elevation angle.
[0055] In this embodiment, the horizontal azimuth angle refers to the angle between the projection of the sound source onto the horizontal plane and the reference direction; the vertical elevation angle refers to the angle between the sound source and the horizontal plane. The sound source information of the device can be calculated using the time delay difference localization method. In an optional embodiment, the time delay difference localization method can be implemented using the Generalized Cross Correlation-Phase Transform (GCC-PHAT) algorithm, the least squares method, and / or the Chan algorithm.
[0056] In some special deployments, the above algorithm can be simplified. Taking an audio input module comprising eight audio input units arranged in a ring at the same height and symmetrically distributed as an example: any four audio input units can be selected to form a positive or rectangular shape. Using the orthogonal time delay obtained from the time these four audio input units receive voiceprint information, the device's sound source information can be calculated. Its expression can be:
[0057]
[0058]
[0059] Where t1 represents the time when the reference audio input unit receives the voiceprint information, t2, t3, and t4 represent the times when the other three audio input units receive the voiceprint information, v represents the speed of sound, θ represents the horizontal azimuth angle, and α represents the vertical elevation angle. It can represent the horizontal axis in the sound source information of the device. It can represent the vertical coordinate in the sound source information of the device. It can represent the height coordinates in the sound source information of the device.
[0060] In one optional embodiment, obtaining the target device based on the device topology, device location information, and device sound source information may include: determining multiple candidate devices from multiple devices based on target feature information; adjusting the positioning algorithm parameters based on the device topology formed by the multiple candidate devices; and obtaining the target device based on the positioning algorithm parameters, device location information, and device sound source information.
[0061] In this embodiment, factors such as environmental noise, multipath effects, and errors in the audio input unit can cause deviations between the device's sound source information and its actual spatial location. The parameters that cause these deviations are referred to as influencing factors. The positioning algorithm parameters refer to the parameters set to suppress or compensate for these influencing factors. Positioning algorithm parameters may include, but are not limited to, location signal thresholds, multipath / reflection suppression coefficients, hyperbola intersection thresholds, Chan / Taylor iteration convergence thresholds, device matching radius, and / or outlier removal thresholds.
[0062] The positioning reliability threshold can refer to the minimum standard for measuring the reliability of the positioning result, and may include a distance error threshold and a coordinate matching threshold. The multipath / reflection suppression coefficient can refer to a parameter used to reduce interference from multipath effects caused by reflections from objects such as walls, the ground, and equipment. The hyperbola intersection judgment threshold can refer to the maximum permissible error range used to determine whether multiple hyperbolas truly intersect at a point when using time delay differences to construct hyperboloids for positioning; exceeding this range indicates that they cannot effectively intersect. In an exemplary embodiment, the hyperbola intersection judgment threshold can be the Chan / Taylor iteration convergence threshold. The Chan / Taylor iteration convergence threshold can refer to a small error standard used to determine whether the calculation is sufficiently accurate and can stop further calculations and output the final result when using iterative algorithms such as Chan or Taylor to back-calculate coordinates. The device matching radius can refer to the radius of a circular area centered on the device's sound source information, within which the corresponding device is searched and matched. The outlier removal threshold can refer to the boundary value of outliers in the data, which can be used to eliminate jumps caused by electromagnetic interference. In the embodiments of this application, different positioning algorithm parameter templates can be constructed for different device topologies.
[0063] By utilizing positioning algorithm parameters, device location information, and device sound source information, the sound source is located to one of multiple candidate devices to obtain the target device. For example, the positioning algorithm parameters may include a device matching radius of 0.3m or 0.5m. Based on the device location information and device sound source information, a distance difference is calculated; devices with a distance difference smaller than the device matching radius can be used as target devices. In this embodiment, by dynamically adapting the positioning algorithm parameters to the device topology, the accuracy of positioning can be improved.
[0064] In one optional embodiment, dynamically adjusting the beamforming direction and position information of the audio input module based on the target person's voice source information and location information may include: obtaining the person's voice source information based on the target person's location information and the audio input module's position information; the voice source information includes the voice source direction angle, actual distance, and horizontal direction; updating the beamforming direction of the audio input module in response to the voice source direction angle when the actual distance does not exceed the pickup distance; obtaining the movement distance based on the actual distance and the pickup distance when the actual distance exceeds the pickup distance; and controlling the audio input module to move towards the target user based on the movement distance and the horizontal direction, thereby updating the position information of the audio input module.
[0065]
[0066]
[0067]
[0068] Where C represents the direction angle of the sound source, X1 represents the x-coordinate of the target person's position information, Y1 represents the y-coordinate of the target person's position information, Z1 represents the height coordinate of the target person's position information, X2 represents the x-coordinate of the audio input module's position information, Y2 represents the y-coordinate of the audio input module's position information, Z2 represents the height coordinate of the audio input module's position information, L represents the actual distance, and D represents the horizontal direction. In an exemplary embodiment, the target person's position information may be the head coordinates of the target person.
[0069] When the target user moves, the camera can update the target user's location information in real time, requiring recalculation of the sound source direction angle C, actual distance L, and horizontal direction D parameters. If the actual distance L ≤ the substation scene's sound pickup distance S, the sound source direction angle C can be directly sent to the audio input module to update the beamforming direction and maintain the voice enhancement effect. If the actual distance L > the sound pickup distance S, the horizontal direction D and the moving distance (i.e., the difference SL between the sound pickup distance and the actual distance) can be sent to the inspection device to control the audio input module to move towards the target user until the actual distance L ≤ the sound pickup distance S, at which point the voice enhancement direction is updated. In an exemplary embodiment, the sound pickup distance can be set to 5-8 meters to adapt to the open environment of a substation.
[0070] In an optional embodiment, when the actual distance does not exceed the pickup distance, updating the beamforming direction of the audio input module in response to the sound source direction angle may include: determining the guide vector of the audio input module in response to the sound source direction angle; calculating the channel weight of each audio input unit based on the guide vector; and updating the beamforming direction of the audio input module based on the channel weight of each audio input unit.
[0071] In this embodiment, the guide vector can refer to the set of time delays generated by each of the other audio input units in the audio input module relative to the reference audio input unit when the voiceprint information arrives at the audio input unit in sequence.
[0072] This application does not limit the implementation method of calculating the channel weights of each audio input unit based on the steering vector. In one exemplary embodiment, corresponding channel weights can be configured for each microphone channel based on the phase difference information represented by the steering vector. These channel weights can be used to compensate for the signal transmission delay between different channels, enabling phase alignment and energy enhancement of the acoustic signature information in the target direction during coherent superposition. In another exemplary embodiment, the weights can be calculated using a minimum variance distortionless response adaptive beamforming algorithm.
[0073] In an optional embodiment, when the target voiceprint information received by the audio input module is an instruction message, and the target device is repaired according to the target fault type, the process may further include: determining a fault priority based on the importance of the target device and the target fault type; generating a fault diagnosis report based on the type of the target device, the target fault type, the device's sound source information, the fault priority, and the fault duration; obtaining historical fault data and corresponding fault handling strategies based on the fault diagnosis report; correspondingly, when the instruction message with the voiceprint type being a target person is received, and the target device is repaired according to the target fault type, the process may include: when the instruction message with the voiceprint type being a target person is received, repairing the target device according to the fault diagnosis report, historical fault data, and fault handling strategies.
[0074] In this embodiment, the target personnel can trigger the audio input module via keywords to activate voice enhancement and voiceprint recognition functions. The storage unit can load parameters of the substation's electrical equipment, historical fault data, and fault handling strategies. The processing module can determine the fault priority based on the importance of the target equipment and the type of the target fault. In an exemplary embodiment, the risk level assessment can be based on the importance of the target equipment (e.g., the main transformer is a core device) and the type of the target fault (e.g., short-circuit faults and discharge faults are high-risk), setting three levels of fault priority: high, medium, and low, to support differentiated early warning responses. The processing module can generate a fault diagnosis report based on the type of the target equipment, the type of the target fault, the equipment's sound source information, the fault priority, and the fault duration. For example, the fault diagnosis report can be: Warning: A winding short-circuit fault has occurred in main transformer No. 3, location coordinates (15.6m, 8.2m, 3.5m), risk level: high. The target equipment type can refer to parameters such as equipment model. The fault diagnosis report can be output through the equipment display screen and / or voice module, and can synchronously trigger the substation monitoring center alarm system. When the fault duration exceeds the preset duration (e.g., 3 seconds), the fault diagnosis report, voiceprint information, historical fault data and fault handling strategy can be automatically sent to the substation operation and maintenance management platform and the terminal equipment associated with the target personnel. The target personnel can retrieve the historical fault data of the equipment and similar fault handling strategies through voice commands or platform operation, compare the voiceprint changes before and after the fault, and assist in quickly formulating maintenance plans.
[0075] To address the challenges of repeated keyword triggering, limited pickup distance, and inability to identify abnormal noises from substation equipment during user movement in substation scenarios, this application provides a method for user tracking, voice enhancement, and fault voiceprint recognition within substations. The method utilizes a camera to track the target user's three-dimensional position in real-time, dynamically updates the voice enhancement direction, and controls equipment movement via distance detection. Furthermore, it incorporates voiceprint recognition and fault classification algorithms to accurately distinguish between human voice semantics and substation equipment operating sounds. This enables dual identification and location of faulty substation equipment types and fault types, providing early warnings of equipment safety risks and meeting the triple operational requirements of continuous voice interaction, environmental safety monitoring, and accurate fault diagnosis within substations.
[0076] In summary, the embodiments of this application can achieve the following beneficial effects: Adapted to specific needs of substations: In response to the characteristics of dense equipment, noisy environment and wide operating range in substations, it can realize continuous voice recognition for mobile users without the need to repeatedly trigger keywords, and improve work efficiency by more than 40%.
[0077] Accurate and efficient fault diagnosis: The fault sound source location error can be no more than 0.3 meters, the equipment matching accuracy rate can exceed 98%, and the fault type identification accuracy rate can exceed 95%. Compared with traditional manual inspection, the fault detection response time is shortened to the second level, which greatly reduces safety risks.
[0078] High degree of functional integration: It integrates five major functions, including voice enhancement, user tracking, fault identification, risk warning, and data traceability, replacing multiple single devices and reducing the cost of intelligent transformation of substations.
[0079] Highly regarded for safety and practicality: It features industrial-grade design with anti-electromagnetic interference and dustproof features, making it suitable for harsh substation environments. It can also provide fault case matching and historical fault data comparison functions to assist personnel in quick repairs and improve the quality of operation and maintenance.
[0080] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0081] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0082] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides an inspection device, which may include: The camera can be used to track target personnel and determine their location information.
[0083] An audio input module can be used to receive target voiceprint information.
[0084] The processing module can be used to adjust the beamforming direction and position information of the audio input module according to the personnel voice source information and location information of the target personnel; in response to the target voiceprint information received by the audio input module being device sound, extract the feature information of the target voiceprint information to obtain target feature information; obtain device sound source information according to the position information of the audio input module and the target voiceprint information; obtain the target device according to the device topology, device location information and the device sound source information; the device topology is constructed based on the candidate devices determined by the target feature information; obtain the target fault type according to the target feature information and the fault feature information of multiple fault types of the target device; and in response to the target voiceprint information received by the audio input module being instruction information of the target personnel, perform maintenance on the target device according to the target fault type.
[0085] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0086] The apparatus described above is used to implement the corresponding inspection method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0087] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the inspection method described in any of the above embodiments.
[0088] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0089] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0090] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0091] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0092] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0093] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0094] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0095] The electronic devices described above are used to implement the corresponding inspection methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0096] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the inspection method as described in any of the above embodiments.
[0097] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0098] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the inspection method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0099] It should be noted that the embodiments of this application can also be further described in the following ways: An inspection method is applied to an inspection device, which includes an audio input module. The beamforming direction and position information of the audio input module are dynamically adjusted based on the voice source information and position information of the target personnel. The inspection method includes: responding to the target voiceprint information received by the audio input module as device sound, extracting feature information of the target voiceprint information to obtain target feature information; obtaining device sound source information based on the position information of the audio input module and the target voiceprint information; obtaining a target device based on the device topology, device position information, and device sound source information; the device topology is constructed based on candidate devices determined by the target feature information; obtaining a target fault type based on the target feature information and fault feature information of multiple fault types of the target device; and responding to the target voiceprint information received by the audio input module as instruction information, performing maintenance on the target device according to the target fault type.
[0100] Optionally, extracting feature information from the target voiceprint information to obtain target feature information includes: converting the target voiceprint information into spectral data; using a first feature extraction module to extract features from the spectral data within a local time-frequency range to obtain the spectral features; using a second feature extraction module to extract temporal dependencies from the spectral features to obtain the temporal features; using an attention mechanism to calculate the correlation degree of each frame of feature data in the temporal features relative to the target device; obtaining weight information based on the correlation degree; and performing weighted summation processing on each frame of feature data in the temporal features based on the weight information to obtain the target feature information.
[0101] Optionally, obtaining device sound source information based on the location information of the audio input module and the target voiceprint information includes: the audio input module comprising multiple audio input units arranged in a ring; determining the location information of each audio input unit; calculating the physical distance and time difference based on the location information of different audio input units and the time of receiving the target voiceprint information; obtaining the distance difference between the sound source and different audio input units based on the time difference and the speed of sound; calculating the horizontal azimuth and vertical elevation angle based on the distance difference and the physical distance; and obtaining the device sound source information based on the physical distance, the horizontal azimuth, and the vertical elevation angle.
[0102] Optionally, obtaining the target device based on the device topology, device location information, and device sound source information includes: determining multiple candidate devices from multiple devices based on the target feature information; adjusting the positioning algorithm parameters based on the device topology formed by the multiple candidate devices; and obtaining the target device based on the positioning algorithm parameters, the device location information, and the device sound source information.
[0103] Optionally, the method further includes: obtaining the person's sound source information based on the location information of the target person and the location information of the audio input module; the person's sound source information includes the sound source direction angle, actual distance, and horizontal direction; if the actual distance does not exceed the pickup distance, updating the beamforming direction of the audio input module in response to the sound source direction angle; if the actual distance exceeds the pickup distance, obtaining the movement distance based on the actual distance and the pickup distance; and controlling the audio input module to move towards the target user based on the movement distance and the horizontal direction, thereby updating the location information of the audio input module.
[0104] Optionally, if the actual distance does not exceed the pickup distance, updating the beamforming direction of the audio input module in response to the sound source direction angle includes: the audio input module comprising multiple audio input units; determining the guide vector of the audio input module in response to the sound source direction angle; calculating the channel weight of each audio input unit based on the guide vector; and updating the beamforming direction of the audio input module based on the channel weight of each audio input unit.
[0105] Optionally, the method further includes: determining a fault priority based on the importance of the target device and the target fault type; generating a fault diagnosis report based on the type of the target device, the target fault type, the sound source information of the device, the fault priority, and the fault duration; obtaining historical fault data and corresponding fault handling strategies based on the fault diagnosis report; and correspondingly, in response to an instruction from a person whose voiceprint type is that of the target personnel, repairing the target device based on the target fault type, including: in response to an instruction from a person whose voiceprint type is that of the target personnel, repairing the target device based on the fault diagnosis report, the historical fault data, and the fault handling strategy.
[0106] An inspection device includes: a camera for tracking a target person and determining the location information of the target person; an audio input module for receiving target voiceprint information; a processing module for adjusting the beamforming direction and position information of the audio input module based on the voice source information and location information of the target person; in response to the target voiceprint information received by the audio input module being device sound, extracting feature information of the target voiceprint information to obtain target feature information; obtaining device sound source information based on the location information of the audio input module and the target voiceprint information; obtaining a target device based on the device topology, device location information, and device sound source information; the device topology is constructed from candidate devices determined based on the target feature information; obtaining a target fault type based on the target feature information and fault feature information of multiple fault types of the target device; and in response to the target voiceprint information received by the audio input module being instruction information of the target person, performing maintenance on the target device according to the target fault type.
[0107] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the inspection method.
[0108] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the inspection method.
[0109] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0110] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0111] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0112] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. An inspection method applied to an inspection device, the inspection device including an audio input module, wherein the beamforming direction and position information of the audio input module are dynamically adjusted based on the sound source information and position information of the target personnel, including: In response to the target voiceprint information received by the audio input module being device sound, feature information of the target voiceprint information is extracted to obtain target feature information; Based on the location information of the audio input module and the target voiceprint information, the device sound source information is obtained; The target device is obtained based on the device topology, device location information, and device sound source information; the device topology is constructed from candidate devices determined based on the target feature information. Based on the target feature information and the fault feature information of multiple fault types of the target device, the target fault type is obtained; In response to the target voiceprint information received by the audio input module as instruction information, the target device is inspected and repaired according to the target fault type.
2. The method according to claim 1, characterized in that, Extracting the feature information of the target voiceprint information to obtain target feature information includes: The target acoustic signature information is converted into spectral data; The first feature extraction module is used to extract features from the local time-frequency range of the spectrum data to obtain the spectrum features; The temporal dependencies of the spectral features are extracted using the second feature extraction module to obtain the temporal features; Using an attention mechanism, the correlation degree of each frame feature data in the temporal features relative to the target device is calculated; weight information is obtained based on the correlation degree. Based on the weight information, the feature data of each frame in the temporal features are weighted and summed to obtain the target feature information.
3. The method according to claim 1, characterized in that, Based on the location information of the audio input module and the target voiceprint information, the device sound source information is obtained, including: The audio input module includes multiple audio input units, which are arranged in a ring. Determine the position information of each audio input unit; The physical distance and time difference are calculated based on the location information of different audio input units and the time of receiving the target voiceprint information; Based on the time difference and the speed of sound, the distance difference between the sound source and different audio input units is obtained; The horizontal azimuth and vertical elevation angles are calculated based on the distance difference and the physical distance. The sound source information of the device is obtained based on the physical distance, the horizontal azimuth angle, and the vertical elevation angle.
4. The method according to claim 1, characterized in that, Based on the device topology, device location information, and the device sound source information, the target device is obtained, including: Based on the target feature information, multiple candidate devices are determined from multiple devices; Based on the device topology structure formed by the multiple candidate devices, adjust the positioning algorithm parameters; The target device is obtained based on the positioning algorithm parameters, the device location information, and the device sound source information.
5. The method according to claim 1, characterized in that, The method further includes: Based on the location information of the target person and the location information of the audio input module, the person's sound source information is obtained; the person's sound source information includes the sound source direction angle, actual distance, and horizontal direction; If the actual distance does not exceed the pickup distance, the beamforming direction of the audio input module is updated in response to the sound source direction angle. If the actual distance exceeds the pickup distance, the movement distance is obtained based on the actual distance and the pickup distance; based on the movement distance and the horizontal direction, the audio input module is controlled to move towards the target user, and the position information of the audio input module is updated.
6. The method according to claim 5, characterized in that, When the actual distance does not exceed the pickup distance, the beamforming direction of the audio input module is updated in response to the sound source direction angle, including: The audio input module includes multiple audio input units; In response to the direction angle of the sound source, the guide vector of the audio input module is determined; Based on the steering vector, calculate the channel weight of each audio input unit; The beamforming direction of the audio input module is updated based on the channel weights of each audio input unit.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Based on the importance of the target equipment and the type of target fault, the fault priority is determined; A fault diagnosis report is generated based on the type of the target device, the type of the target fault, the sound source information of the device, the fault priority, and the fault duration. Based on the fault diagnosis report, historical fault data and corresponding fault handling strategies are obtained; Accordingly, in response to the instruction information whose voiceprint type is the target person, the target equipment is repaired according to the target fault type, including: In response to the instruction information that the voiceprint type is the target person, the target equipment is repaired according to the fault diagnosis report, the historical fault data and the fault handling strategy.
8. An inspection device, comprising: A camera is used to track a target person and determine the target person's location information; An audio input module is used to receive target voiceprint information; The processing module is used to adjust the beamforming direction and position information of the audio input module according to the voice source information and location information of the target person; in response to the target voiceprint information received by the audio input module being device sound, it extracts the feature information of the target voiceprint information to obtain target feature information; it obtains device sound source information according to the position information of the audio input module and the target voiceprint information; and it obtains the target device according to the device topology, device location information, and device sound source information; the device topology is constructed from candidate devices determined based on the target feature information. Based on the target feature information and the fault feature information of multiple fault types of the target device, the target fault type is obtained; in response to the target voiceprint information received by the audio input module being the instruction information of the target personnel, the target device is repaired according to the target fault type.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1 to 7.