Intelligent hidden danger investigation joint inspection method based on dynamic visual calculation

By sharing location information and visual sensors in the mobile inspection unit, trigger indications are generated, and the intensity of extended motion vectors and time changes is calculated and encoded into a three-dimensional hazard structure tensor field. This solves the problems of insufficient time resolution and data redundancy in the existing technology, and realizes efficient capture and adaptive inspection of high-speed transient hazards.

CN122491773APending Publication Date: 2026-07-31NANJING ANTUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING ANTUAN TECH CO LTD
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing visual perception methods, when processing dynamic visual data, lack sufficient temporal resolution to capture high-speed transient hazards, suffer from excessive data redundancy, lack an adaptive collaborative scheduling mechanism for inspection units, and lack a closed-loop correlation between task scheduling and hazard investigation results.

Method used

By deploying multiple mobile inspection units, sharing location information, and equipping them with visual sensors to perceive dynamic visual data, trigger indications are generated, the intensity of extended motion vectors and time changes is calculated, and encoded into a three-dimensional hazard structure tensor field. Iterative observation status updates are performed, and hazard identification and task scheduling are carried out based on the tensor field.

Benefits of technology

It achieves causal retrospective full-time domain capture of high-speed transient hazards, significantly compresses data volume, adaptively distributes inspection units to high-risk areas, and realizes continuous driving and closed-loop feedback of the hazard risk field on task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_112
    Figure SMS_112
  • Figure QLYQS_10
    Figure QLYQS_10
Patent Text Reader

Abstract

This invention discloses a joint inspection method for intelligent hazard investigation based on dynamic visual computing, relating to the fields of intelligent inspection and computer vision. It addresses the problems of insufficient capture capability for high-speed transient hazards and low efficiency of multi-unit collaboration in traditional inspection methods. This invention deploys multiple mobile inspection units. When a preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated, and the triggered dynamic visual fragment is extracted. The extended motion vector and its temporal change intensity are calculated, and the degree of disorder in the local motion direction distribution is statistically analyzed, encoded into a three-dimensional hazard structure tensor field. The extended observation state is iteratively updated using a composite gradient of the observation value metric and the distance factor to neighboring units, and the step size ratio is adjusted using a precondition matrix to form a hazard investigation list. Task switching is driven by continuous potential energy nodes with activation values ​​obtained from the tensor field in the behavior tree, achieving full-time domain capture of high-speed transient hazards, decentralized collaborative detailed inspection by multiple units, and dynamic updating of the list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent inspection and computer vision, specifically a joint inspection method for intelligent hidden danger investigation based on dynamic visual computing. Background Technology

[0002] In industrial production and infrastructure operation and maintenance, inspection is a crucial means of ensuring safe equipment operation and preventing accidents. Traditional inspection methods mainly rely on manual labor to perform inspection tasks according to fixed routes and cycles. Inspectors judge the operating status of equipment for abnormalities through visual observation and handheld testing instruments. With the development of sensing and robotics technologies, mobile inspection units have been gradually applied to various inspection scenarios, replacing manual labor in perceiving the scene and collecting data by being equipped with visual sensors.

[0003] Visual hazard detection is a core component of intelligent inspection. Common hazard types in industrial settings include electrical sparks, insulator flashovers, loose or detached components, pipeline corrosion and leaks, and structural crack propagation. These hazards often exhibit specific abnormal movement patterns during their occurrence, such as the instantaneous flashing of an electric arc, non-rigid sputtering from detached parts, and periodic abnormal vibrations caused by rotating machinery malfunctions. Capturing these hazards places high demands on the temporal resolution of visual perception. Conventional visual acquisition methods based on fixed frame rates present a trade-off between temporal resolution and data redundancy, making it difficult to simultaneously meet the requirements of high-speed transient capture and wide-area coverage.

[0004] The following problems exist in the existing technology: Conventional visual perception methods use fixed frame rate to acquire images, which is insufficient in terms of temporal resolution for microsecond-level high-speed transient hazards such as electrical sparks and parts falling off, and the constant high frame rate acquisition generates massive data redundancy. Existing methods, when processing dynamic visual data, only calculate the optical flow amplitude or inter-frame difference in the spatial plane, without incorporating the rate of change of motion information along the time dimension and the degree of disorder in local motion directions into a unified representation, resulting in insufficient ability to describe the spatiotemporal structure of potential hazards. Existing methods for collaborative scheduling of multiple mobile inspection units mostly adopt centralized path planning or preset task allocation, lacking a decentralized mechanism that adaptively adjusts the observation status and step size ratio of each unit based on the real-time distribution of hidden dangers and risks. Existing methods lack a closed-loop relationship between inspection task scheduling and hazard investigation results driven by a continuous mathematical field, and task switching relies on preset rules rather than real-time changes in the hazard risk field. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes an intelligent joint inspection method for hidden danger investigation based on dynamic visual computing, which is used to solve the above-mentioned technical problem.

[0006] The first aspect of this invention provides a joint inspection method for intelligent hazard detection based on dynamic visual computing, comprising the following steps: S1: Deploy multiple mobile inspection units, and share location information among the mobile inspection units; the mobile inspection units are equipped with vision sensors to continuously perceive the scene and collect dynamic visual data. When a segment that matches the preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated; according to the trigger indication, visual data of the corresponding time period and corresponding area is extracted from the collected dynamic visual data to form a trigger-based dynamic visual segment. S2: Calculate the extended motion vector and time change intensity of each spatiotemporal point in the triggered dynamic visual segment, and statistically analyze the disorder of the local motion direction distribution in the local neighborhood of each spatiotemporal point; encode the extended motion vector and time change intensity, as well as the disorder of the local motion direction distribution, into a three-dimensional hidden danger structure tensor field; S3: Calculate the observation value metric for the three-dimensional hidden danger structure tensor field, calculate the neighbor unit distance factor by combining the location information of the shared neighbor mobile inspection units, and use the composite gradient of the observation value metric and the neighbor unit distance factor as the update direction to iteratively update the extended observation state of the mobile inspection unit. During the update process, the step size ratio of different types of degrees of freedom in the extended observation state is adjusted by the precondition matrix, and the movement of the mobile inspection unit is controlled according to the extended observation state obtained by the iterative update. After the extended observation state changes, the three-dimensional hidden danger structure tensor field is updated by the newly acquired dynamic visual data. The hidden danger category is judged according to the updated three-dimensional hidden danger structure tensor field to form a hidden danger investigation list. S4: The mobile inspection unit performs inspection tasks according to a preset behavior tree, which includes continuous potential energy nodes and investigation task branches. The activation degree of the continuous potential energy nodes is obtained from the updated three-dimensional hidden danger structure tensor field. When the activation degree exceeds a preset priority threshold, the current inspection task is interrupted and the investigation task branch is entered. The dynamic visual data collected during the execution of the investigation task branch is used to update the three-dimensional hidden danger structure tensor field, thereby updating the hidden danger investigation list.

[0007] Preferably, in step S1, the mobile inspection unit is equipped with a vision sensor to continuously perceive the scene and collect dynamic visual data. When a segment matching a preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated, including the following steps: Dynamic visual data includes an asynchronous impulse event stream output from a high temporal resolution dynamic change perception channel and a frame image stream output from a high spatial resolution texture detail perception channel. A spatiotemporal pulse pattern detector corresponding to each preset abnormal motion pattern is run on the asynchronous pulse event stream with a preset sliding time window. When any detector successfully matches within the current sliding time window, it is determined that a segment that conforms to the preset abnormal motion pattern has been detected. A trigger indication is generated based on the abnormal motion pattern type corresponding to the successfully matched detector, the timestamp of the matching time, and the spatial neighborhood coordinate range covered by the successfully matched event pulse cluster.

[0008] Preferably, in step S1, extracting visual data for the corresponding time period and region from the acquired dynamic visual data according to the trigger indication to form a trigger-based dynamic visual segment includes the following steps: Using the timestamp carried by the trigger indication as the retrieval benchmark, extract the frame image sequence from the first preset time point before the trigger time to the second preset time point after the trigger time from the acquired frame image stream; Using the spatial neighborhood coordinate range carried by the trigger indication as a constraint, extract the local asynchronous pulse event stream within the corresponding spatiotemporal range from the collected asynchronous pulse event stream; The local asynchronous pulse event stream is spatiotemporally aligned with the frame image sequence, and motion compensation parameters are calculated using the successfully matched event pulse clusters carried by the trigger indicator. Based on the motion compensation parameters, at least one compensation operation, such as timestamp remapping and spatial coordinate transformation, is performed on the local asynchronous pulse event stream to complete cross-modal compensation registration. The compensated and registered local asynchronous pulse event stream is then combined with the frame image sequence to form a trigger-based dynamic visual segment.

[0009] Preferably, in step S2, calculating the extended motion vector and time change intensity of each spatiotemporal point in the triggered dynamic visual segment includes the following steps: For each spatiotemporal point in the triggered dynamic visual segment, calculate the optical flow component along the spatial x-direction. and optical flow components along the y-direction in space For the optical flow field of adjacent time periods, perform time difference calculation to calculate the rate of change component of optical flow along the time direction. ;by , and The extended motion vector constitutes the spatiotemporal point; the root mean square value of the pixel intensity change between multiple consecutive frames within the spatiotemporal point and its temporal neighborhood is calculated as the temporal variation intensity of the spatiotemporal point. .

[0010] Preferably, in step S2, the disorder of the local motion direction distribution in the local neighborhood of each spatiotemporal point is statistically analyzed, and the extended motion vector, the intensity of time variation, and the disorder of the local motion direction distribution are encoded into a three-dimensional hidden danger structure tensor field, including the following steps: For each spatiotemporal point in the triggered dynamic visual segment, a local spatiotemporal cube neighborhood is formed by a preset spatial radius and a preset time radius centered on the spatiotemporal point, and the directions of all extended motion vectors in the neighborhood are discretized into K direction intervals. Count the number of extended motion vectors falling into each directional interval, and calculate the probability value of the a-th directional interval. Where a = 1, 2, ..., K; calculate the local optical flow direction entropy. As a measure of the degree of disorder in the distribution of local motion directions; The construction of the three-dimensional hidden danger structure tensor field S(x,y,t) is as follows: .

[0011] Preferably, in step S3, the observation value metric is calculated for the tensor field of the three-dimensional hidden danger structure, and the neighbor unit distance factor is calculated by combining the location information of the shared neighbor mobile inspection units. The composite gradient of the observation value metric and the neighbor unit distance factor is used as the update direction to iteratively update the extended observation state of the mobile inspection unit. During the update process, the step size ratio of different types of degrees of freedom in the extended observation state is adjusted by the precondition matrix, including the following steps: Current extended observation status of the mobile inspection unit Calculate the measurement of observational value The calculation formula is: in, This refers to the field of view cone region under the current extended observation state. Let Frobenius norm be the tensor field of the three-dimensional hidden danger structure. For scaling parameters The spatial resolution term that increases with increasing observation distance d and decreases with increasing observation distance d. For parameters that vary with sampling rate Increases with relative velocity The motion fuzzy term that increases and decreases accordingly; Based on the shared location information of neighboring mobile inspection units, and combined with the observation direction parameters and scaling parameters of the mobile inspection units, the neighbor unit distance with each neighboring mobile inspection unit is calculated. The neighbor unit distance of each neighbor is mapped using an exponential decay function, and the mapped values ​​of each neighboring mobile inspection unit are summed to obtain the neighbor unit distance factor. The tensor field of the local three-dimensional hidden danger structure pointed to by the current observation direction is decomposed into features. The eigenvector corresponding to the largest eigenvalue is extracted as the principal direction of the hidden danger, and a preconditioning matrix is ​​constructed accordingly. The precondition matrix is ​​a diagonal matrix. Among its diagonal elements, the element corresponding to the spatial translational degree of freedom along the main direction of the hazard takes the first step length weight, and the element corresponding to the spatial translational degree of freedom perpendicular to the main direction of the hazard takes the second step length weight. The first step length weight is greater than the second step length weight.

[0012] Preferably, in step S3, controlling the movement of the mobile inspection unit based on the extended observation state obtained through iterative updates includes the following steps: Calculate the extended observation state update increment Where N is the distance factor of the neighboring unit, and S is the tensor field of the three-dimensional hidden danger structure. Here, η is the gradient operator with respect to the extended observation state, and η is the step size factor; The extended observation state is updated incrementally based on the extended observation state, and the spatial position, observation direction, scaling parameters, and sampling rate parameters of the mobile inspection unit are controlled according to the updated extended observation state.

[0013] Preferably, in step S4, the activation degree of continuous potential energy nodes is obtained from the updated three-dimensional hidden danger structure tensor field. When the activation degree exceeds a preset priority threshold, the current inspection task is interrupted and the investigation task branch is switched, including the following steps: For continuous potential energy nodes, the area of ​​concern is determined based on the spatial location of the hazards recorded in the hazard investigation list. The activation degree A of the continuous potential energy nodes is calculated using the following formula: in, For the current moment, To preset the time window length, Here, γ is the Frobenius norm of the updated 3D hazard structure tensor field, γ is the gain coefficient, and θ is the preset bias threshold. It is an S-shaped function; When the activation level exceeds the preset priority threshold, the current inspection task is interrupted, and the investigation task branch corresponding to the continuous potential energy node with the highest activation level is selected as the target to be transferred.

[0014] Compared with the prior art, the beneficial effects of the present invention are: This invention generates a trigger indication when a segment matching a preset abnormal motion pattern is detected in dynamic visual data. Based on this, visual data of the corresponding time period and region are extracted from the collected dynamic visual data to form a trigger-based dynamic visual segment. This achieves causal retrospective full-time domain capture of high-speed transient hazards, while significantly compressing the amount of data. This invention calculates the extended motion vector and the intensity of time change at each spatiotemporal point, and statistically analyzes the degree of disorder in the distribution of local motion directions. These three factors are then jointly encoded into a three-dimensional hazard structure tensor field, achieving a unified representation of the motion geometric characteristics and temporal disorder in the spatiotemporal structure of hazards. This provides a computable mathematical basis for distinguishing different hazard forms. This invention uses the composite gradient of the observation value metric and the distance factor of neighboring units as the update direction to perform decentralized iterative updates on the extended observation state of each unit, and adjusts the step size ratio on different types of degrees of freedom through a precondition matrix, so that the multi-inspection unit cluster can be adaptively distributed to high-risk areas while avoiding observation redundancy. This invention obtains the activation degree of continuous potential energy nodes from the updated three-dimensional hidden danger structure tensor field, and interrupts the inspection task and switches to the investigation task branch when the activation degree exceeds the priority threshold. The dynamic visual data collected during the investigation process is fed back to update the tensor field and the list, realizing the continuous driving of the hidden danger risk field on task scheduling and the closed-loop feedback of the investigation results on the cognitive state. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0017] Please see Figure 1 This invention is a joint inspection method for intelligent hazard investigation based on dynamic visual computing, comprising the following steps: S1: Deploy multiple mobile inspection units, and share location information among the mobile inspection units; the mobile inspection units are equipped with vision sensors to continuously perceive the scene and collect dynamic visual data. When a segment that matches the preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated; according to the trigger indication, visual data of the corresponding time period and corresponding area is extracted from the collected dynamic visual data to form a trigger-based dynamic visual segment. S2: Calculate the extended motion vector and time change intensity of each spatiotemporal point in the triggered dynamic visual segment, and statistically analyze the disorder of the local motion direction distribution in the local neighborhood of each spatiotemporal point; encode the extended motion vector and time change intensity, as well as the disorder of the local motion direction distribution, into a three-dimensional hidden danger structure tensor field; S3: Calculate the observation value metric for the three-dimensional hidden danger structure tensor field, calculate the neighbor unit distance factor by combining the location information of the shared neighbor mobile inspection units, and use the composite gradient of the observation value metric and the neighbor unit distance factor as the update direction to iteratively update the extended observation state of the mobile inspection unit. During the update process, the step size ratio of different types of degrees of freedom in the extended observation state is adjusted by the precondition matrix, and the movement of the mobile inspection unit is controlled according to the extended observation state obtained by the iterative update. After the extended observation state changes, the three-dimensional hidden danger structure tensor field is updated by the newly acquired dynamic visual data. The hidden danger category is judged according to the updated three-dimensional hidden danger structure tensor field to form a hidden danger investigation list. S4: The mobile inspection unit performs inspection tasks according to a preset behavior tree, which includes continuous potential energy nodes and investigation task branches. The activation degree of the continuous potential energy nodes is obtained from the updated three-dimensional hidden danger structure tensor field. When the activation degree exceeds a preset priority threshold, the current inspection task is interrupted and the investigation task branch is entered. The dynamic visual data collected during the execution of the investigation task branch is used to update the three-dimensional hidden danger structure tensor field, thereby updating the hidden danger investigation list.

[0018] Specifically, multiple mobile inspection units are first deployed and share location information. Each unit's onboard visual sensors continuously perceive the scene and collect dynamic visual data. When a segment matching a preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated. Based on the trigger indication, visual data for the corresponding time period and region is extracted from the collected dynamic visual data to form a triggered dynamic visual segment. Subsequently, the extended motion vector and time-varying intensity of each spatiotemporal point in the triggered dynamic visual segment are calculated, and the degree of disorder in the distribution of motion directions within the local neighborhood of each spatiotemporal point is statistically analyzed. These three factors are then jointly encoded into a three-dimensional hazard structure tensor field. After obtaining the tensor field, each mobile inspection unit calculates its observation value and, combined with shared neighbor location information, calculates the neighbor unit distance factor. Using the gradient of their product as the update direction, the extended observation state is iteratively updated. During the update process, the step size ratio on different degrees of freedom is adjusted using a precondition matrix, and the unit's movement is controlled based on the updated extended observation state. After the observation state changes, dynamic visual data is re-acquired to update the tensor field, and the hazard category is determined based on the updated tensor field, forming a hazard investigation list. Each hazard record in the hazard investigation list includes the hazard category, hazard spatial location, occurrence time window, and corresponding maximum feature value. During the inspection process, each mobile inspection unit executes inspection tasks according to a preset behavior tree, which contains continuous potential energy nodes. The activation degree of the continuous potential energy nodes is obtained from the updated tensor field. When the activation degree exceeds a preset priority threshold, the current inspection task is interrupted and the investigation task branch is initiated. The dynamic visual data collected during the execution of the investigation task branch is used to update the tensor field and the hazard investigation list.

[0019] In one embodiment of the present invention, in step S1, the mobile inspection unit is equipped with a vision sensor to continuously perceive the scene and collect dynamic visual data. When a segment that matches a preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated, including the following steps: Dynamic visual data includes an asynchronous impulse event stream output from a high temporal resolution dynamic change perception channel and a frame image stream output from a high spatial resolution texture detail perception channel. A spatiotemporal pulse pattern detector corresponding to each preset abnormal motion pattern is run on the asynchronous pulse event stream with a preset sliding time window. When any detector successfully matches within the current sliding time window, it is determined that a segment that conforms to the preset abnormal motion pattern has been detected. A trigger indication is generated based on the abnormal motion pattern type corresponding to the successfully matched detector, the timestamp of the matching time, and the spatial neighborhood coordinate range covered by the successfully matched event pulse cluster.

[0020] Specifically, the visual sensor on the mobile inspection unit consists of a dynamic change sensing component and a texture detail sensing component. The dynamic change sensing component (e.g., using an event camera) continuously monitors the brightness changes of each pixel location in the scene with a microsecond-level time resolution. When the brightness change of a pixel location exceeds a preset sensitivity threshold, it outputs an asynchronous pulse event. The sensitivity threshold is set based on the noise level under the scene's ambient lighting conditions. Under normal and stable operating conditions of the scene to be inspected, the rate of background events caused by sensor noise in the events output by the dynamic change sensing component is statistically analyzed. The sensitivity threshold is then set to the brightness change amplitude that can suppress background noise events to an acceptable level, typically 10% to 15% of the relative change in pixel brightness value. Each asynchronous pulse event carries the two-dimensional spatial coordinates of the pixel location, the time of the event, and polarity information indicating the direction of brightness change. All asynchronous pulse events output during continuous sensing are aggregated in chronological order to form an asynchronous pulse event stream. The texture detail sensing component (e.g., using a global shutter CMOS camera) continuously acquires texture images of the scene at a preset reference frame rate, forming a frame image stream; the typical preset reference frame rate is 30 frames per second.

[0021] A set of spatiotemporal pulse pattern detectors runs on the internal processor of the mobile inspection unit. Each spatiotemporal pulse pattern detector corresponds to a preset abnormal motion pattern. The preset abnormal motion patterns include at least clustering patterns, non-rigid body divergence patterns, and periodic abnormal patterns. A clustering pattern is defined as a sudden increase in the event firing rate within a spatial neighborhood of an asynchronous pulse event stream to exceed a first threshold within a preset sliding time window. The event firing rate is the number of asynchronous pulse events generated in the spatial neighborhood per unit time. The value of the first threshold is related to the average event firing rate of the spatial neighborhood under normal operating conditions, typically set to 5 to 10 times the average event firing rate. The size of the spatial neighborhood is a rectangular area on the image plane, typically 32×32 pixels or 64×64 pixels.

[0022] The construction and judgment logic of the spatiotemporal pulse pattern detector corresponding to the non-rigid body divergence mode is as follows: Within the current sliding time window, extract all asynchronous pulse events, take the mean of their spatial coordinates as the event cluster center C, calculate the Euclidean distance from each event to C, and take the maximum distance as the dispersion radius; divide the circular region with C as the center and the dispersion radius as the radius radially into equal areas. A concentric ring, Take 4 or 5, and number them sequentially from the inside out as ring 1 to ring 5. The area of ​​the innermost ring 1 is... The number of events that fell into it was Then the event density of this ring is Similarly, the outermost ring... The event density is ; Calculate the divergence intensity , where ε is a very small positive number to prevent division by zero; when Greater than the preset threshold When the event cluster matches a non-rigid body divergence pattern, it is determined that the event cluster matches a non-rigid body divergence pattern. It is a positive number, and its typical value can be determined based on historical samples of the application scenario, such as 0.2.

[0023] A periodic anomaly pattern is defined as a specific periodic fluctuation in the event emission frequency of the asynchronous pulse event stream within a preset sliding time window, with the fluctuation frequency falling within a preset fault characteristic frequency range. The preset fault characteristic frequency range is pre-calibrated based on the type of equipment being inspected; for example, for rotating machinery, the fault characteristic frequency range is 0.5 to 3 times the equipment's rotational frequency. Each spatiotemporal pulse pattern detector continuously operates on the asynchronous pulse event stream using a sliding time window. The typical duration of the sliding time window is 50 to 200 milliseconds, and the sliding step size is one-tenth of the sliding time window duration.

[0024] The system extracts all asynchronous pulse events whose timestamps fall within the current sliding time window from the asynchronous pulse event stream. Spatiotemporal clustering analysis is performed on these events to extract their spatial distribution and temporal statistical features. Spatial distribution features include the spatial center location, spatial dispersion radius, and spatial distribution direction of the event pulses. Temporal statistical features include the time series of event emission rates, the distribution statistics of event emission intervals, and the spectral components of event emission frequencies. The extracted spatial distribution and temporal statistical features are then matched and compared with the feature template of the preset abnormal motion pattern corresponding to the detector, and the matching similarity is calculated. The matching similarity is represented by the normalized mapping value of the cosine similarity or Euclidean distance between feature vectors. When the matching similarity of any detector within the current sliding time window exceeds a preset similarity threshold, a segment conforming to the preset abnormal motion pattern is detected. The similarity threshold is set based on the trade-off between the false trigger rate of the detector on historical normal data and the false miss rate on known hazard samples. In the offline stage, the detector is tested using labeled normal operating data and hazard operating data. Curves of false trigger rate and false miss rate as a function of similarity threshold are plotted, and the value is selected near the intersection point where both are at an acceptable level. Typical values ​​are 0.80 to 0.95.

[0025] After a successful detection, the trigger indication generated by the mobile inspection unit includes the abnormal motion mode type identifier corresponding to the successfully matched detector, the timestamp of the matching time, and the spatial neighborhood coordinate range covered by the successfully matched event pulse cluster. This coordinate range is represented by the coordinates of the four vertices of the minimum bounding rectangle of the event pulse cluster on the image plane. The origin of the coordinates is located at the upper left corner of the image plane, the x-axis extends to the right along the image width direction, and the y-axis extends downward along the image height direction. The coordinate values ​​are in pixels.

[0026] In one embodiment of the present invention, step S1 involves extracting visual data for the corresponding time period and region from the acquired dynamic visual data according to a trigger indication to form a trigger-based dynamic visual segment, including the following steps: Using the timestamp carried by the trigger indication as the retrieval benchmark, extract the frame image sequence from the first preset time point before the trigger time to the second preset time point after the trigger time from the acquired frame image stream; Using the spatial neighborhood coordinate range carried by the trigger indication as a constraint, extract the local asynchronous pulse event stream within the corresponding spatiotemporal range from the collected asynchronous pulse event stream; The local asynchronous pulse event stream is spatiotemporally aligned with the frame image sequence, and motion compensation parameters are calculated using the successfully matched event pulse clusters carried by the trigger indicator. Based on the motion compensation parameters, at least one compensation operation, such as timestamp remapping and spatial coordinate transformation, is performed on the local asynchronous pulse event stream to complete cross-modal compensation registration. The compensated and registered local asynchronous pulse event stream is then combined with the frame image sequence to form a trigger-based dynamic visual segment.

[0027] Specifically, using the timestamp of the trigger indication as the retrieval benchmark, a frame image sequence is extracted from the acquired frame image stream. The extraction time range starts from a first preset time point before the trigger moment and ends at a second preset time point after the trigger moment; the first preset time point is typically 300 milliseconds; the second preset time point is typically 100 milliseconds; the first preset duration is longer than the second preset duration to ensure that the extracted frame image sequence fully covers the precursor stage before the hazard occurs, the instant of occurrence, and the brief continuation stage after occurrence. The frame image stream is acquired at a baseline frame rate of 30 frames per second, and approximately 13 images can be obtained within a total extraction time of 400 milliseconds. These frame images are arranged in chronological order according to their timestamps to form a frame image sequence.

[0028] Using the spatial neighborhood coordinate range of the trigger indication as a constraint, a local asynchronous pulse event stream is extracted from the acquired asynchronous pulse event stream. This spatial neighborhood coordinate range is represented by the coordinates of the four vertices of the minimum bounding rectangle of the event pulse cluster on the image plane. The coordinate values ​​are in pixels, with the origin located at the upper left corner of the image plane, the x-axis pointing to the right along the image width direction, and the y-axis pointing downwards along the image height direction. The filtering criteria are that the two-dimensional spatial coordinates of the event fall within this rectangular area, and the timestamp of the event falls within the time range defined by the first preset duration to the second preset duration. All filtered pulse events are arranged in chronological order of their timestamps to form a local asynchronous pulse event stream.

[0029] The local asynchronous pulse event stream is spatiotemporally aligned with the frame image sequence, and motion compensation parameters are calculated using the successfully matched event pulse clusters carried by the trigger indicator. Although the dynamic change sensing unit and the texture detail sensing unit are physically fixed to the same inspection pan-tilt unit, the slight deviation of the optical axis and the difference in internal electronic delay between the two types of sensing units result in a non-naturally strict alignment between the spatial coordinates and timestamps of events and the pixels of the frame image. This requires correction through cross-modal compensation registration.

[0030] The trigger indication's successful matching event cluster is the set of events confirmed to match a preset abnormal motion pattern during the detection phase. For each event in this cluster, the pixel in the frame image sequence that is closest to the event in both spatial coordinates and timestamp is found, establishing a corresponding point pair between the event and the frame image pixel. For events whose timestamps lie between the exposure times of two adjacent frames, an interpolated image of the corresponding time is generated through linear interpolation of the two adjacent frames to improve the accuracy of establishing the corresponding point pair. For all corresponding point pairs, the optimal parameters of the rigid body transformation model are solved using the least squares method. The rigid body transformation model contains two translation parameters and one rotation parameter. The optimal translation and rotation parameters are obtained by minimizing the sum of squares of the reprojection errors of all corresponding point pairs, which together constitute the motion compensation parameters.

[0031] Based on motion compensation parameters, compensation operations are performed on each event in the local asynchronous pulse event stream. These compensation operations include at least one of spatial coordinate transformation and timestamp remapping. Spatial coordinate transformation refers to rotating and translating the spatial coordinates of the event according to the optimal rigid body transformation parameters (i.e., optimal translation and optimal rotation parameters) to obtain the compensated spatial coordinates. Timestamp remapping refers to correcting a fixed time delay between two types of sensing components. This fixed time delay is determined through offline calibration. The calibration method involves placing a pulsed light source or a rapidly flashing light source in the scene to be inspected; the flashing frequency of the pulsed light source is set to 1 kHz, and the pulse width is less than 1 microsecond; simultaneously, the timestamps of the dynamic change sensing component and the texture detail sensing component's response to the light source are recorded, and the average difference between the two component timestamps is calculated. This is repeated multiple times and averaged to obtain the fixed time delay value. The value of the fixed time delay depends on the specific hardware model and signal transmission path, typically ranging from 1 microsecond to 50 microseconds; for example, for configurations using hardware synchronization signal lines, this value is usually less than 10 microseconds. Timestamp remapping subtracts this fixed time delay value from the original timestamp of each event to obtain the corrected timestamp.

[0032] After completing cross-modal compensation registration, the compensated and registered local asynchronous pulse event stream is combined with the frame image sequence to form a triggered dynamic visual segment. This triggered dynamic visual segment includes all frame images in the frame image sequence, all events in the compensated and registered local asynchronous pulse event stream, and their compensated spatial coordinates and timestamps. During the inspection process, the time delay between the dynamic change sensing component and the texture detail sensing component may change slowly due to factors such as temperature drift. To address this, the mobile inspection unit selects frames with contrast higher than a preset threshold in the frame image sequence at a preset period (e.g., every 30 minutes), extracts their edge pixel positions, retrieves the event spatially closest to the edge pixel position in the asynchronous pulse event stream within the corresponding time period, calculates the difference between the edge pixel frame time and the matching event timestamp, and takes the median of multiple frames and multiple edge points as the residual delay correction value for online updates. This correction value is added to the offline calibrated fixed time delay to obtain the current effective time delay, which is used for compensation registration of subsequent triggered dynamic visual segments.

[0033] In one embodiment of the present invention, step S2, calculating the extended motion vector and time change intensity of each spatiotemporal point in the triggered dynamic visual segment, includes the following steps: for each spatiotemporal point in the triggered dynamic visual segment, calculating the optical flow component along the spatial x-direction. and optical flow components along the y-direction in space For the optical flow field of adjacent time periods, perform time difference calculation to calculate the rate of change component of optical flow along the time direction. ;by , and The extended motion vector constitutes the spatiotemporal point; the root mean square value of the pixel intensity change between multiple consecutive frames within the spatiotemporal point and its temporal neighborhood is calculated as the temporal variation intensity of the spatiotemporal point. .

[0034] Specifically, the frame image sequence is denoted as ,in This represents the total number of frames in the frame image sequence, with indices increasing sequentially by timestamp. For each pair of adjacent frames in the frame image sequence... and (j takes values ​​from 1 to...) The variational optical flow method is used to calculate the dense optical flow field between two frames. Based on the assumption of constant image grayscale and joint optimization of data and smoothing terms, the variational optical flow method solves for the optical flow vector at each pixel location, which includes a component along the x-direction of the image plane space. and the component along the y-direction in space The calculation is performed on a pixel-per-frame basis. The calculated optical flow field covers all pixel locations of the frame image, with each pixel location having a corresponding... Value and For each time interval between two adjacent frames in a frame image sequence, an optical flow field is calculated. For example, if the frame image sequence contains 13 frames, a total of 12 optical flow fields are calculated, each corresponding to a time interval.

[0035] The rate of change of optical flow along the time direction is denoted as It is obtained by performing time difference analysis on the optical flow fields of adjacent time periods. For the j-th optical flow field and the (j+1)-th optical flow field, the time difference is calculated separately for the optical flow components at the same pixel location, that is, using the (j+1)-th optical flow field... Component minus the j-th optical flow field The component is obtained by dividing the difference by the difference between the center times of the two optical flow fields, thus obtaining the value at the pixel location. A sample value; for The same procedure applies to the components. and Take the average as the value at that location. The temporal difference is calculated pairwise over adjacent optical flow field pairs. The results of the calculations for multiple consecutive adjacent optical flow field pairs are then averaged over a temporal neighborhood, typically spanning 3 to 5 optical flow field pairs. During the time periods at the beginning and end of the frame image sequence, due to the lack of sufficient preceding or succeeding optical flow field pairs, the results are calculated only from the available optical flow field pairs, and no moving average is performed.

[0036] At each spacetime point, we obtain from... , , An extended motion vector is formed; this vector acts on the spatiotemporal three-dimensional volume covered by the frame image sequence. The spatial dimension of this spatiotemporal three-dimensional volume is the width and height of the frame image, and the temporal dimension is the sequence of frames. Each pixel position falling within this spatiotemporal three-dimensional volume corresponds to a spatiotemporal point and a frame time, and each spatiotemporal point has an extended motion vector ( , , ).

[0037] Intensity of time variation This is the root mean square value of the pixel intensity variation between consecutive frames at this spatiotemporal point and its temporal neighborhood. For a given frame in a frame image sequence... Let a pixel position (x, y) on the frame be a position within the frame. The pixel intensity value on ,in P is the preset temporal neighborhood half-width, typically 1 or 2; for example, if the temporal neighborhood half-width is 1, then a total of 3 frames are involved. The sum of squares of the first-order differences of pixel intensity values ​​with respect to time within this temporal neighborhood is calculated, divided by the number of difference pairs within the neighborhood, and then the square root is taken to obtain the temporal variation intensity at that spatiotemporal point. The first-order difference is the difference in intensity values ​​at the same pixel location between two adjacent frames, with 2P difference pairs. For regions near the ends of a frame image sequence, the portion of the temporal neighborhood that exceeds the available range of the frame image sequence is not included in the calculation; only the available frames are used, and the number of difference pairs is reduced accordingly.

[0038] In one embodiment of the present invention, in step S2, the disorder of the local motion direction distribution in the local neighborhood of each spatiotemporal point is statistically analyzed, and the extended motion vector, the intensity of time change, and the disorder of the local motion direction distribution are encoded into a three-dimensional hidden danger structure tensor field, including the following steps: For each spatiotemporal point in the triggered dynamic visual segment, a local spatiotemporal cube neighborhood is formed by a preset spatial radius and a preset time radius centered on the spatiotemporal point, and the directions of all extended motion vectors in the neighborhood are discretized into K direction intervals. Count the number of extended motion vectors falling into each directional interval, and calculate the probability value of the a-th directional interval. Where a = 1, 2, ..., K; calculate the local optical flow direction entropy. As a measure of the degree of disorder in the distribution of local motion directions; The construction of the three-dimensional hidden danger structure tensor field S(x,y,t) is as follows: .

[0039] Specifically, for each spatiotemporal point in the spatiotemporal 3D volume covered by the triggered dynamic visual segment, the degree of disorder in the distribution of local motion directions at that point is statistically analyzed. A local spatiotemporal cube neighborhood is constructed by taking a preset spatial radius and a preset temporal radius centered on that spatiotemporal point; the typical value of the preset spatial radius is 8 to 16 pixels, and the typical value of the preset temporal radius is 2 to 4 frames; for example, if the spatial radius is 8 pixels and the temporal radius is 3 frames, then the spatial cross-section of the local spatiotemporal cube neighborhood is a 17×17 pixel square area, and the time span is 7 frames. Within this local spatiotemporal cube neighborhood, the directions of the extended motion vectors at all spatiotemporal points are collected. The direction of the extended motion vector is determined by its spatial components. and The direction angle is determined through calculation. and The direction angle is obtained by the arctangent of the ratio, and the range of the direction angle is from 0 to 360 degrees. The range of the direction angle is uniformly discretized into K direction intervals, and the typical value of K is 8, 12 or 16; for example, when K is 8, each direction interval covers 45 degrees.

[0040] The number of extended motion vectors falling into the a-th direction interval is counted, denoted as . Where 'a' takes the value of an integer from 1 to K; the probability value of the 'a'-th direction interval. The number of vectors in this interval Divide by the total number of extended motion vectors in the neighborhood; if the total number of extended motion vectors in the neighborhood is zero, then divide all... The local optical flow direction entropy is set to an equal, uniform distribution value of 1 / K. The local optical flow direction entropy is calculated as the degree of disorder in the local motion direction distribution at that spatiotemporal point. When the motion directions within the neighborhood are highly consistent and concentrated only in a single directional interval, the entropy value approaches 0; when the motion directions are uniformly distributed across all directional intervals, the entropy value approaches logK; when the motion directions are randomly distributed across multiple intervals but not uniformly, the entropy value lies between the two. The typical range of the local optical flow direction entropy H is 0 to logK; for example, when K is 8, it is approximately 0 to 2.08.

[0041] In the 3D hazard structure tensor field S(x,y,t), the matrix part is a 3×3 real symmetric matrix. Its diagonal elements are the sum of the squares of the components of the extended motion vector and the squares of the time gradient energy, while the off-diagonal elements are the pairwise products of the components of the extended motion vector. The symbol ⊙ represents the Hadamard product, which is each element in the matrix multiplied by the local optical flow direction entropy. The elements of the 3×3 real symmetric matrix are denoted as... ,in, , , , , , ; Perform eigenvalue decomposition on the real symmetric matrix and solve the characteristic equation. That is, to solve the equation about λ: Where I is a 3×3 identity matrix, i.e., a square matrix with all elements on the main diagonal being 1 and all other elements being 0; the three roots of the equation are all real numbers, denoted as in descending order. And all three eigenvalues ​​satisfy .

[0042] For each eigenvalue Solve the system of linear equations The corresponding feature vectors are obtained. Let z = 1, 2, and 3. The three eigenvectors are pairwise orthogonal and normalized to unit length. The three components of each eigenvector correspond to the projection weights of the hidden structure along the x-axis, y-axis, and time axis, respectively. For example, if... If the dominant feature vector points almost entirely to the time axis direction, then the dominant feature vector will be in the same direction.

[0043] The magnitude relationship of the three eigenvalues ​​reflects the spatiotemporal structure type of the hidden danger; for example, when and When pointing towards the time axis, it corresponds to a potential hazard of momentary flickering; when... Furthermore, when the three feature vectors are distributed in multiple directions on the spatial plane, it corresponds to a potential hazard of debris splashing; when and When pointing in a specific spatial direction, it corresponds to a potential crack-like hazard that extends along that direction.

[0044] In one embodiment of the present invention, in step S3, the observation value metric of the three-dimensional hidden danger structure tensor field is calculated, and the neighbor unit distance factor is calculated by combining the location information of the shared neighbor mobile inspection units. The extended observation state of the mobile inspection unit is iteratively updated using the composite gradient of the observation value metric and the neighbor unit distance factor as the update direction. During the update process, the step size ratio of different types of degrees of freedom in the extended observation state is adjusted by the precondition matrix, including the following steps: Current extended observation status of the mobile inspection unit Calculate the measurement of observational value The calculation formula is:

[0045] in, This refers to the field of view cone region under the current extended observation state. Let Frobenius norm be the tensor field of the three-dimensional hidden danger structure. For scaling parameters The spatial resolution term that increases with increasing observation distance d and decreases with increasing observation distance d. For parameters that vary with sampling rate Increases with relative velocity The motion fuzzy term that increases and decreases accordingly; Based on the shared location information of neighboring mobile inspection units, and combined with the observation direction parameters and scaling parameters of the mobile inspection units, the neighbor unit distance with each neighboring mobile inspection unit is calculated. The neighbor unit distance of each neighbor is mapped using an exponential decay function, and the mapped values ​​of each neighboring mobile inspection unit are summed to obtain the neighbor unit distance factor. The tensor field of the local three-dimensional hidden danger structure pointed to by the current observation direction is decomposed into features. The eigenvector corresponding to the largest eigenvalue is extracted as the principal direction of the hidden danger, and a preconditioning matrix is ​​constructed accordingly. The precondition matrix is ​​a diagonal matrix. Among its diagonal elements, the element corresponding to the spatial translational degree of freedom along the main direction of the hazard takes the first step length weight, and the element corresponding to the spatial translational degree of freedom perpendicular to the main direction of the hazard takes the second step length weight. The first step length weight is greater than the second step length weight.

[0046] Specifically, each mobile inspection unit i has maintained its own extended observation state vector when performing step S3. ,in, The coordinates of the mobile inspection unit in physical space. The horizontal angle of the camera's optical axis. For zoom lens focal length, The frame rate is used for data collection.

[0047] Measurement of observational value In the calculation formula, For the mobile inspection unit , as well as The spatial field of view (FOV) is jointly determined; the horizontal and vertical field of view of the FVOV are determined by the focal length and the image sensor size, with the near-end truncation distance being 1 to 5 times the focal length and the far-end truncation distance being 20 to 50 times the focal length. Spatial resolution term. With focal length (i.e., scaling parameter, in millimeters) The effect increases with increasing observation distance d (in meters), and decreases with increasing observation distance d (in meters) to penalize image blur caused by excessive distance or insufficient focal length; a typical implementation is as follows: Motion fuzziness penalty item With frame rate (i.e., sampling rate parameter) Increased and improved with the movement speed of the mobile inspection unit relative to the scene. Increase and worsen; a typical realization is Where L is the reference scale, taken as 0.1 meters.

[0048] The mobile inspection unit uses location information shared with its neighbors to calculate the neighbor unit distance factor N. For each neighbor unit (i.e., the neighbor mobile inspection unit b), the neighbor unit distance in the extended observation state space is calculated. The directional difference weight α is used to convert angular differences into a distance contribution comparable to the square of the spatial distance. Its value depends on the spatial scale of the inspection scenario and the desired angular sensitivity. In typical industrial inspection scenarios, the working distance between inspection units is usually 3 to 10 meters. When the orientation of two units differs by 30 degrees, the offset of the center of their observation area at a distance of 10 meters is approximately 5.2 meters, which is on the same order of magnitude as the typical working distance. To ensure that the directional difference contributes to the neighbor distance in a way comparable to the spatial distance, α is set to 0.1 to 1.0. The focal length difference weight β is used to convert the relative change in focal length into a distance contribution. The logarithmic difference of a focal length change from 24 mm to 50 mm is approximately 0.73, and this zoom corresponds to a change in the observation range of approximately 2 times. To ensure that the focal length difference has a perceptible but not dominant contribution to the neighbor unit distance, β is set to 0.1 to 0.5.

[0049] The logarithmic focal length difference is introduced to ensure that relative changes in focal length (such as 2x zoom) produce the same distance contribution at different base focal lengths. Then, the distance to a single neighbor is mapped using an exponential decay function, and the neighbor cell distance factor is obtained by summing over all neighbors. The compactness coefficient 'o' controls the rate at which the repulsive effect decays with increasing distance. Its value depends on the desired effective repulsive radius. The effective repulsive radius can be defined as the spatial distance at which the mapped value of the neighbor distance factor drops to approximately 0.37, and this spatial distance is inversely proportional to the square root of 'o'. When the desired repulsion essentially decays at a spatial distance of approximately 3 meters, a value of 'o' of 0.1 results in a mapped value of approximately 0.41 at 3 meters (contributing approximately 9), and a decrease to approximately 10 at 10 meters. -4 The magnitude of the repulsion effect is negligible; the value of o ranges from 0.05 to 0.2, with smaller values ​​suitable for sparse deployments on a large scale and larger values ​​suitable for dense deployments on a small scale.

[0050] Perform eigenvalue decomposition on the tensor field of the local three-dimensional hidden danger structure pointed to by the current line of sight, and take the maximum eigenvalue. corresponding feature vector Because the mobile inspection unit moves in the horizontal plane, it extracts... Spatial components And normalize it to a two-dimensional unit vector. This serves as the main direction for identifying potential hazards. The precondition matrix M is a diagonal matrix, with its diagonal elements corresponding to... The step size weights for each degree of freedom are in the order (x, y, z, θ, f, r). For the spatial position degrees of freedom x and y, the weights are set according to the main direction of the potential hazard, and the step size weights are set in the direction u. In the vertical direction Set as above ; and satisfy And the ratio Choose 2 to 5, with a typical configuration as follows , For the z-degree of freedom, since the height is constant, the weight is set to 0 or a fixed small value. For the observation direction θ-degree of freedom, the weight can be 1.0 to allow normal rotation; the focal length f weight can be 0.8 to make the zoom action slightly slower to avoid drastic shaking; the frame rate r weight can be 0.5 to make its changes smoother. Therefore... Among them, W X , W y By the current unit position , Projecting onto the global x and y axes yields the result; the remaining diagonal elements are as shown above.

[0051] The mobile inspection unit calculates the product of the observation value metric and the distance factor of its neighboring units, and then applies this product with respect to... Calculate the gradient to obtain the initial update direction. This gradient already reflects the trend of moving towards high-risk, high-value areas while avoiding excessive clustering with neighbors. Finally, multiplying this gradient by the preconditioning matrix yields the final update direction adjusted for step size.

[0052] In one embodiment of the present invention, step S3, controlling the movement of the mobile inspection unit based on the extended observation state obtained through iterative updates, includes the following steps: Calculate the extended observation state update increment Where N is the distance factor of the neighboring unit, and S is the tensor field of the three-dimensional hidden danger structure. Here, η is the gradient operator with respect to the extended observation state, and η is the step size factor; The extended observation state is updated incrementally based on the extended observation state, and the spatial position, observation direction, scaling parameters, and sampling rate parameters of the mobile inspection unit are controlled according to the updated extended observation state.

[0053] Specifically, expand the incremental update of observation status. In the calculation formula, the step size factor η is used to control the magnitude of each iteration update, and its value needs to take into account both the convergence speed and motion stability. In a typical industrial inspection scenario, the upper limit of the inspection unit's moving speed is about 1.5 m / s, the upper limit of the gimbal's rotation speed is about 30 degrees / s, and the rate of change of focal length and frame rate is also limited by hardware. Based on the dimensional differences of each component of the extended observation state and hardware constraints, the value of η ranges from 0.05 to 0.5.

[0054] Update increments based on extended observation status Update extended observation status: The updated extended observation status Includes new spatial locations Observation direction ,focal length and frame rate The mobile inspection unit adjusts its position in physical space according to the spatial position component in the updated extended observation state, adjusts the optical axis pointing of its onboard visual sensor according to the observation direction component, adjusts the zoom lens of its onboard visual sensor according to the focal length component, and adjusts the acquisition frequency of its onboard visual sensor according to the frame rate component.

[0055] The iterative update process continues until a preset termination condition is met. The termination condition is set as follows: the change in the observation value metric is lower than a preset threshold in multiple consecutive iterations, for example, the relative change in the observation value metric is less than 5% in 5 consecutive iterations; or the preset maximum number of iterations is reached, for example, 50 iterations; or an external termination command is received.

[0056] In one embodiment of the present invention, in step S4, the activation degree of continuous potential energy nodes is obtained from the updated three-dimensional hidden danger structure tensor field. When the activation degree exceeds a preset priority threshold, the current inspection task is interrupted and the investigation task branch is switched, including the following steps: For continuous potential energy nodes, the area of ​​concern is determined based on the spatial location of the hazards recorded in the hazard investigation list. The activation degree A of the continuous potential energy nodes is calculated using the following formula: in, For the current moment, To preset the time window length, Here, γ is the Frobenius norm of the updated 3D hazard structure tensor field, γ is the gain coefficient, and θ is the preset bias threshold. It is an S-shaped function; When the activation level exceeds the preset priority threshold, the current inspection task is interrupted, and the investigation task branch corresponding to the continuous potential energy node with the highest activation level is selected as the target to be transferred.

[0057] Specifically, before the inspection task is initiated, a unified behavior tree is pre-configured for each mobile inspection unit. This behavior tree includes inspection task nodes, continuous potential energy nodes, and investigation task branches. The inspection task node defines the default behavior of the mobile inspection unit when no high-risk hazards are triggered, namely, continuously roaming the scene according to the extended observation state iteration update law of step S3 to maintain dynamic coverage of the global area. The continuous potential energy node is a conditional node whose return value is not a Boolean value, but a continuous activation value between 0 and 1. This node is bound to the area of ​​interest determined by the spatial location of the hazard recorded in the hazard investigation list. The investigation task branch is a sequence of specific investigation actions in the behavior tree corresponding to different hazard types or different areas. Each investigation task branch pre-stores specific expected motion parameters to guide the mobile inspection unit to perform specific observation actions, such as fixed-point scanning. Each continuous potential energy node has a pre-defined correspondence with a troubleshooting task branch, which is established when the behavior tree is constructed. When there are multiple continuous potential energy nodes and multiple troubleshooting task branches in the behavior tree, each continuous potential energy node is associated with the troubleshooting task branch that handles the hidden dangers in the area of ​​concern bound to that node. If there are multiple hidden dangers in an area of ​​concern, the troubleshooting task branches that handle these hidden dangers are chained together into a troubleshooting task sequence and then associated with the continuous potential energy node.

[0058] The area of ​​interest is determined by the spatial location of the hazards recorded in the hazard identification checklist. For each hazard recorded in the hazard identification checklist, its spatial coordinates are extracted; a rectangular area of ​​interest is formed by expanding outward from these spatial coordinates with a preset spatial radius. The preset spatial radius is determined based on the radius of the effective observation range of the mobile inspection unit in inspection mode. Under typical inspection conditions with a focal length of 50 mm and an object distance of 5 meters, the visual sensor on the mobile inspection unit covers a scene width of approximately 3.6 meters and a height of approximately 2.4 meters in a single frame image. To ensure that the area of ​​interest completely covers the potential impact range of the hazard and leaves room for observation, the typical value of the spatial radius is 3 to 5 meters. For example, if the spatial coordinates of the hazard are (12.0, 8.0, 1.5), and the spatial radius is 4 meters, then the area of ​​interest R is a rectangular area with (8.0, 4.0) and (16.0, 12.0) as diagonal vertices. If there are multiple hazard records in the list and their areas of concern overlap spatially, the overlapping areas of concern will be merged into a larger area of ​​concern to avoid generating multiple duplicate continuous potential energy nodes for similar hazards.

[0059] In the formula for calculating activation A, the gain coefficient γ is used to scale the integral value to the sensitive range of the sigmoid function. The integral value is typically low under normal inspection conditions, for example, between 0.1 and 1.0. When a high-risk hazard appears within the area of ​​interest, the integral value can rise to 5.0 to 20.0. To map the normal state to the low activation range and the high-risk state to the high activation range, γ is set to 0.5 to 2.0. θ is the bias threshold, which suppresses the activation in low-risk scenarios to avoid frequent false triggers due to background noise; θ is set to 1.0 to 3.0. For sigmoid functions, the integral value is smoothly mapped to the interval between 0 and 1; a typical form is the logistic function. The behavior tree evaluates the activation level of all continuous potential energy nodes in each decision cycle; the duration of the decision cycle is the cycle in which the mobile inspection unit completes one S3 step iteration update, typically ranging from 200 milliseconds to 500 milliseconds.

[0060] During the patrol mission, the behavior tree continuously monitors the activation level of its associated continuous potential energy nodes. When the activation level of one or more continuous potential energy nodes exceeds a preset priority threshold... When this happens, the behavior tree immediately interrupts the currently executing patrol task node. Based on the trade-off between false trigger rate and response latency, the value is set to 0.5 to 0.8. After an interruption occurs, the behavior tree compares all continuous potential energy nodes whose activation exceeds the priority threshold, and selects the investigation task branch corresponding to the continuous potential energy node with the highest activation as the transition target; after the selected investigation task branch is activated, the behavior mode of the mobile inspection unit switches from coverage-driven to investigation-driven, and then executes the various observation actions pre-stored in the investigation task branch in sequence.

[0061] When a patrol mission is interrupted, the patrol status at the time of interruption is saved, including the current extended observation status and the sequence of remaining uncompleted patrol path points. This information is serialized into a task node to be restored. The task node to be restored carries the identifier of the patrol area and the timestamp of the time of generation, and is attached to the task queue to be restored in the behavior tree.

[0062] When a mobile inspection unit completes its investigation task branch, the unit enters an idle state. Idle units query the queue of tasks awaiting recovery, selecting the task node with the earliest timestamp or closest to the current spatial location. Before officially resuming the task, they check whether the inspection area associated with the task node is covered by other mobile inspection units, i.e., whether other units are currently performing inspection or investigation tasks in that area, and whether the overlap of their observation areas exceeds a preset threshold (e.g., 50%). If a coverage conflict exists, the task node is skipped, and the next candidate node in the queue is selected. If all nodes in the queue have coverage conflicts, the idle unit performs a global random roaming inspection, continuously monitoring the state changes of the task queue during the roaming process. Upon resuming execution, the mobile inspection unit restores its extended observation state to the state saved in the task node and continues inspection from the first point in the remaining path point sequence, re-entering the extended observation state iterative update in step S3.

[0063] During the execution of the investigation task branch, the visual sensors of the mobile inspection unit continuously collect dynamic visual data. Based on the newly collected dynamic visual data, a new local three-dimensional hazard structure tensor field sub-block is calculated. The newly calculated local sub-block is then merged and updated with the existing global tensor field. Then, based on the updated three-dimensional hazard structure tensor field, the hazard investigation list is updated according to the hazard category discrimination method in step S3.

[0064] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A joint inspection method for intelligent hazard investigation based on dynamic visual computing, characterized in that, Includes the following steps: S1: Deploy multiple mobile inspection units, and share location information among the mobile inspection units; The mobile inspection unit is equipped with a vision sensor to continuously perceive the scene and collect dynamic visual data. When a segment that matches a preset abnormal motion pattern is detected in the dynamic visual data, a trigger instruction is generated. Based on the trigger instruction, visual data of the corresponding time period and corresponding area is extracted from the collected dynamic visual data to form a trigger-based dynamic visual segment. S2: Calculate the extended motion vector and time change intensity of each spatiotemporal point in the triggered dynamic visual segment, and statistically analyze the disorder of the local motion direction distribution in the local neighborhood of each spatiotemporal point; encode the extended motion vector and time change intensity, as well as the disorder of the local motion direction distribution, into a three-dimensional hidden danger structure tensor field; S3: Calculate the observation value metric for the tensor field of the three-dimensional hidden danger structure, calculate the neighbor unit distance factor by combining the location information of the shared neighbor mobile inspection units, and use the composite gradient of the observation value metric and the neighbor unit distance factor as the update direction to iteratively update the extended observation state of the mobile inspection unit. During the update process, the step size ratio of different types of degrees of freedom in the extended observation state is adjusted by the precondition matrix, and the movement of the mobile inspection unit is controlled according to the extended observation state obtained by the iterative update. After the observation status changes, the three-dimensional hidden danger structure tensor field is updated by the newly acquired dynamic visual data; the hidden danger category is determined based on the updated three-dimensional hidden danger structure tensor field to form a hidden danger investigation list; S4: The mobile inspection unit performs inspection tasks according to a preset behavior tree, which includes continuous potential energy nodes and investigation task branches. The activation degree of continuous potential energy nodes is obtained from the updated three-dimensional hidden danger structure tensor field. When the activation degree exceeds the preset priority threshold, the current inspection task is interrupted and the investigation task branch is switched. The dynamic visual data collected during the execution of the investigation task branch is used to update the three-dimensional hidden danger structure tensor field, and then update the hidden danger investigation list.

2. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 1, characterized in that, In step S1, the mobile inspection unit is equipped with a vision sensor to continuously perceive the scene and collect dynamic visual data. When a segment matching a preset abnormal motion pattern is detected in the dynamic visual data, a trigger indication is generated, including the following steps: Dynamic visual data includes an asynchronous impulse event stream output from a high temporal resolution dynamic change perception channel and a frame image stream output from a high spatial resolution texture detail perception channel. A spatiotemporal pulse pattern detector corresponding to each preset abnormal motion pattern is run on the asynchronous pulse event stream with a preset sliding time window. When any detector successfully matches within the current sliding time window, it is determined that a segment that conforms to the preset abnormal motion pattern has been detected. A trigger indication is generated based on the abnormal motion pattern type corresponding to the successfully matched detector, the timestamp of the matching time, and the spatial neighborhood coordinate range covered by the successfully matched event pulse cluster.

3. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 2, characterized in that, In step S1, visual data corresponding to the time period and region is extracted from the collected dynamic visual data according to the trigger indication to form a trigger-based dynamic visual segment, including the following steps: Using the timestamp carried by the trigger indication as the retrieval benchmark, extract the frame image sequence from the first preset time point before the trigger time to the second preset time point after the trigger time from the acquired frame image stream; Using the spatial neighborhood coordinate range carried by the trigger indication as a constraint, extract the local asynchronous pulse event stream within the corresponding spatiotemporal range from the collected asynchronous pulse event stream; The local asynchronous pulse event stream is spatiotemporally aligned with the frame image sequence, and motion compensation parameters are calculated using the successfully matched event pulse clusters carried by the trigger indicator. Based on the motion compensation parameters, at least one compensation operation, such as timestamp remapping and spatial coordinate transformation, is performed on the local asynchronous pulse event stream to complete cross-modal compensation registration. The compensated and registered local asynchronous pulse event stream is then combined with the frame image sequence to form a trigger-based dynamic visual segment.

4. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 1, characterized in that, In step S2, calculating the extended motion vector and time change intensity of each spatiotemporal point in the triggered dynamic visual segment includes the following steps: For each spatiotemporal point in the triggered dynamic visual segment, calculate the optical flow component along the spatial x-direction. and optical flow components along the y-direction in space For the optical flow field of adjacent time periods, perform time difference calculation to calculate the rate of change component of optical flow along the time direction. ;by , and The extended motion vector constitutes the spatiotemporal point; the root mean square value of the pixel intensity change between multiple consecutive frames within the spatiotemporal point and its temporal neighborhood is calculated as the temporal variation intensity of the spatiotemporal point. .

5. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 4, characterized in that, In step S2, the disorder of the local motion direction distribution in the local neighborhood of each spatiotemporal point is statistically analyzed. The extended motion vector, the intensity of time variation, and the disorder of the local motion direction distribution are encoded into a three-dimensional hidden danger structure tensor field, including the following steps: For each spatiotemporal point in the triggered dynamic visual segment, a local spatiotemporal cube neighborhood is formed by a preset spatial radius and a preset temporal radius centered on that spatiotemporal point. The directions of all extended motion vectors within this neighborhood are discretized into K directional intervals. The number of extended motion vectors falling into each directional interval is counted, and the probability value of the a-th directional interval is calculated. Where a = 1, 2, ..., K; calculate the local optical flow direction entropy. As a measure of the degree of disorder in the distribution of local motion directions; The construction of the three-dimensional hidden danger structure tensor field S(x,y,t) is as follows: 。 6. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 1, characterized in that, In step S3, the observation value metric is calculated for the tensor field of the three-dimensional hidden danger structure. The neighbor unit distance factor is calculated by combining the location information of the shared neighbor mobile inspection units. The composite gradient of the observation value metric and the neighbor unit distance factor is used as the update direction to iteratively update the extended observation state of the mobile inspection units. During the update process, the step size ratio of different types of degrees of freedom in the extended observation state is adjusted through a precondition matrix. This includes the following steps: Current extended observation status of the mobile inspection unit Calculate the measurement of observational value The calculation formula is: in, This refers to the field of view cone region under the current extended observation state. Let Frobenius norm be the tensor field of the three-dimensional hidden danger structure. For scaling parameters The spatial resolution term that increases with increasing observation distance d and decreases with increasing observation distance d. For parameters that vary with sampling rate Increases with relative velocity The motion fuzzy term that increases and decreases accordingly; Based on the shared location information of neighboring mobile inspection units, and combined with the observation direction parameters and scaling parameters of the mobile inspection units, the neighbor unit distance with each neighboring mobile inspection unit is calculated. The neighbor unit distance of each neighbor is mapped using an exponential decay function, and the mapped values ​​of each neighboring mobile inspection unit are summed to obtain the neighbor unit distance factor. The tensor field of the local three-dimensional hidden danger structure pointed to by the current observation direction is decomposed into features. The eigenvector corresponding to the largest eigenvalue is extracted as the principal direction of the hidden danger, and a preconditioning matrix is ​​constructed accordingly. The precondition matrix is ​​a diagonal matrix. Among its diagonal elements, the element corresponding to the spatial translational degree of freedom along the main direction of the hazard takes the first step length weight, and the element corresponding to the spatial translational degree of freedom perpendicular to the main direction of the hazard takes the second step length weight. The first step length weight is greater than the second step length weight.

7. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 6, characterized in that, In step S3, the movement of the mobile inspection unit is controlled based on the extended observation state obtained through iterative updates, including the following steps: Calculate the extended observation state update increment Where N is the distance factor of the neighboring unit, and S is the tensor field of the three-dimensional hidden danger structure. Here, η is the gradient operator with respect to the extended observation state, and η is the step size factor; The extended observation state is updated incrementally based on the extended observation state, and the spatial position, observation direction, scaling parameters, and sampling rate parameters of the mobile inspection unit are controlled according to the updated extended observation state.

8. The intelligent hazard investigation and joint inspection method based on dynamic visual computing according to claim 1, characterized in that, In step S4, the activation degree of continuous potential energy nodes is obtained from the updated three-dimensional hidden danger structure tensor field. When the activation degree exceeds a preset priority threshold, the current inspection task is interrupted and the investigation task branch is switched, including the following steps: For continuous potential energy nodes, the area of ​​concern is determined based on the spatial location of the hazards recorded in the hazard investigation list. The activation degree A of the continuous potential energy nodes is calculated using the following formula: in, For the current moment, To preset the time window length, Here, γ is the Frobenius norm of the updated 3D hazard structure tensor field, γ is the gain coefficient, and θ is the preset bias threshold. It is an S-shaped function; When the activation level exceeds the preset priority threshold, the current inspection task is interrupted, and the investigation task branch corresponding to the continuous potential energy node with the highest activation level is selected as the target to be transferred.