A multi-modal intelligent monitoring device and method for electric power engineering construction
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU SAILAFU ELECTRIC POWER DEV CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-02
Smart Images

Figure CN122134109A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power engineering technology, and in particular to a multimodal intelligent monitoring device and method for power engineering construction. Background Technology
[0002] With the increasing scale and complexity of power engineering construction, safety management at construction sites faces severe challenges. The activities of personnel, equipment, and materials during construction are frequent and complex. If monitoring methods are limited or outdated, it will be difficult to detect potential risks in a timely manner, thus seriously threatening personnel safety and operational efficiency at the construction site.
[0003] Existing power engineering construction site monitoring technologies mostly adopt traditional video surveillance or manual inspection methods. These methods lack the precision of environmental perception and are difficult to accurately capture abnormal equipment conditions and dynamic risk propagation trends in complex environments. At the same time, due to the lack of effective fusion and linkage of multimodal information, traditional methods still have shortcomings in terms of real-time performance, accuracy, and early warning effect.
[0004] Therefore, how to effectively integrate multi-source sensing data to achieve proactive identification and accurate early warning of risks at construction sites has become a key technical problem that urgently needs to be solved in the field of power engineering construction safety. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a multimodal intelligent monitoring device and method for power engineering construction.
[0006] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a multimodal intelligent monitoring device for power engineering construction, the device comprising: Spatial Vision Fusion Module: Used to perform multimodal feature fusion based on 3D spatial perception data and visual monitoring data, extract the spatial location and corresponding category information of the target object, and generate a spatial motion trajectory carrying category information; Voiceprint feature extraction module: used to decompose acoustic perception data and generate local voiceprint features that characterize the abnormal operating conditions of the target object; Spatiotemporal risk association module: used to perform spatiotemporal mapping association between the spatial motion trajectory and the local features of the voiceprint, and generate risk propagation paths between target objects; Risk topology early warning module: It is used to construct a risk topology network with spatiotemporal characteristic memory mechanism based on the risk propagation path, so as to form risk clustering areas at the construction site and output risk early warning information.
[0007] Furthermore, the generation of the spatial motion trajectory carrying category information includes: Based on visual monitoring data, local texture gradient features of the target object surface are extracted, and key texture regions for classifying are determined according to the sparsity of the texture gradient features in spatial distribution. Based on 3D spatial perception data, determine the spatial location of target objects in multiple frames, and establish the spatiotemporal correspondence between the spatial location of target objects and key texture regions; Based on the spatiotemporal correspondence, a spatial motion trajectory carrying category information is generated.
[0008] Furthermore, the process for determining the key texture regions used to distinguish categories includes: Extract local texture gradient baseline features of different categories of target objects under normal working conditions; Based on the matching correlation between local texture gradient features in real-time visual monitoring data and preset multi-type texture gradient benchmark features, key texture regions that can distinguish the categories of target objects are determined.
[0009] Furthermore, the specific process of decomposing the acoustic sensing data based on the time-frequency feature decomposition algorithm that integrates sparse sensing mechanisms is as follows: Based on the local variation characteristics of acoustic sensing data in the time-frequency domain, the initial time-frequency features are extracted by multi-scale wavelet transform, and the sparse sensing constraints are determined based on the degree of sparse variation of the local energy distribution of the time-frequency features. The sparse sensing constraints are used to locally suppress redundant information in the initial time-frequency features and strengthen the sparse structure related to the abnormal mode, so as to generate local voiceprint features that characterize the abnormal working conditions of the target object.
[0010] Furthermore, the determination of sparse sensing constraints based on the degree of sparsity variation of local energy distribution with time-frequency characteristics includes: Extract the stable energy distribution of the acoustic signature features of the target object in different frequency bands under normal operating conditions to determine the initial sparsity baseline of the acoustic signature energy distribution; The deviation between the frequency band energy distribution of acoustic sensing data and the initial sparse baseline is analyzed in real time, and sparse sensing constraints are constructed based on the local abrupt change relationship of the deviation.
[0011] Furthermore, the risk association algorithm based on a local feature self-correction mechanism for spatiotemporal mapping and association of spatial motion trajectories and local voiceprint features is described in the following specific process: Identify the spatial intersection area of the spatial motion trajectories, and by analyzing the differences in local motion patterns of the target objects within the intersection area, preliminarily determine the critical moment when the target objects intersect in space; Extract the initial trigger time of abnormal patterns in the local features of the voiceprint, and perform local feature self-correction on the key moment by using the temporal mapping relationship of the differences in the local motion patterns of the target object in the spatial location; Based on the spatial correlation between the critical moments and the interaction patterns of the target objects after self-correction, a risk propagation path with high confidence is generated between the target objects.
[0012] Furthermore, the process of performing local feature self-correction at key moments is as follows: Extract the local velocity change trend in the spatial motion trajectory corresponding to the initial trigger time of the abnormal mode, and determine the velocity change abrupt point by the spatial distribution gradient of the velocity change trend. Based on the temporal matching relationship between the abrupt change in velocity and the initial triggering time of the abnormal local feature pattern of the voiceprint, local feature self-correction is performed on key moments.
[0013] Furthermore, the formation process of the risk-accumulation zone at the construction site includes: Extract the frequency of spatial motion interactions between target objects in the risk propagation path, and construct the risk topology network by combining the historical spatiotemporal feature weights in the risk topology network; Local clustering analysis is performed on the interaction intensity between target objects in the risk topology network, and the local high-density regions of the risk propagation path are determined by the spatial clustering density of the interaction intensity. Based on the spatial coupling relationship between the local high-density area and the risk propagation path, a risk clustering area with boundaries is generated at the construction site.
[0014] Furthermore, the process of determining the local high-density region includes: Based on the spatial movement trajectory of target objects in the risk propagation path, identify the intersection events where the spatial distance between target objects is continuously less than a preset spatial distance threshold, calculate the intersection frequency and duration, and construct an initial interaction feature matrix between target objects; The initial interaction feature matrix is normalized, and the normalized interaction intensity distribution is calculated. Based on the gradient variation pattern of the normalized interaction intensity distribution in spatial location, local extreme value regions of the interaction intensity gradient are identified to determine the local high-density regions of the risk propagation path at the construction site.
[0015] Secondly, the present invention provides a multimodal intelligent monitoring method for power engineering construction, the method comprising: Based on 3D spatial perception data and visual monitoring data, multimodal feature fusion is performed to extract the spatial location and corresponding category information of the target object and generate a spatial motion trajectory carrying category information. The acoustic sensing data is decomposed to generate local acoustic features that characterize the abnormal operating conditions of the target object. Spatiotemporal mapping and correlation are performed on the spatial motion trajectory and the local features of the voiceprint to generate risk propagation paths between target objects; Based on the risk propagation path, a risk topology network with a spatiotemporal memory mechanism is constructed to form risk clustering areas at the construction site and output risk warning information.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention uses multimodal feature fusion to accurately obtain the real-time spatial location and category information of target objects at the construction site, effectively overcoming the shortcomings of inaccurate target identification and susceptibility to environmental interference under traditional single monitoring methods, and improving the reliability of monitoring in complex environments.
[0017] This invention uses a voiceprint feature decomposition algorithm based on a sparse sensing mechanism to accurately identify abnormal operating patterns of target objects, thereby improving the sensitivity and real-time performance of acoustic anomaly monitoring and effectively avoiding the delayed early warning problem of traditional acoustic monitoring methods.
[0018] This invention constructs a risk topology network with a spatiotemporal memory mechanism to achieve accurate prediction of risk propagation paths and dynamic identification of risk areas at construction sites. This effectively improves the predictability of risk management, reduces the probability of accidents at power engineering construction sites, and ensures the safety and efficiency of on-site construction. Attached Figure Description
[0019] Figure 1 This is an architecture diagram of a multimodal intelligent monitoring device for power engineering construction, as shown in Example 1.
[0020] Figure 2 This is a flowchart of a multimodal intelligent monitoring method for power engineering construction, as shown in Example 2. Detailed Implementation
[0021] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1 Please see Figure 1 This invention provides a multimodal intelligent monitoring device for power engineering construction, comprising: Spatial Vision Fusion Module: Used to perform multimodal feature fusion based on 3D spatial perception data and visual monitoring data, extract the spatial location and corresponding category information of the target object, and generate a spatial motion trajectory carrying category information; In implementation, this embodiment first uses a high-precision three-dimensional laser scanning device deployed around the construction site to acquire three-dimensional spatial perception data of each target object in the construction site in real time. The three-dimensional spatial perception data includes, but is not limited to, spatial coordinate data and point cloud density data.
[0023] It should be understood that spatial coordinate data is in a spatial rectangular coordinate system. Based on the reference, it represents the position of the target object's surface in the spatial coordinate system of the construction site; the point cloud density data is the number of point cloud data per unit volume, used to reflect the difference between the target object and the environmental space.
[0024] Secondly, high-definition camera devices deployed at different locations on the construction site collect visual monitoring data in real time. The visual monitoring data includes: two-dimensional image data sequence of the target object and continuous video frame data of the target object's motion.
[0025] Understandably, the two-dimensional image data sequence of the target object is used to identify the category features of the target object; the continuous video frame data of the target object's motion is used to capture the object's texture changes and motion trajectory information.
[0026] Specifically, generating the spatial motion trajectory carrying category information includes: Based on visual monitoring data, local texture gradient features of the target object surface are extracted, and key texture regions for classifying are determined according to the sparsity of the texture gradient features in spatial distribution. In implementation, this embodiment first uses a local gradient extraction method (such as the Sobel operator) to calculate the texture gradient of the two-dimensional image of the visual monitoring data to obtain the local gradient magnitude at different locations on the surface of the target object; wherein, the local gradient magnitude is determined based on the square root of the sum of the squares of the gray-level change rates of the pixels in the horizontal and vertical directions.
[0027] Furthermore, to fully reflect the texture features of the target object, this embodiment statistically analyzes the following texture gradient features for multiple local regions (e.g., 10×10 pixels) divided in the image space: the mean of the gradient magnitude, used to characterize the overall intensity of gradient changes within the region; the variance of the gradient magnitude, used to characterize the degree of gradient changes within the region; and the sparsity of the gradient magnitude, used to characterize the proportion of pixels with low gradient magnitudes within the region. The sparsity is determined by calculating the proportion of pixels with gradient magnitudes less than a preset magnitude threshold (e.g., 0.05) within the region to the total number of pixels in that region.
[0028] Subsequently, based on the gradient sparsity calculation results, candidate key texture regions with gradient sparsity less than a preset sparsity threshold (e.g., 50%) are initially selected so that the target object category information can be further accurately determined.
[0029] Furthermore, the process for determining the key texture regions used to distinguish categories includes: Extract local texture gradient baseline features of different categories of target objects under normal working conditions; In practice, multiple sample images of different categories of target objects at the construction site under normal working conditions are collected, and the local texture gradient feature vectors of each sample surface (such as mean, variance, and sparsity combination features) are calculated. Then, the texture gradient features of samples of the same category are subjected to arithmetic average or weighted statistical analysis to obtain the local texture gradient baseline feature vectors of different categories of target objects.
[0030] It should be understood that the benchmark features can be represented in the form of multi-dimensional feature vectors to reflect the different texture characteristics of each type of target.
[0031] Based on the matching correlation between local texture gradient features in real-time visual monitoring data and preset multi-class texture gradient benchmark features, key texture regions that can distinguish the target object category are determined. The specific implementation is as follows: Real-time computation of texture gradient feature vectors for each candidate region in visual monitoring data This embodiment will use... Compare with all known categories in the preset feature library respectively Texture gradient baseline feature vector Compare them.
[0032] In implementation, calculation With each category Minimum vector Euclidean distance between If the minimum distance If the value is less than the preset matching and recognition threshold, the candidate region is considered to contain texture information that can characterize a specific category feature, and it is identified as a key texture region, and a corresponding category label is initially assigned to the region. .
[0033] Subsequently, based on a preset deviation threshold, features with excessive fluctuations (i.e., Interference areas (abnormal fluctuation ranges deviating from the preset center value) are identified, thereby ultimately determining the key texture regions containing candidate category features, and using them as regions of interest for subsequent deep learning recognition.
[0034] It should be noted that the preset deviation threshold was determined based on multiple experiments, and its value ranges from 1.5 times to 3 times the standard deviation obtained from the statistical analysis of the local texture gradient baseline features.
[0035] Based on 3D spatial perception data, determine the spatial location of target objects in multiple frames, and establish the spatiotemporal correspondence between the spatial location of target objects and key texture regions; In the specific implementation process, firstly, this embodiment establishes a unified global coordinate system for the construction site by pre-calibrating the high-definition camera device and the 3D laser scanning device; then, using continuous multi-frame 3D point cloud data, a multi-frame registration algorithm (such as the ICP algorithm) is used to determine the spatial position sequence of the target object relative to the global coordinate system at different times. ; Secondly, based on the camera intrinsic parameter matrix obtained from the joint calibration And the extrinsic parameter matrix of the high-definition camera device relative to the three-dimensional laser scanning device A geometric mapping relationship is established between the visual image coordinate system and the 3D point cloud spatial coordinate system. Through this mapping relationship, key texture region feature points in the 2D image are projected onto the 3D point cloud, achieving precise matching between the key texture regions and their 3D spatial positions. The formula for calculating the mapping relationship is: ; In the formula, Represents the coordinates of a two-dimensional image. This represents the three-dimensional spatial coordinates in the corresponding global coordinate system. Indicates the projection scale factor. express Rotation matrix, express Translation vector.
[0036] Based on the spatiotemporal correspondence, a spatial motion trajectory carrying category information is generated.
[0037] It should be understood that, in order to reduce the computational overhead of real-time processing of full-frame images of the construction site, this embodiment uses the key texture region as the region of interest suggestion window of the deep learning model; real-time target detection algorithms (such as YOLO, Faster R-CNN) are used to perform local deep feature extraction and refined target attribute classification only on the region of interest, thereby improving the detection frame rate while ensuring recognition accuracy and obtaining high-confidence category label information; then the category label information obtained by recognition is fused with the above spatial position sequence frame by frame to form the final spatial motion trajectory data carrying category information, specifically represented as follows: ; In the formula, This represents a spatial motion trajectory carrying category information. Indicates the first The three-dimensional spatial coordinates of the target object in the frame. This represents the total number of frames sampled. This indicates the semantic category label of the target object in the corresponding frame, such as construction machinery, construction workers, or safety facilities.
[0038] Voiceprint feature extraction module: used to decompose acoustic perception data and generate local voiceprint features that characterize the abnormal operating conditions of the target object; In implementation, this embodiment first acquires acoustic sensing data generated during the operation of the target object in real time by using high-sensitivity acoustic sensors deployed at key locations on the construction site. The acoustic sensing data specifically includes a continuous time-domain sound pressure amplitude signal s(t), where t represents the sampling time point and s(t) represents the sound pressure amplitude at that moment. In order to effectively identify the abnormal operating conditions of the target object, the acoustic sensing data needs to be analyzed in the time-frequency domain to obtain local acoustic features that can clearly reflect the abnormal characteristics.
[0039] Specifically, the decomposition of acoustic sensing data based on the time-frequency feature decomposition algorithm that integrates sparse sensing mechanisms is carried out as follows: Based on the local variation characteristics of acoustic sensing data in the time-frequency domain, the initial time-frequency features are extracted by multi-scale wavelet transform, and the sparse sensing constraints are determined based on the degree of sparse variation of the local energy distribution of the time-frequency features. In the specific implementation process, this embodiment uses continuous wavelet transform (CWT) to process the original acoustic data s(t) to obtain time-frequency domain wavelet coefficients at different scales and time offsets. The wavelet basis functions can be Morlet or Daubechies wavelet basis functions. Subsequently, by calculating the square of the modulus of the wavelet coefficients, the energy distribution characteristics corresponding to each time-frequency point are obtained.
[0040] To effectively characterize the differences in acoustic signal anomalous patterns in the time-frequency domain, this embodiment divides the time-frequency plane into several local analysis units. Each analysis unit is defined by a specific time interval and frequency interval, and the entropy value of the energy distribution is calculated within each unit to characterize the sparsity of the local energy distribution. The specific calculation method is as follows: ; In the formula, This represents the energy entropy value within a local analysis cell. Indicates the first unit within the cell The proportion of energy at each time-frequency point to the total energy within the unit is specifically expressed as follows: ; In the formula, Indicates the first unit within the local analysis cell Energy value at each time frequency point This represents the total number of time-frequency points within a local analysis unit.
[0041] The determination of sparse sensing constraints based on the degree of sparsity variation of local energy distribution with time-frequency characteristics includes: Extract the stable energy distribution of the acoustic signature features of the target object in different frequency bands under normal operating conditions to determine the initial sparsity baseline of the acoustic signature energy distribution; It should be noted that in this embodiment, multiple acoustic samples of the target object are collected under normal operating conditions, and the wavelet transform and energy entropy value calculation are performed on each sample as described above. The energy entropy value is statistically analyzed in each frequency band, and the average entropy value of multiple normal samples is taken as the initial sparsity reference for that frequency band.
[0042] The deviation between the frequency band energy distribution of acoustic sensing data and the initial sparse baseline is analyzed in real time, and sparse sensing constraints are constructed based on the local abrupt change relationship of the deviation.
[0043] In the specific implementation process, this embodiment uses the same method as under normal operating conditions to calculate the real-time frequency band energy entropy value of the acoustic sensing data collected in real time. Then compare it with the reference entropy value of the corresponding frequency band. The deviation degree is calculated using the following formula: ; In the formula, This indicates the degree of deviation between the real-time entropy value and the baseline entropy value. Indicates real-time acoustic data in the frequency band and time The energy entropy value at that location.
[0044] Furthermore, this embodiment determines the statistical distribution of the degree of entropy deviation in each frequency band through long-term experimental data analysis, using the standard deviation of the statistical distribution of entropy deviation. For reference, the preset deviation threshold is: When real-time acoustic data is in the frequency band ,time The degree of deviation of the entropy value satisfy When the time-frequency region is considered to have deviated, it is then determined to be the time-frequency region of the sparse sensing constraint condition corresponding to the abnormal mode.
[0045] The sparse sensing constraints are used to locally suppress redundant information in the initial time-frequency features and strengthen the sparse structure related to the abnormal mode, so as to generate local voiceprint features that characterize the abnormal working conditions of the target object.
[0046] In specific implementation, this embodiment performs weighted enhancement processing on the energy characteristics of abnormal regions that meet the sparse sensing constraints, while suppressing the energy of normal regions that do not meet the sparse sensing constraints. The specific calculation method is as follows: ; In the formula, This represents the energy characteristics after weighted processing. This represents the weighting coefficient corresponding to the time-frequency position. In this embodiment, the weighting selection principle is determined by pre-testing the signal-to-noise ratio distribution under normal and abnormal operating conditions, i.e., the weighting coefficient based on the time-frequency position. When it belongs to an anomaly region that satisfies the sparse perception constraint. An enhancement factor with a value between 1.5 and 2 is used to amplify abnormal signal components; time-frequency position When it belongs to a normal region that does not meet the sparse perception constraint, An attenuation coefficient between 0.5 and 0.8 is used to suppress ambient noise; finally, local acoustic features are obtained to characterize the abnormal operating conditions of the target object, specifically expressed as follows: ; In the formula, To characterize the set of local voiceprint features representing the abnormal operating conditions of the target object, A set of anomalous time-frequency regions to satisfy sparse sensing constraints.
[0047] Spatiotemporal risk association module: used to perform spatiotemporal mapping association between the spatial motion trajectory and the local features of the voiceprint, and generate risk propagation paths between target objects; In implementation, the risk association algorithm based on a local feature self-correction mechanism for spatiotemporal mapping and association of spatial motion trajectories and local voiceprint features is as follows: Identify the spatial intersection area of the spatial motion trajectories, and by analyzing the differences in local motion patterns of the target objects within the intersection area, preliminarily determine the critical moment when the target objects intersect in space; In specific implementation, this embodiment first performs spatial clustering analysis on the spatial motion trajectory data to identify the trajectory intersection area; the spatial intersection area is determined to be the spatial area where the minimum distance between the motion trajectories of two or more target objects in three-dimensional space is less than a preset distance threshold (e.g., 0.5m); specifically, the minimum spatial distance is determined by calculating the vector displacement magnitude between the spatial position coordinates of any two target objects in real time.
[0048] Taking motion speed as an example, the real-time motion speed of the target object is obtained by calculating the ratio of the displacement of the target object's spatial position between adjacent sampling times to the sampling time interval.
[0049] By statistically analyzing the trend of movement speed changes, the speed difference when the target objects enter the intersection area is identified, and the moment when the speed difference first exceeds or equals a preset threshold (e.g., 0.2 m / s) is determined as the initial key moment of spatial intersection.
[0050] Extract the initial trigger time of abnormal patterns in the local features of the voiceprint, and perform local feature self-correction on the key moment by using the temporal mapping relationship of the differences in the local motion patterns of the target object in the spatial location; In specific implementation, this embodiment uses a set of local voiceprint features. The initial trigger time of the abnormal mode is extracted; the initial trigger time of the abnormal mode is defined as the time point when the voiceprint energy feature first exceeds or equals the normal operating condition threshold, specifically expressed as: ; In the formula, Indicates the initial trigger time of the abnormal mode. Indicates time The energy characteristics corresponding to the anomalous time-frequency region This indicates the preset energy threshold determined by normal operating conditions.
[0051] Furthermore, this embodiment utilizes the difference in motion patterns in spatial location and the initial triggering time of the abnormal voiceprint pattern for timing mapping. The method for determining the timing deviation is as follows: ; in, Indicates timing deviation. This indicates the key moment of the spatial convergence that has been preliminarily determined above.
[0052] It should be noted that when the deviation amount When the time is greater than or equal to the preset allowable range (e.g., greater than or equal to 1 second), the key moment of spatial intersection is precisely adjusted through local feature self-correction to eliminate timing deviations.
[0053] Specifically, the process of performing local feature self-correction at key moments is as follows: Extract the local velocity change trend in the spatial motion trajectory corresponding to the initial trigger time of the abnormal mode, and determine the velocity change abrupt point by the spatial distribution gradient of the velocity change trend. In the implementation process, this embodiment first uses a three-dimensional spline interpolation algorithm to fit the discrete trajectory points of the target object, and constructs a continuous velocity field function that varies with spatial position. Subsequently, spatial gradient analysis is performed on the velocity field to characterize the degree of velocity abrupt change of the target object when passing through a specific spatial region. The specific calculation formula is as follows: ; In the formula, Indicates spatial location The velocity space gradient at that location, This represents a local continuous motion velocity function constructed through fitting.
[0054] Understandably, by calculating and analyzing the distribution of this spatial velocity gradient value, it can be identified that the spatial velocity gradient value is greater than or equal to a preset gradient threshold (e.g., 0.1 m / s). 2 The moment corresponding to the spatial location of a local region is determined as the moment of the abrupt change in velocity.
[0055] Based on the temporal matching relationship between the abrupt change in velocity and the initial triggering time of the abnormal local feature pattern of the voiceprint, local feature self-correction is performed on key moments.
[0056] In specific implementation, this embodiment calculates the moment of the abrupt change in velocity. and the initial trigger time of the abnormal mode The temporal differences between them are used to determine the abrupt change in velocity closest to the initial trigger moment of the abnormal mode, as shown below: ; In the formula, This indicates the critical moment of spatial intersection that is finally determined after local feature self-correction processing.
[0057] The adjusted time As a key moment of spatial convergence after local feature self-correction, it ensures accurate temporal matching between abnormal voiceprint patterns and abnormal spatial motion features.
[0058] Based on the spatial correlation between the critical moments and the interaction patterns of the target objects after self-correction, a risk propagation path with high confidence is generated between the target objects.
[0059] In specific implementation, this embodiment, based on the key moments of spatial intersection after self-correction through the above steps, and combined with the relative motion state between target objects within the spatial trajectory intersection area, further analyzes the interaction risk patterns between target objects. To accurately characterize the collision or interference risks between target objects (especially during relative motion), this embodiment uses an index based on relative velocity magnitude and spatial distance weights to determine the interaction intensity. The specific calculation formula is as follows: ; In the formula, Indicates time target object With the target object The strength of interaction between them Indicates the two target objects at time... The magnitude of the relative velocity vector, Indicates the two target objects at time... Euclidean distance, This represents a preset, extremely small positive constant to avoid the denominator being zero.
[0060] Finally, based on the above interaction intensity and duration, a risk propagation path is constructed, specifically represented as follows: ; In the formula, This indicates the risk propagation path between target objects. This indicates the preset interaction strength threshold. Indicates the allowable time deviation range (e.g., 2s).
[0061] Understandably, the above calculations ensure that when two target objects approach each other at a high relative speed, their interaction intensity increases, reflecting the physical risk trend at the construction site.
[0062] Risk topology early warning module: It is used to construct a risk topology network with spatiotemporal characteristic memory mechanism based on the risk propagation path, so as to form risk clustering areas at the construction site and output risk early warning information.
[0063] It is understood that this embodiment first constructs a risk topology network with a spatiotemporal feature memory mechanism based on the risk propagation path determined above, and identifies risk clustering areas at the construction site by analyzing the network structure characteristics, and finally achieves effective early warning of the risk status at the construction site.
[0064] Specifically, the formation process of the risk-accumulation area at the construction site includes: Extract the frequency of spatial motion interactions between target objects in the risk propagation path, and construct the risk topology network by combining the historical spatiotemporal feature weights in the risk topology network; In practical implementation, this embodiment maps each target object within the monitoring area to a location node in the risk topology network, defining the risk topology network. , where the set of nodes The location node of the target object, and the set of edges. It represents the interactive evolution relationship between nodes; it counts the frequency of spatial motion interactions between any two target objects through risk propagation paths, and defines the interaction frequency. For two target objects , The cumulative number of times the spatial distance is less than a preset distance threshold (e.g., 0.5m) within a preset time period is represented as follows: ; In the formula, This is an indicator function that returns 1 if the condition is true, and 0 otherwise. Represents the target object , At any moment Spatial distance, This indicates the preset distance threshold.
[0065] Furthermore, to effectively reflect the impact of historical characteristics of risk propagation on current risk prediction, this embodiment introduces historical spatiotemporal feature weights. Specifically, it is expressed as: ; In the formula, Indicates the previous time period target object , Historical weight between them This indicates the frequency of interactions between target objects during the current time period. This represents the weighting adjustment coefficient, with a value ranging from 0.5 to 0.9. The specific value is determined based on the stability of the risk characteristics at the construction site.
[0066] By combining the historical weights and current interaction frequencies, the final comprehensive weight of the risky topology network edges is obtained. , is represented as: .
[0067] The above calculations form a complete risk topology network. .
[0068] Local clustering analysis is performed on the interaction intensity between target objects in the risk topology network, and the local high-density regions of the risk propagation path are determined by the spatial clustering density of the interaction intensity. In implementation, this embodiment identifies key clustering nodes in risk evolution by spatiotemporally quantifying the interaction behavior in the risk topology network; specifically, the process of determining the local high-density region includes: Based on the spatial movement trajectory of target objects in the risk propagation path, identify the intersection events where the spatial distance between target objects is continuously less than a preset spatial distance threshold, calculate the intersection frequency and duration, and construct an initial interaction feature matrix between target objects; In the specific implementation process, this embodiment first extracts the spatial motion trajectory data of each target object from the risk propagation path. By traversing the spatial position of each time point in the risk propagation path, it identifies objects whose spatial distance is continuously lower than a preset spatial distance threshold. Intersection events (e.g., 0.5m-1.0m).
[0069] For each pair of target objects, this embodiment records the cumulative frequency of their intersection events within a preset monitoring period. and the average duration of a single intersection event .
[0070] Based on this, an initial interaction feature matrix is constructed using the target object as the row and column coordinates. Specifically, it is expressed as: ; In the formula, Represents the total number of target objects within the monitoring area; matrix elements Represents the target object and target object The original cumulative intensity of spatial interaction behavior between them.
[0071] The initial interaction feature matrix is normalized, and the normalized interaction intensity distribution is calculated. In the specific implementation process, in order to eliminate the influence of sampling frequency on data dimensions under different construction conditions, this embodiment adjusts the initial interactive feature matrix. Normalization is performed; by calculating the difference between each element in the matrix and the minimum value, and dividing by the range of the entire matrix data, the interaction feature data is mapped to the [0,1] interval, resulting in the normalized interaction intensity distribution. .
[0072] It should be understood that the interaction intensity distribution It can intuitively reflect the relative strength of the risk relationship between various nodes within the construction site.
[0073] Based on the gradient variation pattern of the normalized interaction intensity distribution in spatial location, local extreme value regions of the interaction intensity gradient are identified to determine the local high-density regions of the risk propagation path at the construction site.
[0074] In specific implementation, this embodiment is based on the normalized interaction intensity distribution. By combining the real-time spatial coordinates of the corresponding nodes at the construction site, spatial gradient analysis of interaction intensity is performed to accurately identify the local evolution center of interaction intensity in spatial location.
[0075] First, this embodiment extracts each pair of target objects that generate interactive behavior. and Spatial coordinates of the center point at the moment of intersection And the corresponding normalized interaction intensity distribution This serves as the energy weight at that coordinate point; furthermore, a Gaussian kernel density estimation algorithm is employed, using the coordinates of each center point... Spatial interpolation and smoothing are performed on the source point to map the discrete interaction intensity to the monitoring space, thereby constructing a continuous risk interaction intensity potential field. The spatial gradient value is calculated based on this potential energy field, specifically expressed as follows: ; Specifically, the partial derivatives are calculated using the finite difference method.
[0076] Subsequently, this embodiment calculates the spatial gradient magnitude. Specifically, it is expressed as: ; Secondly, statistical analysis is performed on the calculated spatial gradient magnitudes to identify gradient magnitudes greater than or equal to a preset gradient threshold. Local extreme value regions (e.g., 1.5 times the mean of the gradient magnitude) are identified as core candidate points for risk propagation paths.
[0077] Furthermore, this local extreme value area is considered as a local high-density area of the risk propagation path at the construction site.
[0078] Finally, this embodiment uses a density clustering algorithm (such as the DBSCAN algorithm) to cluster the nodes in the aforementioned local extreme value region.
[0079] In specific implementation, taking the extreme value region node as the center, within the radius... Spatial clustering density for calculating interaction strength within the neighborhood Specifically, it is expressed as: ; In the formula, Represented by node Centered on, with radius The set of all nodes within the spatial neighborhood, This indicates the number of nodes in the set.
[0080] When spatial cluster density When the density is greater than or equal to the preset high-density threshold, the connected area is ultimately determined as a local high-density area of the risk propagation path at the construction site.
[0081] It should be noted that the radius... The value range is between 0.5m and 1.5m, and the preset high-density threshold ranges between 0.60 and 0.85.
[0082] Based on the spatial coupling relationship between the local high-density area and the risk propagation path, a risk clustering area with boundaries is generated at the construction site.
[0083] In this embodiment, the local high-density areas determined by density clustering are taken as the core risk sources. By calculating the envelope range of the risk propagation path in space, the influence boundary of risk evolution is determined.
[0084] First, in this embodiment, the node set generated by the density clustering described above is... Spatial projection overlap verification with real-time risk propagation paths; definition of overlap coefficient. Risk transmission path The ratio of the sum of the path segment lengths traversing a local high-density area to the total length of the risk propagation path; when When the value exceeds a preset coupling threshold (e.g., 0.7), the local high-density area is determined to be a valid risk clustering center.
[0085] Secondly, this embodiment employs either the convex hull algorithm or the concave hull algorithm for the selected set of nodes. The risk propagation trajectory points coupled with these points are fitted with circumscribed polygons to generate a construction site risk cluster area with three-dimensional spatial geometric boundaries.
[0086] Finally, based on the generated risk clustering areas, risk warning information is output; the risk warning information includes: the dynamic boundary three-dimensional coordinates of the risk clustering areas, and the risk level within the area (based on interaction intensity). The mean value is determined, along with a list of target object types. This visualized boundary guides on-site safety managers to make targeted interventions to ensure the safety of construction site operations.
[0087] Example 2 like Figure 2 As shown, based on the same inventive concept and referring to Embodiment 1, this embodiment discloses a multimodal intelligent monitoring method for power engineering construction. For details not covered in this embodiment, please refer to the relevant parts of Embodiment 1 above. The method includes: Based on 3D spatial perception data and visual monitoring data, multimodal feature fusion is performed to extract the spatial location and corresponding category information of the target object and generate a spatial motion trajectory carrying category information. The acoustic sensing data is decomposed to generate local acoustic features that characterize the abnormal operating conditions of the target object. Spatiotemporal mapping and correlation are performed on the spatial motion trajectory and the local features of the voiceprint to generate risk propagation paths between target objects; Based on the risk propagation path, a risk topology network with a spatiotemporal memory mechanism is constructed to form risk clustering areas at the construction site and output risk warning information.
[0088] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.
[0089] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A multimodal intelligent monitoring device for power engineering construction, characterized in that, The device includes: Spatial Vision Fusion Module: Used to perform multimodal feature fusion based on 3D spatial perception data and visual monitoring data, extract the spatial location and corresponding category information of the target object, and generate a spatial motion trajectory carrying category information; Voiceprint feature extraction module: used to decompose acoustic perception data and generate local voiceprint features that characterize the abnormal operating conditions of the target object; Spatiotemporal risk association module: used to perform spatiotemporal mapping association between the spatial motion trajectory and the local features of the voiceprint, and generate risk propagation paths between target objects; Risk topology early warning module: It is used to construct a risk topology network with spatiotemporal characteristic memory mechanism based on the risk propagation path, so as to form risk clustering areas at the construction site and output risk early warning information.
2. The multimodal intelligent monitoring device for power engineering construction according to claim 1, characterized in that, The generation of spatial motion trajectories carrying category information includes: Based on visual monitoring data, local texture gradient features of the target object surface are extracted, and key texture regions for classifying are determined according to the sparsity of the texture gradient features in spatial distribution. Based on 3D spatial perception data, determine the spatial location of target objects in multiple frames, and establish the spatiotemporal correspondence between the spatial location of target objects and key texture regions; Based on the spatiotemporal correspondence, a spatial motion trajectory carrying category information is generated.
3. The multimodal intelligent monitoring device for power engineering construction according to claim 2, characterized in that, The process for determining the key texture regions used to distinguish categories includes: Extract local texture gradient baseline features of different categories of target objects under normal working conditions; Based on the matching correlation between local texture gradient features in real-time visual monitoring data and preset multi-type texture gradient benchmark features, key texture regions that can distinguish the categories of target objects are determined.
4. The multimodal intelligent monitoring device for power engineering construction according to claim 1, characterized in that, The specific process of the time-frequency feature decomposition algorithm based on the fusion sparse sensing mechanism for decomposing acoustic sensing data is as follows: Based on the local variation characteristics of acoustic sensing data in the time-frequency domain, the initial time-frequency features are extracted by multi-scale wavelet transform, and the sparse sensing constraints are determined based on the degree of sparse variation of the local energy distribution of the time-frequency features. The sparse sensing constraints are used to locally suppress redundant information in the initial time-frequency features and strengthen the sparse structure related to the abnormal mode, so as to generate local voiceprint features that characterize the abnormal working conditions of the target object.
5. A multimodal intelligent monitoring device for power engineering construction according to claim 4, characterized in that, The determination of sparse sensing constraints based on the degree of sparsity variation of local energy distribution with time-frequency characteristics includes: Extract the stable energy distribution of the acoustic signature features of the target object in different frequency bands under normal operating conditions to determine the initial sparsity baseline of the acoustic signature energy distribution; The deviation between the frequency band energy distribution of acoustic sensing data and the initial sparse baseline is analyzed in real time, and sparse sensing constraints are constructed based on the local abrupt change relationship of the deviation.
6. A multimodal intelligent monitoring device for power engineering construction according to claim 1, characterized in that, The risk association algorithm based on a local feature self-correction mechanism, which performs spatiotemporal mapping and association between spatial motion trajectories and local voiceprint features, is as follows: Identify the spatial intersection area of the spatial motion trajectories, and by analyzing the differences in local motion patterns of the target objects within the intersection area, preliminarily determine the critical moment when the target objects intersect in space; Extract the initial trigger time of abnormal patterns in the local features of the voiceprint, and perform local feature self-correction on the key moment by using the temporal mapping relationship of the differences in the local motion patterns of the target object in the spatial location; Based on the spatial correlation between the critical moments and the interaction patterns of the target objects after self-correction, a risk propagation path with high confidence is generated between the target objects.
7. A multimodal intelligent monitoring device for power engineering construction according to claim 6, characterized in that, The process of performing local feature self-correction at key moments is as follows: Extract the local velocity change trend in the spatial motion trajectory corresponding to the initial trigger time of the abnormal mode, and determine the velocity change abrupt point by the spatial distribution gradient of the velocity change trend. Based on the temporal matching relationship between the abrupt change in velocity and the initial triggering time of the abnormal local feature pattern of the voiceprint, local feature self-correction is performed on key moments.
8. A multimodal intelligent monitoring device for power engineering construction according to claim 1, characterized in that, The formation process of the risk-concentration zone at the construction site includes: Extract the frequency of spatial motion interactions between target objects in the risk propagation path, and construct the risk topology network by combining the historical spatiotemporal feature weights in the risk topology network; Local clustering analysis is performed on the interaction intensity between target objects in the risk topology network, and the local high-density regions of the risk propagation path are determined by the spatial clustering density of the interaction intensity. Based on the spatial coupling relationship between the local high-density area and the risk propagation path, a risk clustering area with boundaries is generated at the construction site.
9. A multimodal intelligent monitoring device for power engineering construction according to claim 8, characterized in that, The process of determining the local high-density region includes: Based on the spatial movement trajectory of target objects in the risk propagation path, identify the intersection events where the spatial distance between target objects is continuously less than a preset spatial distance threshold, calculate the intersection frequency and duration, and construct an initial interaction feature matrix between target objects; The initial interaction feature matrix is normalized, and the normalized interaction intensity distribution is calculated. Based on the gradient variation pattern of the normalized interaction intensity distribution in spatial location, local extreme value regions of the interaction intensity gradient are identified to determine the local high-density regions of the risk propagation path at the construction site.
10. A multimodal intelligent monitoring method for power engineering construction, characterized in that, The method includes: Based on 3D spatial perception data and visual monitoring data, multimodal feature fusion is performed to extract the spatial location and corresponding category information of the target object and generate a spatial motion trajectory carrying category information. The acoustic sensing data is decomposed to generate local acoustic features that characterize the abnormal operating conditions of the target object. Spatiotemporal mapping and correlation are performed on the spatial motion trajectory and the local features of the voiceprint to generate risk propagation paths between target objects; Based on the risk propagation path, a risk topology network with a spatiotemporal memory mechanism is constructed to form risk clustering areas at the construction site and output risk warning information.