Remote online monitoring method and system based on machine vision and artificial intelligence
Multimodal data is collected through edge computing nodes for cross-domain feature alignment and lightweight deep learning, and intelligent early warning is carried out in combination with historical data, solving the problems of insufficient fusion of multimodal data and complexity of deep diagnostic models, and achieving efficient equipment status monitoring and adaptive alarms.
Patent Information
- Application Number
- CN202510571529.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-05-06
AI Technical Summary
When facing multimodal data in complex industrial environments, the prior art has problems such as insufficient feature alignment and fusion, complex and difficult to deploy deep diagnostic models, and lack of dynamic learning capabilities in alarm systems, resulting in low information utilization and frequent false alarms and missed reports.
Multimodal data is collected through edge computing nodes, cross-domain feature alignment and lightweight deep learning, intelligent early warning is carried out in combination with historical data, a composite feature matrix of time and space synchronization is constructed, and an adaptive alarm threshold is generated.
It realizes efficient feature fusion and real-time diagnosis of multimodal data, improves the accuracy of device status perception and the intelligent response capabilities of the system, and reduces the false alarm rate and missed alarm rate.
Smart Images

Figure CN120105312B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial intelligent monitoring, and in particular to a remote online monitoring method and system based on machine vision and artificial intelligence. Background Art
[0002] As industrial equipment safety and maintenance management become increasingly intelligent, remote online monitoring methods based on multimodal perception and artificial intelligence algorithms are becoming a key support tool. With the development of technologies such as machine vision, infrared thermal imaging, vibration sensing, edge computing, and deep learning, traditional equipment monitoring methods, which rely primarily on single-modality, fixed-point inspections, are shifting towards multi-source integration, real-time diagnosis, and intelligent early warning.
[0003] However, when faced with complex operating environments and multimodal data, existing technologies still have the following key technical problems that need to be solved urgently: First, heterogeneous modal data have obvious differences in time granularity, spatial structure and semantic levels, and lack effective feature alignment and fusion mechanisms, resulting in low utilization of multimodal information; Second, deep diagnostic models generally have complex structures, large parameters, and are difficult to deploy on edge devices, which limits their real-time online application capabilities; Third, most alarm systems are based on static threshold triggering, lack the ability to dynamically learn and adaptively adjust the historical operating status of the equipment, and are prone to false alarms or missed alarms. Summary of the Invention
[0004] Based on the above objectives, the present invention provides a remote online monitoring method and system based on machine vision and artificial intelligence, which supports multimodal spatiotemporal feature fusion, has lightweight diagnostic capabilities, and can link historical data for intelligent early warning and response remote online monitoring to meet the high-precision and high-robustness intelligent perception requirements in complex industrial scenarios.
[0005] The remote online monitoring method based on machine vision and artificial intelligence includes the following steps:
[0006] S1: The edge computing node collects visible light video stream, infrared thermal imaging data, and three-dimensional vibration spectrum data of the target device to form a multimodal sensing data set;
[0007] S2: Perform cross-domain feature alignment processing on the multimodal sensor data set to generate a spatiotemporally synchronized composite feature matrix.
[0008] The cross-domain feature alignment process includes: extracting key frames from the visible light video stream, using a spatiotemporal interpolation algorithm to increase the infrared thermal imaging sampling rate to synchronize with the vibration spectrum, and constructing a cross-modal feature correlation map based on the Gram matrix;
[0009] S3: Input the composite feature matrix into the cascade deep learning model, and output a three-dimensional diagnostic vector including equipment health status classification, abnormal area location, and failure probability assessment.
[0010] The cascaded deep learning model sequentially includes an expanded convolutional network, a multi-head cross attention unit, and a lightweight classification unit;
[0011] S4: Dynamically matching the three-dimensional diagnostic vector with the reference feature vector in the pre-built equipment historical operation database to generate a comparison result including a deviation index, and generating an adaptive alarm threshold and maintenance recommendation plan based on the comparison result.
[0012] Optionally, the S1 includes:
[0013] S11: Deploy edge computing nodes in the target device operation area. The edge computing nodes integrate a high-resolution visible light camera module, an infrared thermal imaging unit, and a three-axis vibration sensor array. Each sensor module is synchronized through a unified clock source to form a hardware-level linkage control architecture.
[0014] S12: The edge computing node receives a continuous image frame sequence from the visible light camera module in real time, extracts the visible light video stream based on a multi-threaded image caching mechanism, and adds a timestamp to generate a time-standardized visible light video stream;
[0015] S13: The infrared thermal imaging unit samples a thermal field image sequence at a fixed time interval, combines the temperature gradient enhancement algorithm of the edge node, generates time-scaled infrared thermal imaging data, and performs time tag alignment processing with the time-scaled visible light video stream;
[0016] S14: capturing the three-axis time-domain vibration acceleration signals of the target device along each axis using the three-axis vibration sensor array, converting the three-axis time-domain vibration acceleration signals into a two-dimensional frequency-time spectrum using a short-time Fourier transform, and constructing three-dimensional vibration spectrum data including the three-axis data, with a sampling time label attached;
[0017] S15: performing time tag alignment and modal coding processing on the time-standardized visible light video stream, the time-standardized infrared thermal imaging data, and the three-dimensional vibration spectrum data, and fusing them into a multi-modal sensing data set with a unified structure.
[0018] Optionally, the S2 includes:
[0019] S21: Content-driven keyframe extraction is performed on time-scaled visible light video streams in multimodal sensor datasets. Using the inter-frame structural similarity (SSIM) and entropy change rate as dual criteria, redundant frames are removed and frames with typical visual states are retained to obtain a keyframe sequence.
[0020] S22: Based on the timestamp alignment results, the cubic spline interpolation algorithm is used to enhance the time density of the time-scaled infrared thermal imaging data, thereby increasing its time resolution to be consistent with the three-dimensional vibration spectrum data, and generating interpolation-enhanced infrared thermal imaging data with time dimension alignment.
[0021] S23: Extract feature vectors from keyframe sequences, interpolated enhanced infrared thermal imaging data, and 3D vibration spectrum data. Use convolutional neural networks for intra-modal embedding encoding to obtain a uniform modal feature representation. Compute pairwise correlations between the three modal types based on the Gram matrix and construct a cross-modal feature correlation mapping matrix.
[0022] S24: Perform weighted fusion of the modal feature representation and the cross-modal feature correlation mapping matrix, and generate a spatiotemporal synchronous composite feature matrix including spatial dimension, temporal dimension and modal relationship dimension through multi-scale feature splicing and temporal expansion operations.
[0023] Optionally, the S21 includes:
[0024] S211: For each frame of the time-scaled visible light video stream, the structural similarity index and the image entropy change rate are used as a joint criterion to compare the current frame with the previous frame in turn. degree of similarity;
[0025] The calculation formula of the structural similarity index is: ;
[0026] in: is the structural similarity index, No. Frame image, 、 Frame and The average pixel value, 、 Frame and The standard deviation of for and The covariance between 、 is a stable constant to prevent the denominator from being 0;
[0027] The rate of change of image entropy is defined as: ;
[0028] in: For images The information entropy of represents the complexity of image grayscale distribution. For the The rate of change of information entropy of a frame relative to the previous frame;
[0029] S212: Set two discrimination thresholds, structural similarity index threshold and entropy change rate threshold , when satisfied When the current frame Mark as a keyframe and add it to the keyframe sequence.
[0030] Optionally, the S3 includes:
[0031] S31: The spatiotemporal synchronized composite feature matrix is input into the dilated convolutional network. By introducing the dilation rate parameter, the multi-scale feature tensor including multi-scale context information is extracted while maintaining the expansion of the convolution receptive field without increasing the number of parameters.
[0032] S32: The multi-scale feature tensor is input into the multi-head cross attention mechanism, multiple attention heads are generated through query, key, and value mapping, and the dependencies between modalities are calculated in parallel in different subspaces to obtain a weighted context representation tensor that integrates spatial, temporal, and modal information.
[0033] Optionally, the S3 further includes:
[0034] S33: The weighted context representation tensor is input into the lightweight classification unit. The dimension is compressed through fully connected mapping and activation function, and three types of results are output in sequence, including the equipment health status classification result, the abnormal area location index, and the fault probability value.
[0035] S34: Integrate the equipment health status classification results, abnormal area location index and fault probability value to generate a three-dimensional diagnosis vector.
[0036] Optionally, the S4 includes:
[0037] S41: Extract three-dimensional diagnostic vectors under stable operating conditions from the equipment operation history data. Perform cluster analysis using the K-means clustering algorithm. The cluster centers are used as the baseline feature vector sets representing typical operating conditions. Build an equipment operation history database.
[0038] S42: Perform dynamic similarity matching on the three-dimensional diagnostic vector obtained in S34 and each benchmark feature vector in the historical database to obtain a deviation index.
[0039] Optionally, the S4 further includes:
[0040] S43: combining the deviation index with the system's historical alarm curve to dynamically generate an adaptive alarm threshold and determine the current deviation;
[0041] S44: Output a structured maintenance suggestion plan based on the current deviation index, classification results and abnormal location.
[0042] The remote online monitoring system based on machine vision and artificial intelligence is used to implement the above-mentioned remote online monitoring method based on machine vision and artificial intelligence, and includes the following modules:
[0043] Edge acquisition module: used to collect data from target devices through edge computing nodes. The collected data includes visible light video streams, infrared thermal imaging data, and three-dimensional vibration spectrum data. It also performs time synchronization and structural packaging to generate a multimodal sensing data set.
[0044] Feature alignment module: This module performs cross-domain feature alignment on multimodal sensing datasets, including keyframe extraction for visible light video streams, temporal interpolation enhancement for infrared thermal imaging data, and the construction of a Gram matrix-based cross-modal feature correlation map. It then outputs a spatiotemporally synchronized composite feature matrix.
[0045] Diagnostic reasoning module: used to input the spatiotemporal synchronous composite feature matrix into a pre-built cascaded deep learning model, which sequentially includes a dilated convolutional network, a multi-head cross-attention unit, and a lightweight classification unit, and outputs a three-dimensional diagnostic vector including equipment health status classification, abnormal area location, and fault probability assessment;
[0046] Comparison and evaluation module: used to perform dynamic similarity matching between the three-dimensional diagnostic vector and the benchmark feature vector set in the historical operation database, and generate a comparison result including a deviation index based on a combination of dynamic time warping algorithm and cosine similarity calculation;
[0047] Strategy generation module: used to dynamically adjust the adaptive alarm threshold according to the comparison results and output maintenance recommendation plans, which include inspection recommendations, preventive maintenance recommendations, emergency shutdown recommendations and follow-up observation recommendations.
[0048] Beneficial effects of the present invention:
[0049] This invention deploys visible light video, infrared thermal imaging, and triaxial vibration sensors at edge computing nodes and introduces unified clock control, timestamp alignment, and interpolation enhancement strategies to construct a time-scaled multimodal sensing dataset. Furthermore, through keyframe extraction, cross-modal feature embedding, and Gram matrix correlation modeling, a spatiotemporal synchronized composite feature matrix is constructed. This solution effectively addresses issues such as inconsistent sampling frequencies across heterogeneous multi-source data and non-uniform feature expression spaces across modalities. It provides highly consistent and expressive input features for subsequent deep learning models, significantly improving the accuracy and stability of device state perception.
[0050] This paper constructs a cascaded deep learning model consisting of a dilated convolutional network, a multi-head cross-attention unit, and a lightweight classification unit, balancing the ability to capture multi-scale context with the requirements of lightweight model deployment. The dilated convolutional network expands the feature receptive field, the multi-head attention mechanism models long-range dependencies between modalities and time series, and the lightweight classifier implements multi-task decoding for three diagnostic tasks: health status classification, anomaly localization, and fault probability assessment. This model can run on-device and offers efficient inference, improving real-time diagnostic capabilities and fault prediction accuracy for complex operating conditions.
[0051] This invention implements an adaptive risk identification and early warning mechanism by dynamically matching diagnostic results with baseline feature vectors of typical operating conditions in a pre-built historical database. It also introduces a comparison algorithm that fuses DTW and cosine similarity, quantifies deviation indicators, and generates dynamic alarm thresholds based on historical distributions. Furthermore, it outputs structured maintenance recommendations based on diagnostic results, covering various scenarios such as inspections, preventive maintenance, and emergency shutdowns. This enhances the system's intelligent maintenance response capabilities, reduces false alarm and missed alarm rates, and improves operational reliability and safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;
[0054] Figure 2 Schematic diagram of the system flow of an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It is also noted that, to provide a more detailed description, the following embodiments are best and preferred embodiments, and those skilled in the art may employ alternative methods for implementing certain known technologies. Furthermore, the accompanying drawings are intended only to provide a more detailed description of the embodiments and are not intended to limit the present invention.
[0056] It should be noted that references in the specification to "one embodiment," "an embodiment," "exemplary embodiments," "some embodiments," etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment will include such specific features, structures, or characteristics. Furthermore, when specific features, structures, or characteristics are described in conjunction with an embodiment, it is within the knowledge of persons skilled in the relevant art to implement such features, structures, or characteristics in conjunction with other embodiments (whether or not explicitly described).
[0057] In general, terms can be understood, at least in part, from their use in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending at least in part on the context, allow for the presence of other factors that are not necessarily explicitly described.
[0058] like Figure 1 As shown, the remote online monitoring method based on machine vision and artificial intelligence includes the following steps:
[0059] S1: The edge computing node collects visible light video stream, infrared thermal imaging data, and three-dimensional vibration spectrum data of the target device to form a multimodal sensing data set;
[0060] S2: Perform cross-domain feature alignment on the multimodal sensor dataset to generate a spatiotemporally synchronized composite feature matrix.
[0061] The cross-domain feature alignment process includes: extracting key frames from the visible light video stream, using a spatiotemporal interpolation algorithm to increase the infrared thermal imaging sampling rate to synchronize with the vibration spectrum, and constructing a cross-modal feature correlation map based on the Gram matrix;
[0062] S3: Input the composite feature matrix into the cascade deep learning model, and output a three-dimensional diagnostic vector including equipment health status classification, abnormal area location, and failure probability assessment.
[0063] The cascaded deep learning model consists of a dilated convolutional network, a multi-head cross-attention unit, and a lightweight classification unit.
[0064] S4: Dynamically match the three-dimensional diagnostic vector with the benchmark feature vector in the pre-built equipment historical operation database to generate a comparison result including a deviation index. Based on the comparison result, an adaptive alarm threshold and maintenance recommendation plan are generated.
[0065] S1 includes:
[0066] S11: Deploy edge computing nodes in the target equipment operation area. The edge computing nodes integrate high-resolution visible light camera modules, infrared thermal imaging units, and three-axis vibration sensor arrays. Each sensor module is synchronized through a unified clock source to form a hardware-level linkage control architecture.
[0067] S12: The edge computing node receives a continuous sequence of image frames from the visible light camera module in real time. Based on the multi-threaded image caching mechanism, it extracts the visible light video stream and adds timestamps to generate a time-scaled visible light video stream. The specific steps are as follows:
[0068] S121: Initialize the image receiving channel, start the multi-threaded frame capture unit built into the edge node, and receive the image frame sequence captured by the visible light camera module at a fixed frame rate;
[0069] S122: Build an image frame cache queue and use a producer-consumer mechanism to write and read image frames in real time to avoid data frame loss and thread blocking.
[0070] S123: Bind a timestamp generated by the unified system clock to each frame of the image to form a time-continuous image sequence;
[0071] S124: Combining the frame sequence in timestamp order, encoding it into a structured video data format, and outputting a time-scaled visible light video stream;
[0072] S13: The infrared thermal imaging unit samples the thermal field image sequence at fixed time intervals, combines it with the temperature gradient enhancement algorithm of the edge node, generates time-scaled infrared thermal imaging data, and aligns the time tags with the time-scaled visible light video stream. The specific steps are as follows:
[0073] S131: Control the infrared thermal imaging unit to collect thermal field images and obtain a thermal radiation distribution map of the target device;
[0074] S132: Input the thermal field image into the temperature gradient enhancement algorithm of the edge node, and use local contrast stretching and spatial smoothing to enhance the thermal image details and improve the image feature clarity;
[0075] S133: adding a timestamp to the processed thermal image to generate structured time-scaled infrared thermal imaging data;
[0076] S134: performing time alignment processing based on the timestamp and the time-standardized visible light video stream, and improving the time accuracy of the heat map sequence to be consistent with the visible light frame sequence through interpolation;
[0077] S14: Capture the three-axis time-domain vibration acceleration signals of the target device along each axis using a three-axis vibration sensor array, convert the three-axis time-domain vibration acceleration signals into a two-dimensional frequency-time spectrum using a short-time Fourier transform, and construct three-dimensional vibration spectrum data including the three-axis data with sampling time labels.
[0078] S141: The three-axis vibration sensor collects acceleration signals of the device in the X, Y, and Z directions respectively to form an original three-axis time-domain vibration acceleration signal;
[0079] S142: Performing window function frame processing on the signal of each axis and performing short-time Fourier transform (STFT) to obtain the frequency distribution of each frame signal;
[0080] S143: Splice the STFT results of each axis into a two-dimensional time-frequency diagram, and further combine the X / Y / Z three-axis spectrum to construct a complete three-dimensional vibration spectrum data;
[0081] S144: Add a system time tag to the spectrogram of each time period to ensure consistency with the video stream and heat map data in the time domain;
[0082] S15: Perform time tag alignment and modal encoding on the time-scaled visible light video stream, time-scaled infrared thermal imaging data, and 3D vibration spectrum data to fuse them into a multimodal sensing dataset with a unified structure. This multimodal sensing dataset serves as the input for the subsequent "cross-domain feature alignment processing" in S2.
[0083] By integrating multimodal sensor components within edge computing nodes and implementing unified clock control, time tag alignment, and feature fusion strategies, the efficient construction of a multimodal sensing dataset with standardized structure and consistent time is effectively achieved, providing standardized input for subsequent spatiotemporal synchronization and model reasoning, and improving system response speed and perception accuracy.
[0084] S2 includes:
[0085] S21: Content-driven keyframe extraction is performed on time-scaled visible light video streams in multimodal sensor datasets. Using the inter-frame structural similarity (SSIM) and entropy change rate as dual criteria, redundant frames are removed and frames with typical visual states are retained to obtain a keyframe sequence.
[0086] S22: Based on the timestamp alignment results, a cubic spline interpolation algorithm is used to perform time density enhancement on the time-scaled infrared thermal imaging data, thereby increasing its time resolution to be consistent with the three-dimensional vibration spectrum data, and generating interpolation-enhanced infrared thermal imaging data aligned in the time dimension. The specific steps are as follows:
[0087] S221: Infrared thermal image data By timestamp Arrange in ascending order to form a time series of infrared images;
[0088] S222: Use the cubic spline interpolation method to improve the time density of the sequence. The interpolation function is expressed as follows:
[0089] ;
[0090] in, Indicates a time point The interpolated pixel value, For interval The interpolation coefficient of Interpolate time points for the target;
[0091] S223: Resample the new image sequences obtained by interpolation at all time points so that their timestamp density is consistent with the three-dimensional vibration spectrum data, and generate interpolation-enhanced infrared thermal imaging data.
[0092] S23: Extract feature vectors from key frame sequences, interpolated enhanced infrared thermal imaging data, and 3D vibration spectrum data. Use convolutional neural networks (CNNs) for intra-modal embedding encoding to obtain a uniform modal feature representation. Then, calculate the pairwise correlations between the three modal types based on the Gram matrix and construct a cross-modal feature correlation mapping matrix. The specific steps are as follows:
[0093] S231: Input the key frame sequence, interpolation enhanced infrared thermal imaging data and 3D vibration spectrum data into the pre-trained convolutional neural network, and extract the modal feature vectors respectively, which are recorded as:
[0094] is the visual feature vector of the visible light modality, is the thermal map feature vector of infrared mode, is the spectral eigenvector of the vibration mode;
[0095] S232: Compute the Gram matrix correlation between any two modes:
[0096] , ,in, Indicates time Upper modal With modal The feature correlation matrix of represents the matrix inner product, Represents a transpose operation;
[0097] S233: Combining all modes in pairs Construct a set of cross-modal feature correlation mapping matrices as the input basis for subsequent modal fusion;
[0098] S24: Perform weighted fusion of the modal feature representation and the cross-modal feature correlation mapping matrix, and generate a spatiotemporal synchronous composite feature matrix including spatial dimension, temporal dimension, and modal relationship dimension through multi-scale feature splicing and temporal expansion operations. The specific steps are as follows:
[0099] S241: Every moment The three modal eigenvectors of are concatenated into a unified eigenvector :
[0100] ;
[0101] S242: Based on Gram Matrix and fusion weight coefficient , define the fusion feature vector as follows:
[0102] ,in, is the inter-modal weighted fusion coefficient, satisfying , Represents the fused feature vector;
[0103] S243: All time steps The fusion features on are concatenated into the final spatiotemporal composite feature matrix, which is expressed as:
[0104] , As input to subsequent deep learning models;
[0105] By introducing mechanisms such as key frame extraction, interpolation enhancement, and Gram matrix mapping, the differences in time granularity and semantic space of multimodal data sources are resolved, and a composite feature expression with unified structure and consistent time sequence is achieved, providing highly expressive input feature support for subsequent model reasoning, and improving the system's recognition accuracy and generalization ability for complex working conditions.
[0106] S21 includes:
[0107] S211: For each frame of the time-scaled visible light video stream, the structural similarity index and the image entropy change rate are used as a joint criterion to compare the current frame with the previous frame in turn. degree of similarity;
[0108] The calculation formula of the structural similarity index is: ;
[0109] in: is the structural similarity index, No. Frame image, 、 Frame and The average pixel value, 、 Frame and The standard deviation of for and The covariance between 、 is a stable constant to prevent the denominator from being 0;
[0110] The rate of change of image entropy is defined as: ;
[0111] in: For images The information entropy of represents the complexity of image grayscale distribution. For the The rate of change of information entropy of a frame relative to the previous frame;
[0112] S212: Set two discrimination thresholds, structural similarity index threshold and entropy change rate threshold , when satisfied When the current frame Mark as a keyframe and add it to the keyframe sequence.
[0113] S3 includes:
[0114] S31: Input the spatiotemporal synchronous composite feature matrix into the dilated convolutional network. By introducing the dilation rate parameter, the multi-scale feature tensor including multi-scale context information is extracted while maintaining the expansion of the convolution receptive field without increasing the number of parameters. The specific steps are as follows:
[0115] S311: Input spatiotemporal synchronization composite feature matrix It is regarded as a one-dimensional time series feature map and input into a set of convolutional layers with different dilation rates for multi-scale feature extraction.
[0116] S312: Set multiple groups of one-dimensional dilated convolutional layers with kernel size of , void ratio parameter They are , each set of convolution operations is defined as follows:
[0117] ,in, For the The convolution output at the dilation rate is For the The convolution kernel of the weights, The input time step Composite eigenvectors on ;
[0118] S313: Concatenate the outputs of all dilated convolution channels in the feature dimension direction, expressed as:
[0119] ;
[0120] Finally, we get a multi-scale feature tensor in the time dimension , this tensor serves as the input of the multi-head cross attention unit in S32;
[0121] S32: Input the multi-scale feature tensor into the multi-head cross attention mechanism, generate multiple attention heads through query, key, and value mapping, and parallelly calculate the inter-modal dependencies in different subspaces to obtain a weighted context representation tensor that integrates spatial, temporal, and modal information. The specific steps are as follows:
[0122] S321: Multi-scale feature tensor Enter the multi-head attention mechanism unit. First, the query, key, and value matrices are generated through three sets of linear transformation mappings:
[0123] ,in: Represents query, key and value respectively, is the trainable weight matrix, is the dimension of each attention head;
[0124] S322: Input the attention calculation function and get the attention output, which is expressed as:
[0125] ,in, represents the similarity between each time step, Ensure attention weight normalization;
[0126] S323: In multi-head attention, there is subspaces, projecting the original input into subspaces and calculate the above attention operations in parallel, concatenate the outputs of each subspace and then perform a linear transformation, which can be expressed as:
[0127] ;
[0128] ;
[0129] in, Map weights to the final output;
[0130] Output is the weighted context representation tensor , tensor The joint dependency features of time, modality and context are retained and used as input for subsequent lightweight classification units.
[0131] S3 also includes:
[0132] S33: The weighted context representation tensor is input into the lightweight classification unit. The dimension is compressed through full connection mapping and activation function, and three types of results are output in sequence, including the equipment health status classification result, the abnormal area location index, and the fault probability value. The specific steps are as follows:
[0133] S331: Representing weighted context as a tensor The time dimension is compressed by the average pooling (or maximum pooling) operation to obtain the global context feature vector, which is expressed as:
[0134] ;
[0135] S332: Input the global context feature vector into three independent fully connected sub-units, and output three types of task results respectively:
[0136] (1) Health status classification results (using Softmax) are expressed as:
[0137] ;
[0138] Output three status categories (normal / sub-healthy / fault);
[0139] (2) Abnormal area positioning index (based on the output of the maximum attention score point), expressed as:
[0140] ;
[0141] (3) Failure probability evaluation value (using Sigmoid), expressed as:
[0142] ,in , output probability value ;
[0143] S34: Integrate the equipment health status classification results, abnormal area location index and fault probability value to generate a three-dimensional diagnosis vector, which is expressed as: ;
[0144] By constructing a cascaded deep learning model consisting of an expanded convolutional network, a multi-head cross-attention unit, and a lightweight classification unit, while maintaining inference efficiency, it achieves deep perception, semantic understanding, and structured decoding of high-dimensional spatiotemporal composite features, effectively improving the accuracy of equipment operation status identification, the precision of anomaly location, and the credibility of fault probability estimation.
[0145] S4 includes:
[0146] S41: Extract the three-dimensional diagnostic vector under the stable operating state from the equipment operation history data, and use the K-means clustering algorithm to perform cluster analysis. The cluster center is used as the benchmark feature vector set representing the typical working condition, which is expressed as: , and build a historical operation database of the equipment for subsequent comparison process;
[0147] S42: The three-dimensional diagnostic vector obtained in S34 Each benchmark feature vector in the historical database Perform dynamic similarity matching based on the dynamic time warping algorithm combined with cosine similarity calculation to achieve misalignment alignment and similarity measurement under periodic fluctuations, and obtain the deviation index as follows:
[0148] Cosine similarity calculation: ;
[0149] Dynamic time warping calculates the minimum path distance:
[0150] ;
[0151] Comprehensive similarity score (weighted after normalization):
[0152] ;
[0153] in, The current diagnosis vector and the The combined similarity of the reference vectors, is the preset maximum DTW distance for normalization. is the fusion weight coefficient, Matching path for DTW;
[0154] After calculating all the comparison results, select the one with the lowest similarity to quantify the degree of deviation and define the deviation index for: , deviation index The larger the value, the more significant the difference between the current diagnostic status and the historical benchmark.
[0155] The S4 also includes:
[0156] S43: Combine the deviation index with the system's historical alarm curve to dynamically generate adaptive alarm thresholds , and judge the current deviation:
[0157] like : The status is in the safe range;
[0158] like : Trigger an abnormal alarm and record the trigger time, threshold and deviation vector;
[0159] The alarm threshold is adaptively adjusted based on the historical statistical distribution, expressed as: ,in, is the historical deviation mean, is the historical deviation standard deviation, is the adjustment factor (e.g., 1.96 corresponds to a 95% confidence level);
[0160] S44: Output a structured maintenance suggestion plan based on the current deviation index, classification results, and abnormal location. The suggestion content includes but is not limited to:
[0161] Inspection suggestion: If If the value is low but exceeds the threshold, manual inspection and verification is recommended;
[0162] Preventive maintenance recommendations: If If the condition is high and sub-healthy, it is recommended to replace parts and adjust the working conditions within the planned window;
[0163] Emergency shutdown maintenance suggestions: If If the value is extremely high and the status is faulty, it is recommended to shut down the system immediately and notify the operation and maintenance team.
[0164] Follow-up observation suggestions: When the threshold is approaching and the change trend is significant, it is recommended to continuously monitor in the short term and lower the alarm threshold;
[0165] The final output includes: , , suggestion type, processing time limit, suggestion executor As structured maintenance instructions for the platform or user end.
[0166] like Figure 2 As shown, the remote online monitoring system based on machine vision and artificial intelligence is used to implement the above-mentioned remote online monitoring method based on machine vision and artificial intelligence, and includes the following modules:
[0167] Edge acquisition module: used to collect data from target devices through edge computing nodes. The collected data includes visible light video streams, infrared thermal imaging data, and three-dimensional vibration spectrum data. It also performs time synchronization and structural packaging to generate a multimodal sensing data set.
[0168] Feature alignment module: This module performs cross-domain feature alignment on multimodal sensing datasets, including keyframe extraction for visible light video streams, temporal interpolation enhancement for infrared thermal imaging data, and the construction of a Gram matrix-based cross-modal feature correlation map. It then outputs a spatiotemporally synchronized composite feature matrix.
[0169] Diagnostic reasoning module: This module inputs the spatiotemporal synchronized composite feature matrix into a pre-built cascaded deep learning model. The cascaded deep learning model sequentially comprises a dilated convolutional network, a multi-head cross-attention unit, and a lightweight classification unit. The output is a three-dimensional diagnostic vector that classifies the equipment health status, locates the abnormal area, and assesses the probability of failure.
[0170] Comparison and evaluation module: This module is used to dynamically match the three-dimensional diagnostic vector with the benchmark feature vector set in the historical operation database. Based on the combined calculation of the dynamic time warping algorithm and cosine similarity, it generates a comparison result including a deviation index.
[0171] Strategy generation module: used to dynamically adjust the adaptive alarm threshold according to the comparison results and output maintenance recommendation plans, which include inspection suggestions, preventive maintenance suggestions, emergency shutdown suggestions and follow-up observation suggestions.
[0172] The present invention encompasses any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention. To provide a thorough understanding of the present invention, specific details are described in detail below in connection with the preferred embodiments of the present invention, but those skilled in the art will be able to fully understand the present invention without these detailed descriptions. Furthermore, to avoid unnecessary confusion regarding the essence of the present invention, well-known methods, processes, procedures, components, and circuits have not been described in detail.
[0173] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A remote online monitoring method based on machine vision and artificial intelligence, characterized in that: The following steps are involved: S1: The edge computing node collects visible light video stream, infrared thermal imaging data, and three-dimensional vibration spectrum data of the target device to form a multimodal sensing data set; S2: Perform cross-domain feature alignment processing on the multimodal sensor data set to generate a spatiotemporally synchronized composite feature matrix. The cross-domain feature alignment process includes: extracting key frames from the visible light video stream, using a spatiotemporal interpolation algorithm to increase the infrared thermal imaging sampling rate to synchronize with the vibration spectrum, and constructing a cross-modal feature correlation map based on the Gram matrix; S3: Input the composite feature matrix into the cascade deep learning model, and output a three-dimensional diagnostic vector including equipment health status classification, abnormal area location, and failure probability assessment. The cascaded deep learning model sequentially includes an expanded convolutional network, a multi-head cross attention unit, and a lightweight classification unit; S4: Dynamically matching the three-dimensional diagnostic vector with a reference feature vector in a pre-built equipment historical operation database to generate a comparison result including a deviation index, and generating an adaptive alarm threshold and a maintenance recommendation plan based on the comparison result; The S2 specifically includes: S21: Content-driven keyframe extraction is performed on the time-scaled visible light video stream in the multimodal sensor dataset. The dual criteria of inter-frame structural similarity and entropy change rate are used to remove redundant frames and retain frames with typical visual states to obtain a keyframe sequence. S22: Based on the timestamp alignment results, a cubic spline interpolation algorithm is used to enhance the time density of the time-scaled infrared thermal imaging data, thereby increasing its time resolution to be consistent with the three-dimensional vibration spectrum data, and generating interpolation-enhanced infrared thermal imaging data with time dimension alignment. S23: Extract feature vectors from keyframe sequences, interpolated enhanced infrared thermal imaging data, and 3D vibration spectrum data. Use convolutional neural networks for intra-modal embedding encoding to obtain a uniform modal feature representation. Compute pairwise correlations between the three modal types based on the Gram matrix and construct a cross-modal feature correlation mapping matrix. S24: Perform weighted fusion of the modal feature representation and the cross-modal feature correlation mapping matrix, and generate a spatiotemporal synchronous composite feature matrix including spatial dimension, temporal dimension and modal relationship dimension through multi-scale feature splicing and temporal expansion operations.
2. The remote online monitoring method based on machine vision and artificial intelligence according to claim 1, characterized in that: Said S1 comprises: S11: deploying an edge computing node in the target device operation area, wherein the edge computing node integrates a visible light camera module, an infrared thermal imaging unit, and a three-axis vibration sensor array; S12: The edge computing node receives a continuous image frame sequence from the visible light camera module in real time, extracts the visible light video stream based on a multi-threaded image caching mechanism, and adds a timestamp to generate a time-standardized visible light video stream; S13: The infrared thermal imaging unit samples a thermal field image sequence at a fixed time interval, combines the temperature gradient enhancement algorithm of the edge node, generates time-scaled infrared thermal imaging data, and performs time tag alignment processing with the time-scaled visible light video stream; S14: capturing the three-axis time-domain vibration acceleration signals of the target device along each axis using the three-axis vibration sensor array, converting the three-axis time-domain vibration acceleration signals into a two-dimensional frequency-time spectrum using a short-time Fourier transform, and constructing three-dimensional vibration spectrum data including the three-axis data, with a sampling time label attached; S15: performing time tag alignment and modal coding processing on the time-standardized visible light video stream, the time-standardized infrared thermal imaging data, and the three-dimensional vibration spectrum data, and fusing them into a multi-modal sensing data set with a unified structure.
3. The remote online monitoring method based on machine vision and artificial intelligence according to claim 1, characterized in that: The S21 includes: S211: For each frame of the time-scaled visible light video stream, the structural similarity index and the image entropy change rate are used as a joint criterion to compare the current frame with the previous frame in turn. degree of similarity; The calculation formula of the structural similarity index is: ; in, is the structural similarity index, No. Frame image, 、 Frame and The average pixel value, 、 Frame and The standard deviation of for and The covariance between 、 is a stable constant to prevent the denominator from being 0; The rate of change of image entropy is defined as: ; in, For images The information entropy of represents the complexity of image grayscale distribution. For the The rate of change of information entropy of a frame relative to the previous frame; S212: Set two discrimination thresholds, structural similarity index threshold and entropy change rate threshold , when satisfied When the current frame Mark as a keyframe and add it to the keyframe sequence.
4. The remote online monitoring method based on machine vision and artificial intelligence according to claim 1, characterized in that: The S3 includes: S31: The spatiotemporal synchronized composite feature matrix is input into the dilated convolutional network. By introducing the dilation rate parameter, the multi-scale feature tensor including multi-scale context information is extracted while maintaining the expansion of the convolution receptive field without increasing the number of parameters. S32: The multi-scale feature tensor is input into the multi-head cross attention mechanism, multiple attention heads are generated through query, key, and value mapping, and the dependencies between modalities are calculated in parallel in different subspaces to obtain a weighted context representation tensor that integrates spatial, temporal, and modal information.
5. The remote online monitoring method based on machine vision and artificial intelligence according to claim 4 is characterized in that: Said S3 further comprises: S33: The weighted context representation tensor is input into the lightweight classification unit. The dimension is compressed through fully connected mapping and activation function, and three types of results are output in sequence, including the equipment health status classification result, the abnormal area location index, and the fault probability value. S34: Integrate the equipment health status classification results, abnormal area location index and fault probability value to generate a three-dimensional diagnosis vector.
6. The remote online monitoring method based on machine vision and artificial intelligence according to claim 5, characterized in that: The S4 includes: S41: Extract three-dimensional diagnostic vectors under stable operating conditions from the equipment operation history data. Perform cluster analysis using the K-means clustering algorithm. The cluster centers are used as the baseline feature vector sets representing typical operating conditions. Build an equipment operation history database. S42: Perform dynamic similarity matching on the three-dimensional diagnostic vector obtained in S34 and each benchmark feature vector in the historical database to obtain a deviation index.
7. The remote online monitoring method based on machine vision and artificial intelligence according to claim 6, characterized in that: Said S4 further comprises: S43: combining the deviation index with the system's historical alarm curve to dynamically generate an adaptive alarm threshold and determine the current deviation; S44: Output a structured maintenance suggestion plan based on the current deviation index, classification results and abnormal location.
8. A remote online monitoring system based on machine vision and artificial intelligence, used to implement the remote online monitoring method based on machine vision and artificial intelligence as claimed in any one of claims 1 to 7, characterized in that: Includes the following modules: Edge acquisition module: used to collect data from target devices through edge computing nodes. The collected data includes visible light video streams, infrared thermal imaging data, and three-dimensional vibration spectrum data. It also performs time synchronization and structural packaging to generate a multimodal sensing data set. Feature alignment module: This module performs cross-domain feature alignment on multimodal sensing datasets, including keyframe extraction for visible light video streams, temporal interpolation enhancement for infrared thermal imaging data, and the construction of a Gram matrix-based cross-modal feature correlation map. It then outputs a spatiotemporally synchronized composite feature matrix. Diagnostic reasoning module: used to input the spatiotemporal synchronous composite feature matrix into a pre-built cascaded deep learning model, which sequentially includes a dilated convolutional network, a multi-head cross-attention unit, and a lightweight classification unit, and outputs a three-dimensional diagnostic vector including equipment health status classification, abnormal area location, and fault probability assessment; Comparison and evaluation module: used to perform dynamic similarity matching between the three-dimensional diagnostic vector and the benchmark feature vector set in the historical operation database, and generate a comparison result including a deviation index based on a combination of dynamic time warping algorithm and cosine similarity calculation; Strategy generation module: used to dynamically adjust the adaptive alarm threshold according to the comparison results and output maintenance recommendation plans, which include inspection recommendations, preventive maintenance recommendations, emergency shutdown recommendations and follow-up observation recommendations.
Citation Information
Patent Citations
Power grid health assessment and analysis method based on multiple modes
CN118657404A
Early warning analysis method based on intelligent vision and server
CN119810757A