Remote dynamic monitoring method and system for industrial automation equipment production line

By deploying edge computing nodes and multi-protocol adaptive driver engines on the industrial automation equipment production line, combining digital twin models and deep learning technology, the real-time, adaptability and security problems in existing remote monitoring technologies are solved, and an efficient and secure remote dynamic monitoring system is realized.

CN120196072APending Publication Date: 2025-06-24SHANDONG YAJIE INTELLIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510406610.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The remote monitoring technology of existing industrial automation equipment production lines has problems such as insufficient real-time control, lack of dynamic adaptability, difficulty in compatibility with heterogeneous protocols and weak network security protection, which is difficult to meet the needs of Industry 4.0 and intelligent manufacturing.

Method used

Edge computing nodes are used to collect production line operation status data in real time, the built-in multi-protocol adaptive driver engine supports a variety of industrial protocols, and dynamic physical simulation and predictive control are performed through digital twin models and LSTM neural networks. At the same time, image or video information is processed using an improved convolutional neural network model and weighted averaging algorithm, encrypted control instruction streams are generated, and low-latency transmission is ensured through the blockchain consensus mechanism and the hybrid networking architecture of time-sensitive network TSN and 5G URLLC.

Benefits of technology

It realizes low-latency response, dynamic strategy optimization, multi-protocol adaptation and active security protection, improves the remote dynamic monitoring capabilities of the production line, and meets the evolution needs of industrial automation to distributed intelligent control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196072A_ABST
    Figure CN120196072A_ABST
Patent Text Reader

Abstract

The invention discloses a remote dynamic monitoring method and system for an industrial automation equipment production line, relates to the technical field of monitoring, and solves the problem that the monitoring capability is lagged in the prior art. According to the technical scheme, the method comprises the following steps that an image or video of the operation state of the production line is obtained, so that basic information of the operation state of the production line is obtained; setting information frames of a plurality of preset time points for the image or video information in a preset time period; the acquired images or videos are stored so as to permanently store the operation state information of the production line, and subsequent calling, checking, evidence obtaining and tracking are facilitated; and extracting global time sequence characteristics of the plurality of image or video information time sequence input vectors, and processing the acquired images or videos based on a weighted average algorithm to obtain a comprehensive comparison value so as to further screen out the abnormal operation state information of the production line. The remote dynamic monitoring capability of the industrial automation equipment production line is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial automation equipment, and more specifically to a remote dynamic monitoring method and system for an industrial automation equipment production line. Background Art

[0002] With the in-depth development of Industry 4.0 and intelligent manufacturing, industrial automation production lines have gradually realized equipment networking and data collection functions. Traditional monitoring systems mostly adopt local deployment mode, and realize equipment status monitoring and process parameter adjustment through control systems such as PLC and SCADA. Its control logic is solidified at the equipment end, which is difficult to adapt to flexible production needs.

[0003] The existing remote monitoring technology mainly has the following technical bottlenecks: Insufficient real-time control: Most systems use a polling data transmission mechanism (such as Modbus TCP). Network delays cause a time difference of seconds between control instructions and device status feedback, which cannot meet the millisecond response requirements of scenarios such as precision machining and high-speed assembly lines.

[0004] Lack of dynamic adaptability: Although the remote monitoring solutions proposed in patents such as CN201910235678.X can realize the visualization of equipment status, the control strategy still relies on preset parameter thresholds and lacks dynamic optimization capabilities based on real-time operating data, resulting in an increase in equipment downtime rate under abnormal conditions.

[0005] Difficulty in compatibility of heterogeneous protocols: There are more than 20 industrial protocols in industrial sites, including PROFINET, EtherCAT, OPC UA, etc. Public literature (such as the 2021 study of "IEEE Transactions on Industrial Informatics") points out that existing systems mostly use protocol conversion gateways to achieve data collection, but protocol parsing errors are prone to occur when control instructions are issued, resulting in abnormal equipment execution.

[0006] Weak network security protection: According to the NIST IR 8228 report, existing remote control systems mostly use the static protection mode of VPN+firewall, which is difficult to resist man-in-the-middle attacks (MITM) against the control instruction flow, and there is a risk of production equipment being maliciously manipulated.

[0007] It is worth noting that although international patents such as EP3564883B1 propose a 5G-based remote control architecture, it adopts a centralized data processing mode. When the number of edge nodes exceeds 500, the system response delay increases exponentially (test data see the "5G-ACIA White Paper" 2022 edition), which seriously restricts large-scale production line deployment.

[0008] Therefore, there is an urgent need to build a remote dynamic monitoring system with low-latency response, dynamic policy optimization, multi-protocol adaptability, and active security protection capabilities to break through the technical barriers in the evolution of industrial automation towards distributed intelligent control. Summary of the Invention

[0009] Aiming at the deficiencies of the above technologies, the present invention discloses a method and system for remote dynamic monitoring of an industrial automation equipment production line, which can realize remote dynamic monitoring of the industrial automation equipment production line, enabling the remote control end to dynamically obtain the operating status of the industrial automation equipment production line under unattended conditions.

[0010] The present invention adopts the following technical solutions: A method for remote dynamic monitoring of an industrial automation equipment production line, which includes the following steps: Real-time collect the operating status data and process parameters of the industrial automation equipment production line through edge computing nodes deployed at the industrial automation equipment production line end. The edge computing nodes are built-in with a multi-protocol adaptive drive engine, supporting parallel parsing and data encapsulation of at least three industrial protocols, namely PROFINET, EtherCAT, and OPC UA; dynamically obtain images or videos of the production line operating status to obtain basic information on the production line operating status; and set information frames at multiple predetermined time points within a predetermined time period for the image or video information. Perform dynamic physical simulation on the operating status data of the industrial automation equipment production line based on the digital twin model to generate a dynamic optimization parameter set for the control method of the industrial automation equipment production line. The digital twin model predicts the operating conditions of the industrial automation equipment production line at the future t+Δt moment through an LSTM neural network, where Δt is a preset control period and Δt≤50ms; store the obtained images or videos to permanently store the production line operating status information for subsequent retrieval, viewing, evidence collection, and tracking. Compare the dynamic optimization parameter set in the operating status of the industrial automation equipment production line with the preset process standard parameters in real time. When it is detected that the parameter deviation value exceeds the dynamic tolerance threshold, trigger the dynamic reconstruction of the control instruction, generate an encrypted control instruction stream containing a timestamp and the texture of the industrial automation equipment production line, extract the global temporal features of the multiple image or video information temporal input vectors, and process the obtained images or videos based on the weighted average algorithm to obtain a comprehensive comparison value to further screen out the abnormal operating status information of the production line, enabling the operating status information of the production line to be known under unattended conditions. The information frames that set multiple predetermined time points for the image or video information within a predetermined time period include: arranging multiple information frames in the time dimension within the predetermined time period through an improved convolutional neural network model to obtain the time series input vectors of the multiple image or video information; and using a feature analysis algorithm to obtain the time series feature vectors of the information frames at multiple predetermined time points. The extraction of the global time series features of the multiple time series input vectors of the image or video information includes: annotating the image or video information data through an information frame capture module with a time stamp. The processing of the acquired image or video based on the weighted average algorithm includes image enhancement, image information encoding, image fusion, and image frame analysis.

[0011] The present invention also adopts the following technical solutions: The improved convolutional neural network model includes an input layer, a multi-dimensional convolutional layer, an activation layer, a pooling layer, a classification layer, a diagnostic layer, and an output layer. The data information of the input layer is the image or video of the production line operation status. The method for dynamically acquiring the image or video of the production line operation status includes: Collecting image data through a multi-spectral vision sensor array deployed at key workstations on the production line. The sensor array integrates a visible light camera and an infrared thermal imaging module, and the sampling frequency is not less than 60Hz. Performing spatio-temporal alignment processing on the collected original images, and using an improved 3D convolutional kernel to extract the dynamic features of continuous N frames of images in the time domain, where N≥5 and satisfies N×frame interval≤Δt / 2 with Δt. The output layer performs multi-node verification on the encrypted control instruction stream through a blockchain-based consensus mechanism. After passing the verification, it is sent to the target device end through a low-latency transmission channel. The low-latency transmission channel adopts a hybrid networking architecture of Time-Sensitive Networking (TSN) and 5G Ultra-Reliable Low-Latency Communication (URLLC) to ensure that the end-to-end transmission latency ≤ 10ms.

[0012] The present invention also adopts the following technical solutions: The improved convolutional neural network model is a spatio-temporal two-stream architecture, including: Spatial stream network: Using an atrous convolutional layer to extract the device state features of a single frame image, and the atrous rate is positively correlated with the device movement speed; Temporal stream network: Capturing the abnormal state evolution features between consecutive frames through a three-dimensional convolutional layer. Feature fusion module: Using a gated attention mechanism to dynamically allocate spatio-temporal feature weights, and the abnormal detection sensitivity is increased by ≥21%. The working method of the improved convolutional neural network model is: when inputting data x through the input layer, the output of the model through the output layer is Then the output function of the output layer is: In formula (1), Indicates variable parameters, and the output image timing frame y takes k values. Indicates the timing sorting of frames, and the image information in the output model is denoted as For the input feature map Perform convolution calculation through a multi-dimensional convolution layer, perform one-dimensional convolution calculation through time-domain convolution, then perform two-dimensional calculation on the input feature map and the convolution kernel in the depth direction, improve the calculation ability through an activation function, and then perform global max pooling and global average pooling to obtain two feature vectors, calculate the global max pooling and the global average pooling In formula (2), represents the C-dimensional feature vector of the input feature map F at the spatial position ; In formula (3), represents the C-dimensional feature vector of the input feature map F at the spatial position ; Send these two feature vectors into a shared multi-layer perceptron (MLP) respectively to learn channel-level features, and calculate the number of neurons in the first layer of the MLP In formula (4), represents the activation function of the MLP, is a learnable parameter; In formulas (4) and (5), are learnable parameters; Add the two feature vectors output by the MLP to obtain a 1×1×C feature vector z, and map the z feature vector through the Sigmoid activation function to obtain the final channel attention weight matrix In formula (6), represents the MLP output of the global max pooling, represents the MLP output of the global average pooling, then In formula (7), represents the Sigmoid activation function, and its numerical range is (0, 1), which is used to represent the channel attention weight. The present invention also adopts the following technical solutions: The method for image processing is: Perform image enhancement by improving the Retinex algorithm, perform multi-scale Gaussian filtering on the brightness component V in the HSV color space, and the scale parameter σ = [5, 15, 30]; Compensate the details of the dark area through the adaptive gain adjustment function where λ is dynamically adjusted according to the ambient light intensity; Perform non-local means denoising on the enhanced image, and the search window radius is associated with the device vibration amplitude; Store the production line operation status images for record; extract the production line operation status information model for comparative analysis using the texture feature set corresponding to the production line operation status image database; Compare the features of the stored production line operation status images with those of the retrieved production line operation status images; Calculate the relationship between the features of the production line operation status images and those of the retrieved production line operation status images in the database. Assume that the set similarity threshold for the production line operation status is Y, and the comprehensive comparison value for the image or video is Z. When Z is greater than Y, it is considered that the production line operation status is normal; when Z is less than Y, it is considered that the production line operation status is abnormal; Output the calculation results for the reference of the monitoring center staff. The present invention also adopts the following technical solutions: The processing process of the weighted average algorithm includes: a) Extract the HSV color space histogram features and LBP texture features for each information frame to construct a multi-dimensional feature vector; b) Calculate the Mahalanobis distance of the feature vector based on a sliding time window and dynamically assign the weight coefficients for each frame where Δd is the deviation degree of the current frame from the historical benchmark, and k is the sensitivity adjustment factor; c) When the comprehensive comparison value exceeds the threshold trigger a three-level abnormal alarm mechanism; The calculation method of the weighted average Z in the weighted average algorithm is: In formula (8), Z ij is the comparison value between the i-th captured photo and the j-th production line operation status library; K Ii is the quality coefficient of the production line operation status photo of the i-th captured photo; K Qj is the quality coefficient of the photo in the j-th production line operation status library; The total number of captured photos is m, 1 ≤ i ≤ m; The total number of photos in the production line operation status library is n, 1 ≤ j ≤ m; N is the total number of comparison values between the captured photos and the production line operation status photo library, N = m × n. The present invention also adopts the following technical solutions: The production line operation status similarity threshold is set by the monitoring center staff, and the production line operation status similarity threshold is 8 - 15.

[0013] The present invention also adopts the following technical solutions: Use a weighted softmax loss function to handle the remaining imbalance in the resampled segment in the weighted average algorithm. For the training set { x ’’, y ’’} p , the weights of each performance parameter type in the loss function β are calculated as: In formula (9), P is the total number of weight parameter categories; The weighted softmax loss function loss is calculated as: In formula (10),sum ’ is the sum of samples of performance parameters in the weighted average algorithm. The present invention also adopts the following technical solutions: When enhancing the image, the contrast of the image is enhanced by adjusting the brightness value of the image, and the histogram is made more uniform by redistributing the pixel values in the image; the clarity of the image is improved by enhancing the edges and details in the image, and the local contrast of the image is enhanced by the Retinex method; the image is decomposed into components of different scales and different frequencies by using wavelet transform to facilitate multi-scale analysis and denoising; When encoding the image information, the image data is transformed from the spatial domain to the frequency domain by using orthogonal transform, and then quantized, and encoded according to the frequency of occurrence of pixel values by Huffman coding or arithmetic coding to reduce redundant information; When fusing the images, encoding is performed according to the frequency of occurrence of pixel values, redundant information is reduced by the Bayesian algorithm, and the images are fused at multiple scales to retain information at different scales; When analyzing the image frames, the images are fused at multiple scales to retain information at different scales; deep learning models such as CNN are used to identify the information in the images, track the objects in the video, and associate the targets in consecutive frames. The present invention also adopts the following technical solutions: A system for remotely and dynamically monitoring the home environment, comprising: A monitoring device, configured to obtain images or videos of the operation status of the production line to obtain basic information on the operation status of the production line; A central control unit, configured to process and store the images or videos obtained by the monitoring device; A storage unit, configured to store the obtained images or videos to permanently store the operation status information of the production line, facilitating subsequent retrieval, viewing, evidence collection and tracking; An image processing unit, configured to process the obtained images or videos based on a weighted average algorithm to obtain a comprehensive comparison value, so as to further screen out the abnormal operation status information of the production line, enabling the operation status information of the production line to be known when the production line is unattended; A cloud server, configured to permanently store the processed image or video information to securely store the information and not easily lose it; A handheld terminal, configured to enable the staff in the monitoring center to immediately obtain the processed image or video information to make manual intervention actions according to the dynamic operation status of the production line; wherein: The monitoring device is communicatively connected to the central control unit through a communication unit, the central control unit is integrated with a storage unit and an image processing unit, the output end of the central control unit is connected to the input end of the cloud server, and the output end of the cloud server is connected to the input end of the handheld terminal; The central control unit further includes a production line operation status library for storing the production line operation status images monitored for record. The central control unit communicates with the monitoring device through a communication unit, and the communication unit is a wired communication unit or a wireless communication unit. The monitoring device is arranged at the door, indoors or at the window of the monitoring center staff, and the monitoring device is a wide-angle camera or a 360 camera with a wireless communication interface.

[0014] The present invention also adopts the following technical solutions. The image processing unit further includes: An extraction module for extracting the production line operation status information model for comparative analysis using the texture feature set corresponding to the production line operation status image database. A comparison module for comparing the stored production line operation status image features with the retrieved production line operation status image features. A calculation module for calculating the relationship between the production line operation status image features and the production line operation status image features retrieved from the database. An output module for outputting the calculation result for reference by the monitoring center staff. The extraction module extracts the deep learning features of the production line operation status based on an improved convolutional neural network algorithm model.

[0015] Positive and beneficial effects: The present invention obtains images or videos of the production line operation status, sets information frames at multiple predetermined time points for the image or video information within a predetermined time period, improving the analysis ability. By storing the acquired images or videos, the production line operation status information is permanently stored. Based on the weighted average algorithm, the acquired images or videos are processed to obtain a comprehensive comparison value to further screen out the abnormal operation status information of the production line, enabling the production line operation status information to be known when the production line is unattended. Through the image processing unit, the images can be quickly processed, and the pictures and video information captured by the monitoring device can be recognized relatively quickly. The present invention realizes the wireless upload of monitoring data by adopting remote communication technology. By calculating the relationship between the production line operation status image features and the production line operation status image features retrieved from the database, the efficiency of identifying the production line operation status is rapidly improved. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where: Figure 1 It is a schematic diagram of the process structure of the present invention; Figure 2 It is a schematic diagram of the process of the image processing unit of the present invention; Figure 3 It is a schematic diagram of the system architecture of the present invention. Specific embodiments

[0017] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0018] As Figures 1-3 shown, a method for remote dynamic monitoring of an industrial automation equipment production line is characterized by including the following steps: Through the edge computing nodes deployed at the industrial automation equipment production line end, the operation status data and process parameters of the industrial automation equipment production line are collected in real time. The edge computing nodes are built-in with a multi-protocol adaptive drive engine, supporting parallel parsing and data encapsulation of at least three industrial protocols, namely PROFINET, EtherCAT, and OPC UA; to dynamically obtain images or videos of the production line operation status to obtain basic information on the production line operation status; and set information frames at multiple predetermined time points within a predetermined time period for the image or video information. Based on the digital twin model, dynamic physical simulation is performed on the operation status data of the industrial automation equipment production line to generate a dynamic optimization parameter set for the control method of the industrial automation equipment production line. The digital twin model predicts the working condition status of the industrial automation equipment production line at the future t+Δt moment through the LSTM neural network, where Δt is a preset control period and Δt≤50ms; the obtained images or videos are stored to permanently store the production line operation status information for subsequent retrieval, viewing, evidence collection, and tracking. Compare the dynamic optimization parameter set in the operating state of the industrial automation equipment production line with the preset process standard parameters in real time. When it is detected that the parameter deviation value exceeds the dynamic tolerance threshold, trigger the dynamic reconstruction of the control instruction, generate an encrypted control instruction stream containing the timestamp and the texture of the industrial automation equipment production line, extract the global temporal features of the multiple image or video information temporal input vectors, process the acquired images or videos based on the weighted average algorithm, and obtain a comprehensive comparison value to further screen out the abnormal operating state information of the production line, so that the operating state information of the production line can be known when the production line is unmanned; Among them, setting information frames at multiple predetermined time points within a predetermined time period for the image or video information includes: arranging multiple information frames in the time dimension within a predetermined time period through an improved convolutional neural network model to obtain the multiple image or video information temporal input vectors; and using a feature analysis algorithm to obtain the temporal feature vectors of the information frames at multiple predetermined time points; Among them, extracting the global temporal features of the multiple image or video information temporal input vectors includes: annotating the image or video information data through an information frame capture module with a timestamp; Processing the acquired images or videos based on the weighted average algorithm includes image enhancement, image information encoding, image fusion, and image frame analysis.

[0019] In the above embodiments, an indoor surveillance camera is installed to ensure that the camera covers the area to be monitored. Install an indoor surveillance camera to ensure that the camera covers the area to be monitored. Install an indoor surveillance camera to ensure that the camera covers the area to be monitored. Install an indoor surveillance camera to ensure that the camera covers the area to be monitored. Through feature extraction, basic information on the operating status of the production line is extracted, such as equipment status, process, production efficiency, etc. The extracted information is recorded in a database. During time segmentation and frame extraction, through time planning, according to needs, the video is segmented into multiple predetermined time periods. Within each predetermined time period, video frames are extracted at predetermined time points. Then information frame processing is performed. When frame analysis is to be carried out, each extracted frame is further analyzed, such as motor speed identification, vibration amplitude, conveyor belt tension, tool wear amount, cylinder stroke position, sensor signal strength, process anomaly detection, etc. Then information integration is performed. The extracted information is associated with timestamps to form information frames. During information frame processing, each extracted frame is further analyzed through frame analysis, such as behavior recognition, anomaly detection, etc. Finally, information integration is carried out: the extracted information is associated with timestamps to form information frames. The extracted information frames and related information are stored in a database. Functions such as querying, retrieving, updating, and deleting data are implemented. In the above embodiments, the acquired images or videos are stored to permanently store the operating status information of the production line for subsequent retrieval, viewing, evidence collection, and tracking; after the video stream arrives at the server, it is first decoded, and then video frames are extracted. Image recognition software is used to process the extracted frames to identify the operating status of the production line and extract relevant information. This information is stored together with the original video frames or in a separate database. The videos and the extracted information can be stored on a local hard drive or storage device. To provide higher reliability and accessibility, the data can be stored on a cloud server. Metadata of the videos and images are stored, including timestamps, location information, event descriptions, etc. An index is established to quickly retrieve relevant data for a specific time, location, or event.

[0020] In a further embodiment, the improved convolutional neural network model includes an input layer, a multi-dimensional convolutional layer, an activation layer, a pooling layer, a classification layer, a diagnostic layer, and an output layer, The data information of the input layer is an image or video of the operating status of the production line; The method for dynamically acquiring an image or video of the operating status of the production line includes: Collecting image data through a multi-spectral vision sensor array deployed at key workstations on the production line. The sensor array integrates a visible light camera and an infrared thermal imaging module, and the sampling frequency is not less than 60 Hz; Perform spatio-temporal alignment processing on the collected original images, and use an improved 3D convolution kernel to extract the dynamic features of continuous N frames of images in the time domain dimension, where N≥5 and satisfies N×frame interval≤Δt / 2; The output layer performs multi-node verification on the encrypted control instruction stream through a blockchain-based consensus mechanism. After passing the verification, it is sent to the target device end through a low-latency transmission channel. The low-latency transmission channel adopts a hybrid networking architecture of Time-Sensitive Networking (TSN) and 5G Ultra-Reliable Low-Latency Communication (URLLC) to ensure that the end-to-end transmission latency ≤ 10 ms.

[0021] In the above embodiment, in an industrial production line, accurately obtaining the operation status information of key workstations on the production line is of great significance for ensuring production quality, improving production efficiency, and performing equipment maintenance. The multi-spectral vision perception technology can monitor the production line from different dimensions by integrating multiple types of sensors and obtain rich multi-modal data. This technical solution focuses on the collection and preprocessing of multi-spectral data, aiming to provide high-quality, spatio-temporally aligned multi-spectral image sequences for subsequent data analysis and decision-making.

[0022] Deploy multi-spectral sensor arrays at key workstations on the production line, such as assembly stations and quality inspection stations. This array integrates visible light cameras and infrared thermal imaging modules, and its deployment principle is based on the functional characteristics of different types of sensors and the requirements of production line monitoring. The visible light camera has a high resolution (≥4K) and frame rate (60Hz), and can capture visible state information such as the mechanical movement of equipment and the position of materials. In an industrial production environment, the mechanical movement state of equipment and the position accuracy of materials directly affect the quality and efficiency of production. High-resolution images can clearly present the details of equipment and the position of materials, while a high frame rate can capture fast mechanical movements to ensure that key production links are not missed. The temperature measurement range of the infrared thermal imaging module is -20°C to 500°C, and the accuracy is ±1°C. It is mainly used to monitor the temperature fields of key components such as motor windings and bearings. In industrial production, the temperature changes of components such as motors and bearings are important indicators of the equipment operation status. Excessive temperature can indicate potential equipment faults, such as bearing wear and motor overload. By monitoring the temperature fields of these key components in real time, potential problems can be detected in time for preventive maintenance to avoid production interruptions caused by equipment failures. The sensor array is powered by Power over Ethernet (PoE) technology, which not only simplifies the wiring but also improves the reliability of the system. At the same time, the sensor array is directly connected to a Time-Sensitive Network (TSN) switch, and the characteristics of the TSN network are used to ensure that the time synchronization error ≤1μs. Time synchronization is crucial for the acquisition of multi-spectral data because only when the data collected by different sensors have accurate time correspondence can effective spatio-temporal alignment processing be carried out. The image data collected by the multi-spectral sensor needs to be subjected to spatio-temporal alignment processing to ensure the consistency of the data collected by different sensors in space and time, thereby providing an accurate basis for subsequent analysis.

[0023] Spatial alignment uses the Scale-Invariant Feature Transform (SIFT) feature matching algorithm. The core idea of this algorithm is to extract scale-invariant feature points in visible light images and infrared images, and find the corresponding relationships between them by comparing the descriptors of these feature points. The specific steps are as follows: Feature point detection: Detect points with significant features in visible light images and infrared images respectively. These points have good stability under scale, rotation, and illumination changes in the image.

[0024] Feature descriptor generation: Generate a descriptor for each detected feature point. This descriptor contains local feature information of the feature point, such as gradient direction, scale, etc.

[0025] Feature matching: Find the matching pairs between visible light images and infrared images by comparing the descriptors of feature points.

[0026] Affine transformation matrix calculation: Based on the matched feature point pairs, calculate the affine transformation matrix between the visible light image and the infrared image. The affine transformation matrix can describe the translation, rotation, and scaling relationships between images. By applying this matrix to one of the images, it can be spatially aligned with the other image.

[0027] Time alignment is based on the IEEE 802.1AS clock synchronization protocol of the TSN network. This protocol establishes a high-precision clock synchronization mechanism in the network, enabling the clocks of different sensors to be consistent. The specific implementation process is as follows: Master clock determination: Select a node in the TSN network as the master clock, which has a high-precision clock source. Slave clock synchronization: Other sensor nodes act as slave clocks and adjust their clocks to be synchronized with the master clock by communicating with the master clock. Timestamp alignment: When sensors collect data, add accurate timestamps to each data frame. Through the clock synchronization mechanism, ensure that the data frames collected by different sensors have consistent timestamps, thereby achieving time alignment.

[0028] Principle of generating multi-spectral image sequences After spatio-temporal alignment processing, a spatio-temporally aligned multi-spectral image sequence is generated. The time resolution Δt' = Δt / 2 = 25ms, and this setting satisfies the sampling theorem. The sampling theorem states that in order to accurately reconstruct the original signal, the sampling frequency must be greater than or equal to twice the highest frequency of the original signal. During the multi-spectral image acquisition process, an appropriate time resolution can ensure that the acquired image sequence can accurately reflect the dynamic changes of the production line, providing sufficient information for subsequent analysis and decision-making. This multi-spectral visual perception technology solution realizes the effective acquisition and preprocessing of multi-modal data by reasonably deploying a multi-spectral sensor array and using an advanced spatio-temporal alignment processing algorithm. The generated spatio-temporally aligned multi-spectral image sequence provides an accurate and reliable data basis for the operation status monitoring, fault diagnosis, and quality control of industrial production lines, helping to improve the intelligent level and production efficiency of industrial production.

[0029] In the above further embodiment, the spatio-temporal dual-stream convolutional network constructs a dual-stream architecture including a spatial stream network and a temporal stream network, which respectively process the spatial features and temporal sequence dynamic features of multi-spectral images. For example, when inputting a single-frame multi-spectral fusion image (H×W×4 channels, including RGB three channels and the thermal radiation intensity channel), the dynamic dilated convolution design adopts a strategy of adaptive adjustment of the dilation rate , where v is the device movement speed (unit: mm / s), which is obtained in real time by the encoder. This design enables the effective receptive field to dynamically expand to 3 times the original with the movement speed, solving the problem of feature loss of traditional convolution in high-speed movement scenarios. Feature extraction: Cascade 3 dilated convolutional layers (with dilation rates of 1, 2, and 3 respectively) to form a dilated convolution group to capture multi-scale spatial features. Form a dilated convolution group to capture multi-scale spatial features. The temporal flow network takes as input a temporal cube (H×W×5×4) composed of 5 consecutive frames, designs a separable 3D convolutional kernel (kernel size 3×3×3), which is decomposed into a spatial convolution (3×3×1) and a temporal convolution (1×1×3) in the spatio-temporal dimension, reducing the computational cost by 40% while retaining spatio-temporal correlation. Through the sliding window mechanism in the temporal dimension, dynamic features such as the device movement trajectory and the evolution of the temperature field are extracted. The spatio-temporal feature fusion mechanism is related to the gated attention module, and an adaptive weight allocation mechanism is designed in specific applications. This module dynamically adjusts the contribution degrees of spatio-temporal features through learning, emphasizing spatial features in static scenarios (such as when the device is stationary) and temporal features in dynamic scenarios. In a further embodiment, the improved convolutional neural network model includes an input layer, a multi-dimensional convolutional layer, an activation layer, a pooling layer, a classification layer, a diagnostic layer, and an output layer. In a specific embodiment, the input layer converts the original data into a form that can be processed by the neural network. For images, it is usually a three-dimensional array containing height, width, and color channels (such as RGB). Each image or video frame is passed as input data to the network. Local features in the image or video frame are extracted through the multi-dimensional convolutional layer. The input data is convolved using convolutional kernels (also known as filters), and features such as edges, textures, and shapes are extracted. Each convolutional layer contains multiple convolutional kernels. The convolutional kernels slide over the input data and perform convolutional operations on local regions. The result of the convolutional operation is a feature map that contains the extracted features. Nonlinearity is introduced through the activation layer to increase the expressive power of the network. The ReLU (Rectified Linear Unit) activation function is used, which converts all negative values to 0 and retains positive values, thus introducing nonlinearity. The activation function is applied to the output of each convolutional layer. The size of the feature map is reduced through the pooling layer, reducing the number of parameters and improving the robustness of the model. The spatial dimension of the feature map is reduced through downsampling operations (such as max pooling or average pooling). After the convolutional layer and the activation layer, the pooling operation is applied to the feature map. The extracted features are classified through the classification layer. Usually, fully connected layers are used to map the features to the output classes. The pooled feature map is flattened and then passed to the fully connected layer. The entire model is trained through the backpropagation algorithm, and the weights and biases of the convolutional kernels are adjusted through a large number of labeled data sets so that the model can learn effective feature representations. After training, the model can be used for feature extraction and classification of new image or video data.

[0030] In the above embodiments, the key frames or specific frames in the video sequence are captured by the information frame capture module, and a time stamp is assigned to each frame. The captured information frames are labeled to extract key information. Through time series aggregation or a time series model, the time series features that can represent the entire video sequence are extracted. The extracted time series features are combined into a global time series feature vector. Through the above process, a global time series feature vector that can reflect the characteristics of the entire video sequence can be extracted from multiple image or video information, and these features can be used in various applications such as video understanding, event detection, and behavior analysis.

[0031] In a further embodiment, the working method of the improved convolutional neural network model is as follows: when data x is input through the input layer, the output of the model through the output layer is Then the output function of the output layer is: In formula (1), represents a variable parameter, and the value of the output image time series frame y has k values, represents the time series sorting of the frames, and the image information in the output model is denoted as For the input feature map Convolution calculation is performed through a multi-dimensional convolutional layer, one-dimensional convolution calculation is performed through time domain convolution, and then two-dimensional calculation is performed on the input feature map and the convolution kernel in the depth direction. The calculation ability is improved through an activation function, and then global max pooling and global average pooling are performed to obtain two feature vectors, and calculate the global max pooling and the global average pooling In formula (2), represents the C-dimensional feature vector of the input feature map F at the spatial position ; In formula (3), represents the C-dimensional feature vector of the input feature map F at the spatial position ; These two feature vectors are respectively fed into a shared multi-layer perceptron (MLP) to learn channel-level features, and calculate the number of neurons in the first layer of the MLP In formula (4), represents the activation function of the MLP, is a learnable parameter; In formulas (4) and (5), are learnable parameters; the two feature vectors output by the MLP are added to obtain a 1×1×C feature vector z, and the z feature vector is mapped through a Sigmoid activation function to obtain the final channel attention weight matrix In formula (6), represents the MLP output of the global max pooling, denotes the MLP output of global average pooling, then In Equation (7), denotes the Sigmoid activation function, whose value range is (0, 1), and is used to represent the channel attention weight. In a further embodiment, the method for image processing is as follows: The method for image processing is as follows: Perform image enhancement by improving the Retinex algorithm, perform multi-scale Gaussian filtering on the brightness component V in the HSV color space, and the scale parameter σ = [5, 15, 30]; through the adaptive gain adjustment function Compensate for the details in the dark area, where λ is dynamically adjusted according to the ambient light intensity; Perform non-local means denoising on the enhanced image, and the search window radius is associated with the device vibration amplitude; Store the production line operation status images for record monitoring; Extract the production line operation status information model for comparative analysis using the texture feature set corresponding to the production line operation status image database; Compare the stored production line operation status image features with the retrieved production line operation status image features; Calculate the relationship between the production line operation status image features and the retrieved production line operation status image features in the database. Assume that the set production line operation status similarity threshold is Y, and the comprehensive comparison value of the image or video is Z. When Z is greater than Y, it is considered that the production line operation status is normal. When Z is less than Y, it is considered that the production line operation status is abnormal; Output the calculation results for reference by the monitoring center staff. In a specific embodiment, store the production line operation status images and texture feature sets for record monitoring. First, store the production line operation status images. Store the feature images of the production line operation status for record in the database. These images can be preprocessed, such as normalization, denoising, etc. Then perform texture feature extraction on the production line operation status images for record monitoring. Texture features refer to the repeating patterns or structures in an image, and they can be used to describe skin texture, facial contours, etc. In a specific embodiment, when constructing the gray-level co-occurrence matrix (GLCM), extract texture features by analyzing the spatial relationship between pixels in the image. When performing wavelet transform, use wavelet transform to analyze the texture information of the image at different frequencies. When applying the histogram of oriented gradients (HOG), extract texture features by analyzing the gradient direction and frequency of pixels in the image.

[0032] When extracting the production line operation status information model, a model is constructed to extract the texture features of the production line operation status image. The production line operation status images under record monitoring and the set of texture features corresponding to their rotation operations are stored to form a feature library. Then, the features of the production line operation status images are compared for feature comparison. When comparison is needed, the features of the production line operation status images under record monitoring are retrieved from the database. The same feature extraction process is performed on the retrieved production line operation status images to obtain their texture features. The features of the retrieved production line operation status images are compared with the features of the production line operation status images under record monitoring. Then, similarity calculation is carried out to calculate the relationship between the features of the production line operation status images under record monitoring and the features of the retrieved production line operation status images. This is usually achieved by calculating the distance or similarity between the features. When setting the threshold, a similarity threshold Y is set, and this threshold is based on predefined rules or determined through experiments. When calculating the comprehensive comparison value, a comprehensive comparison value Z is calculated, which can be a certain measure of feature similarity. Z can be the reciprocal of the feature distance, the product of similarities, etc. According to the comparison result of the calculated comprehensive comparison value Z and the threshold Y, it is decided whether the retrieved production line operation status image is considered to have normal production line operation status. The calculation result is output for the reference of the monitoring center staff.

[0033] When extracting the production line operation status features, CNN or texture analysis technology is used to extract the features of the production line operation status image. When comparing features, the similarity or distance between the features is calculated to determine whether two images correspond to the same status.

[0034] When making threshold judgment, a preset threshold is used to determine whether the production line operation status image matches the production line image information in the record database.

[0035] When outputting the decision, the comparison result is output to help users identify and verify their identities. where Δd is the deviation degree of the current frame from the historical benchmark, and k is the sensitivity adjustment factor; c) when the comprehensive comparison value exceeds the threshold a three - level abnormal alarm mechanism is triggered; in the weighted average algorithm, the calculation method of the weighted average value Z is: In formula (8), Z ij is the comparison value between the i - th captured photo and the j - th production line operation status library; K Ii is the production line operation status photo quality coefficient of the i - th captured photo; K Qjis the photo quality coefficient of the j-th production line operation status library; the total number of captured photos is m, 1 ≤ i ≤ m; the total number of photos in the production line operation status library is n, 1 ≤ j ≤ m; N is the total number of comparison values between the captured photos and the production line operation status photo library, N = m × n. In a further specific embodiment, the production line operation status similarity threshold is a value set by the monitoring center staff, and the production line operation status similarity threshold is 8 - 15.

[0036] In a further specific embodiment, a weighted softmax loss function is used to handle the remaining imbalance in the resampled segment of the weighted average algorithm for the training set { x ’’, y ’’} p , the weights of each performance parameter type in the loss function β are calculated as: In Equation (9), P is the total number of weight parameter categories; the weighted softmax loss function loss is calculated as: In Equation (10), sum ’ is the sum of samples of the performance parameter in the weighted average algorithm.

[0037] In a specific embodiment of the present invention, a weighted softmax loss function is used to handle the remaining imbalance in the resampled segment of the weighted average algorithm, and its working principle is as follows: 1. Total number of weight parameter categories (P) First, determine the total number of weight parameter categories P, which usually refers to the number of categories to be classified in the model. For example, in a multi-classification problem, P is the number of categories.

[0038] 2. Calculation of weight β According to Equation (9), calculate the weight β of each performance parameter type. The weight β usually reflects the frequency or importance of each category in the training set. The calculation method can be as follows: where, ( f i ) is the sample frequency of the i-th category, that is, the number of times this category appears in the training set. 3. Weighted softmax loss function The weighted softmax loss function is used to handle the imbalanced data set in the multi-classification problem. Its basic idea is to assign different weights to the predicted probabilities of each category to increase the prediction accuracy of the minority classes.

[0039] 4. Loss function calculation According to Equation (10), the calculation of the weighted softmax loss function is as follows: where: ( y_i' ) is the true label (0 or 1) of the i-th sample.

[0040] (\hat{y}_i) is the probability predicted by the model for the i-th sample, i.e., the softmax output.

[0041] (p) is the total number of samples in the training set. 5. Principle of weighted softmax Weighted average algorithm: In the resampling segment, there are some categories with a small number of samples, making it difficult for the model to learn the features of these categories. Weighted softmax solves this problem by increasing the weights of these minority classes.

[0042] Remaining imbalance in the resampling segment: During the resampling process, there may be situations where the number of samples in some categories is still insufficient. By using weighted softmax, the impact of this imbalance on the model performance can be reduced.

[0043] Sum of samples of performance parameters: When calculating the loss function, the sum of samples of each category needs to be considered to ensure that the contribution of the minority classes is not ignored due to the small number of samples when calculating the prediction probability.

[0044] Workflow 1. Calculate the weight β of each category.

[0045] 2. During the training process, apply the corresponding weight β to the loss contribution of each sample.

[0046] 3. By optimizing the loss function, the model will learn more balanced category features, thereby improving the recognition ability for minority classes.

[0047] In this way, the weighted softmax loss function can help the model better handle the problem of class imbalance in the training set, improving the generalization ability and classification accuracy of the model.

[0048] In a further embodiment of the present invention, during image enhancement, the contrast of the image is enhanced by adjusting the brightness value of the image, and the histogram is made more uniform by redistributing the pixel values in the image; the sharpness of the image is improved by enhancing the edges and details in the image, and the local contrast of the image is enhanced by the Retinex method; the image is decomposed into components of different scales and different frequencies by using wavelet transform for multi-scale analysis and denoising; In a specific embodiment, image enhancement is an image processing technique aimed at improving the visual quality of the image to make it easier to observe and analyze. When adjusting the brightness value of the image to enhance the contrast, by adjusting the brightness value of the image, the overall brightness of the image can be changed, thereby enhancing the contrast of the image. This is achieved through the following steps: Histogram Equalization: Redistributes the pixel values in an image to make the histogram more uniform. This method can improve the local contrast of the image, especially when the brightness of the image is uneven.

[0049] Histogram Specification: Adjusts the brightness of an image according to a specific target histogram distribution, making certain regions of the image more prominent.

[0050] Adaptive Histogram Equalization: Combines local region information to equalize the image, which can reduce the image distortion caused by histogram equalization.

[0051] When enhancing edges and details to improve sharpness, the sharpness of the image can be improved by enhancing the edges and details in the image. Common methods include: Edge Detection: Uses algorithms such as Canny, Sobel, or Prewitt to detect edges in the image.

[0052] Detail Enhancement: Enhances the details in the image through high-pass filters or sharpening techniques.

[0053] Local Contrast Enhancement: Uses the Retinex method or other local contrast enhancement techniques to improve the local contrast of the image, making the image details more obvious.

[0054] When using the Retinex method to enhance local contrast, the Retinex (Retinal Image Exponentiation) method is an image enhancement technique based on the perception principle of the human visual system. It enhances the local contrast of the image through the following steps: Brightness Extraction: Extracts the brightness component from the image, usually achieved through a non-linear filter.

[0055] Contrast Enhancement: Enhances the contrast of the extracted brightness component to make the details of the image more prominent.

[0056] Color Restoration: Restores the color information of the image to keep the enhanced image in its original color.

[0057] Wavelet Transform for Multiscale Analysis and Denoising: Decomposes the image into components of different scales and frequencies for multiscale analysis and denoising. The specific steps are as follows: Decomposition: Decomposes the image into multiple wavelet coefficients, each coefficient representing the features of the image at a specific scale and direction.

[0058] Denoising: Applies threshold processing to the wavelet coefficients to remove noise and retain the detail information of the image.

[0059] Reconstruction: Use the denoised coefficients to reconstruct the image and obtain the denoised image. Through these image enhancement techniques, the quality of the image can be significantly improved, making it more useful in subsequent analysis and applications. For example, in the fields of medical image analysis, satellite image processing, human visual system simulation, etc., image enhancement techniques play an important role.

[0060] When encoding image information, use orthogonal transformation to convert the image data from the spatial domain to the frequency domain, then perform quantization, and encode according to the frequency of pixel values using Huffman coding or arithmetic coding to reduce redundant information. When performing image fusion, encode according to the frequency of pixel values, reduce redundant information through the Bayesian algorithm, and fuse images at multiple scales to retain information at different scales. When analyzing image frames, fuse images at multiple scales to retain information at different scales; use deep learning models such as CNN to identify information in the image, track objects in the video, and associate targets in consecutive frames.

[0061] In a further embodiment, a system for remotely and dynamically monitoring the home environment Comprises: A monitoring device for acquiring images or videos of the production line operation status to obtain basic information on the production line operation status; A central control unit for processing and storing the images or videos acquired by the monitoring device; A storage unit for storing the acquired images or videos to permanently store the production line operation status information for subsequent retrieval, viewing, evidence collection, and tracking; An image processing unit for processing the acquired images or videos based on the weighted average algorithm to obtain a comprehensive comparison value, so as to further screen out the abnormal operation status information of the production line, enabling the production line operation status information to be known in the case of unmanned operation of the production line; A cloud server for permanently storing the processed image or video information to securely store the information and make it not easily lost; A handheld terminal for enabling the monitoring center staff to immediately obtain the processed image or video information to make manual intervention actions according to the dynamic operation status of the production line; wherein: The monitoring device is communicatively connected to the central control unit through a communication unit. The central control unit integrates a storage unit and an image processing unit. The output end of the central control unit is connected to the input end of the cloud server, and the output end of the cloud server is connected to the input end of the handheld terminal; Wherein the central control unit further includes a production line operation status library for storing the production line operation status images for record filing and monitoring; The central control unit communicates with the monitoring device through a communication unit, where the communication unit is a wired communication unit or a wireless communication unit; The monitoring device is arranged at the door, indoors or at the window of the monitoring center staff, and the monitoring device is a wide-angle camera or a 360 camera with a wireless communication interface.

[0062] The image processing unit further includes: An extraction module, configured to extract the production line operation status information model for comparative analysis using the texture feature set corresponding to the production line operation status image database; A comparison module, configured to compare the stored production line operation status image features with the retrieved production line operation status image features; A calculation module, configured to calculate the relationship between the production line operation status image features and the production line operation status image features retrieved from the database; An output module, configured to output the calculation result for reference by the monitoring center staff; The extraction module extracts the deep learning features of the production line operation status based on an improved convolutional neural network algorithm model.

[0063] The following are the principles and implementation processes of these modules: Extraction module Principle: The extraction module extracts the deep learning features of the production line operation status based on an improved convolutional neural network (CNN) algorithm model. CNN is a very effective deep learning model in image recognition tasks, which can automatically learn features from images.

[0064] Implementation process: 1. Preprocessing: First, preprocess the input production line operation status image, including size adjustment, normalization, denoising, etc.

[0065] 2. Feature extraction: Use the improved CNN model to process the preprocessed production line operation status image. This model can include multiple convolutional layers, pooling layers, and fully connected layers.

[0066] 3. Feature learning: In the training stage, CNN learns the features of the production line operation status image through the backpropagation algorithm. These features are usually deep and abstract, and can distinguish different production line operation statuses.

[0067] 4. Feature optimization: According to specific requirements, CNN can be improved, such as introducing new layers, adjusting the network structure, or using a pre-trained model.

[0068] Comparison module Principle: The comparison module is used to compare the extracted production line operation status features with the features stored in the database to determine whether it is the production line operation status under the same state.

[0069] Implementation process: 1. Feature extraction: Use the same CNN model to extract features from the retrieved production line operation status images.

[0070] 2. Feature comparison: Compare the newly extracted features with the features stored in the database, usually using distance metrics (such as Euclidean distance, cosine similarity, etc.).

[0071] Calculation module Principle: The calculation module is used to calculate the relationship between the extracted production line operation status features and the retrieved production line operation status features, which is usually measured by similarity or distance.

[0072] Implementation process: 1. Feature distance calculation: Calculate the distance between the newly extracted features and each feature in the database.

[0073] 2. Similarity threshold: Determine whether to match according to the preset similarity threshold (Y).

[0074] Output module Principle: The output module is responsible for presenting the results of comparison and calculation to the user in an easy-to-understand manner.

[0075] Implementation process: 1. Result evaluation: Judge whether the production line operation status is normal according to the comparison result of the calculated similarity and the threshold (Y).

[0076] 2. Result output: Output the judgment result (familiar or abnormal production line operation status) to the user, which can be through a graphical interface, API call or other communication methods.

[0077] The implementation process of the entire system involves the following steps: 1. Data collection: Collect a large amount of production line operation status image data, including front, side, and images under different operation status conditions.

[0078] 2. Model training: Use the collected data to train an improved CNN model so that it can accurately extract production line operation status features.

[0079] 3. System integration: Integrate the extraction, comparison, calculation, and output modules into a system.

[0080] 4. Testing and Optimization: Test the performance of the system and optimize the model and the system according to the test results.

[0081] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are only illustrative examples. Without departing from the principles and essence of the present invention, those skilled in the art can make various omissions, substitutions, and changes to the details of the above methods and systems. For example, combining the above method steps so as to perform substantially the same function in a substantially the same way to achieve substantially the same result falls within the scope of the present invention. Therefore, the scope of the present invention is only defined by the appended claims.

Claims

1. A remote dynamic monitoring method for an industrial automation equipment production line, characterized in that: The following steps are involved: The edge computing node deployed at the end of the industrial automation equipment production line collects the operation status data and process parameters of the industrial automation equipment production line in real time. The edge computing node has a built-in multi-protocol adaptive drive engine, which supports parallel parsing and data encapsulation of at least three industrial protocols: PROFINET, EtherCAT, and OPC UA. The image or video of the production line operation status is dynamically acquired to obtain basic information about the production line operation status. The image or video information is set as information frames at multiple predetermined time points within a predetermined time period. Based on the digital twin model, dynamic physical simulation is performed on the operating status data of the industrial automation equipment production line to generate a dynamic optimization parameter set of the industrial automation equipment production line control method, wherein the digital twin model predicts the operating status of the industrial automation equipment production line at the future time t+Δt through the LSTM neural network, Δt is a preset control cycle and Δt≤50ms; the acquired images or videos are stored to permanently store the production line operating status information for subsequent retrieval, viewing, evidence collection and tracking; The dynamically optimized parameter set in the operation state of the industrial automation equipment production line is compared with the preset process standard parameters in real time. When it is detected that the parameter deviation value exceeds the dynamic tolerance threshold, the dynamic reconstruction of the control instruction is triggered, and an encrypted control instruction stream containing a timestamp and the texture of the industrial automation equipment production line is generated. The global timing characteristics of the timing input vectors of the multiple images or video information are extracted, and the acquired images or videos are processed based on the weighted average algorithm to obtain a comprehensive comparison value, so as to further screen out the abnormal operation state information of the production line, so that the operation state information of the production line can be known when the production line is unattended; The step of setting the image or video information to a plurality of information frames at a predetermined time point within a predetermined time period comprises: arranging the plurality of information frames according to a time dimension within a predetermined time period by using an improved convolution neural network model to obtain the plurality of image or video information time series input vectors; and obtaining the information frame time series feature vectors at the plurality of predetermined time points by using a feature analysis algorithm; Wherein extracting the global temporal features of the plurality of image or video information temporal input vectors comprises: marking the image or video information data by an information frame capture module having a timestamp; The acquired image or video is processed based on the weighted average algorithm, including image enhancement, image information encoding, image fusion and image frame analysis.

2. The remote dynamic monitoring method of an industrial automation equipment production line according to claim 1, characterized in that: The improved convolutional neural network model includes an input layer, a multi-dimensional convolution layer, an activation layer, a pooling layer, a classification layer, a diagnosis layer and an output layer. The data information of the input layer is an image or video of the operating status of the production line; Methods for dynamically acquiring images or videos of the production line operation status include: Image data is collected by deploying a multi-spectral visual sensor array at key positions on the production line. The sensor array integrates a visible light camera and an infrared thermal imaging module, with a sampling frequency of no less than 60 Hz; Perform spatiotemporal alignment processing on the collected original images, and use the improved 3D convolution kernel to extract the dynamic features of N consecutive frames of images in the time domain dimension, where N ≥ 5 and Δt satisfies N × frame interval ≤ Δt / 2; The output layer performs multi-node verification on the encrypted control instruction stream through a consensus mechanism based on blockchain. After verification, it is sent to the target device via a low-latency transmission channel. The low-latency transmission channel adopts a hybrid networking architecture of time-sensitive network TSN and 5G URLLC to ensure that the end-to-end transmission delay is ≤10ms.

3. The remote dynamic monitoring method of an industrial automation equipment production line according to claim 2 is characterized in that: The improved convolutional neural network model is a spatiotemporal dual-stream architecture, including: Spatial Stream Network: A dilated convolutional layer is used to extract the device status features of a single frame image. The dilation rate is positively correlated with the device movement speed. Temporal Stream Network: Captures the abnormal state evolution characteristics between consecutive frames through 3D convolutional layers; Feature fusion module: Uses gated attention mechanism to dynamically allocate spatiotemporal feature weights, improving anomaly detection sensitivity by ≥21%; The working method of the improved convolution neural network model is: When data x is input through the input layer, the output model is Then the output function of the output layer is: In formula (1), represents a variable parameter, and the value of the output image time sequence frame y is k. Represents the temporal order of frames, and the image information in the output model is recorded as For the input feature map Convolution calculation is performed through multi-dimensional convolution layer, one-dimensional convolution calculation is performed through time domain convolution, and then the feature map and convolution kernel are input in the depth direction for two-dimensional calculation. The activation function is used to improve the computing power, and then global maximum pooling and global average pooling are performed to obtain two The feature vector of , calculate the global maximum pooling and global average pooling In formula (2), Represents the input feature map F at the spatial position The C-dimensional feature vector at ; In formula (3), Represents the input feature map F at the spatial position The C-dimensional feature vector at the location; these two feature vectors are sent to a shared multi-layer perceptron (MLP) to learn channel-level features and calculate the number of neurons in the first layer of MLP In formula (4), represents the activation function of MLP, is a learnable parameter; In formulas (4) and (5), is a learnable parameter; add the two feature vectors output by MLP to obtain a 1×1×C feature vector z, map the z feature vector with the Sigmoid activation function, and obtain the final channel attention weight matrix In formula (6), represents the global maximum pooling MLP output, represents the global average pooled MLP output, then In formula (7), Represents the Sigmoid activation function, whose actual value range is (0, 1), and is used to represent the channel attention weight.

4. The remote dynamic monitoring method for industrial automation equipment production line according to claim 1, characterized in that: The image processing method is: The image is enhanced by improving the Retinex algorithm, and the brightness component is subjected to multi-scale Gaussian filtering in the HSV color space, with the scale parameter σ=[5,15,30]; the adaptive gain adjustment function is used Compensate for dark area details, where λ is dynamically adjusted according to the ambient light intensity; perform non-local mean denoising on the enhanced image, and the search window radius is associated with the device vibration amplitude; Store the production line operation status images for record monitoring; Extracting the production line operation status information model to compare and analyze the texture feature set corresponding to the production line operation status image database; Comparing the stored production line operation status image features with the retrieved production line operation status image features; Calculate the relationship between the production line operation status image features and the production line operation status image features retrieved from the database. Assume that the set production line operation status similarity threshold is Y, and the comprehensive comparison value of the image or video is Z. When Z is greater than Y, the production line operation status is considered normal. When Z is less than Y, the production line operation status is considered abnormal. The calculation results are output for reference by the monitoring center staff.

5. The remote dynamic monitoring method of an industrial automation equipment production line according to claim 1, characterized in that: The processing process of the weighted average algorithm includes: a) extracting HSV color space histogram features and LBP texture features for each information frame to construct a multidimensional feature vector; b) calculating the Mahalanobis distance of the feature vector based on the sliding time window and dynamically allocating the weight coefficient of each frame Where Δd is the deviation between the current frame and the historical benchmark, and k is the sensitivity adjustment factor; c) When the comprehensive comparison value Exceeding the threshold When , the three-level abnormal alarm mechanism is triggered; the weighted average value Z in the weighted average algorithm is calculated as follows: In formula (8), Z ij is the comparison value between the i-th captured photo and the j-th production line operation status library; K Ii K is the quality coefficient of the production line operation status photo of the i-th snapshot; Qj is the photo quality coefficient of the jth production line operation status library; the total number of captured photos is m, 1≤i≤m; the total number of photos in the production line operation status library is n, 1≤j≤m; N is the total number of comparison values ​​between the captured photos and the production line operation status photo library, N=m×n.

6. The remote dynamic monitoring method for industrial automation equipment production line according to claim 5, characterized in that: The production line operation status similarity threshold is a value set by the monitoring center staff, and the production line operation status similarity threshold is 8-15.

7. The remote dynamic monitoring method of an industrial automation equipment production line according to claim 5, characterized in that: The weighted softmax loss function is used to deal with the residual imbalance of the resampling segment in the weighted average algorithm. x '', y ''} p , the weight of each performance parameter type in the loss function β Calculated as: In formula (9), P is the total number of weight parameter categories; weighted softmax loss function loss The calculation is: In formula (10), sum ' is the sample sum of the performance parameters in the weighted average algorithm.

8. The remote dynamic monitoring method for industrial automation equipment production line according to claim 1, characterized in that: When enhancing an image, the contrast of the image is enhanced by adjusting the brightness value of the image, and the histogram is made more uniform by redistributing the pixel values ​​in the image; the clarity of the image is improved by enhancing the edges and details in the image, and the local contrast of the image is enhanced by the Retinex method; the image is decomposed into components of different scales and frequencies by using wavelet transform to facilitate multi-scale analysis and denoising; When encoding image information, an orthogonal transform is used to convert image data from the spatial domain to the frequency domain, and then quantized. Huffman coding or arithmetic coding is used to encode according to the frequency of occurrence of pixel values ​​to reduce redundant information. When images are fused, they are encoded according to the frequency of occurrence of pixel values, and the Bayesian algorithm is used to reduce redundant information. Images are fused at multiple scales to retain information at different scales. When analyzing image frames, images are fused at multiple scales to preserve information at different scales; Use deep learning models such as CNN to recognize information in images and track objects in videos by associating targets in consecutive frames.

9. A system for remote dynamic monitoring of home environment, characterized in that: A remote dynamic monitoring method for an industrial automation equipment production line according to any one of claims 1 to 8, comprising: Monitoring device, used to obtain images or videos of the operating status of the production line to obtain basic information on the operating status of the production line; A central control unit, used for processing and storing images or videos acquired by the monitoring device; A storage unit is used to store the acquired images or videos to permanently store the production line operation status information for subsequent retrieval, viewing, evidence collection and tracking; An image processing unit is used to process the acquired image or video based on a weighted average algorithm to obtain a comprehensive comparison value, so as to further screen out abnormal operation status information of the production line, so that the operation status information of the production line can be obtained when the production line is unattended; Cloud servers are used to store processed image or video information permanently to store information securely and not easily lost; Handheld terminal, used to enable monitoring center staff to instantly obtain processed image or video information, so as to make manual intervention actions according to the dynamic operation status of the production line; among which: The monitoring device is connected to the central control unit through a communication unit, the central control unit is integrated with a storage unit and an image processing unit, the output end of the central control unit is connected to the input end of the cloud server, and the output end of the cloud server is connected to the input end of the handheld terminal; The central control unit further includes a production line operation status library for storing the production line operation status images for record monitoring; wherein the central control unit communicates with the monitoring device via a communication unit, wherein the communication unit is a wired communication unit or a wireless communication unit; The monitoring device is arranged at the door, indoors or at the window of the staff of the monitoring center, and the monitoring device is a wide-angle camera or a 360 camera with a wireless communication interface.

10. A system for remote dynamic monitoring of home environment according to claim 9, characterized in that: The image processing unit also includes: An extraction module is used to extract the production line operation status information model for comparative analysis using a texture feature set corresponding to the production line operation status image database; A comparison module, used for comparing the stored production line operation status image features with the retrieved production line operation status image features; A calculation module, used for calculating the relationship between the production line operation status image features and the production line operation status image features retrieved from the database; The output module is used to output the calculation results for reference by the staff of the monitoring center; The extraction module extracts the deep learning features of the production line operation status based on an improved convolutional neural network algorithm model.

Citation Information

Patent Citations

  • Injection molding machine

    CN110315726A