Automatic equipment inspection method in intelligent manufacturing

By deploying lightweight edge computing and 5G network slicing technologies on the device side, combined with software-defined networking and the TinyML framework, the data transmission path and model update are dynamically optimized, solving the problem of lagging real-time data processing of automatic inspection equipment. This enables low-latency transmission and real-time analysis of key data, improving the accuracy and response speed of equipment status monitoring.

CN121078486APending Publication Date: 2025-12-05南昌职业大学
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511327449.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing automated inspection equipment lags behind in real-time data processing capabilities, resulting in the inability to provide timely warnings of equipment malfunctions and causing production losses.

Method used

Lightweight edge computing nodes are deployed on the device side, and data is cleaned, compressed, and prioritized using 5G network slicing technology. Transmission paths are optimized using software-defined networking and reinforcement learning algorithms. The TinyML framework is deployed for real-time analysis, and model parameters are dynamically updated through an incremental learning framework to build a closed-loop feedback mechanism.

Benefits of technology

It enables low-latency transmission and real-time analysis of critical data, improves the accuracy and response speed of equipment status monitoring, reduces dependence on cloud resources, and optimizes network stability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121078486A_ABST
    Figure CN121078486A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically inspecting equipment in intelligent manufacturing, which comprises the following steps of: 1, deploying edge computing nodes, combining with 5G network slices, cleaning and compressing inspection data in real time, dividing priorities, performing low-delay transmission on key data through a special bandwidth, and locally caching or filtering non-key data; 2, developing a controller based on the software defined network, monitoring a link state, introducing a reinforcement learning algorithm, and dynamically optimizing a transmission path; 3, deploying a TinyML framework at an equipment end, integrating a model compression technology to operate a lightweight AI model, and combining sensor data to perform real-time early warning and support offline analysis; 4, constructing an incremental learning framework, dynamically updating model parameters through knowledge distillation, and only adjusting part of parameters to adapt to state change; and step 5, an analysis result is fed back to the system by using a message queue, and a patrol strategy is dynamically adjusted in combination with a rule engine to form a closed-loop optimization mechanism of data acquisition-transmission-analysis-decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automated inspection technology, and more particularly to an automated inspection device and method in intelligent manufacturing. Background Technology

[0002] As a core direction for the transformation and upgrading of the manufacturing industry, intelligent manufacturing stems from the global manufacturing sector's urgent pursuit of efficiency improvement, cost optimization, and personalized production. With rising labor costs, shorter product iteration cycles, and higher quality requirements, traditional production models relying on manual inspection and centralized data processing are no longer adequate. Intelligent manufacturing, by integrating technologies such as the Internet of Things, big data, and artificial intelligence, has become a key path to achieving flexible manufacturing, precise control, and efficient resource utilization. Against this backdrop, automated inspection equipment, as the "sensory tentacles" of intelligent manufacturing, undertakes core tasks such as equipment status monitoring, environmental parameter collection, and product quality inspection, and its importance is increasingly prominent. These devices can replace manual labor in completing high-risk, high-frequency, and high-precision inspection tasks.

[0003] However, current automated inspection equipment generally faces the problem of lagging real-time data processing capabilities. In the traditional mode, the data collected by the equipment needs to be transmitted to the cloud for centralized analysis. Due to network latency, bandwidth limitations and cloud load, data processing is often significantly delayed. For example, a wafer inspection device in a semiconductor factory collects thousands of temperature data per second, but due to network congestion, the cloud analysis delay is more than 20 seconds. This results in the inability to provide timely warnings about yield decline caused by abnormal temperatures, ultimately causing a single batch of wafers to be scrapped, resulting in a loss of more than one million. Therefore, an automated inspection device method in intelligent manufacturing is proposed. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing an automated inspection device method for intelligent manufacturing.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: A method for automated inspection equipment in intelligent manufacturing, comprising: Step 1: Deploy lightweight edge computing nodes on the device side, and combine them with 5G network slicing technology to clean, compress and prioritize the raw data generated by the inspection equipment in real time; by dynamically adjusting the compression algorithm, such as selecting lossless or lossy compression according to network load, ensure that critical data (such as fault warnings) are transmitted to the cloud with low latency through dedicated 5G bandwidth, while non-critical data is cached or filtered locally to reduce cloud load and network congestion risks. Step 2: Based on software-defined networking technology, develop a network controller to monitor link bandwidth, latency, and node load in real time, and dynamically optimize data transmission paths by combining reinforcement learning algorithms; introduce the "load balance" index to quantitatively evaluate the load balance between subtrees, and avoid overload on a single path through local adjustment mechanisms (such as subtree grafting and node migration) to ensure the stability and real-time performance of data transmission. Step 3: Deploy the TinyML framework on the resource-constrained inspection equipment, integrate model compression and quantization techniques such as knowledge distillation and pruning, run lightweight AI models such as LSTM fault prediction and CNN visual anomaly detection; combine the sensor data (vibration, temperature, sound, etc.) on the equipment to analyze the equipment status in real time and trigger early warnings, support offline caching and local analysis when the network is interrupted, and reduce dependence on cloud resources. Step 4: Construct an incremental learning framework to dynamically update the AI ​​model parameters on the device through online learning; combine historical and new data, and use knowledge distillation or transfer learning techniques to adjust only some parameters of the model, such as the last fully connected layer, to avoid retraining the entire model; the framework can adapt to changes in device status (such as wear and tear, aging), improve the accuracy and real-time performance of anomaly detection, and reduce the consumption of computing resources. Step 5: Feed back the real-time analysis results (anomaly type, location, severity) from the device to the inspection system via message queues such as MQTT and Kafka; combine rule engines or reinforcement learning algorithms to dynamically adjust inspection strategies (such as path, frequency, and method), for example, increase the inspection frequency of high-risk areas or avoid congested areas; the closed-loop feedback mechanism realizes real-time coordination between the inspection process and the device status, improving response speed and intelligence level.

[0006] The above technical solution further includes: Furthermore, the deployment of lightweight edge computing nodes on the device side, combined with 5G network slicing technology, to perform real-time cleaning, compression, and prioritization of the raw data generated by the inspection equipment includes the following steps: Deploy low-power, high-performance edge computing modules, such as embedded devices equipped with ARM Cortex-A series chips, at or near the inspection equipment, with sufficient memory (e.g., 2-4GB RAM) and storage space (e.g., 32GB eMMC), to support real-time data processing; integrate lightweight operating systems (e.g., RTOS or simplified Linux) and edge computing frameworks (e.g., AWS IoT Greengrass, Azure IoT Edge) to perform local data preprocessing, model inference, and communication management functions; In cooperation with telecom operators, dedicated 5G network slices are allocated to inspection equipment. These slices feature high bandwidth, low latency, and isolation guarantees, ensuring that critical data transmission is not interfered with by other service traffic. Based on real-time data priority and network load, slice resource allocation is dynamically adjusted, such as increasing the bandwidth ratio of high-priority data streams to optimize network utilization. By using rule engines (such as regular expression matching) or lightweight machine learning models (such as KNN-based anomaly detection), noise in sensor data, such as random interference in equipment vibration signals, can be identified and filtered; missing values, out-of-range values, or logically contradictory values ​​(such as a temperature sensor reporting -50℃ and 150℃ simultaneously) can be corrected or marked to improve data validity. Based on the data type (such as time series data, image data) and network status, the lossless or lossy compression algorithm is dynamically selected; for example, lossless compression is used for equipment fault codes to retain complete information, while lossy compression is used for environmental temperature and humidity data to reduce the amount of data transmitted. Priority allocation rules: High priority: Data requiring immediate response, such as equipment failure warnings and safety incidents (e.g., gas leaks), is marked as "urgent" and assigned the highest transmission priority; Medium priority: Data that needs to be monitored in real time, such as equipment operating status (e.g., speed, current), are marked as "important" and assigned the second highest priority; Low priority: Environmental monitoring data (such as temperature, humidity, dust concentration) and other data that can be processed later are marked as "normal" and assigned the lowest priority; For low-priority data or data generated during network congestion, a circular buffer or solid-state drive is used for temporary storage to prevent data loss; the data transmission order is dynamically adjusted based on priority and network slice resources. For example, high-priority data is transmitted in real time via dedicated 5G slices, medium-priority data is transmitted when the network is idle, and low-priority data is transmitted in batches periodically. After the edge nodes complete data cleaning, compression, and prioritization, they only transmit high-priority data and compressed medium- and low-priority data to the cloud, reducing the cloud's processing load. The edge nodes periodically report device status to the cloud, such as uptime and fault history. The cloud adjusts the edge nodes' preprocessing rules based on global information (such as device maintenance plans), such as modifying compression algorithms or priority thresholds.

[0007] Furthermore, the method of dynamically adjusting the compression algorithm to enable low-latency transmission of critical data to the cloud via dedicated 5G bandwidth, while non-critical data is cached or filtered locally, includes the following steps: Compression algorithms are dynamically selected based on real-time network load (e.g., 5G base station congestion), data type (e.g., time-series signals, images), and priority (e.g., fault warning, environmental monitoring). Under high network load, lossy compression, such as JPEG-LS for image data and differential compression for time-series data, is prioritized, sacrificing a small amount of non-critical information (e.g., minor fluctuations in ambient temperature and humidity) to improve transmission efficiency. Under low network load, lossless compression, such as LZ4 for text data and FLAC for audio data, is switched to preserve data integrity and is suitable for scenarios requiring precise analysis, such as early fault characteristics in equipment vibration signals. The switching of compression algorithms is automatically triggered by real-time monitoring of network latency or bandwidth utilization at edge nodes, without manual intervention. Dedicated 5G network slices are allocated to critical data (such as device fault codes and security event logs). These slices feature high-priority scheduling, low-latency transmission, and bandwidth guarantees, ensuring they are unaffected by other service traffic. The QUIC protocol is used to replace traditional TCP, reducing handshake delay and packet loss retransmission time, further improving the efficiency of critical data transmission.

[0008] Local caching involves deploying solid-state drives or high-durability flash memory at edge nodes to store non-critical data, such as ambient temperature and humidity, and device runtime. The cache space is managed using a "first-in, first-out" or "least recently used" algorithm. When the storage capacity reaches a threshold, low-priority or timed-out data is deleted first. Data filtering: Time-dimensional filtering: For periodic data, such as ambient temperature and humidity collected every second, only samples with changes exceeding a threshold are retained to reduce redundant transmission; Spatial dimension filtering: For multi-sensor data, such as vibration signals from different locations of the same device, the most representative data is selected through correlation analysis (such as Pearson correlation coefficient) and duplicate information is removed; Edge nodes periodically synchronize their current compression algorithms, cache status, and filtering rules to the cloud (e.g., every 5 minutes). The cloud dynamically adjusts the compression strategies of edge nodes based on global network load (e.g., cross-regional base station congestion) and equipment maintenance plans (e.g., upcoming equipment overhauls). If anomalies are detected in non-critical data cached locally, such as continuous abnormal fluctuations in ambient temperature and humidity, edge nodes automatically mark them as "potentially critical data" and prioritize their transmission to the cloud to avoid missing important information.

[0009] Furthermore, the development of a network controller based on software-defined networking technology to monitor link bandwidth, latency, and node load in real time, and to dynamically optimize data transmission paths using reinforcement learning algorithms, includes the following steps: Deploy SDN controllers, such as OpenDaylight and ONOS, in data centers or core network nodes to serve as the "brain" of network management; the controllers communicate with network devices (such as switches and routers) through southbound interfaces (such as the OpenFlow protocol) to centrally control network traffic; Monitor real-time status: Link bandwidth: Real-time bandwidth utilization (e.g., used bandwidth / total bandwidth) of each physical link (such as fiber optic cable, 5G base station backhaul link) is collected periodically via SNMP or NetFlow protocol. Latency measurement: Periodically probe the end-to-end latency of the link using Ping or ICMP protocols, such as RTT from the edge node to the cloud server, and record historical latency distribution, such as average latency and 95th percentile latency. Node load: Real-time load status of network devices (such as switch CPU and memory) is obtained through Telemetry technology to identify high-load nodes; State space: Defines quantitative indicators of network status, including bandwidth utilization of each link, real-time latency of each link (such as discretized intervals of 10ms, 50ms, and 100ms), and load status of each node, such as three-level labels: "normal", "high load", and "overload". Action space: Defines the path adjustment operations that the controller can perform, including: Path switching: Switching the current data flow from a congested link to an idle link, such as switching from a link of 5G base station A to a link of base station B; Bandwidth adjustment: Dynamically allocate link bandwidth, such as increasing bandwidth for high-priority data streams by 20% and reducing bandwidth for low-priority data streams; Node avoidance: Bypassing high-load nodes, such as rerouting data flows that originally passed through switch X to switch Y; Reward function: Design quantitative indicators to evaluate the effectiveness of actions, including: Transmission success rate: The percentage of data packets that successfully reach the cloud; Latency reduction: The amount by which the backend latency of the action execution is reduced; Load balancing: The variance of bandwidth utilization of each link after the action is executed (the smaller the variance, the higher the reward). The SDN controller updates the network state every second (or less) and inputs the current state into the reinforcement learning model. Based on the current state and historical experience (Q-table), the reinforcement learning model selects the action with the highest expected reward, such as switching to a link with lower latency. The controller sends flow tables to network devices via the OpenFlow protocol, such as "match data flow IP prefix, execute action: output to port 2", to adjust the path. The model updates the Q-table based on the actual reward after the path is executed (such as a 5% increase in transmission success rate) to optimize future decisions.

[0010] Furthermore, the introduction of the "load balancing degree" metric to quantitatively evaluate the load balancing between subtrees, and to avoid overload on a single path through a local adjustment mechanism, includes the following steps: Load balancing is used to quantitatively evaluate the load balancing of each link / node within the same region or subtree (such as the same 5G base station cluster or factory intranet) to avoid overload of a single path. Mathematical expression:

[0011] in, Let be the bandwidth utilization rate of the i-th link, such as the link utilization rate of 5G base station A being 70%; LBD represents the average bandwidth utilization of all links within the subtree; n is the total number of links within the subtree; LBD value range: 0 (completely unbalanced, such as one link being 100% loaded while the others are idle) to 1 (completely balanced, with all links having similar loads). The LBD trigger threshold is set according to business needs, and the threshold is dynamically adjusted based on historical data. For example, it is relaxed to 0.6 during peak hours and tightened to 0.8 during off-peak hours. Real-time monitoring and judgment: The SDN controller calculates the LBD of each link in the subtree every second. When it detects that the LBD is lower than the threshold, it automatically triggers the local adjustment process.

[0012] Local adjustment: Subtree grafting, a link within a subtree (such as the link of 5G base station A) is overloaded ( >80%), while the load on adjacent subtrees (such as the link of 5G base station B) is low ( <50%); migrate some data streams from high-load links to low-load links. For example, redirect patrol equipment data (such as camera video streams) originally transmitted through base station A to base station B, and issue flow table rules through the SDN controller, such as "match device IP prefix and output to the port of base station B"; Node migration: A network node (such as switch X) in a subtree is overloaded (CPU utilization > 80%) and becomes a bottleneck; bypass the high-load node and replan the data flow path; for example, redirect the data flow that originally passed through switch X (such as control commands from the edge node to the cloud) to switch Y, and update the flow table through the SDN controller, such as "match the data flow protocol type and output to the port of switch Y"; After the adjustment is executed, the SDN controller continuously monitors the LBD and node load of each link in the subtree and evaluates the adjustment effect, such as whether the LBD has increased to above the threshold. If the LBD is still below the threshold after the adjustment, the controller will start a new round of local adjustment (such as further migrating data flow or adjusting bandwidth allocation) until a balanced state is reached. Traditional methods (such as round-robin scheduling) only focus on link bandwidth utilization, while LBD assesses load balance through variance quantification, more accurately identifying imbalance states; traditional methods require manual setting of fixed thresholds, while LBD combines dynamic thresholds and real-time adjustments to adapt to changes in network load (such as 5G base station load fluctuations); traditional methods may affect other areas due to global adjustments, while local adjustment mechanisms only target unbalanced subtrees, reducing the impact on the overall network.

[0013] Furthermore, the deployment of the TinyML framework on resource-constrained inspection equipment, integrating model compression and quantization technologies, and running a lightweight AI model includes the following steps: To address resource constraints on the inspection equipment, such as low-power MCUs and limited memory, the TinyML framework, specifically designed for embedded devices, is selected. TensorFlow Lite for Microcontrollers: Supports quantized model deployment with a memory footprint as low as tens of KB, suitable for microcontrollers based on the ARM Cortex-M series; Edge Impulse: Provides an end-to-end development toolchain, supports automatic model optimization and sensor data preprocessing, and simplifies the device deployment process; Hardware adaptation: Adjust the framework parameters according to the hardware configuration of the device (such as STM32F4 series MCU, ESP32-S3 SoC), such as enabling hardware acceleration (such as using the MCU's DSP instruction set or the SoC's NPU module), and configuring memory management strategies (such as paging loading model parameters to avoid occupying too much RAM at once). Model compression: Pruning: Remove redundant neurons or connections in the model, such as pruning neurons with weights close to zero in fully connected layers to reduce the model size; for example, pruning a 1 million parameter LSTM model to 500,000 parameters while maintaining more than 95% accuracy. Knowledge distillation: Transferring knowledge from a large “teacher model” (such as ResNet-50 trained in the cloud) to a small “student model” (such as MobileNetV2 deployed on the device), and guiding the training of the student model through soft labels (output probabilities of the teacher model) to achieve a balance between accuracy and model size; Quantification: Post-training quantization: Converts model parameters from 32-bit floating-point numbers to 8-bit integers, reducing memory usage and accelerating inference; Hybrid quantization: Retains floating-point precision for sensitive layers (such as activation function layers) in the model, while using integer quantization for other layers, balancing accuracy and efficiency; Deploy targeted models based on the needs of the patrol mission, for example: Time series analysis: Lightweight LSTM or GRU (Gated Cyclic Unit) models are used to process time series data such as equipment vibration and temperature, and to detect early faults in real time, such as bearing wear. Image anomaly detection: Lightweight CNN models such as MobileNetV2 or EfficientNet-Lite are used to process images captured by the camera and identify surface defects (such as cracks and rust) or abnormal behavior (such as intrusion). Optimization model: Layer reduction: The standard LSTM's 3-layer structure is reduced to 1-2 layers, reducing computational load; Filter count adjustment: The number of filters in the convolutional layers of the CNN was reduced from 64 / 128 to 16 / 32 to reduce model complexity; Input resolution compression: Reduces the image input resolution from 224x224 to 96x96 or lower, reducing pixel-level computational burden; The system reads sensor data (such as accelerometer and temperature sensor) from the device in real time via SPI, I2C, or UART interfaces; performs data normalization, noise reduction (such as moving average filtering), and feature extraction (such as time-domain features and frequency-domain features) within the TinyML framework; the model performs inference once at a fixed period and outputs classification results (such as "normal", "warning", "fault") or regression values ​​(such as remaining life prediction); when the inference result exceeds a preset threshold, a local warning is immediately triggered via GPIO or CAN bus, such as a buzzer or LED indicator, and key data (such as abnormal timestamps and feature vectors) is simultaneously transmitted to the cloud; Real-time tracking of CPU utilization, memory usage, and power consumption is achieved through device-side system monitoring tools (such as FreeRTOS's memory management module) to ensure that model operation does not exceed hardware limits (e.g., CPU utilization <70%, memory remaining >20%). The model version is dynamically switched according to resource load (e.g., switching to a lighter "simplified" model under high load and switching to a "full" model under low load). Model parameters are dynamically adjusted via cloud commands (e.g., the forget gate threshold of LSTM and the activation function slope of CNN) to adapt to changes in device status (e.g., vibration feature shift caused by aging).

[0014] Furthermore, the support for offline caching and local analysis during network interruptions includes the following steps: The device sends heartbeat packets periodically (e.g., a simple request to the cloud every 5 seconds) and waits for a response. If no response is received after a set threshold is exceeded, it is determined that the network is interrupted. The device obtains the link status (e.g., physical layer connection status, signal strength) in real time through the device's network interface (e.g., Wi-Fi, 5G module). If the link status is "disconnected" or the signal strength is consistently below the threshold, a network interruption warning is triggered. Using SLC or eMMC type flash memory chips, it has the characteristics of shock resistance and high temperature resistance, making it suitable for long-term storage in industrial environments; a fixed area is designated in the memory as a temporary cache, and an "overwrite" mechanism is adopted. When the cache is full, the new data automatically overwrites the oldest data to avoid memory overflow. Assign priority labels, such as "high", "medium", and "low", to cached data based on data type (e.g., device status, environmental monitoring) and urgency level (e.g., fault warning, general logs). Dynamically adjust the caching ratio of each priority data based on the device's remaining storage space and the duration of network interruption. For example, if the network interruption exceeds 1 hour, increase the caching ratio of high-priority data to 80%. During network outages, the device continuously collects data (such as vibration and temperature) through sensors and directly inputs it into the deployed TinyML model (such as an LSTM fault prediction model). The model uses the device's computing power (such as the DSP module of the MCU or the NPU of the SoC) to complete inference and output classification results (such as "normal", "warning", "fault") or regression values ​​(such as remaining lifetime prediction). If the inference result exceeds the preset threshold (e.g., vibration amplitude > 0.5g, temperature > 80℃), a local warning will be triggered immediately; abnormal events (e.g., timestamps, sensor data, inference results) will be stored in a structured format in the cache medium and marked as "needs to be synchronized to the cloud"; The device continuously monitors the network status, and automatically starts the data synchronization process when the network is detected to have recovered (such as receiving a response from the cloud or the link signal strength rising back above the threshold). Incremental synchronization: Prioritize synchronizing high-priority data: First, upload abnormal events and critical status data marked as "needs to be synchronized to the cloud", such as fault warnings and equipment downtime records; Batch synchronization of low-priority data: After completing the synchronization of high-priority data, batch upload low-priority data such as environmental monitoring data, such as temperature and humidity, and equipment runtime. Calculate a hash value (such as SHA-256) for the synchronized data and compare it with the hash value stored in the cloud to ensure that the data has not been tampered with or lost. If the network is interrupted again during the synchronization process, record the position of the uploaded data and continue the transmission from the breakpoint after the network is restored to avoid duplicate uploads.

[0015] Furthermore, the incremental learning framework is constructed to dynamically update the AI ​​model parameters on the device through online learning; combining historical and new data, knowledge distillation or transfer learning techniques are used to adjust only some model parameters, including the following steps: An incremental learning engine is deployed in the cloud to receive new data uploaded from the device and perform model updates. The device retains a lightweight copy of the historical model parameters to support fast local inference. The cloud database stores historical data uploaded from the device, such as device status and fault records for the past 3 months, and organizes the data by device ID and timestamp. New data uploaded from the device (such as sensor data from the last hour) first enters the cloud message queue and is processed by the incremental learning engine in batches (such as 100 data points per batch). When new data accumulates to a preset threshold or the device detects a significant change in data distribution, such as a shift in the mean of vibration characteristics exceeding 20%, the cloud-based incremental learning engine automatically initiates model updates. Parameter tuning: Knowledge distillation: The teacher model uses the full model trained with historical data as the source of knowledge; the student model is a lightweight model deployed on the device, with only the parameters of its last few fully connected layers (such as output layer weights) adjusted, while most of the convolutional layer parameters remain unchanged; knowledge transfer guides the training of the student model through soft labels (output probabilities of the teacher model), so that the adjusted parameters can adapt to new data while retaining historical knowledge; Transfer learning: Pre-trained model: A model pre-trained on a general device dataset to extract common features, such as edge detection and texture analysis.

[0016] Fine-tuning: Only the classification layer (such as the fully connected layer) at the top of the model is fine-tuned to adapt to the specific tasks of the inspection equipment, such as fault type classification, while the parameters of the bottom feature extraction layer remain unchanged; During model updates, L2 regularization is applied to key parameters (such as weights that performed well in historical tasks) to limit their adjustment range and prevent them from being over-modified by new data. After the historical model has been trained, the contribution of each parameter to the loss function is calculated (e.g., by approximation using the second derivative) to identify parameters that were important to the historical tasks. During incremental learning, additional penalties are applied to the adjustment of important parameters (e.g., by adding a loss function term) to make their adjustment range smaller than that of unimportant parameters. After the model is updated, offline testing is performed using retained historical data (such as device status data from the most recent month) to calculate metrics such as accuracy and recall, ensuring that the performance of the updated model is no lower than that of the previous version. The updated model (version B) and the old version (version A) are simultaneously deployed to some devices, and their performance in real-time inference (such as fault detection latency and false positive rate) is compared. The version with better performance is selected as the primary model. If the updated model continues to deteriorate in online performance (such as an accuracy drop of more than 5%), it is automatically rolled back to the previous version, and an alarm is triggered to notify maintenance personnel. The cloud only sends the parameter differences after model adjustment (such as the weight changes of the last few fully connected layers) to the device, instead of the full model, reducing the amount of data transmitted (e.g., from 10MB to 1KB); the device caches the parameter differences of multiple historical model versions (e.g., the three most recent versions) and selects the appropriate version to load based on network conditions; for example, when the network is interrupted, the latest parameter differences cached locally are used to update the model; when the network is restored, the latest version is synchronized to the cloud.

[0017] The present invention has the following beneficial effects: This invention utilizes edge computing and 5G network slicing technology to achieve real-time data cleaning, compression, and prioritization, dynamically adjusting the compression algorithm to ensure low-latency transmission of critical data. It optimizes transmission paths by combining software-defined networking and reinforcement learning algorithms, introducing a "load balancing" metric to avoid path overload and improve network stability. A lightweight AI model is run using the TinyML framework on the device side, supporting real-time status analysis and offline early warning. Model parameters are dynamically updated using an incremental learning framework, adjusting only local parameters to adapt to state changes and reducing computational resource consumption. Finally, results are fed back to the system via a message queue, and the inspection strategy is dynamically adjusted using a rule engine or reinforcement learning algorithm, forming a closed-loop optimization mechanism of data collection, transmission, analysis, and decision-making, effectively solving the problem of lagging real-time data processing capabilities. Attached Figure Description

[0018] Figure 1 This is a flowchart of an automatic inspection equipment method in intelligent manufacturing proposed in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figure 1 As shown, the present invention is an automatic inspection device method in intelligent manufacturing, comprising: Step 1: Deploy lightweight edge computing nodes on the device side, and combine them with 5G network slicing technology to clean, compress and prioritize the raw data generated by the inspection equipment in real time; by dynamically adjusting the compression algorithm, such as selecting lossless or lossy compression according to network load, ensure that critical data (such as fault warnings) are transmitted to the cloud with low latency through dedicated 5G bandwidth, while non-critical data is cached or filtered locally to reduce cloud load and network congestion risks. Step 2: Based on software-defined networking technology, develop a network controller to monitor link bandwidth, latency, and node load in real time, and dynamically optimize data transmission paths by combining reinforcement learning algorithms; introduce the "load balance" index to quantitatively evaluate the load balance between subtrees, and avoid overload on a single path through local adjustment mechanisms (such as subtree grafting and node migration) to ensure the stability and real-time performance of data transmission. Step 3: Deploy the TinyML framework on the resource-constrained inspection equipment, integrate model compression and quantization techniques such as knowledge distillation and pruning, run lightweight AI models such as LSTM fault prediction and CNN visual anomaly detection; combine the sensor data (vibration, temperature, sound, etc.) on the equipment to analyze the equipment status in real time and trigger early warnings, support offline caching and local analysis when the network is interrupted, and reduce dependence on cloud resources. Step 4: Construct an incremental learning framework to dynamically update the AI ​​model parameters on the device through online learning; combine historical and new data, and use knowledge distillation or transfer learning techniques to adjust only some parameters of the model, such as the last fully connected layer, to avoid retraining the entire model; the framework can adapt to changes in device status (such as wear and tear, aging), improve the accuracy and real-time performance of anomaly detection, and reduce the consumption of computing resources. Step 5: Feed back the real-time analysis results (anomaly type, location, severity) from the device to the inspection system via message queues such as MQTT and Kafka; combine rule engines or reinforcement learning algorithms to dynamically adjust inspection strategies (such as path, frequency, and method), for example, increase the inspection frequency of high-risk areas or avoid congested areas; the closed-loop feedback mechanism realizes real-time coordination between the inspection process and the device status, improving response speed and intelligence level.

[0021] In one embodiment, the deployment of lightweight edge computing nodes on the device side, combined with 5G network slicing technology, to perform real-time cleaning, compression, and prioritization of the raw data generated by the inspection equipment includes the following steps: Deploy low-power, high-performance edge computing modules, such as embedded devices equipped with ARM Cortex-A series chips, at or near the inspection equipment, with sufficient memory (e.g., 2-4GB RAM) and storage space (e.g., 32GB eMMC), to support real-time data processing; integrate lightweight operating systems (e.g., RTOS or simplified Linux) and edge computing frameworks (e.g., AWS IoT Greengrass, Azure IoT Edge) to perform local data preprocessing, model inference, and communication management functions; In cooperation with telecom operators, dedicated 5G network slices are allocated to inspection equipment. These slices feature high bandwidth, low latency, and isolation guarantees, ensuring that critical data transmission is not interfered with by other service traffic. Based on real-time data priority and network load, slice resource allocation is dynamically adjusted, such as increasing the bandwidth ratio of high-priority data streams to optimize network utilization. By using rule engines (such as regular expression matching) or lightweight machine learning models (such as KNN-based anomaly detection), noise in sensor data, such as random interference in equipment vibration signals, can be identified and filtered; missing values, out-of-range values, or logically contradictory values ​​(such as a temperature sensor reporting -50℃ and 150℃ simultaneously) can be corrected or marked to improve data validity. Based on the data type (such as time series data, image data) and network status, the lossless or lossy compression algorithm is dynamically selected; for example, lossless compression is used for equipment fault codes to retain complete information, while lossy compression is used for environmental temperature and humidity data to reduce the amount of data transmitted. Priority allocation rules: High priority: Data requiring immediate response, such as equipment failure warnings and safety incidents (e.g., gas leaks), is marked as "urgent" and assigned the highest transmission priority; Medium priority: Data that needs to be monitored in real time, such as equipment operating status (e.g., speed, current), are marked as "important" and assigned the second highest priority; Low priority: Environmental monitoring data (such as temperature, humidity, dust concentration) and other data that can be processed later are marked as "normal" and assigned the lowest priority; For low-priority data or data generated during network congestion, a circular buffer or solid-state drive is used for temporary storage to prevent data loss; the data transmission order is dynamically adjusted based on priority and network slice resources. For example, high-priority data is transmitted in real time via dedicated 5G slices, medium-priority data is transmitted when the network is idle, and low-priority data is transmitted in batches periodically. After the edge nodes complete data cleaning, compression, and prioritization, they only transmit high-priority data and compressed medium- and low-priority data to the cloud, reducing the cloud's processing load. The edge nodes periodically report device status to the cloud, such as uptime and fault history. The cloud adjusts the edge nodes' preprocessing rules based on global information (such as device maintenance plans), such as modifying compression algorithms or priority thresholds.

[0022] In one embodiment, the method of dynamically adjusting the compression algorithm to enable low-latency transmission of critical data to the cloud via dedicated 5G bandwidth, while non-critical data is cached or filtered locally, includes the following steps: Compression algorithms are dynamically selected based on real-time network load (e.g., 5G base station congestion), data type (e.g., time-series signals, images), and priority (e.g., fault warning, environmental monitoring). Under high network load, lossy compression, such as JPEG-LS for image data and differential compression for time-series data, is prioritized, sacrificing a small amount of non-critical information (e.g., minor fluctuations in ambient temperature and humidity) to improve transmission efficiency. Under low network load, lossless compression, such as LZ4 for text data and FLAC for audio data, is switched to preserve data integrity and is suitable for scenarios requiring precise analysis, such as early fault characteristics in equipment vibration signals. The switching of compression algorithms is automatically triggered by real-time monitoring of network latency or bandwidth utilization at edge nodes, without manual intervention. Dedicated 5G network slices are allocated to critical data (such as device fault codes and security event logs). These slices feature high-priority scheduling, low-latency transmission, and bandwidth guarantees, ensuring they are unaffected by other service traffic. The QUIC protocol is used to replace traditional TCP, reducing handshake delay and packet loss retransmission time, further improving the efficiency of critical data transmission.

[0023] Local caching involves deploying solid-state drives or high-durability flash memory at edge nodes to store non-critical data, such as ambient temperature and humidity, and device runtime. The cache space is managed using a "first-in, first-out" or "least recently used" algorithm. When the storage capacity reaches a threshold, low-priority or timed-out data is deleted first. Data filtering: Time-dimensional filtering: For periodic data, such as ambient temperature and humidity collected every second, only samples with changes exceeding a threshold are retained to reduce redundant transmission; Spatial dimension filtering: For multi-sensor data, such as vibration signals from different locations of the same device, the most representative data is selected through correlation analysis (such as Pearson correlation coefficient) and duplicate information is removed; Edge nodes periodically synchronize their current compression algorithms, cache status, and filtering rules to the cloud (e.g., every 5 minutes). The cloud dynamically adjusts the compression strategies of edge nodes based on global network load (e.g., cross-regional base station congestion) and equipment maintenance plans (e.g., upcoming equipment overhauls). If anomalies are detected in non-critical data cached locally, such as continuous abnormal fluctuations in ambient temperature and humidity, edge nodes automatically mark them as "potentially critical data" and prioritize their transmission to the cloud to avoid missing important information.

[0024] In one embodiment, the development of a network controller based on software-defined networking technology to monitor link bandwidth, latency, and node load in real time, and to dynamically optimize data transmission paths using reinforcement learning algorithms, includes the following steps: Deploy SDN controllers, such as OpenDaylight and ONOS, in data centers or core network nodes to serve as the "brain" of network management; the controllers communicate with network devices (such as switches and routers) through southbound interfaces (such as the OpenFlow protocol) to centrally control network traffic; Monitor real-time status: Link bandwidth: Real-time bandwidth utilization (e.g., used bandwidth / total bandwidth) of each physical link (such as fiber optic cable, 5G base station backhaul link) is collected periodically via SNMP or NetFlow protocol. Latency measurement: Periodically probe the end-to-end latency of the link using Ping or ICMP protocols, such as RTT from the edge node to the cloud server, and record historical latency distribution, such as average latency and 95th percentile latency. Node load: Real-time load status of network devices (such as switch CPU and memory) is obtained through Telemetry technology to identify high-load nodes; State space: Defines quantitative indicators of network status, including bandwidth utilization of each link, real-time latency of each link (such as discretized intervals of 10ms, 50ms, and 100ms), and load status of each node, such as three-level labels: "normal", "high load", and "overload". Action space: Defines the path adjustment operations that the controller can perform, including: Path switching: Switching the current data flow from a congested link to an idle link, such as switching from a link of 5G base station A to a link of base station B; Bandwidth adjustment: Dynamically allocate link bandwidth, such as increasing bandwidth for high-priority data streams by 20% and reducing bandwidth for low-priority data streams; Node avoidance: Bypassing high-load nodes, such as rerouting data flows that originally passed through switch X to switch Y; Reward function: Design quantitative indicators to evaluate the effectiveness of actions, including: Transmission success rate: The percentage of data packets that successfully reach the cloud; Latency reduction: The amount by which the backend latency of the action execution is reduced; Load balancing: The variance of bandwidth utilization of each link after the action is executed (the smaller the variance, the higher the reward). The SDN controller updates the network state every second (or less) and inputs the current state into the reinforcement learning model. Based on the current state and historical experience (Q-table), the reinforcement learning model selects the action with the highest expected reward, such as switching to a link with lower latency. The controller sends flow tables to network devices via the OpenFlow protocol, such as "match data flow IP prefix, execute action: output to port 2", to adjust the path. The model updates the Q-table based on the actual reward after the path is executed (such as a 5% increase in transmission success rate) to optimize future decisions.

[0025] In one embodiment, the introduction of the "load balancing degree" metric to quantitatively evaluate the load balancing between subtrees and to avoid overload on a single path through a local adjustment mechanism includes the following steps: Load balancing is used to quantitatively evaluate the load balancing of each link / node within the same region or subtree (such as the same 5G base station cluster or factory intranet) to avoid overload of a single path. Mathematical expression:

[0026] in, Let be the bandwidth utilization rate of the i-th link, such as the link utilization rate of 5G base station A being 70%; LBD represents the average bandwidth utilization of all links within the subtree; n is the total number of links within the subtree; LBD value range: 0 (completely unbalanced, such as one link being 100% loaded while the others are idle) to 1 (completely balanced, with all links having similar loads). The LBD trigger threshold is set according to business needs, and the threshold is dynamically adjusted based on historical data. For example, it is relaxed to 0.6 during peak hours and tightened to 0.8 during off-peak hours. Real-time monitoring and judgment: The SDN controller calculates the LBD of each link in the subtree every second. When it detects that the LBD is lower than the threshold, it automatically triggers the local adjustment process.

[0027] Local adjustment: Subtree grafting, a link within a subtree (such as the link of 5G base station A) is overloaded ( >80%), while the load on adjacent subtrees (such as the link of 5G base station B) is low ( <50%); migrate some data streams from high-load links to low-load links. For example, redirect patrol equipment data (such as camera video streams) originally transmitted through base station A to base station B, and issue flow table rules through the SDN controller, such as "match device IP prefix and output to the port of base station B"; Node migration: A network node (such as switch X) in a subtree is overloaded (CPU utilization > 80%) and becomes a bottleneck; bypass the high-load node and replan the data flow path; for example, redirect the data flow that originally passed through switch X (such as control commands from the edge node to the cloud) to switch Y, and update the flow table through the SDN controller, such as "match the data flow protocol type and output to the port of switch Y"; After the adjustment is executed, the SDN controller continuously monitors the LBD and node load of each link in the subtree and evaluates the adjustment effect, such as whether the LBD has increased to above the threshold. If the LBD is still below the threshold after the adjustment, the controller will start a new round of local adjustment (such as further migrating data flow or adjusting bandwidth allocation) until a balanced state is reached. Traditional methods (such as round-robin scheduling) only focus on link bandwidth utilization, while LBD assesses load balance through variance quantification, more accurately identifying imbalance states; traditional methods require manual setting of fixed thresholds, while LBD combines dynamic thresholds and real-time adjustments to adapt to changes in network load (such as 5G base station load fluctuations); traditional methods may affect other areas due to global adjustments, while local adjustment mechanisms only target unbalanced subtrees, reducing the impact on the overall network.

[0028] In one embodiment, deploying the TinyML framework on a resource-constrained inspection device, integrating model compression and quantization techniques, and running a lightweight AI model includes the following steps: To address resource constraints on the inspection equipment, such as low-power MCUs and limited memory, the TinyML framework, specifically designed for embedded devices, is selected. TensorFlow Lite for Microcontrollers: Supports quantized model deployment with a memory footprint as low as tens of KB, suitable for microcontrollers based on the ARM Cortex-M series; Edge Impulse: Provides an end-to-end development toolchain, supports automatic model optimization and sensor data preprocessing, and simplifies the device deployment process; Hardware adaptation: Adjust the framework parameters according to the hardware configuration of the device (such as STM32F4 series MCU, ESP32-S3 SoC), such as enabling hardware acceleration (such as using the MCU's DSP instruction set or the SoC's NPU module), and configuring memory management strategies (such as paging loading model parameters to avoid occupying too much RAM at once). Model compression: Pruning: Remove redundant neurons or connections in the model, such as pruning neurons with weights close to zero in fully connected layers to reduce the model size; for example, pruning a 1 million parameter LSTM model to 500,000 parameters while maintaining more than 95% accuracy. Knowledge distillation: Transferring knowledge from a large “teacher model” (such as ResNet-50 trained in the cloud) to a small “student model” (such as MobileNetV2 deployed on the device), and guiding the training of the student model through soft labels (output probabilities of the teacher model) to achieve a balance between accuracy and model size; Quantification: Post-training quantization: Converts model parameters from 32-bit floating-point numbers to 8-bit integers, reducing memory usage and accelerating inference; Hybrid quantization: Retains floating-point precision for sensitive layers (such as activation function layers) in the model, while using integer quantization for other layers, balancing accuracy and efficiency; Deploy targeted models based on the needs of the patrol mission, for example: Time series analysis: Lightweight LSTM or GRU (Gated Cyclic Unit) models are used to process time series data such as equipment vibration and temperature, and to detect early faults in real time, such as bearing wear. Image anomaly detection: Lightweight CNN models such as MobileNetV2 or EfficientNet-Lite are used to process images captured by the camera and identify surface defects (such as cracks and rust) or abnormal behavior (such as intrusion). Optimization model: Layer reduction: The standard LSTM's 3-layer structure is reduced to 1-2 layers, reducing computational load; Filter count adjustment: The number of filters in the convolutional layers of the CNN was reduced from 64 / 128 to 16 / 32 to reduce model complexity; Input resolution compression: Reduces the image input resolution from 224x224 to 96x96 or lower, reducing pixel-level computational burden; The system reads sensor data (such as accelerometer and temperature sensor) from the device in real time via SPI, I2C, or UART interfaces; performs data normalization, noise reduction (such as moving average filtering), and feature extraction (such as time-domain features and frequency-domain features) within the TinyML framework; the model performs inference once at a fixed period and outputs classification results (such as "normal", "warning", "fault") or regression values ​​(such as remaining life prediction); when the inference result exceeds a preset threshold, a local warning is immediately triggered via GPIO or CAN bus, such as a buzzer or LED indicator, and key data (such as abnormal timestamps and feature vectors) is simultaneously transmitted to the cloud; Real-time tracking of CPU utilization, memory usage, and power consumption is achieved through device-side system monitoring tools (such as FreeRTOS's memory management module) to ensure that model operation does not exceed hardware limits (e.g., CPU utilization <70%, memory remaining >20%). The model version is dynamically switched according to resource load (e.g., switching to a lighter "simplified" model under high load and switching to a "full" model under low load). Model parameters are dynamically adjusted via cloud commands (e.g., the forget gate threshold of LSTM and the activation function slope of CNN) to adapt to changes in device status (e.g., vibration feature shift caused by aging).

[0029] In one embodiment, the support for offline caching and local analysis during network interruptions includes the following steps: The device sends heartbeat packets periodically (e.g., a simple request to the cloud every 5 seconds) and waits for a response. If no response is received after a set threshold is exceeded, it is determined that the network is interrupted. The device obtains the link status (e.g., physical layer connection status, signal strength) in real time through the device's network interface (e.g., Wi-Fi, 5G module). If the link status is "disconnected" or the signal strength is consistently below the threshold, a network interruption warning is triggered. Using SLC or eMMC type flash memory chips, it has the characteristics of shock resistance and high temperature resistance, making it suitable for long-term storage in industrial environments; a fixed area is designated in the memory as a temporary cache, and an "overwrite" mechanism is adopted. When the cache is full, the new data automatically overwrites the oldest data to avoid memory overflow. Assign priority labels, such as "high", "medium", and "low", to cached data based on data type (e.g., device status, environmental monitoring) and urgency level (e.g., fault warning, general logs). Dynamically adjust the caching ratio of each priority data based on the device's remaining storage space and the duration of network interruption. For example, if the network interruption exceeds 1 hour, increase the caching ratio of high-priority data to 80%. During network outages, the device continuously collects data (such as vibration and temperature) through sensors and directly inputs it into the deployed TinyML model (such as an LSTM fault prediction model). The model uses the device's computing power (such as the DSP module of the MCU or the NPU of the SoC) to complete inference and output classification results (such as "normal", "warning", "fault") or regression values ​​(such as remaining lifetime prediction). If the inference result exceeds the preset threshold (e.g., vibration amplitude > 0.5g, temperature > 80℃), a local warning will be triggered immediately; abnormal events (e.g., timestamps, sensor data, inference results) will be stored in a structured format in the cache medium and marked as "needs to be synchronized to the cloud"; The device continuously monitors the network status, and automatically starts the data synchronization process when the network is detected to have recovered (such as receiving a response from the cloud or the link signal strength rising back above the threshold). Incremental synchronization: Prioritize synchronizing high-priority data: First, upload abnormal events and critical status data marked as "needs to be synchronized to the cloud", such as fault warnings and equipment downtime records; Batch synchronization of low-priority data: After completing the synchronization of high-priority data, batch upload low-priority data such as environmental monitoring data, such as temperature and humidity, and equipment runtime. Calculate a hash value (such as SHA-256) for the synchronized data and compare it with the hash value stored in the cloud to ensure that the data has not been tampered with or lost. If the network is interrupted again during the synchronization process, record the position of the uploaded data and continue the transmission from the breakpoint after the network is restored to avoid duplicate uploads.

[0030] In one embodiment, the construction of the incremental learning framework dynamically updates the device-side AI model parameters through online learning; combining historical and new data, and employing knowledge distillation or transfer learning techniques to adjust only some model parameters, includes the following steps: An incremental learning engine is deployed in the cloud to receive new data uploaded from the device and perform model updates. The device retains a lightweight copy of the historical model parameters to support fast local inference. The cloud database stores historical data uploaded from the device, such as device status and fault records for the past 3 months, and organizes the data by device ID and timestamp. New data uploaded from the device (such as sensor data from the last hour) first enters the cloud message queue and is processed by the incremental learning engine in batches (such as 100 data points per batch). When new data accumulates to a preset threshold or the device detects a significant change in data distribution, such as a shift in the mean of vibration characteristics exceeding 20%, the cloud-based incremental learning engine automatically initiates model updates. Parameter tuning: Knowledge distillation: The teacher model uses the full model trained with historical data as the source of knowledge; the student model is a lightweight model deployed on the device, with only the parameters of its last few fully connected layers (such as output layer weights) adjusted, while most of the convolutional layer parameters remain unchanged; knowledge transfer guides the training of the student model through soft labels (output probabilities of the teacher model), so that the adjusted parameters can adapt to new data while retaining historical knowledge; Transfer learning: Pre-trained model: A model pre-trained on a general device dataset to extract common features, such as edge detection and texture analysis.

[0031] Fine-tuning: Only the classification layer (such as the fully connected layer) at the top of the model is fine-tuned to adapt to the specific tasks of the inspection equipment, such as fault type classification, while the parameters of the bottom feature extraction layer remain unchanged; During model updates, L2 regularization is applied to key parameters (such as weights that performed well in historical tasks) to limit their adjustment range and prevent them from being over-modified by new data. After the historical model has been trained, the contribution of each parameter to the loss function is calculated (e.g., by approximation using the second derivative) to identify parameters that were important to the historical tasks. During incremental learning, additional penalties are applied to the adjustment of important parameters (e.g., by adding a loss function term) to make their adjustment range smaller than that of unimportant parameters. After the model is updated, offline testing is performed using retained historical data (such as device status data from the most recent month) to calculate metrics such as accuracy and recall, ensuring that the performance of the updated model is no lower than that of the previous version. The updated model (version B) and the old version (version A) are simultaneously deployed to some devices, and their performance in real-time inference (such as fault detection latency and false positive rate) is compared. The version with better performance is selected as the primary model. If the updated model continues to deteriorate in online performance (such as an accuracy drop of more than 5%), it is automatically rolled back to the previous version, and an alarm is triggered to notify maintenance personnel. The cloud only sends the parameter differences after model adjustment (such as the weight changes of the last few fully connected layers) to the device, instead of the full model, reducing the amount of data transmitted (e.g., from 10MB to 1KB); the device caches the parameter differences of multiple historical model versions (e.g., the three most recent versions) and selects the appropriate version to load based on network conditions; for example, when the network is interrupted, the latest parameter differences cached locally are used to update the model; when the network is restored, the latest version is synchronized to the cloud.

[0032] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method of an automatic patrol device in smart manufacturing, characterized in that, Comprise: Step one: Deploy lightweight edge computing nodes on the device end, combine 5G network slicing technology, real-time cleaning, compression and priority division of the original data generated by the patrol device; by dynamically adjusting the compression algorithm, make the key data through the 5G dedicated bandwidth low delay transmission to the cloud, while the non-key data is cached or filtered locally; Step two: Based on software defined network technology, develop network controller to monitor link bandwidth, delay and node load in real time, combine reinforcement learning algorithm to dynamically optimize data transmission path; Introduce "load balancing degree" index to quantify the load balancing between sub trees, avoid single path overload through local adjustment mechanism, improve the stability and real-time performance of data transmission; Step three: Deploy TinyML framework on the resource-constrained patrol device, integrate model compression and quantization technology, run lightweight AI model; Combine device sensor data, real-time analysis of device status and trigger early warning, support offline caching and local analysis when the network is interrupted; Step four: Build incremental learning framework, dynamically update device end AI model parameters through online learning; Combine historical data and new data, use knowledge distillation or transfer learning technology, only adjust part of the model parameters, avoid retraining the entire model; The framework adapts to changes in device state, improves the accuracy and real-time performance of anomaly detection, while reducing the consumption of computing resources; Step five: Feedback the real-time analysis results of the device end to the patrol system through the message queue; Combine rule engine or reinforcement learning algorithm to dynamically adjust the patrol strategy.

2. The method of claim 1, wherein, The deployment of lightweight edge computing nodes on the device end, combined with 5G network slicing technology, real-time cleaning, compression and priority division of the original data generated by the patrol device, comprises the following steps: Deploy low-power, high-performance edge computing modules on the patrol device or adjacent location, equipped with memory and storage space, support real-time data processing; Integrate lightweight operating system and edge computing framework for local data preprocessing, model inference and communication management functions; Cooperate with telecom operators to allocate dedicated 5G network slices for patrol devices, which have high bandwidth, low delay and isolation to ensure that critical data transmission is not affected by other traffic; Dynamically adjust slice resource allocation according to real-time data priority and network load to optimize network utilization; Identify and filter noise in sensor data through rule engine or lightweight machine learning model; Correct or mark missing values, out-of-range values or logical contradictions to improve data effectiveness; Dynamically select lossless or lossy compression algorithms according to data type and network status; Priority division: High priority: including device fault warning, security incidents, data that need immediate response, marked as "urgent" and assigned the highest transmission priority; Medium priority: including device running status, data that need real-time monitoring, marked as "important" and assigned the second highest priority; Low priority: including environmental monitoring data, data that can be delayed, marked as "ordinary" and assigned the lowest priority; For low-priority data or data during network congestion, use ring buffer or solid state disk for temporary storage to avoid data loss; dynamically adjust data transmission order according to priority and network slice resources; After edge node performs data cleaning, compression and priority division, only high-priority data and compressed medium and low-priority data are transmitted to the cloud; edge node reports device status to the cloud regularly, and the cloud adjusts the pre-processing rules of edge node according to global information.

3. The method of claim 1, wherein, The key data is transmitted to the cloud through 5G dedicated bandwidth with low delay by dynamically adjusting the compression algorithm, while the non-key data is cached or filtered locally, including the following steps: According to real-time network load, data type and priority, dynamically select compression algorithm, when network load is high, preferentially use lossy compression, sacrifice a small amount of non-key information to improve transmission efficiency; when network load is low, switch to lossless compression to preserve data integrity; real-time monitoring of network delay or bandwidth utilization by edge node automatically triggers compression algorithm switching; Assign independent 5G network slice to key data, the slice has high priority scheduling, low delay transmission and bandwidth guarantee, so it is not affected by other business traffic; use QUIC protocol to reduce handshake delay and packet retransmission time; Deploy solid state disk or high durability flash memory in edge node to store non-key data; use "first in, first out" or "least recently used" algorithm to manage cache space, when storage capacity reaches threshold, delete low-priority or overdue data with priority; For periodic data, only samples with change amplitude exceeding threshold are retained to reduce redundant transmission; for multi-sensor data, the most representative data is selected through correlation analysis to eliminate redundant information; Edge node synchronizes the currently used compression algorithm, cache state and filtering rules to the cloud regularly, and the cloud dynamically adjusts the compression strategy of edge node according to global network load and device maintenance plan; if abnormality is detected in non-key data cached locally, edge node automatically marks it as "potential key data" and transmits it to the cloud with priority.

4. The method of claim 1, wherein, Based on software-defined network technology, develop network controller to monitor link bandwidth, delay and node load in real time, and dynamically optimize data transmission path combined with reinforcement learning algorithm, including the following steps: Deploy SDN controller in data center or core network node as "brain" of network management; controller communicates with network devices through southbound interface for centralized control of network traffic; Real-time state monitoring: Link bandwidth: periodically collect real-time bandwidth utilization of each physical link through SNMP, NetFlow protocol; Delay measurement: use Ping, ICMP protocol to periodically detect end-to-end delay of link and record historical delay distribution; Node load: real-time obtain load state of network device through Telemetry technology to identify high-load nodes; The state space definition includes quantitative indicators of network state of each link bandwidth usage, real-time delay of each link, and load state of each node; the action space definition includes path switching, bandwidth adjustment, and node avoidance, which are path adjustment operations executable by the controller; the reward function design includes quantitative indicators of transmission success rate, delay reduction amount, and load balancing degree to evaluate the effect of actions; The SDN controller updates the network state every second and inputs the current state into the reinforcement learning model; the reinforcement learning model selects the action with the highest expected reward based on the current state and historical experience; the controller issues a flow table to the network device through the OpenFlow protocol to adjust the path; the model updates the Q table based on the actual reward after path execution to optimize future decisions.

5. The automatic patrolling device method in intelligent manufacturing according to claim 1, wherein, The introduction of the "load balancing degree" indicator quantitatively evaluates the load balancing among sub-trees, and avoids overloading a single path through a local adjustment mechanism, including the following steps: The load balancing degree is used to quantitatively evaluate the load balancing of each link and node within the same region or sub-tree, avoiding overloading a single path; Mathematical expression: wherein, is the bandwidth utilization of the ith link; is the average bandwidth utilization of all links within the sub-tree; n is the total number of links within the sub-tree; LBD has a value ranging from 0 to 1; According to the business requirements, set the LBD trigger threshold, which is dynamically adjusted in combination with historical data; the SDN controller calculates the LBD of each link within the sub-tree every second, and automatically triggers the local adjustment process when the LBD is detected to be lower than the threshold; Local adjustment: Sub-tree grafting: the load of a certain link within a sub-tree is too high, while the load of an adjacent sub-tree is relatively low; part of the data flow is migrated from the high-load link to the low-load link; Node migration: a certain network node within a sub-tree has a very high load and becomes a bottleneck; the data flow path is re-planned to bypass the high-load node; After the adjustment is executed, the SDN controller continuously monitors the LBD of each link and the node load within the sub-tree to evaluate the adjustment effect; if the LBD after adjustment is still lower than the threshold, the controller will start a new round of local adjustment until the balanced state is reached.

6. The automatic patrolling device method in intelligent manufacturing of claim 1, wherein, The TinyML framework is deployed on the resource-constrained patrol device end, integrating model compression and quantization technologies, and running lightweight AI models, including the following steps: According to the resource limitations of the patrol device end, a TinyML framework designed specifically for embedded devices is selected, and the framework parameters are adjusted according to the hardware configuration of the device end; Model compression: Pruning: remove redundant neurons or connections in the model to reduce the size of the model; Knowledge distillation: migrate the knowledge of a large "teacher model" to a small "student model", and guide the student model training through soft labels to balance accuracy and model size; Quantization: Post-training quantization: convert model parameters from 32-bit floating-point numbers to 8-bit integers to reduce memory usage and speed up inference; Mixed quantization: retain floating-point precision for sensitive layers in the model and use integer quantization for other layers, balancing accuracy and efficiency; According to the patrol task requirements, deploy targeted models, use lightweight LSTM or GRU models to process time series data including device vibration and temperature, and detect early faults in real time; use lightweight CNN models including MobileNetV2 or EfficientNet-Lite to process camera-captured images and identify surface defects or abnormal behaviors; Optimize the model, reduce the 3-layer structure of the standard LSTM to 1-2 layers to reduce the amount of calculation; reduce the number of convolutional layer filters of CNN from 64 / 128 to 16 / 32 to reduce the complexity of the model; reduce the image input resolution from 224x224 to 96x96 or lower to reduce the pixel-level calculation burden; Real-time reading of device-end sensor data through SPI, I2C or UART interface; data normalization, noise reduction and feature extraction within the TinyML framework; the model performs inference once every fixed period, outputting classification results or regression values; when the inference result exceeds the preset threshold, trigger local warning through GPIO or CAN bus and simultaneously transmit key data to the cloud; Real-time tracking of CPU occupancy, memory usage and power consumption through device-end system monitoring tools to ensure that the model runs within hardware limits; dynamically switch model versions according to resource load; dynamically adjust model parameters through cloud instructions to adapt to device state changes.

7. The automatic patrolling device method in intelligent manufacturing of claim 1, wherein, The offline caching and local analysis during network interruption in the support network include the following steps: The device end sends heartbeat packets at regular intervals and waits for a response; if no response is received beyond the set threshold, it is determined that the network is interrupted; real-time acquisition of link status through the device-end network interface, if the link status is "disconnected" or the signal strength continuously falls below the threshold, the network interruption warning is triggered; Use SLC or eMMC type flash memory chips, and divide a fixed area in the memory as temporary cache, use "overwrite" mechanism, when the cache is full, the new data automatically covers the oldest data; According to the data type and urgency, assign priority labels to the cached data; dynamically adjust the cache proportion of each priority data according to the remaining storage space of the device and the duration of network interruption; During network interruption, the device end continuously collects data through sensors and directly inputs into the deployed TinyML model; the model uses device-end computing power to complete inference and outputs classification results or regression values; If the inference result exceeds the preset threshold, local warning is triggered immediately; store the abnormal event in a structured format to the cache medium and mark it as "to be synchronized to the cloud"; The device end continuously monitors the network status and automatically starts the data synchronization process when network recovery is detected; first upload the abnormal events marked as "to be synchronized to the cloud" and key state data; after completing the synchronization of high-priority data, batch upload low-priority data such as environmental monitoring data; Calculate the hash value of the synchronized data and compare it with the hash value stored in the cloud to check if the data has been tampered with or lost; if network interruption occurs again during synchronization, record the location of the uploaded data and continue transmission from the breakpoint when the network recovers.

8. The automatic patrolling device method in intelligent manufacturing according to claim 1, wherein, The incremental learning framework is constructed to dynamically update the device-end AI model parameters through online learning; combine historical data with new data, use knowledge distillation or transfer learning technology to adjust only part of the model parameters, including the following steps: Incremental learning engine is deployed in the cloud, receiving new data uploaded by devices and performing model updates; devices keep lightweight copies of historical model parameters, supporting fast local inference; the cloud database stores historical data uploaded by devices, organized by device ID and timestamp; new data uploaded by devices is first put into the cloud message queue, and processed by the incremental learning engine in batches; When new data accumulates to a preset threshold or the device detects a significant change in data distribution, the cloud incremental learning engine automatically starts model updates; Parameter adjustment: Knowledge distillation: the teacher model is a full model trained on historical data, serving as the knowledge source; the student model is a lightweight model deployed on devices, with only the last few fully connected layers adjusted, while most of the convolutional layer parameters remain unchanged; knowledge transfer guides the student model training through soft labels, allowing the adjusted parameters to adapt to new data while retaining historical knowledge; Pre-trained models on general device datasets extract general features; only the top classification layers of the model are fine-tuned to adapt to the specific tasks of the inspection device, while the bottom feature extraction layer parameters remain unchanged; When updating the model, apply L2 regularization to key parameters to limit the adjustment range and prevent over-modification of these parameters due to new data; after the historical model is trained, calculate the contribution of each parameter to the loss function to identify important parameters for historical tasks; during incremental learning, impose an additional penalty on important parameters to make their adjustment range smaller than that of non-important parameters; After the model is updated, use the preserved historical data for offline testing to calculate indicators such as accuracy and recall, ensuring that the performance of the updated model is not lower than that of the previous version; deploy the updated model and the old version to some devices simultaneously to compare their performance in real-time inference, and select the better version as the main model; if the updated model continues to perform poorly online, automatically roll back to the previous version and trigger an alert to notify the maintenance personnel; The cloud only sends the parameter difference of the updated model to the device, and the device caches the parameter differences of multiple historical model versions, selecting the appropriate version for loading based on network conditions.

Citation Information

Cited By

  • Eddy current data compression transmission optimization method

    CN121309684A

  • Power operation data monitoring method and system based on cloud network

    CN121663813A