Online updating system and method for AI algorithm of industrial Internet of Things, and medium
By obtaining industrial IoT data, building equipment connection diagrams, identifying performance bottlenecks, performing equipment failure prediction and life analysis, building an AI task scheduling model, and performing online updates, the problem that traditional industrial IoT AI algorithms cannot adapt to equipment changes in a timely manner, realizing accurate fault prediction and equipment life evaluation, and improving the efficiency and stability of the production line.
Patent Information
- Application Number
- CN202510564721.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional industrial IoT AI algorithms cannot adapt to device changes and environmental dynamic changes in time, resulting in inaccurate equipment failure prediction, incomplete equipment life analysis, poor adaptability of AI task scheduling models, making it difficult to deal with emergencies.
By obtaining industrial Internet of Things data, performing device-level information feature extraction, building device connection diagrams, identifying performance bottlenecks, performing device failure prediction and lifespan analysis, generating task scheduling data, building an AI task scheduling model, and performing online updates.
It improves the real-time and flexibility of the system, accurately predicts equipment failures, accurately evaluates equipment life, improves the efficiency and stability of the production line, ensures the adaptability of complex scheduling requirements and timely adjustments to emergencies.
Smart Images

Figure CN120389956A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial Internet of Things, and particularly relates to an online update system, method and medium for an AI algorithm of industrial Internet of Things. Background Art
[0002] The data acquisition of traditional industrial Internet of Things AI algorithms and the extraction of device-level information rely on traditional static modes, resulting in the inability to adapt to device changes and dynamic environmental changes in a timely manner. Although the construction of the device connection diagram reflects the connection relationship between devices, when the network topology changes or the number of devices increases, traditional methods cannot quickly and accurately update the connection diagram, affecting the real-time performance and flexibility of the system. There are delays in the identification of performance bottlenecks and device fault prediction. Traditional methods usually rely on historical data for inference, ignoring the processing of real-time data and the timely response to emergencies, resulting in inaccurate or lagged device fault prediction. Device life analysis is usually based on simplified models, unable to comprehensively consider the behavior of devices under complex working conditions, and prone to underestimating or overestimating the actual life of devices. Although the construction of the AI task scheduling model can schedule according to historical task data, for real-time dynamic task changes and complex scheduling requirements, the adaptability of traditional methods is poor, and it is difficult to make intelligent adjustments for emergencies. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide an online update system, method and medium for an AI algorithm of industrial Internet of Things to solve at least one of the above technical problems.
[0004] To achieve the above object, an online update method for an AI algorithm of industrial Internet of Things includes the following steps:
[0005] Step S1: Obtain industrial Internet of Things data, extract device-level information features to obtain device-level information; construct a device connection diagram based on the device-level information;
[0006] Step S2: Identify performance bottlenecks in the device-level information to obtain performance bottleneck data; predict device faults based on the performance bottleneck data to obtain fault data; perform device life analysis according to the fault data and the performance bottleneck data to generate device life data;
[0007] Step S3: Identify abnormal connections in the device connection diagram according to the device life data to generate device abnormal connection data; perform industrial production task scheduling based on the device abnormal connection data to obtain task scheduling data; construct an AI task scheduling model based on the task scheduling data;
[0008] Step S4: Perform intelligent scheduling of industrial tasks on the industrial Internet of Things data according to the AI task scheduling model to obtain intelligent scheduling tasks, and perform online update of the AI algorithm on the AI task scheduling model and deploy it to the industrial Internet of Things device side.
[0009] Through the real-time acquisition of industrial Internet of Things data and the extraction of device-level information, the present invention solves the problem that it is difficult to cope with device changes and environmental dynamic changes in the traditional static mode. By constructing a device connection graph, the system can quickly reflect the connection relationship between devices and update the connection graph in a timely manner when the number of devices or the network topology changes, thereby improving the real-time performance and flexibility of the system. The identification of performance bottlenecks by combining real-time data avoids the latency problem of traditional methods, enabling more accurate device fault prediction and timely response to emergencies. The device life analysis, with more comprehensive considerations, avoids the problem of underestimation or overestimation of device life under traditional simplified models and more accurately evaluates the actual service life of devices under complex working conditions. The AI task scheduling model improves the flexibility of task scheduling through the analysis and update of real-time dynamic data, can perform intelligent scheduling according to the actual changes of production tasks, ensure the adaptability to complex scheduling requirements and timely adjustment to emergencies, thus significantly improving the efficiency and stability of the production line. Brief Description of the Drawings
[0010] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:
[0011] Figure 1 It is a schematic flowchart of the steps of the online update method of the industrial Internet of Things AI algorithm of the present invention;
[0012] Figure 2 It is a schematic flowchart of the detailed steps of step S1 in the present invention;
[0013] Figure 3 It is a schematic flowchart of the detailed steps of step S3 in the present invention;
[0014] The realization of the object of the present invention, functional features, and advantages will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0015] The technical method of the present invention patent will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0016] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0017] It should be understood that although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0018] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides an online update method for an AI algorithm of an industrial Internet of Things, and the method includes the following steps:
[0019] Step S1: Obtain industrial Internet of Things data, perform feature extraction of device-level information to obtain device-level information; construct a device connection graph based on the device-level information;
[0020] In this embodiment, it is necessary to connect the sensor interfaces of each Internet of Things device. The interfaces of these devices adopt standard industrial communication protocols, such as Modbus, MQTT, and OPC-UA, which can ensure the reliability and real-time nature of data acquisition. The collected data mainly includes multiple dimensions such as the temperature, humidity, current, voltage, CPU load, and memory usage rate of the device. The data acquisition frequency for each dimension is set to once per minute to ensure the real-time monitoring of the device's working status. Specifically, the temperature and humidity sensors can read the environmental data of the device through the Modbus protocol, while the current and voltage sensors transmit the electrical characteristics of the power equipment through the OPC-UA protocol. Data such as the CPU load and memory usage rate are read by the edge computing device to obtain the internal performance information of the device. The data of all sensors will be preliminarily processed at the edge computing device. First, noise reduction processing is performed, and the commonly used methods are median filtering or mean filtering to eliminate abnormal fluctuations in the data. The processed data also needs to be normalized to convert sensor data of different magnitudes into a unified range (e.g., 0 - 1) to ensure the comparability of data. When extracting the information characteristics at the device level, the data will be classified according to fields such as the device ID, device category (such as pumps, fans, sensors, etc.), device location (e.g., the specific location in the workshop or production line), and device working status (such as normal, faulty, standby). Specifically, the device can be identified by its unique ID. The category field helps to distinguish different types of devices, the location field helps to understand the spatial distribution of the devices, and the working status field helps to judge the current working status of the device. The extracted device-level information will be stored in a structured format, such as a table or JSON format, which contains the relevant attributes of each device, such as: device ID, device category, device location, device status, etc. The value of each field will be directly written into the database during the data acquisition process to ensure the integrity and consistency of the information. A device connection diagram is constructed based on the extracted device-level information. First, it is necessary to determine the connection method between devices, which includes wired connections (such as Ethernet, RS-485, etc.) and wireless connections (such as Wi-Fi, ZigBee, etc.). The transmission rates of each connection method are different. The typical wired network transmission rate can be between 10 Mbps and 1000 Mbps, while the rate of wireless connections is usually lower. According to the connection type of the device, the corresponding transmission rate is determined. On this basis, a network topology algorithm (such as the Dijkstra shortest path algorithm) is used to draw the connection relationship diagram between devices. The Dijkstra algorithm will select the optimal path according to the connection weight (i.e., the transmission rate) between devices to ensure the maximization of data transmission efficiency. In the device connection diagram, each device will be used as a node, and the connection between devices is represented as an edge. The weight of the edge is the communication bandwidth or transmission rate between devices, with the unit of Mbps.The specific construction process of the device connection diagram includes identifying the direct connection relationships between devices, checking whether the devices can communicate with each other (through connection quality assessment), and determining the communication latency between devices. To ensure the integrity and accuracy of the connection diagram, it is necessary to confirm the actual connection relationships between devices one by one based on the device hierarchy information. The technical methods involved in this step include connection diagnosis and connectivity check. First, based on characteristics such as the physical location of the device connection and the connection protocol, it is judged whether each device can be directly connected. For wired-connected devices, the connection status of the physical line needs to be checked; for wireless devices, factors such as signal strength and interference need to be evaluated. With the support of the device hierarchy information, it is confirmed whether each device can communicate with other devices. In addition, it is also necessary to calculate the communication latency between devices according to the actual network bandwidth of the devices. Especially in a large-scale device network, the magnitude of the communication latency directly affects the real-time performance of the system. The calculation formula for latency is usually "latency = distance / network bandwidth", and a threshold is set according to the actual application requirements. If the latency exceeds a certain standard (such as 100 ms), then it is necessary to consider optimizing the device connection method or adjusting the device configuration.
[0021] Step S2: Identify the performance bottlenecks of the device hierarchy information to obtain performance bottleneck data; perform device fault prediction based on the performance bottleneck data to obtain fault data; perform device life analysis according to the fault data and the performance bottleneck data to generate device life data;
[0022] In this embodiment, it is necessary to extract the working state characteristics of the device level information, and particularly focus on key performance indicators, such as the CPU load, memory usage, disk I / O, network bandwidth, etc. of the device. These performance indicators reflect the operating efficiency and resource consumption of the device, and help to identify devices with performance bottlenecks. The working state data of the device has usually been collected in real time by sensors and preprocessed in step S1. At this stage, these key performance indicators are first extracted from the device's operation log, and preliminary screening is performed on each piece of data. Statistical methods (such as mean and variance analysis) can be used to evaluate the collected data. Specifically, calculate the mean and standard deviation of each performance indicator, and analyze its fluctuation. If the value of a certain indicator is much higher or lower than the mean, it indicates that the device is abnormal. Conduct a detailed monitoring and inspection of the performance of each device. First, detect the CPU load of the device. According to the usage requirements and working environment of the device, set the monitoring threshold for the CPU load. For example, if the CPU load of the device exceeds 90% continuously for 5 minutes, it is marked as having a high load. Similarly, if the memory usage exceeds 85%, and the response time of the device exceeds 200 ms, it is also regarded as a performance bottleneck. The response time refers to the processing time of the device for input requests. In the industrial Internet of Things, the response time is usually used to judge the real-time performance of the device. To ensure that these thresholds are reasonable, the CPU load and memory usage thresholds of the device can refer to the standards provided by the device manufacturer or be fine-tuned according to the device operating environment (such as temperature, load type, etc.). When the performance indicators of the device exceed these predetermined thresholds, the device is immediately determined to be a low-performance device. For these devices, further performance bottleneck simulation is carried out. During the simulation process, by setting the task load and running resource-intensive operations (such as data transmission, computationally intensive tasks, etc.), the real workload of the device is simulated. During the task execution, the system will continuously collect the response time data of the task, and these data can reflect the actual performance of the device under high load. At this time, the collected response time data needs to be accurately calculated, and special attention is paid to the execution delay of the task. If the simulation results show that the device response time exceeds 300 ms, and the CPU load of the device remains higher than 90% continuously for 10 minutes, then the device can be determined to be in a CPU overload state. Once low-performance devices are identified and performance bottleneck simulation is carried out, the corresponding data, such as task response time, device load, etc., can be collected and sorted out. These data will be further used for cross-device performance bottleneck analysis. Cross-device analysis is to perform correlation analysis on the performance data of different devices through multi-dimensional data integration technology, so as to find out the factors affecting the device performance. For example, if there is a high correlation between the memory usage and CPU load of a certain device, and the network bandwidth bottleneck appears under high load, this device needs to be optimized. Combining all performance bottleneck data can comprehensively identify the bottlenecks and generate a performance bottleneck data report.The report content includes the following aspects: First, the ID of the faulty device, which can uniquely identify the device with performance bottlenecks through the device ID; Second, the generated timestamp, which records the exact time when the performance bottleneck occurred, facilitating subsequent tracking and problem reproduction; Finally, the specific values of each performance indicator are listed in detail (such as CPU load 90%, memory usage rate 85%, response time 300ms, etc.), helping to analyze the root cause of device problems. Ultimately, the generated performance bottleneck data report will provide an important basis for subsequent device optimization and troubleshooting.
[0023] Step S3: Identify abnormal connections in the device connection diagram based on the device life data to generate device abnormal connection data; perform industrial production task scheduling based on the device abnormal connection data to obtain task scheduling data; construct an AI task scheduling model based on the task scheduling data;
[0024] In this embodiment, it is necessary to count the fault time based on the fault data and analyze the fault mode of the device. The fault data comes from various fault records during the operation of the device. Each time the device fails, the system will record the specific time of the fault occurrence and the fault type. The fault type can include communication faults, overheating, overload, etc. The fault data is usually stored in the form of timestamps and fault descriptions. By analyzing the time of multiple fault events, the device can be fault-classified to identify two types of data: intermittent faults and periodic faults. Intermittent faults usually manifest as occasional interruptions or abnormalities in the working state of the device, while periodic faults refer to faults that occur repeatedly within a specific period of the device. By classifying and counting the fault events, it can help analyze the impact of each type of fault on the device life in the subsequent analysis. For intermittent faults, first, analyze the communication interruption of the device. The communication interruption of the device is often due to problems in data transmission between devices, such as insufficient network bandwidth, transmission interference, etc. To accurately judge the occurrence of communication interruption, it is first necessary to analyze by comparing the loss situation of the device communication data packets and the response time of the device. If the loss rate of the device communication data packets exceeds 10%, and the response time exceeds 5 minutes, it can be judged that the device has a communication interruption. Specifically, the judgment criterion for communication interruption can be set as the device communication delay exceeding 5 minutes and the data loss rate exceeding 10%. During this analysis process, the device communication logs and transmission protocols (such as Modbus, MQTT, etc.) will provide the required communication packet loss and response time data. Once a communication interruption is detected, the system will record the fault detection delay and count the communication interruption duration of the device. If the communication interruption time of the device exceeds 5 minutes, it is necessary to record the fault detection delay of the device as the key data for subsequent life prediction. For periodic faults, especially those related to the cooling system of the device, first, it is necessary to evaluate the status of the device cooling system by combining the temperature sensor data inside the device. The role of the device cooling system is to ensure that the device maintains a stable temperature under high load. If the cooling system malfunctions, it can cause the internal temperature of the device to rise, thus affecting the normal operation of the device. The abnormal judgment criterion for the cooling system can be set as the internal temperature of the device exceeding 70°C and the start time delay of the cooling system exceeding 10 seconds. At this time, the temperature sensor will provide real-time temperature data, and the working state of the cooling system will be judged according to the response time when the cooling system starts. If the temperature exceeds the set threshold and the cooling system does not start in time, it can be determined that the cooling system has a fault. In addition, it is also necessary to use the temperature sensor to obtain the temperature data of the electronic components inside the device for overheating evaluation of the electronic components. If the temperature of the electronic components exceeds their working range (usually 65°C to 85°C), it will be judged that the components are overheated, affecting the long-term stable operation of the device.Through the above analysis of the equipment cooling system and overheating of electronic components, the failure types of the equipment are classified to obtain key indicators such as relevant temperature data, failure causes, and the status of the cooling system. These data will be input into the equipment life prediction model for comprehensive analysis. The equipment life prediction model can be trained using regression analysis based on historical data, statistical models, or machine learning algorithms (such as decision trees, support vector machines, etc.). The input data includes the failure time of the equipment, failure types, the workload of the equipment, the status of the cooling system, etc. Through model analysis, the predicted remaining life value of the equipment, failure type prediction, and relevant maintenance suggestions can be obtained. The output equipment life prediction data will include the following parts: the predicted equipment life, that is, the remaining time that the equipment is expected to operate normally under the existing working conditions; the prediction of the equipment failure type, for example, predicting the failure types that the equipment will occur in the future (such as communication failures, overheating, etc.); the system will also give corresponding maintenance suggestions, including when to perform preventive maintenance on the equipment, the parts that need to be replaced, and how to optimize the working state of the equipment. These data will provide a scientific basis for equipment management personnel to help them identify and handle potential problems in advance, thereby extending the service life of the equipment, reducing downtime, and improving production efficiency.
[0025] Step S4: Perform intelligent scheduling of industrial tasks on the industrial Internet of Things data according to the AI task scheduling model to obtain intelligent scheduling tasks, and perform online update of the AI algorithm on the AI task scheduling model and deploy it to the industrial Internet of Things device side.
[0026] In this embodiment, it is necessary to reasonably allocate devices according to device life data and task requirements. At this stage, first obtain the health status information of the devices, including device failure records, performance data, life prediction data, etc. The health status of the devices is one of the key factors determining the task allocation priority. To ensure the efficiency and reliability of task allocation, tasks are preferentially allocated to those devices with good health status. The evaluation criteria for health status can include the device's failure history, performance metrics (such as CPU load, memory usage, etc.), and the device's life prediction results. If a device is in a failure state or is expected to fail soon, the task priority of this device is reduced, and the tasks of this device are preferentially suspended, and the tasks are allocated to healthy devices to ensure the continuity of production tasks. Task priority sorting is the core link in task scheduling, and its criteria include the health status of the devices, the device idle time, the urgency of the tasks, and the difficulty of task execution, etc. Specifically, the evaluation result of the device's health status will directly affect task allocation. If a device is experiencing a failure or is about to fail, according to the device's failure data and life prediction results, the scheduling system adjusts the task priority of this device to the lowest level to ensure that the faulty device does not undertake too many tasks and prevent the failure from getting worse. On the contrary, for devices with good health status, high-priority tasks are preferentially allocated. The urgency and execution time of the tasks will also be used as the basis for sorting. More urgent tasks will be preferentially allocated to idle devices to ensure efficient production. The generation of task scheduling data depends on the optimization of scheduling algorithms. Common task scheduling algorithms include the shortest job first algorithm, genetic algorithm, etc. These algorithms calculate the optimal task scheduling scheme by inputting factors such as the device idle time, task execution time, and device load. For example, when using the shortest job first algorithm, the scheduling system will first select the task with the shortest execution time for scheduling to ensure the efficient use of resources. In the genetic algorithm, the scheduling system simulates the natural selection process, iteratively optimizes the task allocation strategy, and thus realizes the efficient scheduling of multiple devices and multiple tasks. After the AI task scheduling model is generated, it needs to be updated online in real time to adapt to the changing situations in the industrial Internet of Things environment. The goal of online update is to retrain the scheduling model through real-time data to improve the adaptability and accuracy of the model in practical applications. The updated parameters include device status weights, task priority factors, etc., and these parameters affect the decision-making process of the scheduling algorithm. For example, the adjustment of the device status weight will affect the degree to which the device health status affects the task allocation priority, and the adjustment of the task priority factor will directly affect the task sorting method. During the online update process, first obtain real-time data through the data feedback mechanism. These data include the real-time health status of the devices, the actual execution time of the tasks, device load and other information. Based on these real-time data, use optimization algorithms (such as gradient descent method, genetic algorithm, etc.) to dynamically adjust the scheduling algorithm to achieve the optimization of the scheduling strategy.After each online update, it is necessary to verify the effectiveness of the updated model. The purpose of verification is to ensure that the model can effectively perform task scheduling in a real-time environment after being updated, and to avoid the decline of scheduling effectiveness caused by parameter adjustment. The verification process usually includes backtesting of historical data and testing in a simulation environment. During backtesting, the updated model is applied to historical data to compare the deviation between the scheduling effectiveness and the actual results; simulation testing evaluates the performance of the model in different scenarios by simulating different tasks and device conditions. After ensuring that the model performs well during the verification process, it will be deployed to the industrial Internet of Things device side. The updated AI task scheduling model will be deployed to the industrial Internet of Things device side for real-time task scheduling. Through the connection with the central scheduling system, the device side can obtain task scheduling information in a timely manner and perform task allocation and scheduling in real time. On the device side, the AI scheduling model will dynamically adjust the task allocation strategy according to the real-time status of the device and the task requirements to ensure the efficient utilization of device resources and the smooth transition of task execution during the production process.
[0027] Optionally, step S1 is specifically as follows:
[0028] Step S11: Obtain industrial Internet of Things data, and perform feature extraction of device-level information to obtain device-level information;
[0029] In this embodiment, it is necessary to connect to the device through the Internet of Things gateway to ensure that the communication interface has been configured. The communication protocols of the device can include Modbus, MQTT, OPC-UA, etc., and the specific protocol selected depends on the device type and application scenario. The data acquisition period is set to once per minute, and the acquired data includes information such as the temperature, humidity, current, voltage, CPU load, memory usage rate, and device status of the device. After the data acquisition is completed, data preprocessing is performed to remove noise, and median filtering is used for denoising to ensure the accuracy and stability of the data. For the acquired device data, it is classified by fields such as device ID, device category, and device status. The extraction of device-level information includes extracting information such as device ID, device category, working status (such as running, standby, fault, etc.), device location, and device function, and storing it in JSON or CSV format. At this time, the obtained device-level information should include the detailed attributes of each device, such as the name, model, type, location, and device status of the device.
[0030] Step S12: Define the device connection relationship according to the device-level information, and set the maximum transmission rate of each connection relationship to be between 10 Mbps and 1000 Mbps to obtain the device connection relationship;
[0031] In this embodiment, the connection relationship of devices is constructed according to the type, function, and physical connection method of the devices. First, according to the installation location, usage, and working environment of the devices, the connection method between the devices is confirmed. The connection methods are divided into two types: wired and wireless. The specific connection method depends on the on-site network configuration and device interconnection method. For wired-connected devices, the transmission medium (such as Ethernet, optical fiber, etc.) is determined. For wireless devices, the network protocol (such as Wi-Fi, Bluetooth, Zigbee, etc.) is determined. The transmission rate of each device connection is defined by the network hardware support standard and the device interface specification, and the transmission rate range is limited between 10 Mbps and 1000 Mbps. According to the connection method and the maximum transmission rate of the devices, the communication link between the devices is established. For each connection relationship, the data transmission rate between the devices is specified and stored in the database. At this time, the obtained device connection relationship includes the physical connection method and the data transmission rate information between each device.
[0032] Step S13: Construct a device tree-like hierarchical structure based on the device hierarchical information, where the hierarchical structure is set to a maximum of five levels to obtain the device tree-like hierarchical structure;
[0033] In this embodiment, each node in the tree structure represents a device, and the upper-level node of each device represents its upper-level management device or control device. The superior-subordinate relationship between the devices is determined by sorting according to the device ID and device type. For example, sensor devices belong to higher-level monitoring and control devices, and the control devices are further managed by the central management system or cloud platform. The device hierarchical structure is divided into a maximum of five levels to ensure clear levels and facilitate management and operation. During the construction of the device hierarchical structure, the device roles at each level need to be defined. For example, the first level is the data acquisition layer, which includes all sensor devices; the second level is the control layer, which includes execution devices and control devices; the third level is the management layer, which includes devices responsible for scheduling and management, etc. Each device is connected to its upper-level and lower-level devices through the device hierarchical relationship to form a tree structure of the devices. At this time, the obtained device tree-like hierarchical structure is a tree structure with the device ID as the node and the device hierarchical relationship as the edge, and the hierarchical structure should meet the device management requirements of the actual scenario.
[0034] Step S14: Construct a device connection diagram according to the device tree-like hierarchical structure and the device connection relationship to obtain the device connection diagram.
[0035] In this embodiment, the device connection diagram is composed of device nodes and the connection edges between them. The nodes represent devices, and the edges represent the connection relationships between devices. First, according to the device tree - like hierarchical structure, the superior - subordinate relationships of each device are determined. The connection relationships between devices are connected according to the connection relationships defined in step S12, and the weights of the edges are set according to the transmission rates between devices, with the range between 10 Mbps and 1000 Mbps. The construction of the device connection diagram follows the following steps: First, determine the connection relationships between devices through the hierarchical structure of the devices; second, set the transmission rates of the connection edges according to the device connection relationships; third, calculate the optimal connection paths between devices according to factors such as device connection methods and device performance, and select the most suitable network topology for data transmission. Using the shortest - path algorithm in graph theory (such as Dijkstra's algorithm), the shortest data - transmission paths from a certain device to other devices can be calculated, and the data traffic between devices can be determined according to the transmission - rate limits. Finally, the obtained device connection diagram shows the device nodes and their connection relationships in a graphical way and can be displayed through visualization tools (such as Graphviz), clearly presenting the network connection relationships and transmission capabilities between devices.
[0036] Optionally, the performance - bottleneck identification includes:
[0037] Extract the device working - state characteristics from the device - level information to obtain the device working state;
[0038] In this embodiment, devices are connected through the Internet - of - Things gateway to ensure that the working - state data can be collected from the devices in real time. The data - collection process includes obtaining the state information of the devices, such as running, standby, and fault states, and performance metrics such as the device CPU load, memory usage rate, network - bandwidth usage, and disk I / O. The working - state data of each device includes information such as device ID, device type, running state, timestamp, temperature, current, memory, and CPU usage rate. This information is collected in real time through device - monitoring tools (such as Zabbix or Prometheus) and transmitted through protocols (such as SNMP, Modbus). For each device, the collected performance data is arranged according to the timestamp, and data normalization processing is performed using a time window. Finally, the working - state data of all devices is stored in a database (such as MySQL or InfluxDB) for subsequent analysis. Through these data, the real - time working state of each device can be obtained, providing a basis for subsequent performance evaluation.
[0039] Statistical analysis of device performance metrics based on the device working state;
[0040] In this embodiment, the real-time performance data of the device is regularly obtained through the device monitoring system (such as Zabbix), and data processing is carried out using data analysis tools (such as Pandas in Python). For the CPU load, the average CPU usage rate of the device within different time periods (such as every 5 minutes) is statistically calculated. The memory usage rate is also statistically calculated for 5 minutes. The disk I / O is statistically calculated based on the number of read and write operations per second, and the network bandwidth usage is statistically calculated as the total traffic per minute. In order to ensure the accuracy of the data, the sliding window method is used to perform time-weighted averaging on each performance metric, eliminate abnormal data, and calculate the performance status of the device. In addition, for the performance metrics of each device, reasonable thresholds are set (such as CPU load greater than 90%, memory usage rate greater than 85%, etc.) and marked for subsequent screening of low-performance devices.
[0041] Low-performance devices are screened based on device performance metrics. If the CPU load of the device exceeds 90% within 5 consecutive minutes, the device memory usage rate exceeds 85%, and the average response time of the device within 5 minutes exceeds 200 ms, it is determined as a low-performance device;
[0042] In this embodiment, the thresholds of performance metrics are set: the CPU load threshold is 90%, the memory usage rate threshold is 85%, and the response time threshold is 200 ms. For each device, its performance data within 5 consecutive minutes is checked, including CPU load, memory usage rate, and response time. The specific operation is as follows: the data of the CPU load, memory usage rate, and response time of the device in the past 5 minutes are obtained one by one and compared and analyzed with the set thresholds. If the CPU load of the device always exceeds 90% within 5 consecutive minutes, it is considered that the CPU load of the device is too high and marked as a high-load device; if the memory usage rate of the device exceeds 85% within these 5 minutes, it is considered that the device is in a memory bottleneck state and marked as a memory bottleneck device; if the average response time of the device within 5 minutes exceeds 200 ms, it is considered that the device responds slowly and marked as a device with slow response. If a device exceeds the set threshold in any one performance metric, the device can be marked as a low-performance device. All eligible low-performance devices will be recorded in a separate device list, and the recorded information includes the device ID, the type of performance bottleneck (such as high CPU load, memory bottleneck, or long response time), and the corresponding time period. This data set provides a basis for subsequent task simulation and performance bottleneck analysis, ensuring that the subsequent processing process can focus on devices with performance problems and further optimize or adjust resource allocation according to the performance of these devices.
[0043] Resource-intensive task simulation is carried out based on low-performance devices to obtain task simulation data;
[0044] In this embodiment, typical high-load tasks are selected, such as big data processing, file compression, image rendering, etc. These tasks can efficiently consume the computing resources of the device. During the simulation process, specialized testing tools such as Apache JMeter and Locust are used to execute the tasks. These tools are designed to simulate a large number of concurrent users or large-scale data processing and can accurately simulate the resource consumption status of the device when executing high-load tasks. Before the task simulation starts, the execution time of the task is first determined to be 10 minutes, and the task load is distributed to the low-performance device according to the set parameters through the testing tool. Whenever the task starts to execute, the testing tool records the timestamp when the task starts and continuously monitors key performance indicators such as the CPU load, memory usage, and task response time of the device. During the task execution process, the tool real-time collects the resource consumption data of the device, including the CPU load change per second, memory usage, and task response time. After the task execution is completed, the end timestamp is recorded again, and the overall response time of the task is calculated, that is, the time consumed by the task from start to completion. The entire simulation process is executed separately for each low-performance device, and various performance data during the task execution are output, including but not limited to CPU load, memory usage, task response time, etc. Through the results of these simulation tasks, the performance data of the device when performing resource-intensive tasks can be obtained, and these data will provide an important basis for subsequent performance bottleneck analysis and help identify the bottlenecks of the device under high-load tasks.
[0045] Calculate the response time from the task simulation data to obtain the response time;
[0046] In this embodiment, the execution data of each task is collected, including the start time and end time of the task. By calculating the difference between the start time and end time of the task, the execution duration of each task can be obtained, with the unit being seconds. Subsequently, for each low-performance device, the response time during the task execution is calculated. The response time not only includes the execution duration of each task, but also further analyzes the average response time, minimum response time, and maximum response time of the task. Specifically, when implementing, the NumPy library in Python is used to process the response time data to ensure the accuracy of the statistical calculation process. Obtaining the start time and end time of the task can be achieved by recording the system timestamp. The calculation method is: task response time = end time - start time. Then, the execution duration data of all tasks is aggregated, and the average response time, minimum response time, and maximum response time are calculated respectively. For the average response time, the mean() function in NumPy can be used for calculation. For the minimum and maximum response times, the min() and max() functions in NumPy are used for calculation respectively. The task response time data of all devices will be stored in a database, such as MySQL or InfluxDB, which are commonly used databases, for subsequent data analysis. For each low-performance device, it is checked whether its task response time exceeds the set threshold of 300 ms. If the response time of a certain device continuously exceeds 300 ms during the task execution, then there is a performance bottleneck in this device, especially when the CPU load is too high. If the response time lasts too long, further CPU overload judgment is required to enter the next analysis process.
[0047] Based on the response time, CPU overload judgment is carried out. It is set that if the CPU load of the device exceeds 90% within 10 minutes and the response time always exceeds 300 ms, then it is determined that the CPU is overloaded, and the CPU overload data is obtained;
[0048] In this embodiment, for each low-performance device, it is necessary to check its CPU load data within 10 consecutive minutes. The CPU load data of the device is usually collected by the device monitoring system or performance collection tools (such as Prometheus, Nagios, etc.) during real-time monitoring. The CPU load of the device is recorded every minute, and these data are stored in a database (such as MySQL, InfluxDB) for subsequent processing. For each low-performance device, when querying its CPU load data in the past 10 minutes, if the CPU load continuously exceeds 90% within these 10 minutes, then continue to perform the next operation. Next, analyze the response time data of the device during the same time period. This data comes from the records of the simulated task process in the previous step, and the response time of each task is recorded through task simulation tools (such as Apache JMeter, Locust). It is necessary to check whether the response time of the device continuously exceeds 300 ms throughout the 10 minutes. If the CPU load of the device remains above 90% all the time, and the response time always exceeds 300 ms throughout the time period, then it can be determined that the device has a CPU overload problem. If the CPU load of the device exceeds 90% multiple times within 10 minutes, and the response time always remains above 300 ms during the task execution, then the device will be marked as being in a CPU overload state. To ensure the accuracy of the judgment, all judgment results will be recorded, including the device ID, CPU load value, response time, and the specific time period when the overload occurred. These data will be stored in the database for subsequent cross-device performance bottleneck analysis. Finally, through this step, it is possible to accurately identify which low-performance devices have CPU overload problems and provide data support for performance optimization and bottleneck analysis.
[0049] Perform cross-device performance bottleneck analysis on low-performance devices based on CPU overload data to generate performance bottleneck data.
[0050] In this embodiment, a performance bottleneck correlation model between devices is established by combining the connection relationships and working state data of the devices. The device connection relationship data usually comes from the device connection graph constructed in step 2 through the device hierarchical information and device connection relationships. These data provide information such as the physical connections between devices, network bandwidth, and data flow directions. The working state data can include the operating modes of the devices, resource requirements for processing tasks, etc., and comes from device performance monitoring tools. Using these data, a performance bottleneck correlation model between devices is constructed, aiming to identify resource sharing or bandwidth bottlenecks existing between devices during task execution. The task execution situation of low-performance devices is evaluated through multi-dimensional analysis to analyze whether there is resource competition among the devices. Multi-dimensional analysis includes a detailed evaluation of the resource occupancy, communication latency, data transmission speed, etc. between devices. The analysis of resource occupancy can reveal whether a device overly relies on shared resources such as the CPU, memory, storage, etc. The analysis of communication latency and data transmission speed can reveal whether the network bandwidth between devices is sufficient to support data exchange between devices. Based on these analysis results, it can be preliminarily determined whether there is a cross-device performance bottleneck. To further accurately identify the performance bottleneck, statistical methods are used for correlation analysis. Common statistical methods such as chi-square test, regression analysis, etc. can be used to analyze the correlation between resource sharing and performance bottlenecks between devices. Through the chi-square test, it can be evaluated whether there is a significant statistical relationship between resource competition among different devices and the occurrence of performance bottlenecks; while regression analysis can help identify whether the change in device performance is related to the occupancy of specific resources or network bottlenecks. According to these analysis results, performance bottleneck data can be generated, including device IDs, the time when the performance bottleneck occurs, the related devices involved, and the performance data of the related devices, etc. These performance bottleneck data will be stored in a performance monitoring system (such as Prometheus or Elasticsearch, etc.) for subsequent query and analysis. Through cross-device performance bottleneck analysis, it can help accurately identify which devices have bottleneck problems during task execution, especially problems such as resource competition and insufficient bandwidth between devices. These analysis results provide data support and guiding basis for subsequent optimization of task scheduling and resource allocation, and can effectively improve the overall operating efficiency of the devices and avoid performance degradation or latency.
[0051] Optionally, the device fault prediction includes:
[0052] Monitoring the memory usage of the device based on the performance bottleneck data, where the memory usage is collected once per minute.
[0053] In this embodiment, the memory usage data is obtained through a device monitoring tool (such as vmstat for Linux or Performance Monitor for Windows). The memory usage of each device should be read regularly through the API provided by the system, specifically once per minute. The read memory usage includes key data such as the total current memory of the device, the used memory, the free memory, and the cached memory. Each data point will be stored in a database, such as InfluxDB, in the data format of a timestamp and the corresponding memory usage value. Each time of reading, ensure that the recorded memory usage data is not interfered by other system processes. To avoid abnormal fluctuations in the read values due to factors such as memory cleaning and cache recycling, the timestamp accuracy of each data point is set to the second level, so as to ensure the accuracy of the data under high-frequency monitoring. By storing these data, the subsequent analysis of the device memory usage trend and the occurring abnormal situations can be carried out.
[0054] Monitor the memory recovery amount of the device based on the performance bottleneck data, where the memory amount is collected every 30 seconds.
[0055] In this embodiment, it is necessary to clarify the definition of memory recovery, that is, the amount of memory released or recovered by the device system through the memory management mechanism (such as the free command for Linux or the MemoryCleaner tool for Windows) within a certain period of time. These data are obtained through the API of the device monitoring tool, and information such as the total recovered memory, the fragmented memory, and the recovery timestamp is recorded. The detection of the memory recovery amount is performed once every 30 seconds to ensure that the dynamic changes of the device memory can be reflected during data collection. Similar to the memory usage monitoring, each time data is collected, it is necessary to ensure that the system will not cause measurement errors due to the interference of other processes. Therefore, a low-priority process needs to be used to execute this collection task. The collected data also needs to be stored in the performance monitoring system database, along with the device ID, the recovered memory amount, and the corresponding timestamp information. By collecting these data frequently, a detailed memory recovery behavior record can be provided for subsequent memory leak detection.
[0056] Make a preliminary determination of memory leakage based on the device memory usage and the device memory recovery amount to obtain suspected memory leakage data.
[0057] In this embodiment, a threshold for memory leakage is set: When the memory usage of the device increases by more than 10% within 1 hour continuously and the amount of memory reclaimed does not change significantly, it is considered that memory leakage has occurred. By dynamically monitoring the memory usage and the amount of memory reclaimed of each device, combined with the basic principles of computer memory management, the differences in the amount of memory reclaimed and the change in memory usage collected every 30 seconds are compared. If the memory usage of the device increases significantly within 1 hour, but the amount of memory reclaimed does not increase synchronously, or the amount of memory reclaimed is too low, it is considered that the device has a risk of memory leakage. At this time, record this device as a suspected memory leakage device and save the relevant data, including the device ID, memory usage, amount of memory reclaimed, and timestamp. Further analysis of the suspected leakage device can be confirmed through the memory fragmentation detection in the next step.
[0058] Perform fragmentation detection on the suspected memory leakage data to generate memory fragmentation data;
[0059] In this embodiment, the memory management API provided by the operating system (such as / proc / meminfo of Linux or VirtualAlloc of Windows) is used to obtain memory fragmentation data. The memory fragmentation situation of each device needs to be collected in real time through the system scheduling management tool during each memory allocation and recovery process. The memory fragmentation data of the device includes memory blocks with allocation failures, memory areas not cleaned up, and memory areas with serious fragmentation. The specific numerical threshold can be set as: When the degree of memory fragmentation exceeds 10% of the total memory, it is considered that the fragmentation is serious. The records of memory fragmentation include data such as the device ID, fragmented memory amount, number of fragmented memory blocks, and their distribution. Through fragmentation detection, it can be further confirmed whether there is an actual memory fragmentation problem in the device suspected of memory leakage.
[0060] Perform memory fault clustering on the suspected memory leakage data according to the memory fragmentation data to generate fault data.
[0061] In this embodiment, a clustering algorithm for memory faults is set. For example, the K-means clustering algorithm is used to classify the memory fragmentation data to identify devices with similar memory fault characteristics. According to characteristics such as the degree of device memory fragmentation, the number of memory allocation failures, and the size of fragmented memory blocks, the K-means algorithm is used to divide the devices into several groups. Devices within each group have similar memory fault characteristics. Clustering parameters for each group are set, such as the proportion of fragmented memory and the number of memory allocation failures, and thresholds are set to ensure the accuracy of device classification. After clustering is completed, memory fault data is generated, including information such as device ID, clustering group number, fragmented memory size, and clustering analysis metrics. These fault data are stored in the fault management system for subsequent fault troubleshooting, repair, and optimization. Through memory fault clustering, it can help the operation and maintenance team quickly locate and solve the memory leakage and fragmentation problems of devices.
[0062] Optionally, the device life analysis includes:
[0063] Statistical fault time based on the fault data;
[0064] In this embodiment, a fault data source is obtained, including the fault records of each device. Each fault record includes a timestamp indicating the specific occurrence time of the fault. The time information of the fault occurrence is obtained through the built-in time synchronization module of the device to ensure that the recorded time format is standard and consistent. The common time formats are Unix timestamp (unit: second) or ISO 8601 format (e.g., 2025-02-25T14:30:00Z). These timestamp data are extracted from the device fault logs, and each fault event includes two fields: start time and end time. The system automatically collects the status information of the device every minute through a task scheduler (such as Cron or a system-level scheduling tool) to ensure the accuracy and real-time nature of the fault occurrence time. For each fault event, the system calculates the duration of the fault based on the recorded timestamp, specifically by calculating the difference between the fault end time and the start time to obtain the duration. For example, if the start time of a certain fault event is 14:00:00 and the end time is 14:10:00, the fault lasts for 10 minutes. The system sets a threshold. If the fault duration exceeds 10 seconds, it is considered an effective fault and is statistically counted. The start time, end time, and duration of all effective faults are recorded in the fault log database for subsequent analysis and processing of fault data. In addition, when storing data, the transaction management of the database should be followed to ensure that the records are not lost, not repeated, and are organized in chronological order for subsequent data query and analysis.
[0065] Classify the fault data based on the fault time to obtain intermittent fault data and periodic fault data;
[0066] In this embodiment, all fault records are sorted according to the timestamp when the faults occur. After sorting, the time interval between every two consecutive fault records is checked in sequence to calculate the time difference when the faults occur. If the time difference between the fault records is greater than a set threshold (such as 30 minutes), it is considered that there is a long interval between these two faults, belonging to intermittent faults; if the time difference between the fault records is less than 30 minutes, it is considered that these two faults occur periodically, belonging to periodic faults. Specifically, when implementing, first extract the start time of each fault from the fault data table, sort them in ascending order of time, and calculate the interval time according to the timestamps of two adjacent fault records. For each pair of adjacent fault records, determine whether the time interval between them exceeds 30 minutes. If it exceeds, classify this fault record into the intermittent fault category; if the time interval is less than 30 minutes, classify it into the periodic fault category. The threshold needs to be adjusted according to the actual situation of the device. For example, if the device fault mode has a high occurrence frequency, the threshold needs to be adjusted to 15 minutes or shorter. After completing the type classification, all intermittent fault data and periodic fault data are stored in different tables in the database respectively, which is convenient for subsequent analysis and fault mode recognition. In the database, each fault record includes fields such as fault ID, occurrence time, fault duration, device ID, fault type, etc., which is convenient for quick query and retrieval.
[0067] Perform device communication interruption analysis based on the intermittent fault data to obtain device communication interruption data;
[0068] In this embodiment, obtain the network status data of the device from the device network monitoring system, including information such as the communication connection status, transmission rate, and network latency of the device. By using network monitoring protocols (such as SNMP, Modbus, or other custom device communication protocols), regularly check the communication status between devices. The specific operation is to judge whether there is a communication interruption between the device and other devices through a periodic task (such as checking the network status of the device every 5 seconds). At each check, record the communication status data of the device. If the communication interruption between the device and other devices exceeds 5 seconds, it is considered that a communication interruption event has occurred on this device. The start time of the interruption is recorded as the network check time. When the device resumes normal communication, record the end time of the interruption and calculate the duration of the interruption. If the interruption duration exceeds 5 minutes, this event is marked as a communication interruption fault and belongs to a part of the intermittent fault data. All recorded communication interruption fault data, including information such as device ID, start time of communication interruption, duration of communication interruption, and end time of communication interruption, will be stored in a dedicated database table for subsequent query and analysis. During the storage process, the system will format the interruption data to ensure the unity of the data and the efficiency of query. The storage table includes fields such as: device ID, event type (communication interruption), start time of interruption, end time of interruption, duration, etc.
[0069] Based on the device communication interruption data, perform fault detection delay statistics, where the set communication interruption threshold is greater than 5 minutes to obtain the fault detection delay data;
[0070] In this embodiment, extract the information of all communication interruption events from the device communication interruption data, and focus on the events with an interruption duration exceeding 5 minutes. For each communication interruption record that meets the conditions, first record the start timestamp of the interruption. Next, the system will detect the fault status of the device through the fault detection mechanism. The fault detection system will monitor the communication status and performance indicators of the device according to the preset rules to determine whether the device is already in a fault state. The time point of fault detection will also be recorded. Then, calculate the time difference between the start time of the communication interruption and the time point when the fault detection system confirms the fault, which is the fault detection delay time. According to the set threshold (such as 5 minutes), if the time difference of fault detection is greater than 5 minutes, it is considered that the fault detection of the device is delayed, and this event is recorded as a fault detection event with a long delay. At this time, the fault detection delay data includes information such as device ID, delay time, and communication interruption duration, and will be stored in the fault detection delay statistics data table. This data table will contain relevant fields, such as device ID, event type (fault detection delay), fault detection delay time, and communication interruption duration. All delay time data is sorted by time and saved in the database to ensure easy subsequent analysis and query.
[0071] Based on the periodic fault data, perform abnormal analysis of the device cooling system to obtain the abnormal data of the device cooling system;
[0072] In this embodiment, devices at risk of cooling system anomalies are screened from periodic fault data, focusing on those experiencing frequent failures. The cooling system temperatures of these devices are then monitored, with real-time data acquired through built-in temperature sensors. The temperature sensors regularly collect temperature information, typically once every minute, to ensure accurate recording of the cooling system's temperature status. A reasonable temperature range is set as the normal range for each device and its operating environment. For example, if a device's cooling system temperature should be maintained between 40°C and 60°C, exceeding or falling below this range is considered an anomaly. Specifically, if the temperature remains above 60°C or below 40°C for more than 10 minutes, the device's cooling system is considered to have an anomaly. Temperature anomalies are recorded, including information such as the device ID, temperature value, the start time of the anomaly, and the duration of the anomaly. All of this cooling system anomaly data is stored in a cooling system anomaly data table in the database. This table contains fields such as the device ID, temperature data, the time and duration of the anomaly, ensuring that the cooling system status of each device can be tracked and queried. This method effectively identifies devices with cooling system anomalies and provides data support for subsequent fault analysis.
[0073] Evaluate electronic component overheating based on abnormal data from the equipment cooling system to obtain electronic component overheating data;
[0074] In this embodiment, by analyzing device cooling system anomaly data, we determine which devices' cooling systems have experienced temperature deviations from the normal range. At this point, the device's embedded temperature sensor has already provided the device's cooling system temperature data. When the device's cooling system temperature exceeds the set normal range (for example, exceeding 60°C or falling below 40°C), it is considered a cooling system anomaly, and the next step, overheating assessment, is initiated. To assess the overheating risk of electronic components, a thermal model within the device is used. This model combines the device's cooling system temperature, internal thermal conductivity characteristics, and power consumption data to estimate the actual operating temperature of key electronic components within the device (such as the CPU, memory, hard drive, and power module). An overheating threshold is set for each component; typically, the maximum operating temperature for the CPU and memory is 85°C, and the maximum operating temperature for the power module is 90°C. Once the device's cooling system temperature is abnormal, the thermal model calculates the temperature of the core electronic component. If this temperature exceeds the component's set threshold, the component is considered to be at risk of overheating. The overheating assessment results include information such as the device ID, component type, estimated temperature, and the time of overheating. All overheating data is stored in an electronic component overheating data table. This table facilitates subsequent monitoring and maintenance of equipment at risk of overheating and provides a basis for fault prevention.
[0075] Construct a device life prediction model based on the overheating data of electronic components and the fault detection delay data;
[0076] In this embodiment, the overheating data of electronic components and the fault detection delay data are integrated as the input for constructing the device life prediction model. The overheating data of electronic components includes information such as the overheating temperature, overheating duration, and overheating frequency of the key components (such as CPU, memory, power module, etc.) of each device; while the fault detection delay data contains the time difference from the start of device communication interruption to the detection of the fault. These data provide the abnormal conditions that occur during the operation of the device and their impacts. Next, through the regression analysis method, identify the relationships between the degree of device overheating, the fault detection delay, and the device life. First, organize the historical data of each device to ensure the integrity and accuracy of the data, associate the overheating data and the fault detection delay data with the actual life of the device, and find the correlations between them. Then, based on the results of statistical analysis, construct a device life prediction model. Specifically, use the regression model to fit the device life, set the overheating severity and the fault detection delay as independent variables, and the device life as the dependent variable, calculate the regression coefficients, and obtain the influence degree of each variable on the device life. To further improve the accuracy of the model, the weighted summation method can be used to weight each influencing factor to ensure that the contributions of key factors to life prediction are fully reflected. In addition, if conditions permit, machine learning models such as random forests and support vector machines can be introduced, and the historical data is used to train the model to automatically learn the complex relationships in the data and further optimize the accuracy of device life prediction. Finally, the prediction model will output the estimated remaining life of each device, providing a basis for the maintenance and replacement plans of the device according to this data.
[0077] Input the performance bottleneck data into the device life prediction model, and perform device life prediction to obtain device life data.
[0078] In this embodiment, obtaining performance bottleneck data includes information such as load, response time, memory usage, and CPU load encountered by the device during operation. These data are collected in real time through the device's monitoring system and reflect the performance of the device under high load and heavy operating pressure. The performance bottleneck data can reveal the impact of the device's workload on its lifespan. Especially when the device's workload approaches or exceeds its maximum capacity, it will accelerate the occurrence of failures, thereby shortening the device's service life. Combining these performance bottleneck data with the device overheating data and fault detection delay data obtained in the previous steps as the input of the device lifespan prediction model to re-predict the device lifespan. To ensure the accuracy of the prediction results, in the input data, it is necessary to standardize the performance bottleneck data so that its unit and scale are consistent with the overheating data and fault detection delay data. Next, by adjusting each parameter and weight in the device lifespan prediction model, the model can more accurately evaluate the lifespan according to the specific conditions of different devices. For example, if a device has a continuously high CPU load, its memory usage is close to saturation, and overheating occurs simultaneously, the model will give higher weights to these factors, thus obtaining a prediction result that the device has a shorter lifespan. Finally, combining the performance bottleneck data of the device, the predicted remaining lifespan of each device is output, and the results include information such as device ID, estimated lifespan, and the degree of impact of performance bottlenecks on the lifespan. All data will be stored in the device lifespan prediction data table. This information will provide a basis for device maintenance and optimization, helping operation and maintenance personnel determine when to perform device maintenance, replacement, or adjustment work.
[0079] Optionally, step S3 is specifically as follows:
[0080] Step S31: Perform device maintenance prediction based on device lifespan data to obtain device maintenance data;
[0081] In this embodiment, the remaining lifespan data of each device is obtained from the device lifespan prediction system. These data have been analyzed and predicted based on multi-dimensional factors such as device overheating, fault detection delay, and performance bottlenecks. For each device, a predetermined maintenance cycle threshold is set based on its remaining lifespan, historical fault records, and device type. For some high-load or high-failure-rate devices, their maintenance cycles can be shortened. According to the set maintenance cycle threshold, for example, when the device lifespan is less than 30% of the remaining lifespan, mark the device as entering the maintenance prediction cycle and formulate a device maintenance plan. To this end, it is dynamically adjusted according to the device's remaining lifespan and the device's working environment and production line usage. For example, if the actual working time of the device reaches the set usage threshold (such as the device runs for 2000 hours), even if the remaining lifespan is long, preventive maintenance should be carried out. All maintenance prediction data will include device ID, remaining lifespan, estimated maintenance time, device working environment parameters, etc., and are stored in the device maintenance data table to provide a basis for subsequent device maintenance scheduling.
[0082] Particularly importantly, step S31 includes the following steps:
[0083] Step S311: Perform CPU operating load analysis based on device life data to obtain CPU operating load data;
[0084] In this embodiment, the load data during CPU operation is collected through a device monitoring system, and the monitoring system records the CPU usage rate every second. The CPU usage rate refers to the ratio of the computing workload of the CPU within a certain period of time to the total capacity of the CPU, usually expressed as a percentage. To obtain accurate data, the CPU load value is collected regularly (once a minute), and the recorded data should include the maximum load, minimum load, and average load within each period. According to the actual device operation situation, a threshold is set. When the CPU load exceeds 80%, the load is recorded as a high load state, and when it is lower than 20%, it is recorded as a low load state. By analyzing these data, the high-load operation time periods with continuous load exceeding 90% are extracted, providing a data basis for subsequent performance degradation analysis.
[0085] Step S312: Perform performance degradation analysis based on the CPU operating load data to obtain CPU degradation data;
[0086] In this embodiment, through the collected CPU operating load data, combined with the historical usage of the device (such as device working hours, operating cycles, etc.), performance degradation analysis is performed. In this process, first, the load curve of the CPU is fitted to analyze its load fluctuation trend. Then, considering factors such as the temperature and humidity of the device working environment and the degree of device aging, a degradation function is used to model the performance of the CPU. It is assumed that when the CPU load continuously exceeds 90%, the performance degradation rate of the CPU will accelerate. By long-term monitoring of the CPU operation under high load, the degradation rate is calculated. The specific operation is to analyze the monthly load data to obtain the percentage of CPU performance degradation. If the load exceeds the standard continuously within three months and the degradation rate exceeds the set threshold (such as more than 5% degradation per month), it is considered that the CPU begins to show performance degradation, forming CPU degradation data.
[0087] Step S313: Perform temperature monitoring based on the CPU degradation data to obtain the CPU temperature;
[0088] In this embodiment, the temperature monitoring system needs to regularly collect CPU temperature data and use sensors to detect the surface temperature of the CPU in real time. Each monitoring data needs to record the temperature change accurate to each second. The temperature data includes instantaneous temperature, maximum temperature, minimum temperature and average temperature. According to the operating load data of the device, the temperature of the CPU usually rises under high load. A threshold is set. When the CPU temperature exceeds 85 °C, it is regarded as an abnormal temperature and marked as a high-temperature state. The temperature data should be analyzed under different working loads to ensure that the temperature fluctuation under high load matches the load. If it is found during the degradation analysis that the CPU load is too high and the temperature continuously exceeds 85 °C (for example, for more than 15 minutes), then record this period as an overheating state and analyze possible heat dissipation problems in combination with historical temperature data.
[0089] Step S314: Identify the heat dissipation system failure according to the CPU temperature to obtain the heat dissipation system failure data;
[0090] In this embodiment, by analyzing the relationship between the CPU temperature fluctuation and the device operating load, it is determined whether the temperature exceeds the normal working range of the heat dissipation system. The common manifestation of the heat dissipation system failure is that when the CPU temperature remains normal under low load, but when the load increases, the temperature rises sharply and fails to cool down in time. Combining the CPU temperature data, it is set that if the temperature continuously exceeds the set threshold (such as 90 °C) under high load, and the rotation speed of the cooling fan does not increase significantly, or the temperature of the heat dissipation component (such as the heat sink) is too high, it is directly judged that there is a failure in the heat dissipation system. The determination criteria for the heat dissipation failure data include the temperature fluctuation range, whether the temperature drops stably during the device operating load, the fan operation status, etc., and mark the abnormal heat dissipation system components (such as fans, heat sinks, heat pipes, etc.) and their corresponding failures.
[0091] Step S315: Perform device repair based on the heat dissipation system failure data to obtain the device repair data.
[0092] In this embodiment, the specific steps of the repair include disassembling the heat dissipation components of the device and checking whether there is dust accumulation, failure, aging or damage in the hardware components such as fans, heat sinks, heat pipes, and thermal paste. For the fan, check its rotation speed and whether it can rotate normally; for the heat sink, check whether it is completely attached to the surface of the CPU to ensure effective heat dissipation; for the heat pipe, check whether there is leakage or breakage. All the detection results should be recorded in the device repair data. The repair data includes detailed information such as repair items (such as cleaning the radiator, replacing the fan, re-applying thermal paste, etc.), repair time, repaired parts, specifications of the replaced parts, and performance test results after repair. All the repair data will be summarized and archived and used to analyze the repair history of the device, providing a basis for subsequent device operation and maintenance decisions.
[0093] Step S32: Based on the equipment maintenance data, identify abnormal connections in the equipment connection diagram and generate equipment abnormal connection data;
[0094] In this embodiment, through the equipment maintenance records, identify the equipment that has failed or is in a near-failure state, and according to the network connection topology of the equipment, extract the connection relationships between the equipment. The equipment connection diagram is constructed based on the equipment ID and its corresponding connection status. The nodes of the connection diagram represent the equipment, and the edges represent the communication or transmission links between the equipment. For each piece of equipment that has failed in the maintenance records, based on the equipment maintenance history, judge whether the connection status of the equipment is affected, especially whether the physical connection or data communication path with other equipment is stable. For example, if the maintenance record of equipment X indicates that there is a problem with its power supply, then it affects its connection with equipment Y. The system automatically identifies this potential connection anomaly and generates equipment abnormal connection data. Each piece of abnormal connection data includes the equipment ID where the abnormal connection occurs, the connected equipment ID involved, the connection type (such as power connection, data transmission connection, etc.), the type of anomaly, and the occurrence time, etc., and is recorded in the equipment abnormal connection data table for subsequent processing.
[0095] Step S33: Based on the equipment abnormal connection data, conduct statistics on abnormal equipment tasks to obtain abnormal equipment tasks;
[0096] In this embodiment, extract all the equipment with connection problems from the equipment abnormal connection data. Then, according to the task arrangement of the equipment and the production line work process, analyze which tasks are associated with these abnormal equipment. Obtain the current task status of the equipment through the scheduling system, including information such as task type, execution time, and the equipment ID assigned to the task. If a device on which a task depends has a connection anomaly or failure, the system will automatically mark the task as an abnormal task. For example, if equipment A is responsible for performing a welding task and there is a problem with the connection of equipment A, then the welding task is marked as an abnormal task. The system counts all the tasks marked as abnormal, records the detailed information of the abnormal tasks, including task ID, task name, related equipment ID, task start and end times, reasons for anomalies, etc., and stores them in the abnormal equipment task data table for further scheduling and adjustment.
[0097] Step S34: Schedule industrial production tasks according to the abnormal equipment tasks to obtain task scheduling data;
[0098] In this embodiment, all tasks marked as abnormal are extracted from the abnormal device task data table. These tasks cannot be carried out as scheduled due to equipment failures or connection anomalies. The system will perform scheduling based on multiple dimensions such as the urgency of the tasks, the impact of the tasks on the production line, and the resource requirements of the tasks. For each abnormal task, the system will evaluate the availability and capabilities of alternative devices. For example, according to the load, operating status, available resources, etc. of the devices, suitable alternative devices are selected to execute the tasks. During the scheduling process, the system rearranges the tasks according to the priorities of the tasks and the overall production line plan. For example, if the welding task is affected, the system will arrange other idle devices and adjust their task order to ensure the production progress. Finally, through the task scheduling algorithm, a new task scheduling plan is generated. The task scheduling data will include task IDs, adjusted task times, device IDs, alternative device IDs, task statuses, etc., and are stored in the task scheduling data table to provide a basis for subsequent production operations.
[0099] Step S35: Construct an AI task scheduling model based on the task scheduling data.
[0100] In this embodiment, relevant information such as historical task scheduling data, device status data, and production progress is collected. These data usually include the type of tasks, execution order, required resources, load status of devices, maintenance records, task execution time, and production bottlenecks. The preprocessing of the data includes steps such as cleaning missing values, removing abnormal data, and normalization to ensure the high quality and availability of the data. After data cleaning, feature engineering is carried out to extract key features, such as the failure records of devices, the priority of each task, the expected execution time of tasks, and the health status of devices. These features will provide support in the subsequent model training process. Next, the preprocessed data is input into the AI model for training. The core goal of the AI model is to find the dependency relationships and constraint conditions between tasks by learning the task scheduling patterns, device load conditions, and production progress in historical data, and then optimize the execution order of tasks and the allocation of devices. The optimization process usually uses heuristic-based optimization algorithms, such as genetic algorithms (GA) or particle swarm optimization (PSO). These algorithms can efficiently search the solution space of task scheduling and find the optimal or near-optimal task scheduling scheme. In the genetic algorithm, each scheme of task scheduling is regarded as an individual, and each individual will be optimized through operations such as selection, crossover, and mutation. In particle swarm optimization, each scheduling scheme is regarded as a particle, and the optimal solution is searched by moving in the solution space. Through these algorithms, the model can reasonably allocate between tasks and devices, maximizing resource utilization and production efficiency. To improve the accuracy of task scheduling, the model uses the backpropagation algorithm for update. After each round of task scheduling is completed, the system will adjust the parameters of the model according to the deviation between the actual result and the expected target, so that the prediction and scheduling strategy of the model are gradually optimized. This process is an iterative process. After each task scheduling, the model will evaluate the scheduling effect according to the actual scheduling effect, calculate key indicators such as task completion time, resource utilization rate, and device load, and continuously update the parameters of the model according to these evaluation results. Finally, through this continuous optimization process, the AI model can adaptively adjust the task scheduling strategy in different production environments, can adjust the load of devices, the priority of tasks, the allocation of resources, etc. according to real-time data, so as to generate the optimal task scheduling scheme. This adaptive adjustment mechanism enables the model to cope with complex and changeable situations in actual production, such as device failures and production delays, and ensures the smooth operation of the production process.
[0101] Optionally, step S32 is specifically as follows:
[0102] Step S321: When the following situations occur simultaneously, it is determined that there is an abnormal device connection interruption, and abnormal device connection interruption data is obtained: frequent device failures appear in the device maintenance data, the device communication packet loss rate exceeds the set threshold of 10%, the device connection interruption exceeds the set number of times, and the device maintenance time exceeds the set threshold of 30 minutes;
[0103] In this embodiment, the equipment maintenance data records the frequency of equipment failures. Assume that the set threshold is that the number of failures per month exceeds 3 times. When the failure frequency in the equipment maintenance record exceeds this threshold, it indicates that the equipment has a tendency of frequent failures and further failure analysis is required. At this time, the equipment is determined to have the characteristic of frequent failures and enters the abnormal monitoring range. The communication packet loss rate of the equipment is also included in the determination conditions. When the communication packet loss rate of the equipment exceeds the set threshold of 10%, it indicates that there are major problems in the data transmission process of the equipment, resulting in connection interruptions or data transmission errors. The monitoring of the packet loss rate is through real-time data collection at the network layer, and network monitoring tools (such as SNMP monitoring tools, network traffic analysis tools, etc.) are used to collect the real-time packet loss situation of the equipment communication. If the packet loss rate continues to be higher than the set threshold, it is marked as unstable connection quality. The number of equipment connection interruptions is another important parameter for determining connection interruption abnormalities. Set a maximum number threshold (for example, 3 times). If the equipment has more than 3 connection interruptions within a specified time period (such as one day, three days, etc.), it is considered that there is a problem with the connection of the equipment. The record of equipment connection interruptions is usually provided by the communication module or monitoring system of the equipment. By monitoring the connection status of the equipment and recording the interruption events, the cumulative occurrence times are calculated. The equipment maintenance time exceeding the set threshold of 30 minutes is also used as the determination criterion for connection interruption abnormalities. If the equipment maintenance time exceeds 30 minutes, it indicates that the equipment repair process is slow, affecting production efficiency or equipment reliability. These data can be obtained through the equipment management system or the maintenance record system. Record the start and end times of the maintenance, calculate the maintenance duration, and compare it with the set threshold of 30 minutes. When the equipment maintenance data shows frequent failures, the communication packet loss rate exceeds 10%, the number of connection interruptions exceeds 3 times, and the maintenance time exceeds 30 minutes, when all these conditions are met, it is determined that the equipment has a connection interruption abnormality. Then, record relevant information such as the equipment ID, failure type, communication packet loss rate, number of connection interruptions, and maintenance duration, and store these data in the equipment connection interruption abnormality data table for subsequent data analysis and equipment management use.
[0104] Step S322: When the following situations occur simultaneously, it is determined that the equipment connection quality is abnormal and equipment connection quality abnormal data is obtained: the equipment maintenance data shows that the equipment hardware replacement frequency is relatively high, the equipment connection delay exceeds the set threshold by 50%, the equipment connection recovery time after maintenance is extended beyond the set threshold, and the data transmission failure between equipment exceeds the set number of times;
[0105] In this embodiment, the hardware replacement frequency is one of the key parameters. If the hardware replacement frequency is high in the device's maintenance record, it indicates that there are relatively serious fault problems with the device's hardware. A threshold is set (for example, the number of hardware replacements per month exceeds 2 times) to identify frequent hardware replacement situations. If the number of hardware replacements in the device maintenance record exceeds the set threshold, it indicates that there are repeated fault problems with the device's hardware, resulting in a decline in the device connection quality. In actual operation, the hardware replacement record is usually obtained through the device maintenance system and corresponds to specific maintenance items through the device ID. The device connection delay is an important indicator for evaluating the device connection quality. The device connection delay threshold is set to 50 ms. When the connection delay of the device exceeds this threshold, it indicates that there are significant delay problems in the communication between the device and other devices or the network, affecting the overall performance of the system. The delay monitoring can use network testing tools (such as Ping test, Traceroute, etc.) to detect the connection response time of the device in real time and compare it with the set 50 ms threshold. If the connection delay continuously exceeds the standard, it is determined that the connection quality is abnormal. The connection recovery time after device maintenance is also an important factor in determining the device connection quality. If the time required to restore the connection after device maintenance exceeds the set threshold of 30 minutes, it indicates that the connection recovery process of the device is slow, which will affect the production efficiency. This data can be obtained through the device's maintenance record by recording the start and end times of the maintenance and determining whether it exceeds 30 minutes by calculating the connection recovery time after maintenance. If the recovery time is long, it means that the connection recovery efficiency of the device is low, resulting in too long downtime during the production process and affecting the production stability. The number of data transmission failures between devices is also closely related to the connection quality. A threshold is set (for example, the number of data transmission failures per week exceeds 5 times). When the number of data transmission failures that occur to the device within a certain time range (such as within a week) exceeds the set threshold, it is determined that the device has serious connection quality problems. The record of data transmission failures can be obtained through network monitoring tools or the device's data transmission log. Each time a data transmission failure event occurs, relevant information should be recorded and the failure times should be accumulated for analysis. Combining these parameters, if the device meets the set threshold under any one or more conditions, it is determined that the device connection quality is abnormal. Information such as device ID, hardware replacement frequency, connection delay, recovery time, and number of data transmission failures will be recorded and stored in the device connection quality abnormal data table for subsequent maintenance and fault analysis. These data provide a detailed basis for further analyzing and processing device connection quality problems, helping to timely discover and repair potential problems of the device.
[0106] Step S323: Integrate the device connection interruption abnormal data and the device connection quality abnormal data to obtain the device abnormal connection data.
[0107] In this embodiment, it is necessary to match the device connection interruption exception data and the device connection quality exception data through the device ID. During this process, the system will compare according to the device ID in each piece of data to ensure that the exception data of each device can be correctly associated. For example, if device A has a record in the device connection interruption exception data, the record with the device ID of A will be merged with the relevant records of device A in the device connection quality exception data. For the situation where a device simultaneously experiences device connection interruption exception and device connection quality exception within the same cycle, it is necessary to summarize this data to form a comprehensive exception record. The merged data will include the following information: the maintenance history record of the device, including the fault occurrence time, maintenance type, and hardware replacement frequency; the communication packet loss rate and connection delay of the device; the number of device connection interruptions and maintenance time, etc. In this way, the exception information in multiple dimensions can be merged together to provide more comprehensive data support for subsequent analysis. The integrated data will form comprehensive exception data including device connection interruption exception and device connection quality exception, and these data will be stored in a unified database. The record of each device will include information such as its exception type, exception occurrence frequency, and specific parameters of the fault. This information will support subsequent analysis, device maintenance decision-making, and production optimization to ensure that problems with device connection and quality can be identified and solved in a timely manner.
[0108] Optionally, step S34 is specifically as follows:
[0109] Step S341: Divide the priorities according to the exception device tasks, where the priorities are set from 1 to 10, to obtain the task priority data;
[0110] In this embodiment, the failure frequency of the device is one of the key factors for dividing priorities. Specifically, if the device has more than 3 failures in the past 30 days, the task priority will be set to levels 9 to 10, indicating that this task is very urgent and needs to be processed immediately. If the failure frequency of the device is low, that is, 1 to 2 failures occur in the past 30 days, the task priority will be set to levels 5 to 7, which are tasks of medium priority and need to be processed within a short time but do not need to be executed immediately. If the device has no failures in the past 30 days, the priority of the task is levels 4 to 6, indicating that the urgency of this task is relatively low. Secondly, the urgency of the task will also directly affect the task priority. For example, if the task involves the repair of the core functions of the device, such as host failures, production line shutdowns, etc., the priority of such tasks will be set to levels 8 to 10 because they are crucial for the normal operation of production. Routine equipment maintenance, inspections, or other non-urgent tasks will be set to levels 4 to 6, which are tasks of relatively low urgency and need to be processed regularly but do not require an immediate response. Finally, the load situation of the device is also an important basis for dividing priorities. If the current load of the device exceeds 70%, the priority of the task will be automatically increased to ensure that the device is repaired and maintained in a timely manner and to avoid device failures caused by high loads. If the device load is low (for example, less than 50%), the priority of the task can be set to a lower value because the device has high stability at this time and will not have a great impact on production. Considering the above factors comprehensively, through weighted sum calculation, the priority data of each task can be finally obtained, and the reasonable allocation and timely processing of resources can be ensured.
[0111] Step S342: Perform device task allocation based on the task priority data to obtain task allocation data;
[0112] In this embodiment, the priority data of tasks will be used as the basis for allocating device resources. Tasks with higher priorities will be allocated to devices first to ensure that urgent tasks can be completed in a timely manner. During the process of device task allocation, the available time of each device needs to be considered first. The available time range of the device is from 0 to 480 minutes, and this data is sourced from the device's historical operation records and real-time monitoring data to ensure that the system can understand the remaining available working time of each device. When the device load is high, the system will preferentially allocate high-priority tasks to devices with lower loads to prevent overloaded devices from failing to complete tasks in a timely manner and affecting production efficiency. Secondly, the execution time of tasks (ranging from 0 to 120 minutes) is another key factor. The execution time of tasks must match the execution capabilities of the devices to ensure that the devices can complete tasks within the available time. If the task execution time is long, the system will evaluate whether there are other idle devices available for allocation to ensure the rational use of resources. On this basis, the load and available time of the devices will be compared with the execution requirements of the tasks to ensure that the tasks can be completed on time. For example, if a certain task requires a long execution time and has a heavy load, the system will allocate it to a device with sufficient available time and a light load to ensure the smooth completion of the task. The evaluation of device capabilities is based on its specifications (such as processing power, working hours, etc.), operation time (working hour data of the device), and historical task data (task execution records of the device in the past). The system will select the most suitable device to execute the task based on these evaluation data. By integrating task priorities, device available time, load, and task execution requirements, task allocation data will finally be generated to clarify which tasks are allocated to which devices, ensuring that tasks are completed on time without overloading the devices.
[0113] Step S343: Perform production line resource scheduling according to the task allocation data to obtain resource scheduling data;
[0114] In this embodiment, according to the type and specific requirements of each task, the system will analyze the resources required for each task, such as equipment, tools, and operators. For example, complex tasks may require more equipment and operators, while simpler tasks have relatively fewer resource requirements. The number of equipment resources will be reasonably allocated according to the execution complexity of the task and the expected execution time, usually between 1 and 10 pieces of equipment. When scheduling resources, factors such as the execution time of the task, the types of equipment required, and the available status of the equipment will be taken into account. For the allocation of operators, the system will decide according to the complexity of the task, usually 1 to 5 operators. The number of operators is calculated based on the actual requirements of the task and the work process to ensure that each task can be successfully completed within the specified time. The number of tools will be set according to the specific requirements of the task, usually between 1 and 20 tools, and the system will adjust according to the nature of the task and the operations required in the task. In addition, the system will analyze the availability of equipment and personnel to ensure that there are no resource conflicts. The evaluation range of resource conflicts is from 0 to 10, where the higher the conflict degree, the greater the conflict degree of resource allocation. If the evaluation result shows that the resource conflict degree exceeds 5, the system will adjust the resources, and may assign the task to equipment with lower load or idle, or rearrange the use of operators and tools to ensure that the task can be executed smoothly without delay. Finally, the generated resource scheduling data will include the allocation of equipment, operators, and tools, ensuring that each resource is most effectively allocated, the task can be completed on time, and at the same time, waste or over-occupation of resources is avoided.
[0115] Step S344: Perform scheduling simulation based on the resource scheduling data, where the task execution time is set to 0 - 120 minutes and the resource usage time is set to 0 - 480 minutes to obtain the scheduling simulation data;
[0116] In this embodiment, the system sets the execution time of the task between 0 and 120 minutes, and the specific time is determined according to the complexity of the task. For example, a more complex welding task may require a longer time, while a simple inspection task may be completed in a shorter time. The range of resource usage time is set from 0 to 480 minutes, and the resource time occupied by each task is calculated based on the operation cycle of the production line. During the scheduling simulation process, the system combines the execution time of the task with the resource usage time, calculates the amount of resources required for each task, and simulates the resource usage situation, especially the usage time of equipment, tools, and operators. Then, the system simulates the possible resource conflicts during the task execution process, especially the probability of equipment failure. The equipment failure probability is set from 0 to 100%, which means that the equipment may fail during the task execution, resulting in the unavailability of resources. When the equipment failure probability exceeds the set 80%, the system will automatically reschedule the task to avoid the occurrence of failures and ensure that the task can be completed on time. If the probability of equipment failure is low, the resources can be arranged according to the actual task execution situation. During the simulation process, the system will also predict possible resource bottlenecks. For example, the resources of equipment or operators are tense during a certain period, which may lead to task delays. Through simulation calculations, the resource usage time of each task and the occurrence of resource bottlenecks, the system will optimize the resource scheduling to avoid task delays due to insufficient resources. Finally, the simulation results will generate scheduling simulation data, including the resource usage situation during the task execution process, possible resource conflicts and bottlenecks, and the optimized scheduling plan, providing data support for subsequent task execution monitoring and task scheduling optimization.
[0117] Step S345: Monitor the task execution for the scheduling simulation data to obtain task execution data;
[0118] In this embodiment, the core parameters monitored include task execution delay and resource conflict situation. Task execution delay refers to the gap between the actual execution time and the expected execution time. The execution time of each task has a set range during the simulation phase (for example, 0 to 120 minutes). Therefore, when the actual execution time exceeds the predetermined time by 30 minutes, the alarm mechanism will be triggered. At this time, the system will record the specific delay time and automatically generate a warning to remind the dispatcher to make adjustments to ensure that the task can be completed on time. In addition to delay, the resource conflict situation is also a key parameter for monitoring. The resource conflict degree represents the usage of devices or tools. When multiple tasks require the same resources within the same time period, the resource conflict degree will increase, and the range is set from 0 to 10. If the resource conflict degree exceeds 5, the system will automatically identify the conflict and make adjustments, reallocating devices and tools to ensure that task execution is not affected by resource overload. The system also monitors the execution status of tasks in real time, including whether the task is running, completed, or interrupted, etc. By comparing the actual execution time of each task with the predetermined time, the system can determine whether the progress of the task is normal and identify progress deviations in a timely manner. The monitoring process also records the execution process of tasks, such as the usage of devices and the requirements for tools, to ensure the maximization of resource utilization efficiency. Through real-time monitoring, the system will obtain task execution data and continuously feedback it to the scheduling system, providing accurate data support for subsequent task scheduling optimization to ensure the efficient and stable operation of the production line.
[0119] Step S346: Optimize task scheduling according to task execution data, where the task execution efficiency is set to 0 - 100%, the task allocation adjustment range is 0 - 50%, and the task delay optimization threshold is 0 - 30 minutes, so as to obtain task scheduling data.
[0120] In this embodiment, the task execution efficiency is the core parameter to be optimized, and its range is set from 0 to 100%. The calculation of the execution efficiency is based on the ratio between the actual completion time and the scheduled time of the task. If the execution efficiency of a certain task is lower than 80%, the system will automatically identify that there is an efficiency problem with this task and initiate optimization and adjustment. The specific ways of optimization include reallocating tasks, preferentially allocating them to devices with higher execution efficiency, or adjusting the resource configuration of the devices to ensure the most effective utilization of resources. For example, if a certain device has low efficiency when executing multiple tasks, the system can reallocate some tasks to other devices with lighter loads and higher efficiency. Secondly, the adjustment range of task allocation is set from 0 to 50%, that is, the system can flexibly adjust task allocation within this range. Task allocation adjustment is usually based on the changing demands of the production line, equipment load, and task urgency. For example, if the equipment load is too high or the task priority changes, the system will rearrange task allocation according to the adjustment range to ensure that the task load of each device remains within a reasonable range. When the task delay exceeds the set threshold (usually 30 minutes), the system will initiate the optimization mechanism. Task delay is mainly caused by equipment failures, resource conflicts, or production line bottlenecks. When the task delay exceeds the set threshold, the system will perform scheduling optimization according to the specific situation of the task delay. The optimization measures include adjusting the priority of tasks to make tasks with higher delays execute first, adjusting the equipment resource configuration, reallocating the tasks of high-load devices to other devices, or adjusting the execution time of tasks, such as shortening the operation time of certain links to make up for the delay. The system comprehensively analyzes task execution data, combines the load of the devices, the execution time and priority of the tasks, optimizes the execution order of the tasks, ensures the maximization of the overall efficiency of the production line, and ultimately improves the task completion rate and the operation efficiency of the production line.
[0121] Optionally, this specification also provides an online update system for an industrial Internet of Things AI algorithm, which is used to execute the online update method of the industrial Internet of Things AI algorithm as described above. The online update system for the industrial Internet of Things AI algorithm includes:
[0122] A device connection graph construction module, which is used to obtain industrial Internet of Things data, extract device-level information features to obtain device-level information, and construct a device connection graph based on the device-level information;
[0123] A device life analysis module, which is used to identify performance bottlenecks in the device-level information to obtain performance bottleneck data, predict device failures based on the performance bottleneck data to obtain failure data, and perform device life analysis according to the failure data and the performance bottleneck data to generate device life data;
[0124] The AI task scheduling model construction module is used to identify abnormal connections in the device connection diagram based on device lifespan data, generate device abnormal connection data; perform industrial production task scheduling based on the device abnormal connection data to obtain task scheduling data; construct an AI task scheduling model based on the task scheduling data;
[0125] The algorithm online update module is used to perform intelligent scheduling of industrial tasks on industrial Internet of Things data according to the AI task scheduling model to obtain intelligent scheduling tasks, perform online update of the AI algorithm on the AI task scheduling model, and deploy it to the industrial Internet of Things device side.
[0126] Optionally, this specification also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the online update method of any one of the industrial Internet of Things AI algorithms.
[0127] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0128] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. An online update method for an AI algorithm in the industrial Internet of Things, characterized in that, It includes the following steps: Step S1: Obtain industrial Internet of Things data, extract device-level information features, and obtain device-level information; Construct a device connection graph based on the device-level information; Step S2: Identify performance bottlenecks in the device-level information to obtain performance bottleneck data; Perform device failure prediction based on the performance bottleneck data to obtain failure data; conduct device life analysis based on the failure data and the performance bottleneck data to generate device life data; Step S3: Identify abnormal connections in the device connection graph according to the device life data to generate device abnormal connection data; Perform industrial production task scheduling based on the device abnormal connection data to obtain task scheduling data; construct an AI task scheduling model based on the task scheduling data; Step S4: Perform intelligent scheduling of industrial tasks on the industrial Internet of Things data according to the AI task scheduling model to obtain intelligent scheduling tasks, perform online update of the AI algorithm on the AI task scheduling model, and deploy it to the industrial Internet of Things device side.
2. The online update method of the industrial Internet of Things AI algorithm according to claim 1, characterized in that, Specifically, Step S1 is as follows: Step S11: Obtain industrial Internet of Things data, extract device-level information features, and obtain device-level information; Step S12: Define device connection relationships according to the device-level information, where the maximum transmission rate of each connection relationship is set at 10 Mbps - 1000 Mbps, to obtain device connection relationships; Step S13: Construct a device tree-like hierarchical structure based on the device-level information, where the hierarchical structure is set to a maximum of five levels, to obtain a device tree-like hierarchical structure; Step S14: Construct a device connection graph according to the device tree-like hierarchical structure and the device connection relationships, so as to obtain a device connection graph.
3. The online update method of the industrial Internet of Things AI algorithm according to claim 1, characterized in that The performance bottleneck identification includes: Extract device working state features from the device-level information to obtain the device working state; Statistical device performance indicators based on the device working state; Screen for low-performance devices based on the device performance indicators. If the CPU load of a device exceeds 90% within 5 consecutive minutes, the device memory usage rate exceeds 85%, and the average response time of the device within 5 minutes exceeds 200 ms, it is determined as a low-performance device; Conduct resource-intensive task simulation based on the low-performance devices to obtain task simulation data; Calculate the response time for the task simulation data to obtain the response time; Judge CPU overload based on the response time. If the CPU load of the device exceeds 90% within 10 minutes and the response time always exceeds 300 ms, it is determined that the CPU is overloaded, to obtain CPU overload data; Perform cross-device performance bottleneck analysis on the low-performance devices according to the CPU overload data to generate performance bottleneck data.
4. The online update method of the industrial Internet of Things AI algorithm according to claim 1, wherein, The device failure prediction includes: Monitor the device memory usage amount based on the performance bottleneck data, where the memory usage amount is collected once per minute; Monitor the device memory recovery amount based on the performance bottleneck data, where the memory amount is collected every 30 seconds; Conduct a preliminary determination of memory leakage according to the device memory usage amount and the device memory recovery amount to obtain suspected memory leakage data; Perform fragmentation detection on the suspected memory leakage data to generate memory fragmentation data; Perform memory failure clustering on the suspected memory leakage data according to the memory fragmentation data to generate failure data.
5. The online update method of the industrial Internet of Things AI algorithm according to claim 1, characterized in that The device life analysis mentioned above includes: Statistical analysis of the failure time based on the failure data; Classifying the failure data according to the failure time to obtain intermittent failure data and periodic failure data; Analyzing the device communication interruption based on the intermittent failure data to obtain device communication interruption data; Statistical analysis of the failure detection delay based on the device communication interruption data, where the communication interruption threshold is set to be greater than 5 minutes to obtain failure detection delay data; Analyzing the abnormality of the device cooling system according to the periodic failure data to obtain device cooling system abnormality data; Evaluating the overheating of electronic components based on the device cooling system abnormality data to obtain electronic component overheating data; Constructing a device life prediction model according to the electronic component overheating data and the failure detection delay data; Inputting the performance bottleneck data into the device life prediction model and predicting the device life to obtain device life data.
6. The online update method of the industrial Internet of Things AI algorithm according to claim 1, characterized in that Step S3 is specifically as follows: Step S31: Predicting the device maintenance based on the device life data to obtain device maintenance data; Step S32: Identifying abnormal connections in the device connection diagram based on the device maintenance data to generate device abnormal connection data; Step S33: Counting the abnormal device tasks based on the device abnormal connection data to obtain abnormal device tasks; Step S34: Scheduling the industrial production tasks according to the abnormal device tasks to obtain task scheduling data; Step S35: Constructing an AI task scheduling model based on the task scheduling data.
7. The online update method of the industrial Internet of Things AI algorithm according to claim 6, characterized in that Step S32 is specifically as follows: Step S321: When the following conditions occur simultaneously, it is determined as an abnormal device connection interruption and device connection interruption abnormal data is obtained: frequent device failures appear in the device maintenance data, the device communication packet loss rate exceeds the set threshold of 10%, the device connection interruption exceeds the set number of times, and the device maintenance time exceeds the set threshold of 30 minutes; Step S322: When the following conditions occur simultaneously, it is determined as an abnormal device connection quality and device connection quality abnormal data is obtained: the device hardware replacement frequency is relatively high in the device maintenance data, the device connection delay exceeds the set threshold of 50%, the connection recovery time after device maintenance is extended beyond the set threshold, and the data transmission failure between devices exceeds the set number of times; Step S323: Integrating the device connection interruption abnormal data and the device connection quality abnormal data to obtain device abnormal connection data.
8. The online update method of the industrial Internet of Things AI algorithm according to claim 6, characterized in that, Step S34 is specifically as follows: Step S341: Performing priority classification according to the abnormal device tasks, where the priority is set from 1 to 10, to obtain task priority data; Step S342: Allocating device tasks based on the task priority data to obtain task allocation data; Step S343: Scheduling the production line resources according to the task allocation data to obtain resource scheduling data; Step S344: Performing scheduling simulation based on the resource scheduling data, where the task execution time is set from 0 to 120 minutes and the resource usage time is set from 0 to 480 minutes, to obtain scheduling simulation data; Step S345: Monitoring the task execution for the scheduling simulation data to obtain task execution data; Step S346: Optimize task scheduling based on task execution data, where the task execution efficiency is set to 0 - 100%, the task allocation adjustment range is 0 - 50%, and the task delay optimization threshold is 0 - 30 minutes, so as to obtain task scheduling data.
9. An online update system for an AI algorithm in the industrial Internet of Things, characterized in that, An online update method for implementing the industrial Internet of Things AI algorithm as described in claim 1, the online update system of the industrial Internet of Things AI algorithm includes: A device connection graph construction module, configured to obtain industrial Internet of Things data, extract device hierarchical information features, and obtain device hierarchical information; construct a device connection graph based on the device hierarchical information; A device life analysis module, configured to identify performance bottlenecks in the device hierarchical information to obtain performance bottleneck data; predict device failures based on the performance bottleneck data to obtain failure data; analyze the device life based on the failure data and the performance bottleneck data to generate device life data; An AI task scheduling model construction module, configured to identify abnormal connections in the device connection graph based on the device life data to generate device abnormal connection data; perform industrial production task scheduling based on the device abnormal connection data to obtain task scheduling data; construct an AI task scheduling model based on the task scheduling data; An algorithm online update module, configured to perform intelligent scheduling of industrial tasks on the industrial Internet of Things data according to the AI task scheduling model to obtain intelligent scheduling tasks, perform online update of the AI algorithm on the AI task scheduling model, and deploy it to the industrial Internet of Things device side.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the online update method of the industrial Internet of Things AI algorithm as described in any one of claims 1 to 8.