An agent-based data stream real-time scheduling method

By using a reinforcement learning-driven intelligent scheduling framework, the intelligent agent can perceive and predict data flow characteristics and system status in real time, solving the dynamic adaptability problem of data flow scheduling in smart factories, realizing efficient resource utilization and multi-objective optimization, and ensuring the stability and efficiency of production.

CN121000684BActive Publication Date: 2026-05-01CHINA NAT BUILDING MATERIALS TECH CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA NAT BUILDING MATERIALS TECH CO LTD
Filing Date
2025-08-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies for data flow scheduling in smart factories lack dynamic adaptability, making it difficult to meet the needs of multi-objective optimization and resource pre-allocation, resulting in uneven resource utilization and processing delays.

Method used

A reinforcement learning-driven intelligent scheduling framework is adopted. By enabling agents to perceive data flow characteristics and system status in real time, and combining prediction models and multi-objective scheduling optimization algorithms, dynamic scheduling strategies are generated to achieve pre-allocation and adjustment of resources.

Benefits of technology

It improves the efficiency and stability of data stream processing, avoids resource waste and delays, achieves multi-objective optimization, and ensures stable and efficient production in smart factories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121000684B_ABST
    Figure CN121000684B_ABST
Patent Text Reader

Abstract

The present disclosure provides an agent-based data flow real-time scheduling method, which comprises: according to the factory production environment, building an intelligent scheduling framework driven by reinforcement learning, which comprises an environment, an agent and a reward function; converting data flow characteristics and system state into feature vectors to construct a multi-dimensional feature reinforcement learning state space; the agent perceives the reinforcement learning state space and calls a prediction model to predict the data flow characteristic change trend and the system state change trend in the future preset time period, and generates a resource pre-allocation scheme based on the prediction result; the agent generates at least one scheduling action based on a multi-objective scheduling optimization algorithm to form a scheduling strategy according to the current production target, the resource pre-allocation scheme and the reinforcement learning state space. The present disclosure effectively avoids resource waste and data processing delay, significantly improves resource utilization, and ensures stable and efficient production of the intelligent factory production environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing methods, specifically to a real-time data stream scheduling method, apparatus, storage medium, and electronic device based on intelligent agents. Background Technology

[0002] With the advent of Industry 4.0, data flows in smart factories are characterized by rapid growth, diversification, and complexity, placing higher demands on real-time data flow scheduling. In modern industrial production environments such as smart cement plants, massive amounts of data generated by numerous production equipment, sensors, and control systems require efficient processing to ensure the stability and efficiency of the production process. Currently, data flow scheduling in industrial production environments mainly employs two methods: static scheduling and simple rule-based scheduling.

[0003] Static scheduling assigns specific types of data to pre-defined servers or processing units. While simple to implement, it lacks flexibility and cannot adapt to dynamic changes in system load and data flow characteristics. When data traffic surges or system resources fluctuate, static scheduling can easily overload some processing units while leaving others idle, leading to uneven resource utilization and increased processing latency. Rule-based scheduling, on the other hand, dynamically allocates data flows by pre-setting a series of rules, such as first-come, first-served or shortest-job-first. While more flexible than static scheduling, these rules often fail to cover all scenarios in complex and ever-changing industrial environments, and conflicts between rules can result in poor scheduling performance.

[0004] In recent years, with the development of artificial intelligence technology, reinforcement learning-based scheduling methods have been gradually applied to industrial production environments. However, existing data flow scheduling methods have significant shortcomings in terms of dynamic adaptability, multi-objective optimization, resource pre-allocation, and data flow priority processing, making it difficult to meet the demands of modern industrial production environments such as smart cement plants for real-time and efficient data flow processing. Therefore, there is an urgent need for an intelligent scheduling method that can perceive environmental changes in real time, predict data flow trends, and achieve multi-objective optimization to improve the efficiency and stability of data flow processing in industrial production environments. Summary of the Invention

[0005] The purpose of this disclosure is to provide a real-time data flow scheduling method, apparatus, storage medium, and electronic device based on intelligent agents, so as to solve the problems existing in the prior art.

[0006] The embodiments of this disclosure adopt the following technical solution: a real-time data flow scheduling method based on an intelligent agent, comprising: building a reinforcement learning-driven intelligent scheduling framework according to the factory production environment, the framework including an environment, an intelligent agent, and a reward function; wherein, the environment includes production equipment, production resources, a data transmission network, and a data processing server; the intelligent agent is deployed in a control center, and perceives the data flow characteristics and system status in the environment in real time through the Internet of Things, and generates a scheduling strategy; the reward function gives the intelligent agent a positive or negative reward according to the impact of the scheduling strategy on production; the data flow characteristics and the system status are converted into feature vectors to construct a multi-dimensional feature reinforcement learning state space; the intelligent agent perceives the reinforcement learning state space and calls a prediction model to predict the data flow characteristic change trend and the system status change trend within a future preset time period, and generates a resource pre-allocation scheme based on the prediction results; the intelligent agent generates at least one scheduling action based on a multi-objective scheduling optimization algorithm to form a scheduling strategy according to the current production target, the resource pre-allocation scheme, and the reinforcement learning state space.

[0007] This disclosure also provides a real-time data flow scheduling device based on an intelligent agent, comprising: a framework building module for building a reinforcement learning-driven intelligent scheduling framework according to a factory production environment, the framework including an environment, an intelligent agent, and a reward function; wherein, the environment includes production equipment, production resources, a data transmission network, and a data processing server; the intelligent agent is deployed in a control center, and perceives the data flow characteristics and system state in the environment in real time through the Internet of Things, and generates a scheduling strategy; the reward function gives the intelligent agent a positive or negative reward according to the impact of the scheduling strategy on production; a vector transformation module for converting the data flow characteristics and the system state into feature vectors to construct a multi-dimensional feature reinforcement learning state space; an intelligent scheduling module for calling the intelligent agent to perceive the reinforcement learning state space, and calling a prediction model to predict the data flow characteristic change trend and the system state change trend within a preset time period in the future, and generating a resource pre-allocation scheme based on the prediction results; the intelligent agent generates at least one scheduling action based on a multi-objective scheduling optimization algorithm to form a scheduling strategy according to the current production target, the resource pre-allocation scheme, and the reinforcement learning state space.

[0008] This disclosure also provides a storage medium storing a computer program, characterized in that the computer program, when executed by a processor, implements the steps of the above-described agent-based real-time data flow scheduling method.

[0009] This disclosure also provides an electronic device, including at least a memory and a processor, wherein the memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the above-described agent-based real-time data flow scheduling method.

[0010] The beneficial effects of the embodiments disclosed herein are as follows: Through a reinforcement learning-driven intelligent scheduling framework, dynamic adaptability of data flow scheduling is achieved. Compared with static scheduling or simple rule scheduling in the prior art, it can perceive changes in data flow characteristics and system state in real time, and adjust the scheduling strategy in a timely manner according to these changes, effectively avoiding resource waste and data processing delays. By constructing a reinforcement learning state space containing data flow characteristics, it provides a comprehensive and accurate basis for intelligent agent decision-making, enabling it to grasp the overall situation of production data and the system, thus improving the accuracy of decision-making. The resource pre-allocation strategy and resource elastic allocation mechanism based on prediction models can allocate resources in advance according to prediction results and quickly adjust when the data volume changes suddenly, significantly improving resource utilization. The multi-objective scheduling optimization algorithm achieves collaborative optimization of multiple production objectives, ensuring stable and efficient production in the intelligent factory production environment. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A path diagram for one or more embodiments of the agent-based real-time data stream scheduling method provided in this specification;

[0013] Figure 2 This specification provides a schematic diagram of the structure of a real-time data flow scheduling device based on an intelligent agent for one or more embodiments. Detailed Implementation

[0014] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0015] To address the problems in the prior art, the first embodiment of this disclosure provides a real-time data flow scheduling method based on intelligent agents. This embodiment will describe the specific implementation process of the method in conjunction with the production environment of an intelligent cement plant. Figure 1 The flowchart of the real-time data stream scheduling method of this embodiment is shown, mainly including steps S10 to S40:

[0016] S10 builds a reinforcement learning-driven intelligent scheduling framework based on the factory production environment.

[0017] A scheduling framework is constructed within the intelligent cement plant production environment. This framework comprises three core components: environment, agents, and reward function. The environment includes production equipment such as raw material silos, crushers, raw meal mills, rotary kilns, and cement mills, as well as data transmission networks, data processing servers, and various production resources. Agents are deployed in the data processing center, using an IoT system to perceive the data flow characteristics and system status in the environment in real time. The reward function assigns positive or negative rewards to the agents based on the impact of the scheduling strategy on production. For example, a positive reward is given when the scheduling strategy improves the clinker grade compliance rate, reduces energy consumption, and decreases equipment downtime; a negative reward is given if the scheduling strategy causes data processing delays leading to quality fluctuations or equipment overload shutdowns.

[0018] This embodiment utilizes the LangChain framework or similar frameworks to construct intelligent agents. The agent uses a large deep learning model (LLM) pre-trained on massive amounts of data as its brain, typically used to understand user input, generate responses, and make decisions or provide suggestions in certain situations. In agents built using LangChain, the LLM can be combined with a ReAct mechanism to adjust the scheduling strategy, i.e., through a cycle of Thought→Action→Observation→Though updates until the problem is solved or the task termination condition is met. Tools are an important component of LangChain's agent construction, enabling the agent to perform actual operations, call external services (APIs or custom tools), and interact with the system, thereby increasing the agent's functionality and applicability. By integrating different types of Tools, the agent can complete more complex tasks, such as data acquisition and calling other models. Furthermore, tools can be created according to specific needs during actual operation to achieve the execution of specific tasks. The agent also includes short-term memory, enabling context preservation and personalized responses, allowing the agent to optimize decisions based on historical scheduling data during actual scheduling.

[0019] S20 transforms data flow features and system states into feature vectors to construct a reinforcement learning state space with multi-dimensional features.

[0020] Specifically, the data flow characteristics in this embodiment include: instantaneous data flow generation, data priority coefficient, standard deviation of data arrival interval, data type identifier, temporal correlation of multi-source data, proportion of abnormal data, and data transmission rate; the system status includes at least: peak CPU utilization of the data processing server, slope of memory usage trend, disk I / O response latency, bandwidth utilization of the backbone network, task queue length, idle time of production resources, network transmission latency, data cache size, power consumption, and resource load balancing. In this embodiment, the intelligent agent acquires data flow characteristics and system status in real time through the Internet of Things system. For example, the sensing module of the intelligent agent collects equipment data such as the vibration frequency of the crusher, the motor current of the raw material vertical mill, the kiln tail temperature of the rotary kiln, and the grinding pressure of the cement ball mill every 2 seconds. At the same time, it collects system data such as CPU utilization, memory usage, and network bandwidth usage of the data processing server every 5 seconds. These data can be preprocessed (noise filtering and format conversion) by the edge nodes and then transmitted to the decision module of the intelligent agent.

[0021] After acquiring the aforementioned data, the agent transforms the data flow features and system state into feature vectors to construct a multi-dimensional feature reinforcement learning state space. First, the parameters in the data flow features and system state at the current moment are standardized, including but not limited to outlier cleaning and normalization. Then, according to a preset feature weighting rule, corresponding weight values ​​are assigned to each standardized parameter. The weight values ​​are determined based on the parameter's influence on scheduling decisions. For example, for data from various production equipment in a smart cement plant, the standard deviation of the arrival interval of the rotary kiln temperature data has the greatest impact on scheduling, so it is assigned a high weight of 0.2; the priority coefficient of the raw material mill power consumption data has a weight of 0.15; the peak CPU utilization of the first data server has a weight of 0.12; and other parameters have weights ranging from 0.05 to 0.1. After weighting, the parameters are arranged in a preset order, ultimately forming the feature vector S = [D1, D2, ..., D...] at the current moment. m S1, S2, ..., S n ], where D1, D2, ..., D m The feature vectors corresponding to the features of each data stream are S1, S2, ..., S. nThese correspond to the feature vectors of various system states. Finally, feature vectors from different times are acquired at specific time intervals and integrated to form a reinforcement learning space. It should be noted that in actual implementation, feature vectors can be extracted every 10 minutes, and six feature vectors from one hour can be used to form a reinforcement learning state space, allowing the agent to perceive the state changes of the factory's production environment within one hour.

[0022] S30, the intelligent agent perceives the reinforcement learning state space, and calls the prediction model to predict the trend of data flow characteristics and system state changes within a preset time period in the future, and generates a resource pre-allocation scheme based on the prediction results.

[0023] The intelligent agent perceives the constantly updated reinforcement learning state space and invokes a prediction model to predict the changing trends of data flow characteristics and system state within a preset future time period based on the perception. Specifically, the prediction model can utilize the correlation data of raw material composition, production load, and data processing resource requirements of a smart cement plant over the past N months. Using an LSTM model, it predicts the changing trends of data flow and system state, helping the agent predict changes in the factory's production environment within a specific future time period. Based on the prediction results, the agent can identify all possible data flow changes and system state changes within the factory during the specific future time period, and select data processing tasks with different priorities. Subsequently, based on the predicted data volume and processing requirements of the data processing tasks with different priorities, the agent allocates corresponding processing resources.

[0024] In this embodiment, the length of the preset time period should not be set too long, typically between 5 minutes and 180 minutes, sufficient to meet production forecasts after a certain period. Forecasts with excessively long time periods may lead to decreased forecast accuracy and wasted computing resources. The agent analyzes the changing trends of data stream characteristics within the preset time period and filters out high-priority and non-critical data processing tasks from all data processing tasks based on data priority. This distinction can be made by setting priority thresholds. For example, if the priority of real-time raw material crushing data is 0.7, the priority of historical raw material crushing data is 0.3, and the preset priority threshold is 0.6, then the data processing task for real-time raw material crushing data can be identified as a high-priority task, and the data processing task for historical raw material crushing data can be identified as a non-critical task.

[0025] In some embodiments, the classification of high-priority data processing tasks and non-critical data processing tasks can also be combined with the current production goals. Specifically, firstly, based on the data priority coefficients of data flow features in the predicted data flow feature change trend, the preliminary task priority of the data processing task corresponding to each data flow feature is determined, which is the priority directly obtained from the prediction; then, the preliminary task priority is adjusted in combination with the current production goals. If the current data processing task has a high correlation with the current production goals, or has a significant impact on the current production goals, the preliminary priority is corrected to increase its priority value. For data processing tasks with a low correlation with the current production goals, their preliminary priority value can be reduced, that is, the higher the correlation, the higher the priority after correction; finally, based on the comparison between the corrected priority and a preset threshold, tasks whose corrected priority reaches the preset threshold are determined as high-priority data processing tasks, and the rest are determined as non-critical data processing tasks. For example, for the raw material crushing real-time data processing task, its initial task priority is 0.7. When the current production goal is to ensure the accuracy of raw material proportioning, the correlation between the raw material crushing real-time data processing task and ensuring the accuracy of raw material proportioning is high. After correction, the priority of the raw material crushing real-time data processing task is corrected to 0.9, and it is determined to be a high-priority data processing task. If the current production goal is to improve cement packaging efficiency, the correlation between the raw material crushing real-time data processing task and improving cement packaging efficiency is low, and its priority will be corrected to 0.5, which is lower than the preset threshold. At this time, the raw material crushing real-time data processing task will be identified as a non-critical data processing task.

[0026] In some embodiments, the agent can determine whether the currently predicted data processing task is related to the current production target based on a preset list of data processing tasks corresponding to different production targets; or, the agent can analyze the current production target, determine the operation process required to achieve the production target, and then evaluate whether the current data processing task affects the execution of any of the above operation processes; or, the agent can analyze the degree of correlation between the data processing task and the production target through a pre-trained neural network model, and assign different priority correction values ​​to different degrees of correlation based on the magnitude of the correlation.

[0027] Subsequently, based on the predicted data volume and processing requirements of high-priority data processing tasks, a resource scope is defined to ensure the processing of these tasks. Due to the immediacy and accuracy requirements of high-priority data processing tasks, the agent needs to allocate sufficient system resources for processing. For example, for the real-time data processing task of raw material crushing, considering its real-time data volume (800MB / h) and processing requirements (requiring 2 CPU cores and 4GB of memory), the resource scope is defined as follows: 3 CPU cores are reserved for Server1 (to ensure peak processing requirements) and 6GB of memory (including redundancy), while Server3 is designated as a backup node (1 CPU core is reserved), automatically switching when Server1 is overloaded.

[0028] For non-critical data processing tasks, their corresponding processing priorities and delay conditions are set and integrated into a delay processing queue. In practice, multiple non-critical data processing tasks can be integrated into one delay processing queue. At this time, they can be sorted according to their processing priorities. The delay processing conditions should at least include the conditions that the non-critical data processing task can be started now, and the conditions that ensure the task will not be delayed indefinitely. For example, for the processing task of raw material crushing historical data, its delay processing conditions can be set to allow processing only when the CPU utilization of Server1-Server4 is below 50% and there are no high-priority tasks waiting. At the same time, the delay time of the task is counted. If the delay time exceeds 24 hours (exceeding the daily report statistics requirements), forced processing is triggered.

[0029] While ensuring resource availability for data processing tasks, the agent also needs to assess system state trends to determine whether the system's state will meet the execution requirements of the corresponding data processing tasks in the future. This is especially important for high-priority data processing tasks, ensuring the actual data server resources meet the allocated requirements. If not, the agent needs to determine whether resource migration is necessary and establish triggering conditions to satisfy the data task's processing needs. For example, if the agent predicts that Server1's CPU utilization will rise to 85%, failing to meet the processing requirements of the real-time data processing task for raw material crushing, the agent sets the following triggering conditions: when Server1's CPU utilization is ≥80% for 3 consecutive minutes, or its memory usage is ≥85%, resource migration is triggered. The migration target is Server3's backup resources, and a dual-caching mechanism ensures uninterrupted data transmission during the migration process.

[0030] Finally, the agent integrates the overall resource scope, delayed processing queue, and triggering conditions to form a resource pre-allocation scheme, which is stored in the agent's decision database as the basis for subsequent scheduling.

[0031] S40, the agent generates at least one scheduling action based on the current production target, resource pre-allocation scheme and reinforcement learning state space to form a scheduling strategy.

[0032] After the aforementioned processing, the agent obtains the real-time state of the current factory production environment through reinforcement learning state space. Based on the resource pre-allocation scheme, the agent makes a preliminary decision on the resource scheduling method in a specific time period in the future. In actual scheduling, the agent obtains a scheduling strategy that meets the current production goals and satisfies the optimization of the most goals through a multi-objective scheduling optimization algorithm.

[0033] Specifically, the multi-objective scheduling optimization algorithm is configured in conjunction with the actual production goals of the factory. The optimization goals include, but are not limited to, production efficiency, resource utilization, data processing efficiency, clinker quality stability, unit product energy consumption, and equipment fault early warning response time. Furthermore, the multi-objective scheduling optimization algorithm in this embodiment is configured with a preset action space based on the actual production data flow scheduling process. The preset action space is a set of all scheduling actions that the intelligent agent can select when generating the scheduling strategy. The types of actions included are predefined based on the actual production process and data flow scheduling requirements, covering key links such as data processing, resource allocation, and task scheduling. For example, in a smart cement plant, the specific actions in the preset action space include: data diversion actions, resource adjustment actions, task sorting actions, resource migration actions, and delay processing actions, and the subject and magnitude of the actions are determined according to actual needs.

[0034] In this embodiment, the agent uses the resource pre-allocation scheme and the feature vector of the current moment in the reinforcement learning state space as input parameters for the multi-objective scheduling optimization algorithm. According to the rules of multi-objective scheduling optimization, it obtains at least one candidate scheduling action in the preset action space based on the feature vector of the current moment to meet the multi-objective scheduling optimization requirements at the current moment. Then, it filters all the obtained candidate scheduling actions according to the resource pre-allocation scheme as a constraint, and removes the candidate scheduling actions that do not meet the constraints, so that all the filtered candidate scheduling actions can adapt to the data changes in the future predetermined time period. Finally, the remaining candidate scheduling actions after filtering are executed in order according to the execution priority rule corresponding to the current production goal to form the final scheduling strategy. It should be noted that the execution priority rule is formulated based on the degree of influence of the scheduling action on the multi-objective optimization index and the dependency relationship between actions under the current production goal. In actual formulation, it is advisable to satisfy as many goals as possible to achieve the optimal effect while conforming to the execution logic between actions.

[0035] Furthermore, after the agent generates the scheduling strategy, it drives the real-time scheduling of data flows within the production environment, including various production resources, production equipment, data transmission networks, and data processing servers. The reward function determines whether to award the agent a positive or negative reward by acquiring various production indicators and their changes after the scheduling strategy is executed. Specifically, a positive reward is given when the indicator changes align with a preset optimization direction, such as improved resource utilization or production efficiency; conversely, a negative reward is given when resource utilization or production efficiency decreases. The specific value of the positive or negative reward is determined by the degree of impact of the scheduling strategy on the production objective. After receiving the numerical reward (positive / negative reward) output by the reward function, the agent optimizes the scheduling strategy through a closed-loop logic of "experience storage - value update - strategy adjustment - iterative reinforcement." The core principle is to determine the effectiveness of the current action based on the reward value, reinforcing effective actions and weakening ineffective ones.

[0036] This embodiment achieves dynamic adaptability of data flow scheduling through a reinforcement learning-driven intelligent scheduling framework. Compared with static scheduling or simple rule scheduling in existing technologies, it can perceive changes in data flow characteristics and system state in real time and adjust the scheduling strategy in a timely manner according to these changes, effectively avoiding resource waste and data processing delays. By constructing a reinforcement learning state space containing data flow characteristics, it provides a comprehensive and accurate basis for agent decision-making, enabling it to grasp the overall situation of production data and improve the accuracy of decision-making. The resource pre-allocation strategy and resource elastic allocation mechanism based on prediction models can allocate resources in advance according to prediction results and adjust rapidly when the data volume changes suddenly, significantly improving resource utilization. The multi-objective scheduling optimization algorithm realizes the collaborative optimization of multiple production objectives, ensuring stable and efficient production in the smart factory production environment.

[0037] Based on the same inventive concept, the second embodiment of this disclosure provides a real-time data flow scheduling device based on intelligent agents. This scheduling device can be configured at the upper-level control center of an actual factory, and its structural schematic diagram is shown below. Figure 2As shown, it includes at least: a framework building module 10, used to build a reinforcement learning-driven intelligent scheduling framework based on the factory production environment. The framework includes an environment, an agent, and a reward function. The environment includes production equipment, production resources, data transmission networks, and data processing servers. The agent is deployed in the control center and uses the Internet of Things to perceive the data flow characteristics and system status in the environment in real time, and generates a scheduling strategy. The reward function gives the agent a positive or negative reward based on the impact of the scheduling strategy on production. A vector transformation module 20 is used to transform the data flow characteristics and system status into feature vectors to construct a multi-dimensional reinforcement learning state space. An intelligent scheduling module 30 is used to call the agent to perceive the reinforcement learning state space and call the prediction model to predict the data flow characteristic change trend and system status change trend within a preset time period in the future, and generate a resource pre-allocation plan based on the prediction results. The agent generates at least one scheduling action based on a multi-objective scheduling optimization algorithm to form a scheduling strategy, according to the current production goal, the resource pre-allocation plan, and the reinforcement learning state space.

[0038] Specifically, data flow characteristics include at least: instantaneous data flow generation, data priority coefficient, standard deviation of data arrival interval, data type identifier, temporal correlation of multi-source data, proportion of abnormal data, and data transmission rate; system status includes at least: peak CPU utilization of data processing server, slope of memory usage trend, disk I / O response latency, bandwidth utilization of backbone network, task queue length, idle time of production resources, network transmission latency, data cache size, power consumption, and resource load balancing.

[0039] In some embodiments, the vector transformation module 20 is specifically used to standardize the parameters in the data flow features and system state at the current time; assign corresponding weight values ​​to the standardized parameters according to the preset feature weight allocation rules, wherein the weight values ​​are determined according to the degree of influence of the parameters on the scheduling decision; arrange the weighted parameters in a preset order to form the feature vector at the current time; and form a reinforcement learning state space based on the feature vectors at different times.

[0040] In some embodiments, the intelligent scheduling module 30 is specifically used to: analyze the predicted trend of data flow characteristics within a future preset time period, identify high-priority data processing tasks and non-critical data processing tasks; delineate the resource range to ensure the processing of high-priority data processing tasks based on the predicted data volume and processing requirements of the high-priority data processing tasks; set processing priorities and delayed processing conditions for non-critical data processing tasks to form a delayed processing queue; determine the triggering conditions for resource migration based on the trend of system state changes; and integrate the resource range, delayed processing queue, and triggering conditions to form a resource pre-allocation scheme.

[0041] In some embodiments, the intelligent scheduling module 30 is specifically used to: determine the preliminary task priority of the data processing task corresponding to each data flow feature based on the data priority coefficient of the data flow feature in the data flow feature change trend; correct the preliminary priority according to the degree of correlation between the data processing task and the current production target, with the task with higher correlation having a higher priority after correction; determine the task whose corrected priority reaches a preset threshold as a high-priority data processing task, and the rest as non-critical data processing tasks.

[0042] In some embodiments, the intelligent scheduling module 30 is specifically used for: the agent using the resource pre-allocation scheme and the feature vector of the current moment in the reinforcement learning state space as input parameters of the multi-objective scheduling optimization algorithm to obtain at least one candidate scheduling action in the preset action space; filtering each candidate scheduling action according to the resource pre-allocation scheme as a constraint condition, and eliminating candidate scheduling actions that do not meet the constraint conditions; for the remaining candidate scheduling actions after filtering, determining the execution order according to the execution priority rule corresponding to the current production target to form a scheduling strategy, wherein the execution priority rule is formulated based on the degree of influence of the scheduling action on the multi-objective optimization index under the current production target and the dependency relationship between actions.

[0043] In some embodiments, the intelligent scheduling module 30 is also used to acquire various production indicators before and after the implementation of the scheduling strategy during the production process, and calculate the change in indicators; call the reward function to determine whether to give positive or negative rewards to the agent based on the change in indicators, and give positive rewards when the change in indicators conforms to the preset optimization direction, and give negative rewards otherwise.

[0044] This embodiment achieves dynamic adaptability of data flow scheduling through a reinforcement learning-driven intelligent scheduling framework. Compared with static scheduling or simple rule scheduling in existing technologies, it can perceive changes in data flow characteristics and system state in real time and adjust the scheduling strategy in a timely manner according to these changes, effectively avoiding resource waste and data processing delays. By constructing a reinforcement learning state space containing data flow characteristics, it provides a comprehensive and accurate basis for agent decision-making, enabling it to grasp the overall situation of production data and improve the accuracy of decision-making. The resource pre-allocation strategy and resource elastic allocation mechanism based on prediction models can allocate resources in advance according to prediction results and adjust rapidly when the data volume changes suddenly, significantly improving resource utilization. The multi-objective scheduling optimization algorithm realizes the collaborative optimization of multiple production objectives, ensuring stable and efficient production in the smart factory production environment.

[0045] Based on the same inventive concept, the third embodiment of this disclosure provides a storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the real-time data flow scheduling method based on intelligent agents described in the first embodiment of this disclosure.

[0046] Based on the same inventive concept, the fourth embodiment of this disclosure provides an electronic device, including at least a memory and a processor, wherein the memory stores a computer program, characterized in that the processor, when executing the computer program in the memory, implements the steps of the real-time data flow scheduling method based on intelligent agents described in the first embodiment of this disclosure.

[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A real-time data stream scheduling method based on intelligent agents, characterized in that, include: Based on the factory production environment, a reinforcement learning-driven intelligent scheduling framework is built. The framework includes an environment, an agent, and a reward function. The environment includes production equipment, production resources, data transmission networks, and data processing servers. The agent is deployed in the control center and uses the Internet of Things to perceive the data flow characteristics and system status of the environment in real time, and generates scheduling strategies. The reward function provides positive or negative rewards to the agent based on the impact of the scheduling strategy on production. The data stream features and the system state are transformed into feature vectors to construct a reinforcement learning state space with multi-dimensional features; The agent perceives the reinforcement learning state space and calls the prediction model to predict the trend of data flow characteristics and system state changes within a preset time period in the future, and generates a resource pre-allocation scheme based on the prediction results. The agent generates at least one scheduling action based on a multi-objective scheduling optimization algorithm to form a scheduling strategy, according to the current production target, the resource pre-allocation scheme, and the reinforcement learning state space.

2. The real-time data stream scheduling method according to claim 1, characterized in that, The data stream characteristics include at least: instantaneous generation of data stream, data priority coefficient, standard deviation of data arrival interval, data type identifier, temporal correlation of multi-source data, proportion of abnormal data, and data transmission rate; The system status includes at least: peak CPU utilization of the data processing server, slope of memory usage trend, disk I / O response latency, bandwidth utilization of the backbone network, task queue length, idle time of production resources, network transmission latency, data cache size, power consumption, and resource load balancing.

3. The real-time data stream scheduling method according to claim 1, characterized in that, The step of converting the data stream features and the system state into feature vectors to construct a multi-dimensional feature reinforcement learning state space includes: Standardize the data stream characteristics and various parameters in the system state at the current moment; According to the preset feature weight allocation rules, each standardized parameter is assigned a corresponding weight value, where the weight value is determined based on the degree of influence of the parameter on the scheduling decision. The weighted parameters are arranged in a preset order to form the feature vector at the current moment; The reinforcement learning state space is formed based on the feature vectors at different times.

4. The real-time data stream scheduling method according to claim 3, characterized in that, The process of generating a resource pre-allocation scheme based on the prediction results includes: Analyze the changing trends of data stream characteristics within the predicted future time period to identify high-priority data processing tasks and non-critical data processing tasks; Based on the predicted data volume and processing requirements of high-priority data processing tasks, a resource scope is defined to ensure the processing of the high-priority data processing tasks. Set processing priorities and delay conditions for non-critical data processing tasks to form a delay processing queue; Based on the trend of system state changes, determine the triggering conditions for resource migration; The resource range, delayed processing queue, and triggering conditions are integrated to form the resource pre-allocation scheme.

5. The real-time data stream scheduling method according to claim 4, characterized in that, The process of identifying high-priority data processing tasks and non-critical data processing tasks based on the predicted data stream characteristic change trends within a future preset time period includes: Based on the data priority coefficient of the data flow feature in the trend of data flow feature change, the preliminary task priority of the data processing task corresponding to each data flow feature is determined; The initial priority is adjusted based on the degree of correlation between the data processing tasks and the current production goals. Tasks with higher correlation have higher priority after adjustment. Tasks whose priority reaches the preset threshold after correction are identified as high-priority data processing tasks, while the rest are identified as non-critical data processing tasks.

6. The real-time data stream scheduling method according to claim 4, characterized in that, The agent generates at least one scheduling action to form a scheduling strategy based on a multi-objective scheduling optimization algorithm, according to the current production target, the resource pre-allocation scheme, and the reinforcement learning state space, including: The agent uses the resource pre-allocation scheme and the feature vector of the current moment in the reinforcement learning state space as input parameters of the multi-objective scheduling optimization algorithm to obtain at least one alternative scheduling action in the preset action space. Each candidate scheduling action is screened based on the resource pre-allocation scheme as a constraint, and candidate scheduling actions that do not meet the constraints are eliminated. For the remaining candidate scheduling actions after screening, the execution order is determined according to the execution priority rule corresponding to the current production target to form the scheduling strategy. The execution priority rule is formulated based on the degree of influence of the scheduling actions on the multi-objective optimization index under the current production target and the dependency relationship between actions.

7. The real-time data stream scheduling method according to any one of claims 1 to 6, characterized in that, Also includes: Obtain various production indicators before and after implementing the scheduling strategy during the production process, and calculate the changes in these indicators; The reward function determines whether to give the agent a positive or negative reward based on the change in the indicator. A positive reward is given when the change in the indicator conforms to the preset optimization direction, and a negative reward is given otherwise.

8. A real-time data stream scheduling device based on intelligent agents, characterized in that, include: The framework building module is used to build a reinforcement learning-driven intelligent scheduling framework based on the factory production environment. The framework includes an environment, an agent, and a reward function. The environment includes production equipment, production resources, data transmission networks, and data processing servers. The agent is deployed in the control center and uses the Internet of Things to perceive the data flow characteristics and system status of the environment in real time, and generates scheduling strategies. The reward function gives the agent a positive or negative reward based on the impact of the scheduling strategy on production. The vector transformation module is used to transform the data stream features and the system state into feature vectors to construct a reinforcement learning state space with multi-dimensional features. The intelligent scheduling module is used to invoke the intelligent agent to perceive the reinforcement learning state space and invoke the prediction model to predict the data flow characteristic change trend and system state change trend within a future preset time period, and generate a resource pre-allocation scheme based on the prediction results; the intelligent agent generates at least one scheduling action to form a scheduling strategy based on the current production target, the resource pre-allocation scheme and the reinforcement learning state space and a multi-objective scheduling optimization algorithm.

9. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the real-time data stream scheduling method based on intelligent agents as described in any one of claims 1 to 7.

10. An electronic device, comprising at least a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program on the memory, it implements the steps of the agent-based real-time data flow scheduling method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Workshop scheduling method of multi-objective weight learning based on reinforcement learning and device and application thereof

    CN116307440A

  • Two-stage assembly flow shop dynamic scheduling method based on deep reinforcement learning

    CN117369393A