Self-adaptive memory management system and method for intelligent agent in steelmaking field
By introducing a hierarchical memory system and a multimodal data fusion module into the steelmaking field, and combining it with closed-loop optimization of reinforcement learning, the shortcomings of existing systems in time series modeling and multimodal data processing in steelmaking are solved, enabling rapid response and adaptive decision-making, and improving the efficiency and interpretability of steelmaking production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI BAOSIGHT SOFTWARE CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-28
AI Technical Summary
Existing artificial intelligence systems lack effective time-series modeling capabilities, have a single memory structure, insufficient multimodal data processing capabilities, and lack closed-loop optimization mechanisms and interpretability in the highly complex and time-series-dependent industrial scenarios of steelmaking. This results in low retrieval efficiency, chaotic knowledge updates, inability to adapt to dynamic changes, and opaque decision-making processes in steelmaking production.
A hierarchical memory system, including a dynamic memory module and a static memory module, is adopted, combined with a multimodal data fusion module and a reinforcement learning-based closed-loop optimization module, to achieve multi-heterogeneous data processing and adaptive decision-making in the steelmaking process.
It enables rapid response and self-optimization of the steelmaking process, improves the accuracy and interpretability of decision-making, enhances operator confidence, and adapts to dynamic changes such as steel grade switching and equipment aging.
Smart Images

Figure CN121935243A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to an adaptive memory management system and method for intelligent agents in the steelmaking field. Background Technology
[0002] In the field of artificial intelligence, some existing AI systems have adopted hierarchical memory structures, such as dividing memory into long-term memory (static memory) and short-term memory (dynamic memory), and possessing the ability to integrate and process various types of information, such as text and images. These systems have demonstrated certain capabilities in tasks such as general cognition or spatial exploration.
[0003] However, when such systems are applied to highly complex and time-sensitive industrial scenarios such as steelmaking, their inherent shortcomings become apparent. First, existing systems are designed primarily for general tasks, lacking the ability to effectively model key time-series parameters such as temperature and composition in industrial processes, making them ill-suited to the dynamic characteristics of steelmaking. Second, the memory structures of existing systems are relatively simple, failing to effectively distinguish and coordinate the management of static knowledge such as process specifications and metallurgical principles (as persistent knowledge) with dynamic situations such as real-time changes in furnace conditions and sensor data. This leads to low retrieval efficiency in the face of massive amounts of data and potential confusion in knowledge updates. More critically, many existing systems lack an effective closed-loop feedback optimization mechanism. Their memory management and decision-making strategies tend to become rigid after deployment, unable to self-adjust and evolve based on actual production results (such as product quality and energy consumption), making it difficult to adapt to changes in actual operating conditions such as raw material fluctuations and equipment aging. Furthermore, the decision-making process of such systems is often opaque, like a "black box," failing to provide traceable evidence for decisions. In the high-risk, high-requirement environment of steelmaking production, this makes it difficult to gain the trust and adoption of operators.
[0004] Chinese patent document CN117854059A discloses a large-model-driven intelligent agent scene exploration and memory management method and system. This type of existing system proposes a large-model-driven intelligent agent that, in a vision-dominated exploration scenario, generates "object semantic memory" and "contextual semantic memory" through image recognition and updates the memory bank using a two-dimensional grid storage structure to support goal-oriented autonomous exploration behavior. While this technology has some innovation in general scene exploration tasks, its architectural design is fundamentally incompatible with industrial application scenarios, especially in highly complex, time-series-dependent, and multimodal coupled heavy industrial environments such as steelmaking. Specific drawbacks are as follows: 1. Limited applicability and lack of industrial time-series modeling capabilities: This paper focuses on "spatial exploration" tasks (such as robot indoor navigation and game scene traversal), with its memory management centered on visual images and static object recognition, using a two-dimensional grid to store spatial location information. However, steelmaking is a continuous industrial process with strong time dependence and multivariate coupling. Key decisions rely on high-dimensional time-series signals such as temperature curves, composition changes, and equipment status, rather than spatial coordinates. This system completely fails to consider time-dimensional modeling and cannot handle process parameters fluctuating at the minute or even second level, causing it to fail in industrial dynamic control scenarios.
[0005] 2. The memory structure is simplistic and lacks a static-dynamic hierarchical mechanism. This paper only constructs a unified "semantic memory bank" without distinguishing between long-term knowledge accumulation (such as process specifications and metallurgical principles) and short-term contextual caching (such as current furnace conditions and sensor anomalies). Its "two-dimensional grid storage" is essentially a spatial index structure, which cannot support knowledge graph-style relational reasoning or semantic retrieval of vector databases. Furthermore, it lacks dynamic management mechanisms such as cache eviction and priority scheduling, resulting in low retrieval efficiency and chaotic knowledge updates when faced with massive amounts of historical data and real-time streaming data.
[0006] 3. The modal processing is one-sided, ignoring multi-source heterogeneous industrial data. This paper only processes the visual image modality, relying on "object recognition" and "scene perception" to generate memory. Steelmaking field data includes multi-modal information such as text (work orders, procedures), time series (sensors), images (furnace condition monitoring), audio (equipment abnormalities), and structured parameters (composition, energy consumption), and there are strong semantic relationships between the modalities (e.g., "flame color + temperature curve → decarburization state"). The system lacks cross-modal fusion capabilities and cannot construct a unified semantic space, which severely restricts its perception and decision-making completeness in complex industrial scenarios.
[0007] 4. Lack of a closed-loop optimization mechanism leads to a rigid and stagnant memory system. The memory update in this paper is merely a one-way process of "recognition → storage," without introducing any feedback mechanism based on decision-making effects (such as reinforcement learning). The system cannot automatically optimize its memory retrieval strategy, adjust knowledge weights, or eliminate low-value items based on actual production results (such as the accuracy of predictions or the reduction of energy consumption). This causes the system performance to become rigid over time, unable to adapt to industrial realities such as raw material fluctuations, equipment aging, and process iterations.
[0008] 5. Lack of explainability and industrial credibility assurance. This document does not provide any decision tracing or memory retrieval visualization mechanism. In high-risk industrial scenarios such as steelmaking, operators must understand the basis for the AI's suggestions (e.g., "Why is it recommended to reduce oxygen at this time?"). The system's output is a black box "target item selection," unable to demonstrate the historical cases, expert rules, or real-time data characteristics upon which its reasoning relies, making it difficult to gain the trust of on-site engineers and hindering actual deployment. Summary of the Invention
[0009] To address the shortcomings of existing technologies, the purpose of this invention is to provide an adaptive memory management system and method for intelligent agents in the steelmaking field.
[0010] An adaptive memory management system for an intelligent agent in the steelmaking field, provided by the present invention, includes: The hierarchical memory system consists of a dynamic memory module for processing real-time streaming data from the steelmaking process, and a static memory module for structured storage of persistent domain knowledge. Multimodal data fusion module: connected to the hierarchical memory system, used to extract features from various heterogeneous data sources including text, images, and time-series sensor data during the steelmaking process, map the features to a unified semantic vector space, and store the mapped features in the dynamic memory module; Memory fusion module: Connected to the dynamic memory module and the static memory module respectively, it is used to retrieve relevant information from the dynamic memory module and the static memory module according to the external input request, and integrate the retrieved information based on a preset strategy to generate unified context information; The reinforcement learning-based closed-loop optimization module is connected to the memory fusion module and generates decision actions based on the unified context information. It calculates reward signals based on the actual execution results of the decision actions in the steelmaking process and uses the reward signals to update the system's decision generation strategy and memory management strategy.
[0011] Preferably, the preset strategy upon which the memory fusion module is based includes: Identify the current technological stage of the steelmaking process; And, based on the process stage, adaptively adjust the fusion weights of the information retrieved from the dynamic memory module and the static memory module.
[0012] Preferably, the method by which the reinforcement learning-based closed-loop optimization module calculates the reward signal includes: Based on the product quality indicators, cost indicators, and efficiency indicators after the decision-making action is executed, a multi-objective comprehensive reward value is calculated as the reward signal.
[0013] Preferably, the multimodal data fusion module employs a cross-attention fusion network to jointly analyze data from at least two different modalities, thereby inferring key process parameters in the steelmaking process; The multimodal data fusion module is also used to store the inferred key process parameters as features in the dynamic memory module.
[0014] Preferably, the static memory module is used to store historical production data, expert experience, or process procedures as knowledge triples containing conditions, operations, and effects, and to manage the knowledge triples in a versioned manner.
[0015] An adaptive memory management method for an intelligent agent in the steelmaking field, provided by the present invention, includes: Features are extracted from multiple heterogeneous data sources in the steelmaking process, and these features are mapped to a unified semantic vector space. The real-time streaming data and mapped features of the steelmaking process are stored in a dynamic memory module; and persistent domain knowledge is stored in a structured manner in a static memory module. Based on external input requests, relevant information is retrieved from the dynamic memory module and the static memory module, and the retrieved information is integrated based on a preset strategy to generate unified context information. Based on the unified context information, a decision action is generated; Calculate the reward signal based on the actual execution results of the decision-making action in the steelmaking process; And the decision generation strategy and memory management strategy are updated using the reward signal.
[0016] Preferably, the step of integrating the retrieved information based on a preset strategy includes: Identify the current technological stage of the steelmaking process; And, based on the process stage, adaptively adjust the fusion weights of the information retrieved from the dynamic memory module and the static memory module.
[0017] Preferably, the step of calculating the reward signal includes: Based on the product quality indicators, cost indicators, and efficiency indicators after the decision-making action is executed, a multi-objective comprehensive reward value is calculated as the reward signal.
[0018] Preferably, the step of extracting features from multiple heterogeneous data sources includes: By employing a cross-attention fusion network, data from at least two different modalities are jointly analyzed to infer key process parameters in the steelmaking process. The inferred key process parameters are stored as features in the dynamic memory module.
[0019] Preferably, the step of storing persistent domain knowledge in a structured manner in a static memory module includes: Historical production data, expert experience, or process specifications are stored as knowledge triples containing conditions, operations, and effects, and these knowledge triples are version-managed.
[0020] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention establishes a hierarchical memory system and a dynamic memory module to ensure rapid response to real-time operating conditions. At the same time, by setting up a closed-loop optimization module based on reinforcement learning, the system can continuously optimize itself according to actual production results, adapt to dynamic changes such as steel grade switching and equipment aging, and solve the problems of slow response and rigid performance of traditional systems.
[0021] 2. By setting up a static memory module, this invention can structure and persistently store scattered expert experience, historical data and other tacit knowledge, effectively addressing the risks of knowledge silos and experience loss, and realizing the inheritance and efficient utilization of core process knowledge.
[0022] 3. By setting up a multimodal data fusion module, this invention associates information such as text, images, and time-series data in a unified semantic space, enabling intelligent agents to understand the production status more comprehensively and deeply, thereby improving the accuracy of decision-making.
[0023] 4. Through the memory fusion module, the system's decision-making can be traced back to whether it is based on real-time data or historical successful cases. Combined with the knowledge accumulation mechanism of reinforcement learning, the decision-making reasoning chain can be clearly displayed, which significantly enhances the operator's trust in the system. Attached Figure Description
[0024] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart illustrating an adaptive memory management method for an intelligent agent in the steelmaking field, provided as an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an adaptive memory management system for an intelligent agent in the steelmaking field, provided by an embodiment of the present invention. Figure 3 This is a schematic diagram of an adaptive memory fusion strategy for the process stage in one embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the calculation of a multi-objective reward function in one embodiment of the present invention; Figure 5 This is a diagram of a cross-modal fusion network structure based on a cross-attention mechanism in one embodiment of the present invention.
[0025] Explanation of reference numerals in the attached figures: Detailed Implementation
[0026] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0027] Example 1 This invention provides an intelligent agent adaptive memory management system and its implementation method for use in the steelmaking field, aiming to achieve real-time perception, long-term knowledge accumulation, adaptive decision-making, and closed-loop optimization of the steelmaking process.
[0028] Please see Figure 2 This is a schematic diagram of an adaptive memory management system provided in an embodiment of the present invention. The system can be deployed on computing devices such as servers, cloud computing platforms, or edge computing devices. At the hardware level, the system may include a processor, memory, and a network interface for data communication. At the logical function level, the system includes a hierarchical memory system composed of a dynamic memory module 30 and a static memory module 40. Furthermore, the system also includes a multimodal data fusion module 50, a memory fusion module 60, and a reinforcement learning-based closed-loop optimization module 70. These modules work collaboratively to achieve intelligent management of the steelmaking process.
[0029] Specifically, refer to Figure 1 The flowchart shown illustrates that the system's workflow can form a closed loop, starting from the input of unprocessed raw data and ending with the optimized decision action. The system updates itself based on the feedback of the execution results, thereby achieving continuous learning and evolution.
[0030] In the data access step S10, the system obtains heterogeneous data from multiple data source layers 10 during the steelmaking process through preset interfaces. These data sources may include, but are not limited to: process databases, such as manufacturing execution systems or secondary automation systems, which provide structured time-series data, such as molten steel temperature, composition analysis results, oxygen flow rate, blowing time, converter angle, material addition amount, etc., and the data format is usually relational data tables or time-series database records; image sensors, such as industrial cameras installed above the furnace mouth or ladle, which provide real-time video streams or image sequences to capture visual information such as furnace mouth flame shape, color, and slag state; and text knowledge bases, which include unstructured or semi-structured text data such as electronic process specification documents, historical production scheduling logs, expert operation manuals, and accident analysis reports.
[0031] Subsequently, in the data cleaning step S20 and the knowledge modeling step S30, the data acquisition and modeling module located in the data processing layer 20 preprocesses and structures the incoming raw data. The data cleaning step S20 is used to improve data quality, and specific operations may include: for time-series sensor data, using Lagrange interpolation or spline interpolation to fill missing values, and using the 3-sigma principle or local outlier algorithm to remove outliers; for image data, performing preprocessing such as denoising, contrast enhancement, and region of interest cropping; and for text data, performing natural language processing operations such as word segmentation, stop word removal, and part-of-speech tagging.
[0032] In the knowledge modeling step S30, the system performs in-depth processing and organization on the cleaned data to facilitate subsequent retrieval and utilization. In one embodiment of the present invention, knowledge modeling is mainly achieved in two ways. The first is to construct a domain knowledge graph. The system identifies the core entities in the steelmaking field (such as "converter", "ladle", "molten iron", "scrap steel", "oxygen lance") and the relationships between them (such as "contains", "transported to", "composes of"), and stores these entities and relationships in the form of a graph structure. For example, a database table structure such as CREATE TABLE knowledge_graph (subject VARCHAR(255), predicateVARCHAR(255), object VARCHAR(255)) can be established to represent the triples of the knowledge graph, thereby explicitly expressing the complex relationships between equipment, materials, and processes. The second is to construct a unified semantic vector space. The system uses a pre-trained language model based on the Transformer architecture (for example, a model fine-tuned on a large number of metallurgical literatures) as a unified encoder to encode all modal data, including text, image features, and time-series data features, into high-dimensional semantic vectors. For example, all data can be uniformly mapped to a 1024-dimensional vector space. These vectors and their metadata (such as timestamps, data sources, and original values) are stored in a vector database that supports efficient similarity search. In this embodiment, a vector database architecture based on the FAISS library can be adopted, and its inverted file index and other technologies can be used to accelerate retrieval. At the same time, a timestamp index is established for all vector data to facilitate time-based range queries.
[0033] In the memory category determination step S40, the modeled data is fed to different memory modules according to its properties. This step can be executed based on preset rules. For example, sensor data streams and image streams with real-time timestamps are judged as dynamic information and enter the short-term memory processing step S50; while knowledge with lasting value extracted from historical databases or text knowledge bases is judged as static knowledge and enters the long-term memory processing step S60.
[0034] The short-term memory processing step S50 is executed by the dynamic memory module 30. The function of the dynamic memory module 30 is to cache and process real-time streaming data from the steelmaking process to support rapid response to the current operating conditions and context tracking. In this embodiment, the dynamic memory module 30 can be implemented as a sliding window cache with a fixed capacity. For example, the cache only retains all real-time sensor data and image extraction features from the most recent hour. This design ensures that system decisions are based on the latest on-site situation. To track context in multi-round interactions with the operator, the system generates a unique session ID for each new dialogue or query and associates all real-time data related to that round with that ID, thereby constructing a temporary, task-specific context. The dynamic memory module 30 also includes an intelligent forgetting mechanism. This mechanism is not a simple first-in-first-out (FIFO) approach but is controlled by a reinforcement learning scheduler. This scheduler can be a small neural network that calculates an importance score for all data items in the cache at preset time intervals (e.g., 10 minutes). The calculation of this score can comprehensively consider multiple factors: the time decay of the data, the semantic relevance of the data to the current or recent query, and the contribution of the data to past decisions. Accordingly, the system will discard the lowest-scoring data to make room for new real-time data.
[0035] The long-term memory processing step S60 is executed by the static memory module 40. This static memory module 40 is used to structurally store persistent domain knowledge, such as historical production data, expert experience, and process specifications, to achieve long-term accumulation and inheritance of knowledge. In this embodiment, the static memory module 40 organizes this knowledge into a standardized "knowledge triplet" format, namely {condition, operation, effect}. For example, a historically successful process adjustment record can be represented as: {condition: "Steel temperature is above 1650℃, [C] content is 0.05%", operation: "Add 50kg of cooling scrap steel", effect: "Temperature drops by 10℃ after 5 minutes, [C] content shows no significant change"}. The text in these triplets is converted into semantic vectors by the aforementioned unified encoder and stored in the vector database along with metadata such as the original text, source, and time. To track the evolution of knowledge and adapt to changes in operating conditions, the static memory module 40 also implements version management for these knowledge triplets. Whenever a new, validated triplet is stored, the system assigns it a new version number. If the triple represents a modification of existing knowledge, the system establishes a link between the old and new versions. This approach not only preserves the evolutionary path of knowledge but also allows the system to revert to valid knowledge from a specific historical period when needed.
[0036] The multimodal memory processing step S70 is executed by the multimodal data fusion module 50. This module is connected to the hierarchical memory system and is responsible for extracting deep features from various heterogeneous data sources and mapping these features to the aforementioned unified semantic vector space. Specifically, for text data, the system directly uses a large domain model for encoding; for slag images, the system uses a pre-trained convolutional neural network as a feature extractor to convert the image into a feature vector; for time-series data such as temperature curves and composition change curves, the system uses a long short-term memory network model to capture their dynamic change trends and extract their final hidden states as feature vectors. All these feature vectors from different modalities are ultimately projected into a shared semantic space through a linear mapping layer, thereby breaking down the barriers between different data types and realizing the association and understanding of cross-modal information. The real-time features generated after the above fusion processing will be fed to the dynamic memory module 30.
[0037] In the memory fusion step S80, the memory fusion module 60 plays a crucial role. This module is connected to both the dynamic memory module 30 and the static memory module 40. When the system receives an external request, the memory fusion module 60 is responsible for retrieving relevant information from the two memory modules and integrating them based on a preset strategy to generate unified context information. In a preferred embodiment, the preset strategy specifically includes a dynamic weight adjustment strategy based on the process stage. Since the steelmaking process (taking converter steelmaking as an example) has significantly different reliance on knowledge at different smelting stages, the system identifies the current process stage by real-time monitoring of key process parameters (such as oxygen blowing time, exhaust gas composition, furnace vibration, etc.) and dynamically adjusts the fusion weights "α" and β (where "α" + β = 1) of the dynamic memory (short-term) and static memory (long-term) accordingly.
[0038] The specific strategy and its reference binding relationship with the steelmaking process are as follows: 1. Early stage of refining: Focus on dynamic perception Process characteristics: The reaction is violent, prone to splashing, and the furnace condition fluctuates greatly.
[0039] Fusion strategy: The system automatically increases the dynamic memory weight α (e.g., "α" = 0.8, β = 0.2).
[0040] Technical Logic: The primary tasks at this stage are safe production and prevention of splashing. The system mainly relies on real-time sonar data, vibration data, and image stream characteristics in the dynamic memory module 30 to monitor splashing signs with a millisecond-level response speed. Static historical operation records have low reference value at this time because the differences in the initial conditions of each batch of raw materials have the greatest impact in the early stages of the reaction.
[0041] 2. Mid-stage of refining: Dynamic and static equilibrium Process characteristics: The reaction is relatively stable, mainly involving decarbonization and temperature increase.
[0042] Fusion strategy: The system is adjusted to a balanced weight (e.g., α=0.5, β=0.5).
[0043] Technical Logic: The system combines the real-time decarbonization rate (calculated through exhaust gas analysis) stored in dynamic memory with the temperature rise curves of similar furnaces at the same time point stored in static memory. If the real-time decarbonization rate deviates significantly from the "standard curve" in static memory, the system will trigger an anomaly warning.
[0044] 3. Late stage of refining: Emphasis on static experience Process characteristics: Requires precise control of the final carbon temperature, has a low tolerance for error, and requires a "one-shot hit".
[0045] Fusion strategy: The system automatically increases the static memory weight β (e.g., α=0.3, β=0.7).
[0046] Technical Logic: At this point, relying solely on real-time sensors (which may drift or lag) is often insufficient to accurately determine the endpoint. The system heavily relies on the "historical golden furnace" data and expert experience database (such as the "post-blowing empirical formula") in the static memory module 40. The memory fusion module retrieves the most similar successful cases in history to the current furnace conditions (weight, temperature, oxygen consumption) and uses historical experience to correct the current endpoint prediction value.
[0047] 4. Physical constraint verification in conflict resolution In the aforementioned conflict resolution mechanism, a metallurgical principle constraint layer is introduced. When a conclusion derived from dynamic memory (e.g., judging extremely high temperatures based on images) severely conflicts with static memory (temperature range calculated based on material balance): Strategy: Instead of simply comparing confidence levels, the system introduces physicochemical constants (such as theoretical maximum temperature rise rate and maximum decarbonization rate limit) from the static memory module 40 as hard constraints.
[0048] Execution: If the inference from the dynamic data violates physical constraints (e.g., the rate of temperature rise exceeds the theoretical physical limit), the system will determine that the dynamic sensor data may be "illusion" or faulty, force the adoption of a conservative operating strategy from the static memory, and issue a sensor check alarm.
[0049] In this embodiment, the workflow is as follows: First, a request is received, such as "How to quickly reduce the temperature of molten steel under the current furnace conditions?". Then, semantic retrieval is performed in the dynamic memory module 30 to find the most relevant real-time data. Simultaneously, semantic retrieval is performed in the static memory module 40 to find historical successful cases related to "rapid cooling." The system evaluates the semantic similarity between the dynamic memory retrieval results and the query. If the similarity is lower than a preset threshold, it means that real-time data alone may not be sufficient to answer the question. In this case, the system assigns higher weight to the static memory retrieval results. Furthermore, the system employs a two-layer attention mechanism to determine and weight the relevance of the retrieved information. The first layer of attention is calculated within each memory source to determine which specific data points or knowledge fragments are most relevant to the query. The second layer of attention is calculated between dynamic and static memory information to evaluate their relative importance in answering the current question. In this embodiment, the preset strategy can use fixed weights for fusion, for example, adding the weighted vector representations of dynamic and static memory information to form a unified context vector. In addition, this module also includes a conflict resolution mechanism. If the conclusions inferred from dynamic memory conflict with the cases retrieved from static memory, the system will compare the confidence scores of the two and prioritize the one with higher confidence.
[0050] Finally, in reinforcement learning optimization step S90, the reinforcement learning-based closed-loop optimization module 70 receives unified context information generated by the memory fusion module 60 and completes decision-making and self-optimization. The core of this module is a decision network employing a proximal policy optimization algorithm. In the decision generation phase, the fused unified context vector is input as the state into the decision network, which outputs a probability distribution of an action. Based on this, the system generates a specific decision output 80, such as "Recommendation: Add 50 kg of grade XX cooled scrap steel to the furnace." In the reward calculation phase, after the decision action is executed, the system tracks its actual effect and calculates a scalar reward signal. For example, if the temperature decrease rate matches the expectation and does not cause other negative effects, a positive reward is given; otherwise, a negative reward is given. In the policy update phase, the calculated reward signal is used to update multiple policies in the system, primarily by updating the parameters of the decision network through backpropagation, so that when encountering similar states in the future, the system is more inclined to generate actions that bring high rewards. Simultaneously, this reward signal can also be used to adjust the fusion weights in the memory fusion module 60, or to update the importance scoring model of the forgetting mechanism in the dynamic memory module 30. Through this continuous "state-action-reward-update" cycle, the entire system achieves adaptive optimization.
[0051] Through the collaborative work of the above modules, the system in this embodiment can effectively integrate real-time data and historical knowledge, make decisions that balance rapid response and reliability, and learn from the consequences of each decision to continuously improve its performance in complex steelmaking environments.
[0052] Example 2 This embodiment is a preferred variant of Embodiment 1, with its main improvement lying in the "preset strategy" of the memory fusion module 60. Compared to the fixed fusion weights used in Embodiment 1, this embodiment further considers the stage-specific characteristics of the steelmaking process. It is understood that the steelmaking process is typically divided into different stages such as the early, middle, and late blowing stages, and the tapping stage. At different stages, the types of knowledge relied upon for decision-making also differ. The core idea of this embodiment is to enable the memory fusion strategy to adapt to the current process stage.
[0053] Please see Figure 3 The figure illustrates in detail the adaptive memory fusion strategy for the process stage in this embodiment. Two key sub-modules have been added to the memory fusion module 60: the process stage identification sub-module 61 and the dynamic weight adjuster 62.
[0054] The function of the process stage identification submodule 61 is to determine the specific stage of the current steelmaking process in real time. The input to this submodule is real-time time-series data from the dynamic memory module 30, such as oxygen blowing time, converter tilt angle, oxygen lance position, real-time oxygen flow rate, and exhaust gas composition detected by the furnace gas analyzer. As an optional implementation, the process stage identification submodule 61 can employ a combination of rule-based and machine learning methods. The rule-based part can predefine clear stage division criteria, such as: "When the oxygen blowing time is less than 5 minutes and the converter angle is vertical, it is judged as 'early stage of oxygen blowing'"; the machine learning part can train a classification model (such as a gradient boosting decision tree or a small recurrent neural network), taking the real-time data sequence as input and directly outputting the classification result of the current stage.
[0055] The dynamic weight adjuster 62 receives the output from the process stage identification submodule 61 and dynamically adjusts the dynamic weight α and static weight β used for information fusion accordingly. Internally, this adjuster can be implemented as a predefined strategy library, which stores the optimal weight combinations corresponding to different process stages in key-value pairs. These weight combinations can be pre-set by domain experts or obtained through offline reinforcement learning training. A specific example of a strategy library is as follows: When the process stage is "early blowing": dynamic weight α = 0.8, static weight β = 0.2. The principle is that the early blowing stage is mainly a period of intense oxidation of elements such as silicon and manganese, with complex and rapidly changing reactions within the furnace. At this time, decisions should rely more on the latest sensor data provided by the dynamic memory module 30. When the process stage is "mid-blowing": dynamic weight α = 0.6, static weight β = 0.4. This stage is the main decarburization period, with relatively stable operating conditions, but real-time changes still need to be monitored, and operational experience from similar historical furnace runs can be referenced. When the process stage is "late blowing stage": dynamic weight α = 0.4, static weight β = 0.6. This stage is close to the end of smelting, and the control precision requirements for steel temperature and composition are extremely high. Decisions should rely more on historical successful cases stored in the static memory module 40 to achieve accurate endpoint targeting. When the process stage is "tapping stage": dynamic weight α = 0.9, static weight β = 0.1. This stage mainly focuses on alloying operations and temperature control, with a short operating window and the highest real-time requirements.
[0056] For example, suppose the agent is currently in the "late stage of blowing" and needs to decide whether to add supplementary blowing or a carbon raiser. The process stage identification submodule 61 determines that the current stage is "late stage of blowing" based on real-time furnace gas analysis data and blowing time. This determination is sent to the dynamic weight adjuster 62, which retrieves the corresponding weights from the strategy library and sets the dynamic weight α to 0.4 and the static weight β to 0.6. Subsequently, the memory fusion module 60, when integrating real-time temperature information from the dynamic memory module 30 and historical endpoint control success cases from the static memory module 40, assigns higher weights to the latter. Therefore, the decisions generated by the system will be more inclined to reproduce operations that successfully hit the endpoint component under similar conditions in the past, thereby significantly improving the accuracy of endpoint control.
[0057] In this way, this embodiment deeply couples the memory management strategy with the physical and chemical processes of steelmaking, making the agent's decision-making logic more consistent with domain knowledge and achieving more refined and intelligent decision support.
[0058] Example 3 This embodiment is another preferred variant of Embodiment 1, with its main improvement lying in the calculation method of the reward signal in the reinforcement learning-based closed-loop optimization module 70. It is understood that the optimization objectives of steelmaking production are usually multi-dimensional, often requiring trade-offs between multiple objectives such as product quality, production cost, smelting efficiency, and production safety. A single reward indicator is insufficient to guide the agent to learn such complex, trade-off-based comprehensive strategies. The core idea of this embodiment is to design a multi-objective reward function tailored to the steelmaking scenario.
[0059] Please see Figure 4 The figure illustrates in detail the calculation of the multi-objective reward function in this embodiment. After a complete smelting batch is completed, the system collects multi-dimensional production indicators related to decision-making actions during that period and uses them to calculate a comprehensive reward signal R. This comprehensive reward signal R is defined as a weighted combination function:
[0060] in, These are the weighting coefficients for quality, cost, and efficiency, which are hyperparameters that can be adjusted by the user according to the current production priorities, and satisfy... .
[0061] The calculation methods for each sub-reward item are as follows: Quality Reward Calculation 71: This unit is responsible for calculating quality rewards. The inputs are the key quality indicators of the product after the decision is implemented, typically the deviations between the actual and target values of the final steel temperature and key chemical composition. To encourage precise hits, a Gaussian function inversely proportional to the error can be used to calculate the reward. For example, for temperature, the reward can be calculated as follows: in, This is the actual endpoint temperature. It is the target temperature. It is a parameter representing the acceptable temperature deviation range. Carbon content bonus. A similar calculation can be performed. The final quality reward can be the product or weighted sum of the two: .
[0062] Cost Incentive Calculation 72: This unit is responsible for calculating cost incentives. The inputs are cost indicators relevant to the decision, such as the consumption and yield of auxiliary materials like alloys, coolants, and slag-forming materials. The reward function should be designed to encourage achieving the same effect with less consumption. For example, for a certain alloy, the cost reward could be positively correlated with the yield and negatively correlated with the absolute consumption. in, and This refers to the actual yield and consumption of this operation. and This represents the historical average or standard value.
[0063] Efficiency Bonus Calculation 73: This unit is responsible for calculating efficiency bonuses. Its inputs are efficiency indicators related to decision-making, such as the time spent completing a specific task, the duration of the entire smelting cycle, and electricity consumption per ton of steel. The reward function should be designed to encourage minimizing time and reducing energy consumption while ensuring quality and cost. For example, the efficiency reward for the smelting cycle could be designed as follows: in, For the standard smelting cycle, The actual period. Positive rewards are only awarded when the actual period is shorter than the standard period.
[0064] Weighted Summation Unit 74: This unit calculates the outputs of the above-mentioned sub-reward calculation units. The weighted summation yields the final comprehensive reward signal R, which is then used in the policy update process of the reinforcement learning-based closed-loop optimization module 70.
[0065] For example, suppose an agent performs a series of decision-making actions in a smelting batch. After the batch ends, the system collects data: the final temperature is 2°C higher than the target, the carbon content is perfectly matched, alloy consumption is 5% lower than the standard, but the smelting time is 3 minutes longer than the standard. The reward calculation units within the system calculate the following: It is a positive value close to 1 but slightly smaller. It is a significantly positive value, and It is a negative value. The weighted summation unit 74 calculates the final comprehensive reward signal R based on the set weights (for example, if the current factory emphasizes cost reduction, w_c is set relatively high). If R is positive, the reinforcement learning algorithm will increase the probability of making this series of decisions; if it is negative, it will decrease the corresponding probability.
[0066] By designing this multi-objective reward function, this embodiment enables the agent to learn complex strategies for balancing conflicting production objectives. Its decision-making behavior is no longer about pursuing the optimality of a single indicator, but rather is closer to the complex real-world needs of achieving the overall optimality in steelmaking production.
[0067] Example 4 This embodiment is a further preferred variant of Embodiment 1, with its main improvement lying in the internal implementation of the multimodal data fusion module 50. While the multimodal fusion method in Embodiment 1 achieves semantic alignment, it may overlook the deep-seated, intrinsic physical relationships between different data modalities. For example, in the converter steelmaking process, there is a close physicochemical correspondence between the shape, color, and brightness of the furnace flame and the decarburization reaction rate within the furnace, and the decarburization reaction rate is directly reflected in the changes in CO and CO2 concentrations detected by the exhaust gas analyzer. The core idea of this embodiment is to design a specific fusion network structure that utilizes this intrinsic physical relationship to achieve deeper and more accurate state perception.
[0068] Please see Figure 5 The figure shows in detail the cross-modal fusion network structure based on the cross-attention mechanism in this embodiment. This network is used to jointly analyze furnace flame images and exhaust gas concentration time-series data to jointly infer the key process parameter that is difficult to measure accurately by a single sensor—the "decarbonization rate".
[0069] The network structure mainly includes the following parts: CNN Feature Extractor 51: A convolutional neural network branch specifically designed to process the input flame image. In this embodiment, a lightweight convolutional neural network pre-trained on a general image dataset and fine-tuned using flame image data from the steelmaking field can be employed. For each frame of the input flame image, CNN Feature Extractor 51 outputs a vector representing the high-level visual features of that image. .
[0070] LSTM Feature Extractor 52: A Long Short-Term Memory (LSTM) branch specifically designed for processing input time-series exhaust gas concentration data. This network receives CO and CO2 concentration sequences within a time window as input and outputs a vector that summarizes the dynamic characteristics of that time series. .
[0071] Cross-attention fusion module 53: This is the core of achieving deep fusion. It receives feature vectors from two feature extractors. and As input, it internally implements a cross-attention mechanism. Specifically, this module performs two attention calculations: in the first calculation, the image feature vector is... As a query, the time series feature vector As keys and values, this is equivalent to making visual information "focus" on the most relevant temporal dynamics; in the second calculation, the reverse operation is performed, transforming the temporal feature vector... As a query, the image feature vector As keys and values, this is equivalent to directing temporal dynamic information to "focus" on the most relevant visual features of the flame. The mathematical expressions for both attention calculations are: The results of the two calculations can be concatenated or added to generate a deeply fused feature representation. .
[0072] MLP Predictor 54: A multilayer perceptron network. It receives the fused feature vectors. As input, and through several fully connected layers, the final output is a scalar value, which is an accurate prediction of the current "decarbonization rate".
[0073] For example, during the blowing process, the system collects real-time video streams of the flame at the furnace opening and exhaust gas analysis data from the flue. A cross-attention fusion network within the multimodal data fusion module 50 continuously processes these two data streams. When the decarburization reaction in the furnace is intense, the flame becomes bright white and towering, while the CO concentration in the exhaust gas rises rapidly. The CNN feature extractor 51 captures the visual feature of the bright white and towering flame, while the LSTM feature extractor 52 captures the dynamic feature of the rapidly rising CO concentration. The cross-attention fusion module 53 discovers the high correlation between these two features and fuses them into a strong representation. Based on this representation, the MLP predictor 54 outputs a high predicted decarburization rate. This decarburization rate value inferred from deep fusion of multimodal data, as a high-quality, robust real-time feature, is immediately stored in the dynamic memory module 30. When the agent needs to make a decision to adjust the oxygen blowing rate, it can directly retrieve this accurate decarburization rate from the dynamic memory, thereby making more stable, timely, and realistic oxygen blowing control decisions.
[0074] Through this deep fusion approach based on physical association, this embodiment significantly improves the system's accuracy and robustness in perceiving key but difficult-to-measure process parameters, providing higher-quality decision-making input information for the intelligent agent.
[0075] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0076] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0077] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features of the present invention can be arbitrarily combined with each other.
Claims
1. An adaptive memory management system for intelligent agents in the steelmaking field, characterized in that, include: Hierarchical memory system: A dynamic memory module used to process real-time streaming data from the steelmaking process; And a static memory module for structured storage of persistent domain knowledge; Multimodal data fusion module: connected to the hierarchical memory system, used to extract features from various heterogeneous data sources including text, images, and time-series sensor data during the steelmaking process, map the features to a unified semantic vector space, and store the mapped features in the dynamic memory module; Memory fusion module: Connected to the dynamic memory module and the static memory module respectively, it is used to retrieve relevant information from the dynamic memory module and the static memory module according to the external input request, and integrate the retrieved information based on a preset strategy to generate unified context information; The closed-loop optimization module based on reinforcement learning is connected to the memory fusion module and generates decision actions based on the unified context information. Based on the actual execution results of the decision-making action in the steelmaking process, a reward signal is calculated; and the reward signal is used to update the system's decision generation strategy and memory management strategy.
2. The adaptive memory management system for intelligent agents in the steelmaking field according to claim 1, characterized in that, The preset strategies upon which the memory fusion module is based include: Identify the current technological stage of the steelmaking process; And, based on the process stage, adaptively adjust the fusion weights of the information retrieved from the dynamic memory module and the static memory module.
3. The adaptive memory management system for intelligent agents in the steelmaking field according to claim 1, characterized in that, The method by which the reinforcement learning-based closed-loop optimization module calculates the reward signal includes: Based on the product quality indicators, cost indicators, and efficiency indicators after the decision-making action is executed, a multi-objective comprehensive reward value is calculated as the reward signal.
4. The adaptive memory management system for intelligent agents in the steelmaking field according to claim 1, characterized in that, The multimodal data fusion module employs a cross-attention fusion network to jointly analyze data from at least two different modalities, thereby inferring key process parameters in the steelmaking process. The multimodal data fusion module is also used to store the inferred key process parameters as features in the dynamic memory module.
5. The adaptive memory management system for intelligent agents in the steelmaking field according to claim 1, characterized in that, The static memory module is used to store historical production data, expert experience, or process procedures as knowledge triples containing conditions, operations, and effects, and to manage the knowledge triples in a versioned manner.
6. An adaptive memory management method for intelligent agents in the steelmaking field, characterized in that, include: Features are extracted from multiple heterogeneous data sources in the steelmaking process, and these features are mapped to a unified semantic vector space. The real-time streaming data of the steelmaking process and the mapped features are stored in the dynamic memory module; It also stores persistent domain knowledge in a structured manner in a static memory module; according to external input requests, it retrieves relevant information from the dynamic memory module and the static memory module, and integrates the retrieved information based on a preset strategy to generate unified context information; Based on the unified context information, a decision action is generated; Calculate the reward signal based on the actual execution results of the decision-making action in the steelmaking process; And the decision generation strategy and memory management strategy are updated using the reward signal.
7. The adaptive memory management method for intelligent agents in the steelmaking field according to claim 6, characterized in that, The step of integrating the retrieved information based on a preset strategy includes: Identify the current technological stage of the steelmaking process; And, based on the process stage, adaptively adjust the fusion weights of the information retrieved from the dynamic memory module and the static memory module.
8. The adaptive memory management method for intelligent agents in the steelmaking field according to claim 6, characterized in that, The steps for calculating the reward signal include: Based on the product quality indicators, cost indicators, and efficiency indicators after the decision-making action is executed, a multi-objective comprehensive reward value is calculated as the reward signal.
9. The adaptive memory management method for intelligent agents in the steelmaking field according to claim 6, characterized in that, The steps for extracting features from multiple heterogeneous data sources include: By employing a cross-attention fusion network, data from at least two different modalities are jointly analyzed to infer key process parameters in the steelmaking process. The inferred key process parameters are stored as features in the dynamic memory module.
10. The adaptive memory management method for intelligent agents in the steelmaking field according to claim 6, characterized in that, The step of storing persistent domain knowledge in a structured manner in a static memory module includes: Historical production data, expert experience, or process specifications are stored as knowledge triples containing conditions, operations, and effects, and these knowledge triples are version-managed.
Citation Information
Patent Citations
Large-model-driven intelligent agent scene exploration and memory management method and system
CN117854059A