An aqueous pen ink diluent ratio control method, device and system
By using closed-loop control based on multimodal perception and reinforcement learning decision-making, the consistency and efficiency issues in the production of water-based pen ink thinner were resolved, enabling real-time adaptation and refined control to raw material fluctuations.
Patent Information
- Application Number
- CN202511441308.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-10
AI Technical Summary
The current production process of water-based pen ink thinner relies on manual experience and rigid processes, resulting in poor adaptability to raw material fluctuations, delayed control response, and low batch consistency of the final product.
A multimodal online sensing unit is used to collect spectral and visual data in real time. A state representation vector is generated through a cross-modal fusion neural network, and a reinforcement learning decision agent is used to output refined physical operation commands to achieve closed-loop control.
It enables deep, real-time insights into the mixing process, ensuring high consistency in product performance across batches, reducing material waste, and improving the flexibility and efficiency of the production line.
Smart Images

Figure CN120909138B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of ink preparation equipment, in particular to a water-based pen ink diluent proportioning control method, device and system. BACKGROUND
[0002] As a material widely used in daily writing and office field, the performance of water-based pen ink directly affects the user experience. Among them, the water-based pen ink diluent is one of the key components that determines the final flowability, stability, drying speed and writing smoothness of the ink. The diluent is usually mixed and reacted in a stirred tank according to specific proportions by a variety of solvents, cosolvents, surfactants, pH regulators and stabilizers, etc. Therefore, in the production process of the diluent, accurate control of the proportioning and mixing process of each component is the core link to ensure the stability of the quality of the final product.
[0003] In the existing production practice, the proportioning control process of the water-based pen ink diluent largely depends on the pre-set process procedures and the personal experience of the operators. The production usually follows the standard operating procedures (SOP) and is operated according to the fixed time sequence and the amount of feed. However, this open-loop control method based on fixed parameters is particularly vulnerable when facing various disturbances in actual production. For example, there are inevitable slight differences in the physical and chemical properties of different batches of chemical raw materials such as resins and surfactants. When a fixed process formula is used to process these fluctuating raw materials, it is easy to cause deviations in the key performance indicators such as viscosity and pH value of the final product, thereby causing poor consistency between product batches.
[0004] In order to make up for the shortcomings of this rigid process, experienced technical personnel are often needed to intervene. The operators observe the color, transparency, wall-hanging phenomenon and other macroscopic manifestations of the mixed solution, or take offline sampling detection at irregular times, to subjectively judge the current reaction progress, and accordingly adjust the amount of feed or stirring speed. This method not only introduces strong subjectivity and uncertainty, but also the significant time lag of offline detection itself, so that when the detection result shows an abnormality, the best adjustment opportunity has often been missed, which may lead to the degradation processing or even scrapping of the whole batch of materials.
[0005] Even in some production lines with automatic control, usually only a single online sensor such as an online viscometer or pH meter is used to monitor the mixing process. Although this single-parameter feedback control is an improvement over purely manual operation, the information dimension obtained is extremely limited. The mixing of water-based pen ink diluent is a complex multi-component physicochemical process, and the overall state is the result of the mutual coupling and nonlinear action of multiple parameters. It is difficult to fully and accurately characterize the current real state of the mixed liquid by relying on only one or two isolated physical quantities, and it is also impossible to capture subtle but critical transition characteristics in the state evolution process, so it is impossible to provide sufficient decision basis for truly fine and optimized control.
[0006] Therefore, there is a lack of an intelligent control means in the prior art that can perceive the internal state of the mixing process in real time and comprehensively, and make dynamic and forward-looking decisions based on this, which has become a technical bottleneck restricting the further improvement of the product quality and production efficiency of water-based pen ink diluent. SUMMARY
[0007] In view of the deficiencies of the prior art, the present application provides a water-based pen ink diluent proportioning control method, device and system, which solves the problem that the existing water-based pen ink diluent proportioning process has poor adaptability to raw material fluctuations, delayed control response and low batch consistency of the final product due to reliance on manual experience and rigid processes.
[0008] To achieve the above purpose, the present application is implemented by the following technical solutions: a water-based pen ink diluent proportioning control method, comprising the following steps:
[0009] S1. Real-time perception: at a preset time step, a multi-modal online perception unit deployed on a stirring tank or its circulating pipeline synchronously collects multi-modal raw data of the mixed liquid in the stirring tank, the multi-modal raw data at least including spectral data reflecting molecular information and visual data reflecting macroscopic apparent state;
[0010] S2. State representation: input the multi-modal raw data into a preset cross-modal fusion neural network model for automatically extracting and fusing cross-modal implicit features, and map the multi-modal raw data by the cross-modal fusion neural network model to generate a unified, low-dimensional state representation vector, which is used to comprehensively describe the physicochemical state of the current mixed liquid;
[0011] S3. Intelligent decision-making: input the state representation vector as input to a preset reinforcement learning decision-making agent, and output fine physical operation instructions for guiding the evolution of the mixed liquid state to the target state according to the optimal strategy learned by the reinforcement learning decision-making agent for the current state;
[0012] S4. Closed-loop execution: controlling the actuator connected to the stirred tank to execute the refined physical operation instruction to change the physicochemical state of the mixed solution;
[0013] S5. Loop iteration: repeating steps S1 to S4 until the state characterization vector characterizing the state of the mixed solution meets the preset production termination condition.
[0014] Preferably, the state characterization in step S2 specifically includes:
[0015] Preprocessing the spectral data and visual data;
[0016] Inputting the preprocessed spectral data and visual data into the modality-specific encoder in the cross-modal fusion neural network model to extract high-level features;
[0017] Through the cross-modal fusion in the cross-modal fusion neural network model, the high-level features are subjected to self-attention weighting calculation and fusion to learn cross-modal implicit correlation features and generate a state characterization vector.
[0018] Preferably, the method further comprises:
[0019] After executing the refined physical operation instruction, an instant reward value is calculated according to the change in the physicochemical state of the mixed solution;
[0020] Using the transition sample containing the state characterization vector at the current time, the refined physical operation instruction, the instant reward value, and the state characterization vector at the next time, the optimal policy of the reinforcement learning decision-making agent is updated;
[0021] Wherein, the instant reward value is determined by a progress reward based on the gap between the current physicochemical state of the mixed solution and the final target state and a cost penalty based on the consumption generated by the executed operation.
[0022] Preferably, the progress reward is calculated as:
[0023] Using an auxiliary performance prediction model, the corresponding final product performance is predicted according to the state characterization vector at the current time and the state characterization vector at the next time, respectively;
[0024] The gap between the two final product performances and the preset performance target vector is calculated, and the reduction of the gap is taken as the progress reward.
[0025] Preferably, the multi-modal raw data further includes rheological data and physicochemical data; the rheological data is collected by an online viscometer or a stirring shaft torque sensor; the physicochemical data is collected by a pH electrode, a conductivity meter or a temperature sensor.
[0026] Preferably, an aqueous pen ink diluent proportioning device comprises:
[0027] stirred tank reactor
[0028] a high-precision material conveying unit connected to the stirred tank reactor and configured to convey material to the stirred tank reactor
[0029] a multi-modal online sensing unit configured to collect multi-modal raw data in the stirred tank reactor, the multi-modal raw data comprising at least spectral data and visual data, the multi-modal online sensing unit comprising a spectral analysis module and a machine vision module
[0030] an intelligent control unit electrically connected to the multi-modal online sensing unit and the high-precision material conveying unit, the intelligent control unit being configured to implement a method.
[0031] Preferably, the intelligent control unit specifically comprises:
[0032] an edge computing controller configured to run a cross-modal fusion neural network model and a reinforcement learning decision-making agent to generate refined physical operation instructions
[0033] a programmable logic controller configured to receive the refined physical operation instructions and convert them into bottom-layer control electrical signals for actuators in the high-precision material conveying unit.
[0034] Preferably, the multi-modal online sensing unit further comprises an online rheological / physical parameter module, the online rheological / physical parameter module comprising at least one of an online viscometer, a pH / conductivity electrode, and a temperature sensor.
[0035] Preferably, the intelligent control unit is further configured to:
[0036] determine the uncertainty of the physical and chemical state of the current mixed solution according to the current generated state representation vector
[0037] based on the uncertainty, dynamically adjust the working parameters of the spectral analysis module or the machine vision module in the multi-modal online sensing unit, the working parameters including sampling frequency or resolution.
[0038] An aqueous pen ink diluent proportioning control system, comprising a processor and a memory, the memory storing a computer program, and the processor implementing a method when executing the computer program.
[0039] The present application provides an aqueous pen ink diluent proportioning control method, device and system. It has the following advantages:
[0040] 1、The present application realizes the deep and real-time insight into the physical and chemical state of the mixing process by collecting the spectral data reflecting the molecular information and the visual data of the macroscopic state in real time through the multi-modal online sensing unit, and generating a comprehensive state representation vector using the cross-modal fusion neural network model. Through the state sensing capability, the control system can accurately compensate for the disturbance caused by the batch difference of raw materials or the fluctuation of environmental factors, ensuring the high consistency of the performance of the final product batches.
[0041] 2、The present application adopts a reinforcement learning decision-making agent to replace the traditional fixed program control, which can autonomously decide and output the optimal fine physical operation instruction based on the deep understanding of the current state. It realizes the fundamental change from "open-loop execution" to "closed-loop intelligent guidance", effectively shortens the process debugging time, reduces the material waste and rework caused by unqualified final inspection, and converts the complex adjustment relying on the intuition and experience of experienced technical personnel into an autonomous and data-driven automatic strategy.
[0042] 3、The present application combines deep state representation with the continuous optimization mechanism of reinforcement learning to build a control system that can learn and evolve from production practice. When facing changes such as new product formulations or replacement of raw material suppliers, the system can quickly adapt to new process conditions through online learning without the need for complex reprogramming, and solidify the accumulated process knowledge in the updated model parameters, realizing the digitalization of process knowledge and inheritance, and greatly enhancing the flexibility of the production line. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 is a schematic diagram of the proportioning device structure of the present application;
[0044] Figure 2 is a flow chart of the proportioning control method of the embodiment of the present application;
[0045] Figure 3 is a block diagram of the control system structure of the embodiment of the present application;
[0046] Figure 4 is a schematic diagram of the cross-modal fusion neural network model structure of the embodiment of the present application;
[0047] Figure 5 is a schematic diagram of the reinforcement learning decision-making and training process of the embodiment of the present application;
[0048] Figure 6 is a schematic diagram of the pipeline and instrument connection of the water-based pen ink diluent proportioning device in the embodiment of the present application.
[0049] Wherein, 100, reaction and mixing unit; 110, stirred tank; 120, stirring device; 200, high-precision material conveying unit; 210, raw material tank A; 220, raw material tank B; 300, multi-modal online sensing unit; 310, spectral analysis module; 320, machine vision module; 330, online rheological / physical parameter module; 400, intelligent control unit; 410, edge computing controller; 420, programmable logic controller; 500, water-based pen ink diluent proportioning control system; 510, processor; 520, memory; 530, communication bus; V-201, feed valve one; V-202, feed valve one; V-203, feed valve two; V-204, feed valve two; V-111, electric valve one; V-112, electric valve two; V-113, electric valve three; P-201, metering pump A; P-202, metering pump B; P-110, circulating pump. DETAILED DESCRIPTION
[0050] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the specification of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application. Embodiment 1
[0051] Please refer to the drawings in the specification of the present application Figure 1 and the drawings Figure 6 The present application provides a water-based pen ink diluent proportioning device, which is a physical carrier for realizing the proportioning control method, and the device comprises
[0052] The reaction and mixing unit 100 comprises a stirred tank 110 for containing and mixing various components of the water-based pen ink diluent. The stirred tank 110 is provided with a stirring device 120, which can be but is not limited to a stirring paddle driven by a variable frequency motor, for physically stirring the mixed liquid in the tank;
[0053] The high-precision material conveying unit 200 is connected to the stirred tank 110 for conveying materials thereto, and the high-precision material conveying unit 200 is connected to the stirred tank 110 through a pipeline. Please refer to the drawings in the specification of the present application Figure 6 , the drawings Figure 6A schematic diagram showing the connection of pipelines and instruments in the device, wherein the high-precision material conveying unit 200 includes a plurality of storage tanks for storing different kinds of raw materials (including but not limited to raw material tank A 210 and raw material tank B 220), and high-precision metering pumps and electromagnetic control valves (including but not limited to feed valve one V-201 and feed valve two V-203) connected to each storage tank. The metering pump includes metering pump A P-201 and metering pump B P-202, which can be but are not limited to peristaltic pumps or diaphragm pumps, for accurately conveying a specified volume of raw material from the storage tank to the stirred tank 110;
[0054] A multi-modal online sensing unit 300 for collecting multi-modal raw data in the stirred tank 110, the multi-modal raw data including at least spectral data and visual data, the multi-modal online sensing unit 300 including a spectral analysis module 310 and a machine vision module 320. The spectral analysis module 310 includes an immersion spectral probe mounted on the tank cover of the stirred tank 110, with the sensing end directly in contact with the mixed liquid in the tank. The machine vision module 320 includes an industrial camera and a light source mounted outside the sight window of the stirred tank 110, for collecting images of the mixed liquid.
[0055] In an optional embodiment, the multi-modal online sensing unit 300 further includes an online rheological / physical parameter module 330, which can include at least one of a torque sensor mounted on the drive shaft of the stirring device 120, or a pH electrode, a conductivity meter, and a temperature sensor immersed in the mixed liquid;
[0056] An intelligent control unit 400, which is electrically and communicatively connected to the high-precision material conveying unit 200 and the multi-modal online sensing unit 300. The intelligent control unit 400 is the technical core of the device, which is different from the traditional PLC controller, and has a cross-modal fusion neural network model and a reinforcement learning decision-making agent deployed inside. The intelligent control unit 400 is responsible for receiving and processing multi-modal data, generating state representation vectors, making intelligent decisions, and outputting fine physical operation instructions to control the high-precision material conveying unit 200 and other actuators.
[0057] In a specific embodiment, the intelligent control unit 400 includes an edge computing controller 410 and a programmable logic controller (PLC) 420 in structure.
[0058] Edge computing controller 410, industrial personal computer (IPC), used to perform computationally intensive tasks. It deploys and runs a cross-modal fusion neural network model and reinforcement learning decision-making agent. The edge computing controller 410 receives all raw data from the multi-modal online perception unit 300 through the data interface, performs state representation generation (step S2) and intelligent decision-making (step S3), and finally generates refined physical operation instructions;
[0059] Programmable logic controller 420, programmable logic controller (PLC) 420 communicates with edge computing controller 410 through industrial fieldbus or Ethernet protocol. Its function is to receive the refined physical operation instructions issued by the edge computing controller 410, and convert them into bottom-layer driving electrical signals for each metering pump, electromagnetic valve in the high-precision material conveying unit 200, and the stirring device 120 in the reaction and mixing unit 100. PLC is responsible for performing hardware control tasks with high reliability and high real-time performance.
[0060] In a preferred embodiment, the intelligent control unit 400 adopts a layered architecture, specifically including an edge computing controller 410 for running AI algorithm models, and a programmable logic controller (PLC) 420 for executing bottom-layer hardware control, both of which work together to ensure the intelligence of decision-making and the reliability of execution.
[0061] The multi-modal online perception unit 300 also includes an online rheological / physical parameter module, which includes at least one of an online viscometer, a pH / conductivity electrode, and a temperature sensor.
[0062] In addition, the edge computing controller 410 is also configured to determine the uncertainty of the current mixed liquid physical and chemical state according to the current generated state representation vector. When the uncertainty is higher than a preset threshold, the edge computing controller 410 can generate and send parameter adjustment instructions to the spectral analysis module 310 or the machine vision module 320 in the multi-modal online perception unit 300 to dynamically change its sampling frequency or resolution and other working parameters, so as to achieve more intensive monitoring of the key state transition stage;
[0063] In another preferred embodiment, the intelligent control unit 400 also has the ability of active perception. It can dynamically adjust the working parameters of the spectral analysis module 310 or the machine vision module 320 in the multi-modal online perception unit 300 based on the state uncertainty reflected by the current state representation vector, including sampling frequency or resolution. Example 2
[0064] Please refer to the attached Figure 2 , attached Figure 4 and attached Figure 5This invention provides a method for controlling the mixing ratio of water-based pen ink thinner, comprising the following steps:
[0065] S1. Real-time Sensing: At a preset time step, multimodal online sensing units 300 deployed in the stirred tank 110 or its circulation pipeline synchronously collect multimodal raw data of the mixture within the stirred tank 110. Unlike existing technologies that rely on only a few isolated physicochemical parameters, the multimodal raw data collected in this invention includes at least spectral data reflecting the molecular-level composition and structure of the mixture, as well as visual data reflecting the macroscopic clarity, color, flow morphology, and other apparent states of the mixture. This multidimensional, cross-level data acquisition lays the foundation for a deep understanding of the complex state of the mixing process. The multimodal raw data includes at least spectral data reflecting molecular information and visual data reflecting the macroscopic appearance; it also includes rheological and physicochemical data; rheological data is collected by an online viscometer or a stirring shaft torque sensor; physicochemical data is collected by a pH electrode, conductivity meter, or temperature sensor.
[0066] In step S1, real-time sensing is performed within a preset discrete time step. The process is performed periodically. The sensor modules of the multimodal online sensing unit 300 are physically deployed on the vessel body, lid, or external circulation pipeline of the stirred tank 110 to achieve non-invasive or online continuous monitoring of the mixture.
[0067] In one embodiment, the acquisition of spectral data is achieved through a spectral analysis module 310. The spectral analysis module 310 includes an immersion Raman spectral probe, which is installed via a standard flange interface on the lid of the stirred tank 110, with its sensing end directly immersed in the mixture inside the tank. The immersion Raman spectral probe emits an excitation beam of a preset wavelength into the mixture and collects the Raman scattered light generated by the mixture. The collected scattered light is transmitted to a spectrometer for dispersion and detection, thereby generating spectral data at time t. Spectral data is a one-dimensional vector whose elements are the scattered light intensities corresponding to different Raman shift values. Spectral data directly reflects the molecular vibrational information of each chemical component in the mixture.
[0068] In one embodiment, visual data acquisition is achieved through a machine vision module 320. The machine vision module 320 includes an industrial camera and a programmable light source. The industrial camera is mounted outside a sight glass window on the wall of the stirred tank 110, and its lens focuses on the mixture inside the sight glass. The programmable light source, such as an LED ring light, is arranged coaxially or paraaxially around the camera lens to provide stable and uniform illumination to the camera acquisition area (see attached diagram). Figure 1At time t, the light source is turned on at a preset brightness, and the industrial camera simultaneously exposes and acquires a frame of digital image, generating visual data at time t. This visual data It is a two-dimensional or three-dimensional pixel matrix that records macroscopic appearance information of the mixture within the collection area, such as color, clarity, flow texture, and the presence of bubbles or undissolved particles.
[0069] In an optional embodiment, the multimodal raw data further includes rheological data. and physicochemical data Rheological data can be acquired by a torque sensor mounted on the agitator shaft drive system. This reflects the resistance experienced by the agitator in the mixture. Physicochemical data can be collected by pH electrodes, conductivity electrodes, and PT100 temperature sensors, all immersed in the mixture, corresponding to scalar values of the mixture's acidity / alkalinity, ion concentration, and temperature, respectively.
[0070] To ensure temporal consistency of data across different modalities, all sensors in the multimodal e-mode online sensing unit are connected to a central data acquisition system. The central data acquisition system uses time steps... Based on this, a synchronization trigger signal is sent to each sensor, or a high-precision timestamp is appended to the data packets received from each sensor. In this way, at each time t, the system can obtain a set of time-precisely aligned multimodal raw data. This provides effective data input for subsequent state characterization steps.
[0071] S2. State Representation: Input the multimodal raw data into a pre-defined cross-modal fusion neural network model for automatically extracting and fusing cross-modal latent features. , by neural network model The neural network model maps and generates a unified, low-dimensional state representation vector from the original multimodal data. It's not simply about piecing together data; it's about automatically extracting and fusing deep correlations and subtle, imperceptible features between different data modalities. This is achieved through neural network models. The high-dimensional, heterogeneous original data stream is mapped and a unified, low-dimensional state representation vector is generated. .
[0072] In step S2, the system performs a state representation generation process, converting the multi-modal raw data collected in step S1 into a structured, information-dense state representation vector. The process of generating state representations first involves a series of preprocessing operations on the raw data to eliminate noise, standardize the scale, and make it suitable for use as input to a neural network model.
[0073] In one specific embodiment, spectral data The preprocessing included: applying a Savitzky-Golay filter for smoothing and denoising, and using asymmetric least squares for baseline correction to eliminate fluorescence background interference, resulting in preprocessed spectral data. For visual data The preprocessing includes: scaling the size to a preset fixed size, such as 224x224 pixels, and normalizing the pixel values, for example, scaling them to the range of [-1, 1], to obtain the preprocessed visual data. For other data such as rheological and physicochemical data, moving average filtering and Z-score standardization are applied.
[0074] After preprocessing, all data is input into a pre-defined cross-modal fusion neural network model. Neural network model Structurally, it is designed to include two main parts: a modality-specific encoder and a cross-modality fusion unit.
[0075] The role of a modality-specific encoder is to extract high-level abstract features from the intrinsic structure and characteristics of each data mode. Specifically, for one-dimensional preprocessed spectral data... It uses a one-dimensional convolutional neural network (1D-CNN) as its encoder, extracting feature peaks and combination patterns from the spectrum through multiple convolution and pooling operations. For two-dimensional preprocessed visual data... A two-dimensional convolutional neural network (2D-CNN), such as the backbone of a pre-trained residual network (ResNet), is used as its encoder to extract texture, color distribution, and spatial structure features from the image. Other normalized scalar or vector data is encoded using a multilayer perceptron (MLP). The output of each encoder is a high-level feature vector corresponding to the modality.
[0076] Subsequently, the high-level feature vectors from all modalities are passed to a cross-modal fusion layer. In one embodiment, the cross-modal fusion layer is a Transformer encoder layer based on a self-attention mechanism. These feature vectors from different modalities are treated as a single input sequence. The self-attention mechanism generates a weighted sum for each feature vector by calculating the correlation score between any two feature vectors in the sequence, enabling the model to dynamically and selectively aggregate information from different modalities. This mechanism allows the model to capture, for example, the nonlinear correlation between a specific spectral peak and a particular level of turbidity in an image.
[0077] The output of the Transformer encoder layer is passed through a global average pooling layer, which aggregates the serialized feature information into a single, fixed-dimensional vector. This vector is the final generated state representation vector that comprehensively describes the current physicochemical state of the mixture at time t. The generation process can be formally described by the following expression:
[0078] ;
[0079] in: Let be the state representation vector at time t; A cross-modal fusion neural network model; The spectral data at time t are the preprocessed data. The visual data at time t is the preprocessed data. For the model The learnable network parameters are obtained by pre-training on historical production data or by fine-tuning them online during the control process.
[0080] This state representation vector Used to comprehensively describe the physicochemical state of the current mixture; state characterization specifically includes:
[0081] Preprocessing of spectral and visual data;
[0082] In a preferred embodiment, the cross-modal fusion neural network model includes a modality-specific encoder and a cross-modal fusion unit. The former is used to extract high-level features from spectral, visual, and other data respectively, while the latter performs weighted calculation and fusion of these high-level features through a self-attention mechanism, thereby efficiently learning implicit associations across modalities. Preprocessed spectral and visual data are input into the modality-specific encoder in the cross-modal fusion neural network model to extract high-level features.
[0083] By using the cross-modal fusion unit in the cross-modal fusion neural network model, self-attention weighted calculation and fusion of high-level features are performed to learn cross-modal implicit correlation features and generate state representation vectors.
[0084] S3. Intelligent Decision-Making: Taking the state representation vector as input, receive the state representation vector generated in step S2, which represents the current physicochemical state of the mixture. The system provides a pre-defined reinforcement learning decision-making agent, which models the entire proportioning control process as a Markov decision process. Its goal is to learn an optimal policy to maximize long-term cumulative rewards. Based on this learned optimal policy, the reinforcement learning decision-making agent outputs refined physical operation instructions to guide the mixture state towards the target state, tailored to the current state. The optimal action value function is the basis for its decision-making. Following the Bellman optimal equation: This makes decision-making no longer based on fixed rules, but on a dynamic assessment of the value of future states.
[0085] First, a set of refined physical operation instructions is defined as a preset, discrete action space A. In this embodiment, each action 'a' in action space A corresponds to a specific physical operation that can be executed by the underlying hardware. For example, an action can be defined as an instruction tuple, such as (material number, feed volume, conveying rate) or (stirring speed, duration). By discretizing the continuous control parameters, the complex control problem is transformed into a sequential decision problem.
[0086] The core of a reinforcement learning decision agent is a deep Q-network (DQN), denoted as This network is a deep neural network whose function is to take a state s as input and output the predicted Q-value for each action 'a' in the action space A. The Q-value is an estimate of the expected cumulative reward that can be obtained in the future after performing the corresponding action in that state. The network parameters are... express.
[0087] When the reinforcement learning decision agent receives the state representation vector at time t After that, it will The deep Q-network inside it is passed as input. The network performs a forward propagation computation to generate a corresponding Q-value for each optional action 'a' in the action space A.
[0088] Subsequently, the agent applies the optimal strategy it has learned. Select an action from action space A as the final output refined physics operation instruction. In one specific embodiment, the strategy for Strategy. The specific execution method of the strategy is as follows: The system generates a random number between 0 and 1; if the random number is greater than a preset value... The agent then selects the action 'a' that maximizes the Q-value as its action. This process is called exploitation; if the random number is not greater than... Then the agent randomly selects an action from the action space A as... This process is called exploration.
[0089] Default value It can remain constant throughout the control flow, or it can be designed to be a value that gradually decreases over time or with increasing training steps. In the early stages of the process, a larger value... The value motivates the agent to explore further in order to discover a better sequence of operations; in the later stages of the process, smaller values... The value then prompts the agent to make more use of the learned optimal strategy in order to stably achieve the control objective.
[0090] Ultimately, through Actions selected by the strategy This is the output of step S3, and the instruction will be passed to step S4 for physical execution.
[0091] S4. Closed-loop execution: Controls the actuator connected to the mixing vessel 110 to execute refined physical operation commands. The above steps are repeated in a loop to change the physicochemical state of the mixture until the state of the mixture meets the preset production termination conditions.
[0092] Refined physical operation instructions output from step S3 The instructions are sent to the programmable logic controller (PLC) 420 in the intelligent control unit 400. In one embodiment, the instructions... It is issued by the edge computing controller 410 that runs the reinforcement learning decision agent and sent to the PLC via the internal bus or industrial Ethernet communication protocol.
[0093] The function of a PLC is to process high-level, symbolic instructions. The signals are parsed and converted into precise timing low-level control electrical signals that directly drive the underlying actuators. These actuators are components of the high-precision material conveying unit 200 and the reaction and mixing unit.
[0094] In a specific embodiment, if the instruction Given (material number = A, feed volume = 15.2 mL), and the known flow characteristics of the metering pump for material A, the PLC calculates the operating time required to drive the metering pump to achieve a feed volume of 15.2 mL. Subsequently, within the calculated time period, the PLC outputs a start signal to the relay or frequency converter controlling the metering pump, simultaneously controlling the solenoid valve on the corresponding pipeline to open. After the time period ends, the start signal is canceled and the valve is closed.
[0095] In another embodiment, if the instruction If the stirring speed is 300 rpm, the PLC will output an analog voltage signal or a digital communication message containing specific frequency parameters to the frequency converter controlling the stirring motor, so as to stabilize the stirring paddle speed at the preset value.
[0096] After the actuator performs the physical operation, it directly changes one or more physicochemical properties of the mixture in the stirred tank 110, such as component concentration, temperature, and viscosity. The change in state will be collected by the multimodal online sensing unit 300 in step S1 at the beginning of the next time step, thus forming a closed-loop feedback path for the control method.
[0097] S5. Iteration: Repeat steps S1 to S4 to form a continuous "perception-representation-decision-execution" closed loop until the state of the mixture represented by the state representation vector meets the preset production termination conditions.
[0098] In step S5, the system executes steps S1 to S4 iteratively. The complete sequence of steps S1 to S4 constitutes a discrete-time step... Each control cycle operates independently. After step S4 of each control cycle is completed, the system immediately returns to step S1 and begins multimodal data acquisition for the next cycle, thus forming a continuous closed-loop control process of "perception-representation-decision-execution".
[0099] At the beginning or end of each control cycle, the system determines a preset production termination condition. The purpose of this determination is to ascertain whether the physicochemical state of the current mixture has reached or is sufficiently close to the final production target.
[0100] In one specific embodiment, the termination condition is based on the current state representation vector. It is defined by the difference between the predicted final product performance and the preset performance target vector G. Specifically, the system uses the current state representation vector... Input into the aforementioned auxiliary performance prediction model This yields a predicted performance vector. The system calculates the distance between this predicted performance vector and the performance target vector G, for example, the L2 norm. When this distance is less than a preset threshold... At this point, the system determines that the production termination condition has been met. This condition can be expressed by the following inequality:
[0101] ;
[0102] in: To assist the performance prediction model based on the current state The output predicted performance vector; G is the preset performance target vector; It is a preset positive real number threshold that represents the acceptable error range.
[0103] Once the termination condition is met, the iterative process stops. The system then outputs a production completion signal and can selectively save all state tables, execution instructions, and reward value sequences from the production process for that batch, for subsequent process analysis or offline model optimization. If the termination condition is not met, the system continues to execute the next control cycle.
[0104] In another preferred embodiment, to effectively train the reinforcement learning decision agent, this method further includes a reward calculation and policy update mechanism. After each operation, the system calculates an immediate reward value, which consists of a progress reward and a cost penalty. The progress reward is determined using an auxiliary performance prediction model. To quantify this, the model predicts the final product performance based on the current state, and the progress reward is the reduction in the gap between the predicted performance and the target performance in the current state. This provides dense and explicit guiding signals for the agent's learning. After executing refined physical operation instructions, an immediate reward value is calculated based on the changes in the physicochemical state of the mixture.
[0105] The optimal policy of the reinforcement learning decision agent is updated using a transformation sample that includes the current state representation vector, refined physical operation instructions, immediate reward value, and the state representation vector of the next time step.
[0106] The immediate reward value is determined by a combination of a progress reward based on the difference between the current physicochemical state of the mixture and the final target state, and a cost penalty based on the cost incurred in performing the operation.
[0107] The progress reward is calculated as follows:
[0108] Using an auxiliary performance prediction model, the corresponding final product performance is predicted based on the state representation vector at the current moment and the state representation vector at the next moment.
[0109] Calculate the gap between the performance of the two final products and the preset performance target vector, and use the reduction in the gap as the progress reward. Example 3
[0110] Please see the appendix Figure 3 This invention provides a water-based pen ink thinner mixing control system 500, which can be the intelligent control unit 400 in the aforementioned device embodiments, or an independent server or computer device. The water-based pen ink thinner mixing control system 500 may include at least: a processor 510, a memory 520, and a communication bus 530 for connecting the processor 510 and the memory 520.
[0111] The processor 510 may be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), or one or more other types of processing units.
[0112] Memory 520 may include non-persistent memory, such as random access memory (RAM), and / or persistent memory, such as read-only memory (ROM) or flash memory, hard disk drive (HDD), solid-state drive (SSD). Computer programs are stored in memory 520.
[0113] When the processor 510 is configured to execute a computer program stored in the memory 520, it implements the steps of the water-based pen ink thinner ratio control method as described in any of the foregoing method embodiments.
[0114] In one specific embodiment, the computer program stored in memory 520, when executed by processor 510, can be functionally divided into the following modules:
[0115] Data acquisition and preprocessing module: The data acquisition and preprocessing module is configured to receive synchronous multimodal raw data from the multimodal online sensing unit 300 through a communication interface, and perform preset preprocessing operations on the raw data, such as baseline correction, image enhancement and data standardization.
[0116] State Representation Generation Module: This module contains and executes the logic for the cross-modal fusion neural network model. It receives data processed by the data acquisition and preprocessing module and generates a unified, low-dimensional state representation vector through forward propagation computation of the model.
[0117] Intelligent Decision Module: The intelligent decision module contains and executes the logic of the reinforcement learning decision agent. It receives the state representation vector output by the state representation generation module, performs calculations through its internal deep Q-network, and outputs refined physical operation instructions according to a preset strategy.
[0118] Instruction Execution and Communication Module: The instruction execution and communication module is configured to send refined physical operation instructions generated by the intelligent decision module to one or more programmable logic controllers (PLCs) 420 for execution via an industrial bus or Ethernet interface.
[0119] Model Update Module: The model update module is configured to store transition samples (state, action, reward, next state) generated by system interactions, and calculate and update the network parameters of the deep Q network in the intelligent decision module based on these samples by minimizing a preset loss function.
[0120] Through the coordinated operation of the above modules, the control system of the present invention can fully realize all the functions of the method.
[0121] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An aqueous pen ink diluent formulation control method, characterized by, The method comprises the following steps: S1. Real-time sensing: at a preset time step, synchronously collecting multi-modal raw data of the mixed liquid in the stirred tank (110) by deploying a multi-modal online sensing unit (300) on the stirred tank (110) or its circulating pipeline, wherein the multi-modal raw data at least includes spectral data reflecting molecular information and visual data reflecting macroscopic apparent state, and the multi-modal raw data further includes rheological data and physicochemical data; the rheological data is collected by an online viscometer or a stirring shaft torque sensor; the physicochemical data is collected by a pH electrode, a conductivity meter or a temperature sensor; S2. State representation: inputting the multi-modal raw data into a preset cross-modal fusion neural network model for automatically extracting and fusing cross-modal implicit features, and mapping the multi-modal raw data by the cross-modal fusion neural network model to generate a unified low-dimensional state representation vector, wherein the state representation vector is used for comprehensively describing the physical and chemical state of the current mixed liquid, and the state representation specifically includes: preprocessing the spectral data and the visual data; inputting the preprocessed spectral data and visual data into the modal-specific encoder in the cross-modal fusion neural network model to extract high-level features; performing self-attention weighting calculation and fusion on the high-level features by the cross-modal fusion unit in the cross-modal fusion neural network model to learn cross-modal implicit correlation features and generate the state representation vector; S3. Intelligent decision-making: inputting the state representation vector as an input into a preset reinforcement learning decision-making agent, and outputting a refined physical operation instruction for guiding the evolution of the mixed liquid state to a target state from the reinforcement learning decision-making agent according to the optimal strategy learned by the reinforcement learning decision-making agent; S4. Closed-loop execution: controlling an execution mechanism connected to the stirred tank (110) to execute the refined physical operation instruction to change the physical and chemical state of the mixed liquid; S5. Iterative cycle: repeating steps S1 to S4 until the state of the mixed liquid represented by the state representation vector meets a preset production termination condition.
2. The water-based pen ink diluent ratio control method according to claim 1, characterized by, The method further comprises: calculating an immediate reward value according to the change in the physical and chemical state of the mixed liquid after executing the refined physical operation instruction; updating the optimal strategy of the reinforcement learning decision-making agent by using a transition sample comprising the state representation vector at the current time, the refined physical operation instruction, the immediate reward value and the state representation vector at the next time; wherein the immediate reward value is determined by a progress reward based on the gap between the current physical and chemical state of the mixed liquid and the final target state and a cost penalty based on the consumption generated by the executed operation.
3. The water-based pen ink diluent ratio control method according to claim 2, characterized by, The calculation method of the progress reward is: predicting the final product performance according to the state representation vector at the current time and the state representation vector at the next time, respectively, by using an auxiliary performance prediction model; calculating the gap between the final product performance and a preset performance target vector, and taking the reduction of the gap as the progress reward.
4. An aqueous pen ink diluent proportioning device applied to the aqueous pen ink diluent proportioning control method of any one of claims 1-3, characterized in that, The method comprises: a stirred tank (110); A high-precision material conveying unit (200) connected with the stirred tank (110) and configured to convey material to the stirred tank (110); A multi-modal online sensing unit (300) configured to collect multi-modal raw data in the stirred tank (110), the multi-modal raw data comprising at least spectral data and visual data, the multi-modal online sensing unit (300) comprising a spectral analysis module (310) and a machine vision module (320); An intelligent control unit (400) electrically connected with the multi-modal online sensing unit (300) and the high-precision material conveying unit (200), the intelligent control unit (400) being configured to implement the method of claim 1.
5. The aqueous pen ink diluent compounding device according to claim 4, wherein The intelligent control unit (400) specifically comprises: An edge computing controller (410) configured to run a cross-modal fusion neural network model and a reinforcement learning decision-making agent to generate refined physical operation instructions; A programmable logic controller (420) configured to receive the refined physical operation instructions and convert the refined physical operation instructions into bottom-layer control electrical signals for actuators in the high-precision material conveying unit (200).
6. The aqueous pen ink diluent compounding device according to claim 4, wherein The multi-modal online sensing unit (300) further comprises an online rheological / physical parameter module (330), the online rheological / physical parameter module (330) comprising at least one of an online viscometer, a pH / conductivity electrode, and a temperature sensor.
7. The aqueous pen ink diluent compounding device according to claim 4, wherein The intelligent control unit (400) is further configured to: determine the uncertainty of the current mixed liquid physical and chemical state according to the current generated state representation vector; based on the uncertainty, dynamically adjust the working parameters of the spectral analysis module (310) or the machine vision module (320) in the multi-modal online sensing unit (300), the working parameters including sampling frequency or resolution.
8. An aqueous pen ink diluent formulation control system characterized by, A processor (510) and a memory (520), the memory (520) storing a computer program, the processor (510) executing the computer program to implement the method of any one of claims 1-4.
Citation Information
Patent Citations
Intelligent preparation method and system of water-based propylene ink
CN116189078A
Predictive control system and method for water-based paint production
CN118409552A