Method, device and system for controlling ratio of ink diluent of gel ink pen
By combining multimodal perception and reinforcement learning, the production process of water-based pen ink thinner is monitored and optimized in real time, solving the problems of product consistency and efficiency caused by reliance on human experience, and achieving efficient and flexible production control.
Patent Information
- Application Number
- CN202511441308.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-10
AI Technical Summary
The current production process of water-based pen ink thinner relies on manual experience and rigid processes, resulting in poor adaptability to raw material fluctuations, delayed control response, and low batch consistency of the final product.
A multimodal online sensing unit is used to collect spectral and visual data in real time. A state representation vector is generated through a cross-modal fusion neural network, and a reinforcement learning decision agent is used to output refined physical operation commands to achieve closed-loop control.
It enables deep, real-time insights into the mixing process, ensuring high performance consistency between product batches, reducing waste, and improving production line flexibility and efficiency.
Smart Images

Figure CN120909138A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of ink preparation equipment, in particular to a water-based pen ink diluent proportioning control method, device and system. BACKGROUND
[0002] As a material widely used in daily writing and office field, the performance of water-based pen ink directly affects the user experience. Among them, the water-based pen ink diluent is one of the key components that determines the final flowability, stability, drying speed and writing smoothness of the ink. The diluent is usually mixed and reacted in a stirred tank according to specific proportions by a variety of solvents, cosolvents, surfactants, pH regulators and stabilizers, etc. Therefore, in the production process of the diluent, accurate control of the proportioning and mixing process of each component is the core link to ensure the stability of the quality of the final product.
[0003] In the existing production practice, the proportioning control process of the water-based pen ink diluent largely depends on the pre-set process procedures and the personal experience of the operators. The production usually follows the standard operating procedures (SOP) and is operated according to the fixed time sequence and the amount of feed. However, this open-loop control method based on fixed parameters is particularly vulnerable when facing various disturbances in actual production. For example, there are inevitable slight differences in the physical and chemical properties of different batches of chemical raw materials such as resins and surfactants. When a fixed process formula is used to process these fluctuating raw materials, it is easy to cause deviations in the key performance indicators such as viscosity and pH value of the final product, thereby causing poor consistency between product batches.
[0004] In order to make up for the shortcomings of this rigid process, experienced technical personnel are often needed to intervene. The operators observe the color, transparency, wall-hanging phenomenon and other macroscopic manifestations of the mixed solution, or take offline sampling detection at irregular times, to subjectively judge the current reaction progress, and accordingly adjust the amount of feed or stirring speed. This method not only introduces strong subjectivity and uncertainty, but also the significant time lag of offline detection itself, so that when the detection result shows an abnormality, the best adjustment opportunity has often been missed, which may lead to the degradation processing or even scrapping of the whole batch of materials.
[0005] Even in some production lines with automatic control, usually only a single online sensor such as an online viscometer or a pH meter is used to monitor the mixing process. Although this single-parameter feedback control is an improvement over purely manual operation, the information dimension obtained is extremely limited. The mixing of water-based pen ink diluent is a complex multi-component physicochemical process, and the overall state is the result of the mutual coupling and nonlinear action of multiple parameters. It is difficult to fully and accurately characterize the real state of the mixed liquid at present by relying on only one or two isolated physical quantities, and it is also difficult to capture subtle but critical transition characteristics in the state evolution process, so it is impossible to provide sufficient decision basis for realizing truly fine and optimized control.
[0006] Therefore, there is a lack of an intelligent control means in the prior art that can perceive the internal state of the mixing process in real time and comprehensively, and make dynamic and forward-looking decisions based on this, which has become a technical bottleneck restricting the further improvement of the product quality and production efficiency of water-based pen ink diluent. SUMMARY
[0007] In view of the deficiencies of the prior art, the present application provides a water-based pen ink diluent proportioning control method, device and system, which solves the problems of poor adaptability to raw material fluctuations, lagging control response and low batch consistency of the final product caused by relying on manual experience and rigid processes in the existing water-based pen ink diluent proportioning process.
[0008] To achieve the above purpose, the present application is implemented by the following technical solutions: a water-based pen ink diluent proportioning control method, comprising the following steps: S1. Real-time perception: at a preset time step, a multi-modal online perception unit deployed on a stirring tank or its circulating pipeline is used to synchronously collect multi-modal raw data of the mixed liquid in the stirring tank, and the multi-modal raw data at least includes spectral data reflecting molecular information and visual data reflecting macroscopic apparent state; S2. State representation: the multi-modal raw data is input into a preset cross-modal fusion neural network model for automatically extracting and fusing cross-modal implicit features, and the cross-modal fusion neural network model maps and generates a unified, low-dimensional state representation vector from the multi-modal raw data, and the state representation vector is used to comprehensively describe the physicochemical state of the current mixed liquid; S3. Intelligent decision-making: the state representation vector is input into a preset reinforcement learning decision-making agent, and the reinforcement learning decision-making agent outputs fine physical operation instructions for guiding the evolution of the mixed liquid state to the target state according to the optimal strategy learned by the reinforcement learning decision-making agent; S4. Closed-loop execution: the execution mechanism connected to the stirring tank is controlled to execute the fine physical operation instructions to change the physicochemical state of the mixed liquid; S5. Loop iteration: repeat steps S1 to S4 until the state of the mixed liquor characterized by the state characterization vector meets the preset production termination condition.
[0009] Preferably, the state characterization of step S2 specifically comprises: preprocessing the spectral data and the visual data; inputting the preprocessed spectral data and visual data into the modality-specific encoders in the cross-modal fusion neural network model to extract high-level features; performing self-attention weighting calculation and fusion on the high-level features by the cross-modal fusioner in the cross-modal fusion neural network model to learn cross-modal hidden correlation features and generate a state characterization vector.
[0010] Preferably, the method further comprises: calculating an immediate reward value according to the change in the physical and chemical state of the mixed liquor after executing the refined physical operation instruction; updating the optimal policy of the reinforcement learning decision-making agent using the transition sample containing the state characterization vector of the current time, the refined physical operation instruction, the immediate reward value, and the state characterization vector of the next time; wherein the immediate reward value is determined by a progress reward based on the gap between the current physical and chemical state of the mixed liquor and the final target state and a cost penalty based on the consumption generated by the executed operation.
[0011] Preferably, the progress reward is calculated in the following manner: using an auxiliary performance prediction model to predict the corresponding final product performance according to the state characterization vector of the current time and the state characterization vector of the next time, respectively; calculating the gap between the two final product performances and the preset performance target vector, and taking the reduction of the gap as the progress reward.
[0012] Preferably, the multi-modal raw data further includes rheological data and physicochemical data; the rheological data is collected by an online viscometer or a stirring shaft torque sensor; and the physicochemical data is collected by a pH electrode, a conductivity meter, or a temperature sensor.
[0013] Preferably, an aqueous pen ink diluent proportioning device comprises: a stirred tank; a high-precision material conveying unit connected to the stirred tank for conveying materials thereto; a multi-modal online sensing unit for collecting multi-modal raw data in the stirred tank, the multi-modal raw data at least including spectral data and visual data, the multi-modal online sensing unit including a spectral analysis module and a machine vision module; The intelligent control unit is electrically connected with the multi-modal online perception unit and the high-precision material conveying unit, and the intelligent control unit is configured to implement a method.
[0014] Preferably, the intelligent control unit specifically comprises: An edge computing controller is configured to run a cross-modal fusion neural network model and a reinforcement learning decision-making agent to generate refined physical operation instructions. A programmable logic controller is configured to receive the refined physical operation instructions and convert them into bottom-layer control electrical signals for actuators in the high-precision material conveying unit.
[0015] Preferably, the multi-modal online perception unit further comprises an online rheological / physical parameter module, and the online rheological / physical parameter module comprises at least one of an online viscometer, a pH / conductivity electrode and a temperature sensor.
[0016] Preferably, the intelligent control unit is further configured to: determine the uncertainty of the current mixed liquid physical and chemical state according to the current generated state representation vector; based on the uncertainty, dynamically adjust the working parameters of the spectral analysis module or the machine vision module in the multi-modal online perception unit, and the working parameters include sampling frequency or resolution.
[0017] An aqueous pen ink diluent proportioning control system comprises a processor and a memory, and the memory stores a computer program.
[0018] The present application provides an aqueous pen ink diluent proportioning control method, device and system. 1、The present application realizes deep and real-time insight into the physical and chemical state of the mixing process by using the multi-modal online perception unit to collect spectral data reflecting molecular information and visual data reflecting macroscopic state in real time, and using a cross-modal fusion neural network model to generate a comprehensive state representation vector.
[0019] 2、The present application uses a reinforcement learning decision-making agent to replace the traditional fixed program control, and the agent can autonomously decide and output optimal refined physical operation instructions based on a deep understanding of the current state.
[0020] 3、The application combines deep state representation with the continuous optimization mechanism of reinforcement learning, and constructs a control system that can learn and evolve from production practice. When facing changes such as new product formula or replacement of raw material suppliers, the system can quickly adapt to new process conditions through online learning without complex reprogramming, and the accumulated process knowledge is solidified in the updated model parameters, realizing the digitalization of process knowledge and inheritance, and greatly enhancing the flexibility of the production line. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 It is a schematic diagram of the proportioning device structure of the application; Figure 2 It is a flow chart of the proportioning control method of the embodiment of the application; Figure 3 It is a structure block diagram of the control system of the embodiment of the application; Figure 4 It is a schematic diagram of the cross-modal fusion neural network model structure of the embodiment of the application; Figure 5 It is a schematic diagram of the reinforcement learning decision and training process of the embodiment of the application; Figure 6 It is a schematic diagram of the pipeline and instrument connection of the water-based pen ink diluent proportioning device in the embodiment of the application.
[0022] Among them, 100, reaction and mixing unit; 110, stirred tank; 120, stirring device; 200, high-precision material conveying unit; 210, raw material tank A; 220, raw material tank B; 300, multi-modal online sensing unit; 310, spectrum analysis module; 320, machine vision module; 330, online rheological / physical parameter module; 400, intelligent control unit; 410, edge computing controller; 420, programmable logic controller; 500, water-based pen ink diluent proportioning control system; 510, processor; 520, memory; 530, communication bus; V-201, feed valve one; V-202, feed valve one; V-203, feed valve two; V-204, feed valve two; V-111, electric valve one; V-112, electric valve two; V-113, electric valve three; P-201, metering pump A; P-202, metering pump B; P-110, circulating pump. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the specification of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application. Embodiment 1
[0024] Please refer to the attached Figure 1 and attached Figure 6 , the embodiment of the present application provides a kind of water-based pen ink diluent proportioning device, the device is the physical carrier of realization proportioning control method, device includes Reaction and mixing unit 100, reaction and mixing unit 100 includes stirred tank 110, and stirred tank 110 is used to accommodate and mix the various components of water-based pen ink diluent.The agitator 120 is provided on the stirred tank 110, which can be but not limited to the stirring paddle driven by variable frequency motor, for the physical stirring of mixed liquid in tank; High-precision material conveying unit 200, which is connected to the stirred tank 110, is used to convey materials to it, and the high-precision material conveying unit 200 is connected to the stirred tank 110 by pipeline. Please refer to the attached Figure 6 , attached Figure 6 is the schematic diagram of the connection of pipeline and instrument in the device, wherein the high-precision material conveying unit 200 includes a plurality of storage tanks for storing different kinds of raw materials (including but not limited to raw material storage tank A 210, raw material storage tank B 220), and high-precision metering pump and electromagnetic control valve (including but not limited to feed valve one V-201, feed valve one V-202, feed valve two V-203, feed valve two V-204) connected with each storage tank. Metering pump includes metering pump A P-201, metering pump B P-202, which can be but not limited to peristaltic pump or diaphragm pump, for accurately conveying a specified volume of raw materials from the storage tank to the stirred tank 110; Multi-modal online sensing unit 300 is used to collect multi-modal raw data in stirred tank 110, and the multi-modal raw data at least includes spectral data and visual data, and the multi-modal online sensing unit 300 includes spectral analysis module 310 and machine vision module 320. The spectral analysis module 310 includes an immersion spectral probe, which is installed on the cover of the stirred tank 110, and the sensing end is directly in contact with the mixed liquid in the tank. The machine vision module 320 includes an industrial camera and a light source, which are installed outside the sight window of the stirred tank 110, for collecting images of the mixed liquid.
[0025] In an optional embodiment, the multi-modal online sensing unit 300 further includes an online rheological / physical parameter module 330, which can include at least one of a torque sensor installed on the driving shaft of the agitator 120, or a pH electrode, a conductivity meter and a temperature sensor immersed in the mixed liquid; The intelligent control unit 400 is electrically and data-communicatively connected with the high-precision material conveying unit 200 and the multi-modal online sensing unit 300. The intelligent control unit 400 is the technical core of the device, which is different from the traditional PLC controller. The intelligent control unit 400 is internally disposed with a cross-modal fusion neural network model and a reinforcement learning decision-making agent. The intelligent control unit 400 is responsible for receiving and processing multi-modal data, generating a state representation vector, making intelligent decisions, and outputting fine physical operation instructions to control the high-precision material conveying unit 200 and other actuators.
[0026] In a specific embodiment, the intelligent control unit 400 includes an edge computing controller 410 and a programmable logic controller (PLC) 420 in structure.
[0027] The edge computing controller 410 is an industrial personal computer (IPC) used for performing computationally intensive tasks. The edge computing controller 410 is internally disposed and runs a cross-modal fusion neural network model and a reinforcement learning decision-making agent. The edge computing controller 410 receives all raw data from the multi-modal online sensing unit 300 through a data interface, performs state representation generation (step S2) and intelligent decision-making (step S3), and finally generates fine physical operation instructions. The programmable logic controller 420 is a programmable logic controller (PLC) 420 that communicates with the edge computing controller 410 through an industrial fieldbus or Ethernet protocol. The function of the programmable logic controller 420 is to receive the fine physical operation instructions issued by the edge computing controller 410 and convert them into bottom-layer driving electrical signals for the metering pumps, electromagnetic valves in the high-precision material conveying unit 200, and the stirring device 120 in the reaction and mixing unit 100. The PLC is responsible for performing hardware control tasks with high reliability and high real-time performance.
[0028] In a preferred embodiment, the intelligent control unit 400 adopts a layered architecture, specifically including an edge computing controller 410 for running AI algorithm models and a programmable logic controller (PLC) 420 for executing bottom-layer hardware control. The two work together to ensure the intelligence of decision-making and the reliability of execution.
[0029] The multi-modal online sensing unit 300 further includes an online rheological / physical parameter module including at least one of an online viscometer, a pH / conductivity electrode, and a temperature sensor.
[0030] In addition, the edge computing controller 410 is also configured to determine the uncertainty of the current physical and chemical state of the mixed solution according to the current generated state representation vector. When the uncertainty is higher than a preset threshold, the edge computing controller 410 can generate and send parameter adjustment instructions to the spectral analysis module 310 or the machine vision module 320 in the multi-modal online perception unit 300 to dynamically change the working parameters such as the sampling frequency or resolution, so as to achieve more intensive monitoring of the key state transition stage. In another preferred embodiment, the intelligent control unit 400 also has the ability of active perception. It can dynamically adjust the working parameters of the spectral analysis module 310 or the machine vision module 320 in the multi-modal online perception unit 300 based on the uncertainty of the state reflected by the current state representation vector, including the sampling frequency or resolution. Example 2
[0031] Please refer to the accompanying Figure 2 , the accompanying Figure 4 and the accompanying Figure 5 , the embodiment of the present application provides a water-based pen ink diluent proportioning control method, which comprises the following steps: S1. Real-time perception: at a preset time step, the multi-modal online perception unit 300 deployed on the stirred tank 110 or its circulating pipeline synchronously collects multi-modal raw data of the mixed solution in the stirred tank 110. Unlike the prior art which only relies on a few isolated physical and chemical parameters, the multi-modal raw data collected by the present application at least includes spectral data reflecting the molecular level composition and structure information of the mixed solution, and visual data reflecting the apparent state such as the macroscopic clarity, color and flow pattern of the mixed solution. This multi-dimensional and cross-level data collection lays a foundation for deep understanding of the complex state of the mixing process. The multi-modal raw data at least includes spectral data reflecting molecular information and visual data reflecting macroscopic apparent state; the multi-modal raw data also includes rheological data and physicochemical data; the rheological data is collected by an online viscometer or a stirring shaft torque sensor; the physicochemical data is collected by a pH electrode, a conductivity meter or a temperature sensor.
[0032] In step S1, real-time perception is periodically performed at a preset discrete time step The sensor modules of the multi-modal online perception unit 300 are physically deployed on the tank body, tank cover or external circulating pipeline of the stirred tank 110 to realize non-invasive or online continuous monitoring of the mixed solution.
[0033] In one embodiment, the acquisition of spectral data is achieved through a spectral analysis module 310. The spectral analysis module 310 includes an immersion Raman spectral probe, which is installed via a standard flange interface on the lid of the stirred tank 110, with its sensing end directly immersed in the mixture inside the tank. The immersion Raman spectral probe emits an excitation beam of a preset wavelength into the mixture and collects the Raman scattered light generated by the mixture. The collected scattered light is transmitted to a spectrometer for dispersion and detection, thereby generating spectral data at time t. Spectral data is a one-dimensional vector whose elements are the scattered light intensities corresponding to different Raman shift values. Spectral data directly reflects the molecular vibrational information of each chemical component in the mixture.
[0034] In one embodiment, visual data acquisition is achieved through a machine vision module 320. The machine vision module 320 includes an industrial camera and a programmable light source. The industrial camera is mounted outside a sight glass window on the wall of the stirred tank 110, and its lens focuses on the mixture inside the sight glass. The programmable light source, such as an LED ring light, is arranged coaxially or paraaxially around the camera lens to provide stable and uniform illumination to the camera acquisition area (see attached diagram). Figure 1 At time t, the light source is turned on at a preset brightness, and the industrial camera simultaneously exposes and acquires a frame of digital image, generating visual data at time t. This visual data It is a two-dimensional or three-dimensional pixel matrix that records macroscopic appearance information of the mixture within the collection area, such as color, clarity, flow texture, and the presence of bubbles or undissolved particles.
[0035] In an optional embodiment, the multimodal raw data further includes rheological data. and physicochemical data Rheological data can be acquired by a torque sensor mounted on the agitator shaft drive system. This reflects the resistance experienced by the agitator in the mixture. Physicochemical data can be collected by pH electrodes, conductivity electrodes, and PT100 temperature sensors, all immersed in the mixture, corresponding to scalar values of the mixture's acidity / alkalinity, ion concentration, and temperature, respectively.
[0036] To ensure temporal consistency of data across different modalities, all sensors in the multimodal e-mode online sensing unit are connected to a central data acquisition system. The central data acquisition system uses time steps... Based on this, a synchronization trigger signal is sent to each sensor, or a high-precision timestamp is appended to the data packets received from each sensor. In this way, at each time t, the system can obtain a set of time-precisely aligned multimodal raw data. This provides effective data input for subsequent state characterization steps.
[0037] S2. State representation: inputting the multi-modal raw data into a pre-set cross-modal fusion neural network model for automatic extraction and fusion of cross-modal latent features , the neural network model maps and generates a unified, low-dimensional state representation vector, the neural network model is not simply concatenating data, but is used to automatically extract and fuse the deep correlation and human-undetectable latent features between different data modalities. Through the neural network model , the high-dimensional, heterogeneous raw data stream is mapped and generates a unified, low-dimensional state representation vector .
[0038] In step S2, the system performs the state representation generation process, which converts the multi-modal raw data collected in step S1 into a structured, information-intensive state representation vector . The state representation generation process first performs a series of preprocessing operations on the raw data to eliminate noise, unify the scale, and make it suitable for input to the neural network model.
[0039] In a specific embodiment, the preprocessing of the spectral data includes applying a Savitzky-Golay filter for smoothing and denoising, and using an Asymmetric Least Squares method for baseline correction to eliminate fluorescence background interference, to obtain the preprocessed spectral data . The preprocessing of the visual data includes scaling its size to a pre-set fixed size, such as 224x224 pixels, and normalizing its pixel values, such as scaling to the interval [-1, 1], to obtain the preprocessed visual data . For other data such as rheological data and physicochemical data, moving average filtering and Z-score standardization are performed.
[0040] After preprocessing, all data is input into the pre-set cross-modal fusion neural network model . The neural network model is structurally designed to include two main parts: a modality-specific encoder and a cross-modal fusioner.
[0041] The role of the modality-specific encoder is to extract high-level abstract features for the internal structure and characteristics of each data modality. Specifically, for one-dimensional preprocessed spectral data , a one-dimensional convolutional neural network (1D-CNN) is used as its encoder to extract the characteristic peaks and combined patterns in the spectrum through multiple convolution and pooling operations. For the two-dimensional preprocessed visual data , a two-dimensional convolutional neural network (2D-CNN), e.g., the backbone of a pre-trained residual network (ResNet), is used as its encoder to extract the texture, color distribution, and spatial structure features in the image. For other standardized scalar or vector data, a multi-layer perceptron (MLP) is used for encoding. The output of each encoder is a high-level feature vector for the corresponding modality.
[0042] Subsequently, the high-level feature vectors of all modalities are passed to a cross-modality fusioner. In one embodiment, the cross-modality fusioner is a Transformer encoder layer based on self-attention mechanism. The feature vectors from different modalities are treated as an input sequence. The self-attention mechanism generates a weighted sum for each feature vector by computing the correlation scores between any two feature vectors in the sequence, enabling the model to dynamically and selectively aggregate information from different modalities. This mechanism enables the model to capture nonlinear associations between, e.g., a specific spectral peak and a certain degree of turbidity in the image.
[0043] The output of the Transformer encoder layer is passed through a global average pooling layer to aggregate the sequential feature information into a single, fixed-dimension vector. This vector is the final generated state representation vector at time t that comprehensively describes the physical-chemical state of the current mixture. The generation process can be formally described by the following expression: ; where: is the state representation vector at time t; is the cross-modality fusion neural network model; is the preprocessed spectral data at time t; is the preprocessed visual data at time t; are the model learnable network parameters, whose values are obtained through pre-training on historical production data or online fine-tuning during the control process.
[0044] The state representation vector is used to comprehensively describe the physical-chemical state of the current mixture; the state representation specifically includes: preprocessing of spectral data and visual data; In a preferred embodiment, the cross-modal fusion neural network model comprises a modal-specific encoder and a cross-modal fusioner. The former is used to extract high-level features from spectral, visual, etc. data respectively, and the latter is used to perform weighted calculation and fusion on these high-level features through self-attention mechanism, so as to efficiently learn the cross-modal implicit correlation. The preprocessed spectral data and visual data are input into the modal-specific encoder in the cross-modal fusion neural network model to extract high-level features; The cross-modal fusioner in the cross-modal fusion neural network model performs self-attention weighted calculation and fusion on the high-level features to learn cross-modal implicit correlation features and generate a state representation vector.
[0045] S3. Intelligent decision-making: taking the state representation vector as input, receiving the state representation vector generated by step S2, which represents the current physical and chemical state of the mixed solution , and providing it to a preset reinforcement learning decision-making agent. The reinforcement learning decision-making agent models the entire proportioning control process as a Markov decision process, and its goal is to learn an optimal policy to maximize the long-term cumulative reward. According to the optimal policy learned by the reinforcement learning decision-making agent, the fine physical operation instructions for guiding the evolution of the mixed solution state to the target state are output , whose optimal action value function follows the Bellman optimality equation: This makes the decision no longer based on fixed rules, but on dynamic evaluation of future state values.
[0046] First, define the set of fine physical operation instructions as a preset discrete action space A. In this embodiment, each action a in the action space A corresponds to a specific physical operation that can be executed by the underlying hardware. For example, an action can be defined as an instruction tuple such as (material number, feeding volume, conveying rate) or (stirring speed, duration). By discretizing continuous control parameters, the complex control problem is transformed into a sequential decision-making problem.
[0047] The core of the reinforcement learning decision-making agent is a deep Q-network (Deep Q-Network, DQN), denoted as . The network is a deep neural network whose function is to input a state s and output the predicted Q value of each action a in the action space A. The Q value is an estimate of the expected cumulative reward that can be obtained in the future after performing the corresponding action in the state. The parameters of the network are denoted by .
[0048] When the reinforcement learning decision-making agent receives the state representation vector at time t, it will a deep Q-network inside it The network performs one forward propagation computation to generate a corresponding Q-value for each of the optional actions a in the action space A.
[0049] Subsequently, the agent selects an action from the action space A as the final output of the refined physical operation instruction according to the optimal policy it has learned In one specific embodiment, the policy is a deterministic policy. The specific implementation of the policy is as follows: the system generates a random number between 0 and 1; if the random number is greater than a preset value , the agent selects the action a that maximizes the Q-value as , a process known as exploitation; if the random number is not greater than , the agent randomly selects an action from the action space A as , a process known as exploration.
[0050] The preset value may remain unchanged throughout the control process, or it can be designed as a value that gradually decreases over time or training steps. At the beginning of the process, a larger value encourages the agent to explore more to discover better operation sequences; at the end of the process, a smaller value encourages the agent to exploit the optimal policy it has learned to achieve stable control.
[0051] Finally, the action selected by the policy is the output of step S3, which will be passed to step S4 for physical execution.
[0052] S4. Closed-loop execution: control the actuator connected to the stirred tank 110 to execute the refined physical operation instruction to change the physical and chemical state of the mixed solution, and iterate the above steps until the state of the mixed solution meets the preset production termination condition. The refined physical operation instruction output by step S3 is sent to the programmable logic controller (PLC) 420 in the intelligent control unit 400. In one embodiment, the instruction is issued by the edge computing controller 410 running the reinforcement learning decision-making agent and sent to the PLC through an internal bus or industrial Ethernet communication protocol.
[0053] The function of the PLC is to convert high-level, symbolic instructions The analysis and conversion into precise timing of the underlying control electrical signals directly driving the underlying actuators. The actuators are the high precision material delivery units 200 and the components in the reaction and mixing units.
[0054] In a specific embodiment, if the instruction is (material code = A, dosing volume = 15.2 mL), and the flow characteristics of the dosing pump corresponding to material A are known, the PLC calculates the running time required to drive the dosing pump to achieve a dosing volume of 15.2 mL. Subsequently, the PLC outputs a start electrical signal to the relay or frequency converter controlling the dosing pump within the calculated time period, while controlling the solenoid valve on the corresponding pipeline to open, and revokes the start signal and closes the valve at the end of the time period.
[0055] In another embodiment, if the instruction is (stirring speed = 300 rpm), the PLC outputs an analog voltage signal corresponding to the 300 rpm speed or a digital communication message containing specific frequency parameters to the frequency converter controlling the stirring motor, so that the stirring speed of the stirring paddle is stabilized at the preset value.
[0056] After the actuators perform physical operations, they directly change one or more of the physical and chemical properties of the mixture in the stirred tank 110, such as the concentration of components, temperature, and viscosity. The change in state will be collected by the multi-modal online perception unit 300 in step S1 at the beginning of the next time step, thus forming a closed-loop feedback path of the control method.
[0057] S5. Loop iteration: repeat steps S1 to S4 to form a continuous "perception-characterization-decision-execution" closed loop until the state characterization vector characterizing the state of the mixture meets the preset production termination condition.
[0058] In step S5, the system executes the aforementioned steps S1 to S4 in a loop iteration manner. The complete sequence of steps S1 to S4 constitutes a discrete-time step running, independent control cycle. After the execution of step S4 in each control cycle, the system immediately returns to step S1 to start the multi-modal data collection of the next cycle, thus forming a continuous "perception-characterization-decision-execution" closed-loop control process.
[0059] At the beginning or end of each control cycle, the system judges a preset production termination condition. The purpose of this judgment is to determine whether the current physical and chemical state of the mixture has reached or is close enough to the final production target.
[0060] In a specific embodiment, the termination condition is based on the current state characterization vector The distance between the predicted end-product performance and the pre-set performance target vector G defines the termination condition. Specifically, the system calculates the distance between the current state representation vector to the aforementioned auxiliary performance prediction model to obtain a predicted performance vector. The system then calculates the distance between the predicted performance vector and the performance target vector G, for example, the L2 norm. When the distance is smaller than a pre-set positive real threshold , the system determines that the termination condition has been reached. This condition can be represented by the following inequality: ; wherein: is the predicted performance vector output by the auxiliary performance prediction model according to the current state ; G is the pre-set performance target vector; is a pre-set positive real threshold representing the acceptable error range.
[0061] Once the termination condition is reached, the iterative process stops. The system then outputs a signal indicating that the production is complete, and optionally saves all the state representation vectors, action sequences, and reward values for subsequent process analysis or model offline optimization. If the termination condition is not reached, the system proceeds to the next control cycle.
[0062] In another preferred embodiment, to effectively train the reinforcement learning decision-making agent, the method further comprises a reward calculation and policy update mechanism. After each operation, the system calculates an immediate reward value, which is composed of a progress reward and a cost penalty. The progress reward is quantified by an auxiliary performance prediction model , which predicts the end-product performance based on the current state, and the progress reward is the reduction of the gap between the predicted performance of the current state and the target performance. This provides a dense and explicit guidance signal for the agent's learning. After executing the refined physical operation instructions, an immediate reward value is calculated based on the changes in the physical and chemical state of the mixed solution; The optimal policy of the reinforcement learning decision-making agent is updated using the transition samples containing the state representation vector at the current time, the refined physical operation instructions, the immediate reward value, and the state representation vector at the next time; wherein the immediate reward value is determined by the progress reward based on the gap between the current mixed solution physical and chemical state and the final target state, and the cost penalty based on the cost of the executed operation.
[0063] The progress reward is calculated as follows: The auxiliary performance prediction model is used to predict the corresponding end-product performance based on the state representation vector at the current time and the state representation vector at the next time, respectively; The difference between the two final product performances and the preset performance target vector is calculated, and the reduction of the difference is taken as a progress reward. Embodiment 3
[0064] Please refer to the attached Figure 3 The water-based pen ink diluent proportioning control system 500 can be the intelligent control unit 400 in the foregoing device embodiments, or a server or computer device independent of the foregoing device embodiments. The water-based pen ink diluent proportioning control system 500 can at least include a processor 510, a memory 520, and a communication bus 530 for connecting the processor 510 and the memory 520.
[0065] The processor 510 can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), or one or more other types of processing units.
[0066] The memory 520 can include a non-persistent memory, such as a random access memory (RAM), and / or a persistent memory, such as a read-only memory (ROM) or a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD). The memory 520 stores a computer program.
[0067] When the processor 510 is configured to execute the computer program stored in the memory 520, the steps of the water-based pen ink diluent proportioning control method in any of the foregoing method embodiments are implemented.
[0068] In one specific embodiment, the computer program stored in the memory 520, when executed by the processor 510, can be functionally divided into the following modules: A data acquisition and preprocessing module: The data acquisition and preprocessing module is configured to receive the synchronized multi-modal raw data from the multi-modal online perception unit 300 through the communication interface, and perform a preset preprocessing operation, such as baseline correction, image enhancement, and data standardization, on the raw data.
[0069] A state representation generation module: The state representation generation module contains and executes the logic of a cross-modal fusion neural network model. It receives the data processed by the data acquisition and preprocessing module, and generates a unified, low-dimensional state representation vector through forward propagation calculation of the model.
[0070] An intelligent decision-making module: The intelligent decision-making module contains and executes the logic of a reinforcement learning decision-making agent. It receives the state representation vector output by the state representation generation module, and outputs refined physical operation instructions according to a preset strategy through calculation of its internal deep Q network.
[0071] The instruction execution and communication module is configured to send the refined physical operation instructions generated by the intelligent decision module to one or more programmable logic controllers (PLC) 420 through an industrial bus or an Ethernet interface for execution.
[0072] The model updating module is configured to store the transition samples (state, action, reward, next state) generated by the system interaction, and calculate and update the network parameters of the deep Q network in the intelligent decision module by minimizing a preset loss function according to the samples.
[0073] Through the cooperation of the above modules, the control system of the present application can completely realize all the functions of the method.
[0074] Although embodiments of the present application have been shown and described, it would be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made thereto without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An aqueous pen ink diluent formulation control method, characterized by, The method comprises the following steps: S1. Real-time sensing: synchronously collecting multi-modal raw data of the mixed liquid in the stirred tank (110) at a preset time step by deploying a multi-modal online sensing unit (300) on the stirred tank (110) or its circulating pipeline, wherein the multi-modal raw data at least includes spectral data reflecting molecular information and visual data reflecting macroscopic appearance state; S2. State representation: inputting the multi-modal raw data into a preset cross-modal fusion neural network model for automatically extracting and fusing cross-modal implicit features, and mapping the multi-modal raw data by the cross-modal fusion neural network model to generate a unified low-dimensional state representation vector, which is used to comprehensively describe the physical and chemical state of the current mixed liquid; S3. Intelligent decision-making: inputting the state representation vector as input into a preset reinforcement learning decision-making agent, and outputting a refined physical operation instruction for guiding the evolution of the mixed liquid state to a target state according to the optimal strategy learned by the reinforcement learning decision-making agent; S4. Closed-loop execution: controlling an execution mechanism connected to the stirred tank (110) to execute the refined physical operation instruction to change the physical and chemical state of the mixed liquid; S5. Iterative cycle: repeating steps S1 to S4 until the state of the mixed liquid represented by the state representation vector meets the preset production termination condition.
2. The water-based pen ink diluent ratio control method according to claim 1, characterized by, The state representation of step S2 specifically comprises: preprocessing the spectral data and the visual data; inputting the preprocessed spectral data and visual data into the modal-specific encoder in the cross-modal fusion neural network model to extract high-level features; performing self-attention weighting calculation and fusion on the high-level features by the cross-modal fusion unit in the cross-modal fusion neural network model to learn cross-modal implicit correlation features and generate the state representation vector.
3. The water-based pen ink diluent ratio control method according to claim 1, characterized by, The method further comprises: calculating an immediate reward value according to the change in the physical and chemical state of the mixed liquid after executing the refined physical operation instruction; updating the optimal strategy of the reinforcement learning decision-making agent using the transition sample containing the state representation vector at the current time, the refined physical operation instruction, the immediate reward value, and the state representation vector at the next time; wherein the immediate reward value is determined by a progress reward based on the gap between the current physical and chemical state of the mixed liquid and the final target state and a cost penalty based on the consumption generated by the executed operation.
4. The water-based pen ink diluent ratio control method according to claim 3, characterized by, The calculation method of the progress reward is: using an auxiliary performance prediction model to predict the corresponding final product performance according to the state representation vector at the current time and the state representation vector at the next time, respectively; calculating the gap between the final product performance and the preset performance target vector, and taking the reduction of the gap as the progress reward.
5. The water-based pen ink diluent ratio control method according to claim 1, characterized by, The multi-modal raw data further includes rheological data and physicochemical data; the rheological data is collected by an online viscometer or a stirring shaft torque sensor; and the physicochemical data is collected by a pH electrode, a conductivity meter or a temperature sensor.
6. An aqueous pen ink diluent proportioning device applied to the aqueous pen ink diluent proportioning control method of any one of claims 1-5, characterized in that, It comprises: a stirred tank (110); A high-precision material conveying unit (200) connected with the stirred tank (110) and configured to convey material to the stirred tank (110); A multi-modal online sensing unit (300) configured to collect multi-modal raw data in the stirred tank (110), the multi-modal raw data comprising at least spectral data and visual data, the multi-modal online sensing unit (300) comprising a spectral analysis module (310) and a machine vision module (320); An intelligent control unit (400) electrically connected with the multi-modal online sensing unit (300) and the high-precision material conveying unit (200), the intelligent control unit (400) being configured to implement the method of claim 1.
7. The aqueous pen ink diluent compounding device according to claim 6, wherein The intelligent control unit (400) specifically comprises: An edge computing controller (410) configured to run a cross-modal fusion neural network model and a reinforcement learning decision-making agent to generate refined physical operation instructions; A programmable logic controller (420) configured to receive the refined physical operation instructions and convert the refined physical operation instructions into bottom-layer control electrical signals for actuators in the high-precision material conveying unit (200).
8. The aqueous pen ink diluent compounding device according to claim 6, wherein The multi-modal online sensing unit (300) further comprises an online rheological / physical parameter module (330), the online rheological / physical parameter module (330) comprising at least one of an online viscometer, a pH / conductivity electrode, and a temperature sensor.
9. The aqueous pen ink diluent compounding device according to claim 6, wherein The intelligent control unit (400) is further configured to: determine the uncertainty of the current mixed liquid physical and chemical state according to the current generated state representation vector; based on the uncertainty, dynamically adjust the working parameters of the spectral analysis module (310) or the machine vision module (320) in the multi-modal online sensing unit (300), the working parameters including sampling frequency or resolution.
10. An aqueous pen ink diluent formulation control system characterized by, A processor (510) and a memory (520), the memory (520) storing a computer program, the processor (510) executing the computer program to implement the method of any one of claims 1-5.
Citation Information
Patent Citations
Ink viscosity control system and control method based on neural network
CN114407526A
Intelligent preparation method and system of water-based propylene ink
CN116189078A
Automatic control system and method for preparing ultraviolet curing type printing ink
CN116694129A
Predictive control system and method for water-based paint production
CN118409552A
Printing ink quality online monitoring method and system based on multi-sensor fusion
CN119534813A