A method for point micro-environmental control spraying for bridge support maintenance
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]现有技术中,应用于桥梁支座维护的喷涂装置通常由压力容器、泵送系统、喷枪及简单控制系统构成;部分高端设备集成了基础的环境传感器,能够依据监测到的环境温度或空气湿度,通过预置的查表法或简单补偿公式,对涂料流量或压力进行有限调整;然而,此类技术存在根本性局限;其一,其感知能力不足,通常仅能获取喷涂点附近空气的单一或少数宏观环境参数,无法对工件表面本身的微观状态进行高频、多点的实时探测,例如表面实际温度梯度、冷凝水膜存在与否、污染物附着情况与类型等关键信息严重缺失;其二,其控制逻辑僵化,依赖于基于经典控制理论的线性或简单非线性模型,难以刻画和应对涂料雾化、飞行、撞击、铺展、固化这一复杂物理化学过程中,多种动态微环境变量与最终成膜质量之间高度非线性的耦合关系;其三,其决策与执行过程本质上是开环或弱闭环的,因为对涂层质量如湿膜均匀性、附着强度的有效评估通常滞后于喷涂动作,无法在本次喷涂过程中提供实时反馈以修正参数;因此,现有技术在面对复杂、非均匀的支座微环境时,无法实现喷涂参数的实时动态自优化,导致防腐介质难以在金属表面形成均匀且附着性能最佳的涂层,进而为支座的长周期安全服役埋下隐患
[0015] The beneficial effects of this invention are as follows: by multimodal fusion perception of the microenvironment on the surface of bridge bearings, a precise digital state representation is constructed, and based on this, a deep reinforcement learning model is driven to generate optimized spraying parameters. Then, robustness verification and reinforcement are carried out through digital twin simulation. Finally, a feedforward compensation closed loop based on real-time online detection is introduced in the physical execution, thereby realizing intelligent adaptation and active control of the complex, non-uniform, and dynamically changing microenvironment of bridge bearings. This ensures that the anti-corrosion coating can still obtain uniform, stable and highly adhesive film quality under various harsh field conditions, significantly improving the reliability, economy and long-term protective effectiveness of maintenance operations.
Smart Images

Figure CN122386726B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bridge maintenance construction, and more specifically, to a fixed-point microenvironment control spraying method for bridge bearing maintenance. Background Technology
[0002] As a key force-transmitting component connecting the superstructure and substructure of a bridge, the long-term reliable operation of bridge bearings is crucial to the overall safety of the bridge. Steel bearings, such as pot bearings and spherical bearings, are exposed to complex natural environments during long-term service, bearing the combined effects of vehicle loads, temperature changes, and corrosive media. The integrity of the anti-corrosion coating on their metal surfaces is core to ensuring the function and lifespan of the bearings. The on-site environment for bearing maintenance is extremely complex, not only because the bearing structure itself includes various geometric features and materials such as steel pots, intermediate liners, sealing components, and fastening bolts, but also because their installation locations are diverse. For example, the top surface of bridge piers, the interior of box girders, or near expansion joints; these locations typically have significantly different ventilation conditions and uneven sunlight distribution, resulting in a non-uniform dynamic microenvironment on the bearing surface at the microscale; for example, there is a significant temperature difference between the sun-facing and shaded sides, the humidity is high in areas where air stagnates, and salt and water vapor may accumulate near the drainage holes on the bridge deck; this spatiotemporal difference in the microenvironment caused by structural complexity, location specificity, and external climate places extremely high demands on the coating spraying process in anti-corrosion maintenance; currently, on-site maintenance mostly relies on the experience of operators or preset programs to control the spraying equipment.
[0003] In existing technologies, spraying devices used for bridge bearing maintenance typically consist of a pressure vessel, a pumping system, a spray gun, and a simple control system. Some high-end devices integrate basic environmental sensors, enabling limited adjustments to paint flow rate or pressure based on monitored ambient temperature or humidity using preset lookup tables or simple compensation formulas. However, such technologies have fundamental limitations. First, their sensing capabilities are insufficient, usually only acquiring single or a few macroscopic environmental parameters of the air near the spraying point, failing to perform high-frequency, multi-point real-time detection of the workpiece surface's microscopic state. For example, crucial information such as the actual surface temperature gradient, the presence or absence of condensation film, and the adhesion and type of contaminants is severely lacking. Second, their control logic is rigid, relying on... Classical control theory's linear or simple nonlinear models are insufficient to characterize and address the highly nonlinear coupling between various dynamic microenvironmental variables and the final film quality in the complex physicochemical process of coating atomization, flight, impact, spreading, and curing. Thirdly, the decision-making and execution processes are essentially open-loop or weakly closed-loop because effective assessment of coating quality, such as wet film uniformity and adhesion strength, usually lags behind the spraying action and cannot provide real-time feedback to correct parameters during the spraying process. Therefore, existing technologies cannot achieve real-time dynamic self-optimization of spraying parameters when facing complex and non-uniform support microenvironments, making it difficult for the anti-corrosion medium to form a uniform coating with optimal adhesion on the metal surface, thus posing a hidden danger to the long-term safe service of the support. Summary of the Invention
[0004] This invention addresses the technical problems existing in the prior art by providing a fixed-point microenvironment control spraying method for bridge bearing maintenance, thereby solving the problems mentioned in the background art.
[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: specifically, it includes the following steps: Step S1: The target spraying point of the bridge support is scanned synchronously by a multi-sensor array deployed at the end of the spraying equipment to collect multi-modal raw data including infrared thermal radiation data, hyperspectral reflectance data, macro image data and contact temperature and humidity data; the multi-modal raw data is fused and processed by a preset deep learning feature fusion network to generate a high-dimensional microenvironment state vector that characterizes the comprehensive surface state of the target spraying point. Step S2: Input the high-dimensional microenvironment state vector and the preset target coating specification parameters into a pre-trained deep reinforcement learning decision model; The deep reinforcement learning decision model is trained based on a digital twin simulation module that integrates computational fluid dynamics and coating film formation dynamics. The deep reinforcement learning decision model outputs an initial spraying parameter command vector, which includes spraying pressure, coating flow rate, atomization angle and spray gun movement speed. Step S3: Load the high-dimensional microenvironment state vector and the initial spraying parameter command vector into the digital twin simulation module; drive the digital twin simulation module to perform rapid simulation calculations on the spraying process, predict the coating quality prediction data, and evaluate the robustness of the initial spraying parameter command vector under preset perturbations; output the final spraying parameter command verified by simulation. Step S4: Decouple the final spraying parameter command into a sequence of coordinated control commands for the pressure regulating valve, flow pump, atomizer, and motion actuator, and issue and execute them; simultaneously start a multi-sensor array to perform high-frequency online detection on the sprayed area to obtain real-time microscopic state data; compare the real-time microscopic state data with the coating quality prediction data, generate a dynamic feedforward compensation amount based on the deviation signal generated by the comparison, and use the dynamic feedforward compensation amount to perform real-time fine-tuning compensation on the currently executing coordinated control command sequence.
[0006] In a preferred embodiment, the process of generating the high-dimensional microenvironment state vector in step S1 is as follows: The multimodal raw data, namely infrared thermal radiation data, hyperspectral reflectance data, macro image data, and contact temperature and humidity data, are respectively input into the corresponding branch coding subnetworks in the deep learning feature fusion network for processing. Specifically, the branch coding subnetwork for processing infrared thermal radiation data and hyperspectral reflectance data is configured with a convolutional neural network structure, the branch coding subnetwork for processing macro image data is configured with a convolutional neural network structure, and the branch coding subnetwork for processing contact temperature and humidity data adopts a fully connected network structure; each branch coding subnetwork maps the input multimodal raw data into a high-dimensional feature vector of a unified dimension. Subsequently, the high-dimensional feature vectors output by each branch encoding subnetwork are input into the pre-built cross-modal attention module in the deep learning feature fusion network; the cross-modal attention module linearly transforms each high-dimensional feature vector to a common feature subspace, and calculates the original attention score of each high-dimensional feature vector in the common feature subspace by combining the learnable context vector; Based on the original attention score, the attention weight coefficients corresponding to each high-dimensional feature vector are calculated using a normalized exponential function; based on the calculated attention weight coefficients, each high-dimensional feature vector is weighted, summed, and finally transformed to generate a high-dimensional microenvironment state vector.
[0007] In a preferred embodiment, the high-dimensional microenvironment state vector is a structured multidimensional data object, in which comprehensive state information of the surface of the target spraying point is implicitly encoded. The components of a high-dimensional microenvironment state vector include encoding the surface temperature field distribution pattern, encoding the surface humidity gradient variation trend, and encoding the types and concentration distribution of surface-attached pollutants. The surface temperature field distribution pattern is encoded from the features extracted from infrared thermal radiation data by a branch coding subnetwork and involved in attention-weighted fusion. The encoding of the surface humidity gradient trend comes from the specified spectral band features related to moisture absorption in hyperspectral reflectance data, as well as the features extracted from contact temperature and humidity data by a branch coding subnetwork and fused together. The encoding of surface-attached contaminant types and their concentration distributions is derived from the spectral features characterizing different chemical components in hyperspectral reflectance data, as well as the texture and morphological features extracted from macro image data, which are then combined into a joint feature representation after attention-weighted fusion.
[0008] In a preferred embodiment, step S2, which involves inputting the high-dimensional microenvironment state vector and preset target coating specification parameters into a pre-trained deep reinforcement learning decision model, is as follows: A decision state vector is formed by concatenating a high-dimensional microenvironment state vector with preset target coating specification parameters; the decision state vector is then input into the policy network of a deep reinforcement learning decision model, which is a parameterized deep neural network trained using a proximal policy optimization algorithm. The objective function for near-end policy optimization is constrained by a pruning mechanism. The calculation process is as follows: First, the ratio of the action probability of the new and old policy networks under a given state is calculated during the training of the deep reinforcement learning decision model; then, the action probability ratio is constrained within a preset upper and lower limit range. The optimal network weight parameters are obtained by multiplying the ratio of the action probabilities after pruning with the advantage function, taking the expectation of the product and updating it using gradient ascent; whereby the advantage function is estimated by the value network trained in conjunction with the policy network. The policy network performs multi-layer nonlinear transformation and feature extraction on the input decision state vector, and finally generates the initial spraying parameter command vector directly through the output layer of the policy network. The spraying pressure, paint flow rate, atomization angle and spray gun movement speed in the initial spraying parameter command vector are all continuous real values. The range of their real values is predefined and constrained according to the physical execution capability of the spraying equipment during the training stage of the deep reinforcement learning decision model.
[0009] In a preferred embodiment, the pre-training process of the deep reinforcement learning decision model is completed in an offline virtual environment constructed by a digital twin simulation module, and its training objective is driven by a multi-objective reward function. The multi-objective reward function is calculated based on the coating quality result and resource consumption result predicted by the digital twin simulation module after simulating the given decision state vector and the corresponding initial spraying parameter command vector; The multi-objective reward function consists of three sub-items: coating thickness uniformity reward, coating adhesion compliance reward, and spraying resource consumption penalty, which are linearly weighted and summed according to preset weight coefficients. The coating thickness uniformity bonus is calculated based on the coating thickness distribution map obtained from simulation prediction, and the uniformity is evaluated by quantifying the severity of the spatial gradient change in the thickness distribution. The coating adhesion achievement bonus is evaluated based on the simulated prediction of the coating film formation dynamics. When the film formation state meets the preset adhesion formation conditions, a positive bonus is given. The penalty for spraying resource consumption is proportional to the weighted sum of paint flow rate and spraying pressure in the initial spraying parameter command vector; During the training of the deep reinforcement learning decision model, the output value of the calculated multi-objective reward function is used as a feedback signal and propagated back to the policy network and value network. The weight parameters of the policy network are continuously adjusted so that the policy network learns to output the optimal initial spraying parameter instruction vector after comprehensively balancing coating quality and operating cost.
[0010] In a preferred embodiment, step S3, which involves loading the high-dimensional microenvironment state vector and the initial spraying parameter command vector into the digital twin simulation module for rapid simulation calculation and robustness evaluation, is as follows: Using the high-dimensional microenvironment state vector as the environmental boundary condition and the initial spraying parameter command vector as the nominal action, a set of parallel disturbance actions is generated in the digital twin simulation module. The method for generating the parallel perturbation action set is as follows: based on the parameter value of the nominal action, according to a small fluctuation percentage range defined by a preset relative perturbation intensity vector, multiple parameter perturbation values are randomly generated. Each parameter perturbation value is added to the nominal action parameter value item by item to obtain multiple perturbed spraying parameter combinations. These perturbed parameter combinations together with the original nominal action constitute the parallel perturbation action set. The driving digital twin simulation module uses a high-dimensional microenvironment state vector as a unified initial field to perform fast simulation calculations of coupled computational fluid dynamics and coating film formation dynamics for each combination of spraying parameters in the set of parallel disturbance actions. Each simulation calculation outputs a corresponding subset of coating quality prediction data, which includes the predicted spatial distribution matrix of coating thickness, the spatial distribution matrix of coating curing degree, and the maximum stress value at the interface between the coating and the substrate.
[0011] In a preferred embodiment, the process of evaluating the robustness of the initial spraying parameter command vector under a preset perturbation and outputting the final spraying parameter command verified by simulation is as follows: A robust loss value is calculated based on a subset of all coating quality prediction data obtained from parallel simulation. The calculation logic for the robustness loss value is as follows: First, calculate the values of the thickness field structural sensitivity index and the solidification state overall distribution offset index; multiply the thickness field structural sensitivity index by a preset first weighting coefficient, and multiply the solidification state overall distribution offset index by a preset second weighting coefficient; then add the two product results together, and finally use the sum as the final robustness loss value. The thickness field structure sensitivity index is calculated by comparing the difference between the thickness distribution matrix obtained from the nominal simulation and the thickness distribution matrix obtained from the simulation of each perturbation action. Specifically, it is to calculate the average change intensity of the thickness field spatial gradient structure caused by parameter perturbation. The overall distribution offset index of the solidified state is calculated by comparing the overall probability distribution difference between the set of solidified state values obtained from the simulation of nominal actions and the set of solidified state values obtained from the simulation of each disturbance action. Specifically, the Hellinger distance is used to measure the difference between the two probability distributions. The robustness loss value is compared with a preset robustness threshold. If the robustness loss value is less than or equal to the robustness threshold, the initial spraying parameter command vector is directly used as the final spraying parameter command output verified by simulation. If the robustness loss value is greater than the robustness threshold, a local parameter optimization process running in the digital twin simulation module is triggered to minimize the robustness loss value. The initial spraying parameter command vector is fine-tuned, and the optimized parameter vector is used as the final spraying parameter command output verified by simulation.
[0012] In a preferred embodiment, step S4, which involves decoupling the simulation-verified final spraying parameter commands into a sequence of collaborative control commands and issuing them for execution, is as follows: Receive the final spraying parameter command verified by simulation; The final spraying parameter commands verified by simulation are input into a pre-built collaborative command generator. The collaborative command generator performs a multi-objective optimization calculation based on the physical response dynamics model and electrical characteristic model of the pressure regulating valve, flow pump, atomizer and each motion actuator. The physical response dynamics model and electrical characteristic model are collectively referred to as the actuator dynamic model. The objective function of the multi-objective optimization aims to simultaneously minimize the drastic changes in the control commands of each actuator in the generated sequence of cooperative control commands, and minimize the tracking error between the actual output pressure, actual output flow, actual output atomization angle and actual spray gun pose predicted based on the actuator dynamic model and the time-varying reference trajectory generated by the final spraying parameter commands. By solving a multi-objective optimization problem, a set of coordinated control command sequences that are synchronized in time and coupled in terms of control variables is calculated. The coordinated control command sequence is synchronously sent to the corresponding pressure regulating valve, flow pump, atomizer and motion actuator in chronological order.
[0013] In a preferred embodiment, the process of synchronously activating the multi-sensor array to perform high-frequency online detection on the coated area, obtaining real-time microscopic state data, and comparing the real-time microscopic state data with coating quality prediction data to generate dynamic feedforward compensation is as follows: After the collaborative control command sequence begins to execute, the multi-sensor array is immediately driven to continuously and synchronously scan the sprayed micro area within a fixed lag distance behind the spray gun at a sampling rate higher than the scanning frequency in step S1, and collect real-time microscopic state data of the sprayed micro area. The collected real-time microstate data is input into a preset lightweight online feature extraction network to quickly extract a real-time microstate feature vector that characterizes the current transient properties of the wet film. Meanwhile, from the coating quality prediction data in step S3, the predicted micro-state feature vectors on the sprayed micro-areas are extracted as a reference benchmark for comparison. Calculate the deviation signal between the real-time microstate feature vector and the predicted microstate feature vector; The deviation signals from the current moment and multiple consecutive moments within a short time window in the past, along with the control commands in the current sequence of coordinated control commands, are input into a pre-trained lightweight temporal attention compensation network. The lightweight temporal attention compensation network analyzes the temporal pattern characteristics of the deviation signal sequence through its internal attention mechanism, and infers and generates a dynamic feedforward compensation amount by combining the current control command status.
[0014] In a preferred embodiment, the process of real-time fine-tuning and compensating the dynamically fedforward compensation amount for the currently executed cooperative control command sequence is as follows: Within each control cycle, receive dynamic feedforward compensation generated by a lightweight temporal attention compensation network; The dynamic feedforward compensation is added to the original control commands from the coordinated control command sequence that are planned to be issued in the current control cycle, and a set of coordinated control commands that have been compensated and corrected is generated. The compensated and corrected coordinated control commands are issued in real time to the pressure regulating valve, flow pump, atomizer and motion actuator, covering and replacing the original control commands that were originally planned to be issued. Subsequently, in the next control cycle, the steps of high-frequency online detection, feature extraction, bias calculation, compensation network inference, and instruction correction are repeated. The cycle of execution, perception, and compensation runs continuously at millisecond intervals until the entire spraying process on the current target spraying point is completed.
[0015] The beneficial effects of this invention are as follows: by multimodal fusion perception of the microenvironment on the surface of bridge bearings, a precise digital state representation is constructed, and based on this, a deep reinforcement learning model is driven to generate optimized spraying parameters. Then, robustness verification and reinforcement are carried out through digital twin simulation. Finally, a feedforward compensation closed loop based on real-time online detection is introduced in the physical execution, thereby realizing intelligent adaptation and active control of the complex, non-uniform, and dynamically changing microenvironment of bridge bearings. This ensures that the anti-corrosion coating can still obtain uniform, stable and highly adhesive film quality under various harsh field conditions, significantly improving the reliability, economy and long-term protective effectiveness of maintenance operations. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0019] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0020] This embodiment provides, for example Figure 1The method for targeted microenvironment control spraying for bridge bearing maintenance, shown below, specifically includes the following steps: Step S1: The target spraying point of the bridge support is scanned synchronously by a multi-sensor array deployed at the end of the spraying equipment to collect multi-modal raw data including infrared thermal radiation data, hyperspectral reflectance data, macro image data and contact temperature and humidity data; the multi-modal raw data is fused and processed by a preset deep learning feature fusion network to generate a high-dimensional microenvironment state vector that characterizes the comprehensive surface state of the target spraying point. Step S2: Input the high-dimensional microenvironment state vector and the preset target coating specification parameters into a pre-trained deep reinforcement learning decision model; The deep reinforcement learning decision model is trained based on a digital twin simulation module that integrates computational fluid dynamics and coating film formation dynamics. The deep reinforcement learning decision model outputs an initial spraying parameter command vector, which includes at least spraying pressure, coating flow rate, atomization angle and spray gun movement speed. Step S3: Load the high-dimensional microenvironment state vector and the initial spraying parameter command vector into the digital twin simulation module; drive the digital twin simulation module to perform rapid simulation calculations on the spraying process, predict the coating quality prediction data, and evaluate the robustness of the initial spraying parameter command vector under preset perturbations; output the final spraying parameter command verified by simulation. Step S4: Decouple the final spraying parameter command into a sequence of coordinated control commands for the pressure regulating valve, flow pump, atomizer, and motion actuator, and issue and execute them; simultaneously start a multi-sensor array to perform high-frequency online detection on the sprayed area to obtain real-time microscopic state data; compare the real-time microscopic state data with the coating quality prediction data, generate a dynamic feedforward compensation amount based on the deviation signal generated by the comparison, and use the dynamic feedforward compensation amount to perform real-time fine-tuning compensation on the currently executing coordinated control command sequence.
[0021] In this embodiment, it is specifically necessary to explain that the process of generating a high-dimensional microenvironment state vector characterizing the comprehensive surface state of the target spraying point in step S1 is as follows: The multimodal raw data, namely infrared thermal radiation data, hyperspectral reflectance data, macro image data, and contact temperature and humidity data, are respectively input into the corresponding branch coding subnetworks in the deep learning feature fusion network for processing. The branch coding subnetwork used to process infrared thermal radiation data and hyperspectral reflectance data is configured as a convolutional neural network structure to extract two-dimensional spatial temperature distribution features in infrared thermal radiation data and spatial spectral joint features in hyperspectral reflectance data. The convolutional neural network structure can contain multiple sequentially connected convolutional layers, batch normalization layers and modified linear activation function layers. The branch coding subnetwork used to process hyperspectral reflectance data also contains three-dimensional convolutional layers to extract features in both spatial and spectral dimensions simultaneously. The branch coding subnetwork used to process macro image data is configured as a convolutional neural network structure to extract surface texture features and visual morphological features of contaminants in macro image data; the branch coding subnetwork may contain shallow convolutional layers for extracting edges and contours, and deep convolutional layers for extracting complex textures and morphological semantics. The branch coding subnetwork used to process contact temperature and humidity data adopts a fully connected network structure to extract local temperature and humidity scalar features. The fully connected network structure contains at least one hidden layer, which maps the average temperature and humidity scalar values of a single point or a small area to feature vectors that match the output dimensions of other branches. Each branch coding subnetwork maps the input multimodal raw data to a high-dimensional feature vector of the same dimension. Subsequently, the high-dimensional feature vectors output by each branch encoding subnetwork are input into a pre-constructed cross-modal attention module in the deep learning feature fusion network. This cross-modal attention module first linearly transforms each high-dimensional feature vector to a common feature subspace based on a learnable weight matrix and bias vector, and calculates the original attention score of each high-dimensional feature vector in the common feature subspace in combination with the learnable context vector. The original attention score reflects the importance of the corresponding infrared thermal radiation, hyperspectral reflectance, macro image, and contact temperature and humidity data in the current microenvironment context. The learnable context vector is used to capture cross-modal global context information and interacts with the transformed high-dimensional feature vectors to obtain a scalar score characterizing the relevance of each modality feature to the current task, i.e., the original attention score. Next, based on the raw attention scores, the attention weight coefficients corresponding to each high-dimensional feature vector are calculated using a normalized exponential function. Specifically, for each high-dimensional feature vector's raw attention score, an exponential value with a base of the natural constant is first calculated. This exponential value is then divided by the sum of the exponential values of all the raw attention scores corresponding to the high-dimensional feature vectors, thus obtaining the attention weight coefficient for that high-dimensional feature vector. In actual calculations, to enhance numerical stability, the maximum value among all raw attention scores can be subtracted before performing the exponential operation. The sum of the calculated attention weight coefficients is 1, and each coefficient is a value greater than 0 and less than 1, intuitively representing the contribution ratio of the corresponding modality feature. Based on the calculated attention weight coefficients, each high-dimensional feature vector is weighted, summed, and finally transformed. The weighted summation and final transformation process is as follows: each attention weight coefficient is multiplied by the result of a linear transformation of its corresponding high-dimensional feature vector, then all products are summed, and the sum is linearly transformed through a learnable weight matrix and superimposed with a bias vector to finally generate a high-dimensional micro-environment state vector. This process realizes adaptive feature selection and fusion based on the environment. "The result of a linear transformation of the corresponding high-dimensional feature vector" refers to the independent linear projection of the original high-dimensional feature vectors of each branch coding sub-network through another set of learnable weight matrices before weighting, which aims to align features of different modalities to a more suitable representation space for fusion. The final generated high-dimensional micro-environment state vector is a dense vector of fixed length, such as 256-dimensional or 512-dimensional, serving as a unified environmental state representation for subsequent steps. The high-dimensional microenvironment state vector is a structured multidimensional data object with fixed data dimensions. The multidimensional data object implicitly encodes the comprehensive state information of the target spraying point surface. Although the multidimensional data structure is a flat vector, after network training, its different dimensions or combinations of dimensions form a strong correlation mapping with the specified physical environment attributes. The components of the high-dimensional microenvironment state vector are determined by the features fused by the cross-modal attention module, including at least the encoding of the surface temperature field distribution pattern, the encoding of the surface humidity gradient change trend, and the encoding of the types and concentration distribution of surface-attached pollutants. The surface temperature field distribution pattern is encoded from the features extracted from infrared thermal radiation data through a branch coding subnetwork and involved in attention-weighted fusion. It is used to characterize the non-uniform thermal state of the surface at the target point. This encoding can distinguish the temperature difference between areas under direct sunlight, shaded areas, and metal joints, as well as areas with uneven thermal conduction caused by thickness and corrosion. Its numerical pattern reflects the spatial characteristics of the temperature gradient. The encoding of the surface humidity gradient trend is derived from the specified spectral band features related to moisture absorption in hyperspectral reflectance data, as well as the features extracted and fused from contact temperature and humidity data through a branch coding subnetwork. This encoding is used to characterize the distribution and changes of moisture on and near the surface. The specified spectral bands related to moisture absorption are usually located in the near-infrared region, such as the water molecule absorption peaks near 1450 nm or 1950 nm. This encoding combines contact humidity measurement at a "point" with spectral humidity inversion information on a "surface", and can comprehensively characterize the gradient changes such as the presence of condensation film and the distribution of humidity under corrosion products. The encoding of surface-attached contaminant types and their concentration distributions is derived from the spectral characteristics of different chemical components in hyperspectral reflectance data and the texture and morphological features extracted from macro image data. The resulting joint feature representation, formed by attention-weighted fusion, is used to distinguish and quantify the presence and extent of different contaminants such as salt, grease, and dust. For example, chloride ion contaminants exhibit absorption characteristics in specific spectral ranges, grease shows specific spectral reflectance curves and oily gloss textures in images, while dust appears as a granular coating and may have smoothed spectral features in images. This joint feature representation, by fusing the chemical recognition capabilities of spectra and the morphological recognition capabilities of images, enables more accurate identification and semi-quantitative assessment of mixed contaminants.
[0022] In this embodiment, it is specifically necessary to explain that in step S2, the process of inputting the high-dimensional microenvironment state vector and the preset target coating specification parameters into a pre-trained deep reinforcement learning decision model is as follows: The high-dimensional microenvironment state vector is spliced with the preset target coating specifications in the data dimension to form a decision state vector. This decision state vector integrates information representing the overall surface state of the target spraying point and the performance requirements of the coating process. The splicing operation specifically involves connecting the two vectors end to end to form a new vector with a length equal to the sum of their lengths. The target coating specifications include at least the target dry film thickness and the minimum adhesion strength level. As a joint digital representation of the "environment-task", the decision state vector is the sole input basis for the policy network to make decisions. The decision state vector is input into the policy network of the deep reinforcement learning decision model. The policy network is a parameterized deep neural network trained using a proximal policy optimization algorithm. The policy network typically consists of an input layer, multiple fully connected hidden layers, and an output layer. The hidden layers use activation functions such as modified linear units to introduce nonlinearity. The output layer has the same dimension as the initial spraying parameter command vector and uses activation functions such as Tanh to constrain the output value within a specified range. The objective function of the near-end policy optimization is constrained by a pruning mechanism. The calculation process is as follows: First, calculate the ratio of the action probabilities of the new and old policy networks in a given state during the training of the deep reinforcement learning decision model. That is, divide the action probability output by the new policy network in the current iteration period by the action probability output by the old policy network in the same state in the previous iteration period. The action probability in the continuous action space is usually represented by the Gaussian distribution probability density corresponding to the action value output by the policy network. Then, the probability ratio of the action is limited to a preset upper and lower limit range. The preset upper and lower limit range is a small neighborhood range centered on the value 1, such as between 80% and 120%. This range, such as [0.8, 1.2], is called the pruning range. Its core function is to limit the step size of a single policy update and prevent training instability or collapse due to excessive policy changes. Next, the product of the pruned action probability ratio and the advantage function is calculated. Finally, the expectation of this product is calculated and gradient ascent is performed to update the network weight parameters, thus obtaining a set of optimal network weight parameters. Here, the advantage function reflects the degree of advantage of taking a certain action relative to the average level in a given state, and is estimated by an independent value network trained in conjunction with the policy network. The value network is also a deep neural network, whose input is the decision state vector and whose output is an estimate of the long-term cumulative reward for that state. The advantage function is calculated by subtracting the estimate of the value network from the actual cumulative reward obtained, reflecting the additional value of a specific action. Gradient ascent update refers to adjusting the network weight parameters in the direction that increases the value of the objective function. The specific iterative process of gradient ascent update based on the objective function constrained by the pruning mechanism is as follows: Calculate the ratio of the action probabilities of the old and new policy networks after pruning, multiply it by the advantage function, then calculate the expected value of this result over all possible actions, and finally adjust the weight parameters of the policy network using the stochastic gradient ascent method along the direction that increases the expected value until the deep reinforcement learning decision model converges, thereby obtaining a set of optimal deep reinforcement learning decision model weight parameters. This process is carried out through a large number of trial-and-error interactive iterations in a virtual environment constructed by a digital twin simulation module. Each iteration includes steps such as collecting a batch of interactive data, calculating the advantage function, and updating the policy network and value network parameters until the policy performance tends to stabilize. The policy network performs multi-layer nonlinear transformations and feature extraction on the input decision state vector, and finally directly generates the initial spraying parameter command vector through the output layer of the policy network. The spraying pressure, paint flow rate, atomization angle, and spray gun movement speed in the initial spraying parameter command vector are all continuous real values. The range of these values is predefined and constrained during the training phase of the deep reinforcement learning decision model according to the physical execution capabilities of the spraying equipment. For example, the spraying pressure range can be defined as 0.2 MPa to 0.8 MPa, the paint flow rate as 50 ml / s to 200 ml / s, the atomization angle as 30 degrees to 80 degrees, and the spray gun movement speed as 50 mm / s to 300 mm / s. At the output of the policy network, a scaling function maps the network output to these actual physical ranges. The pre-training process of the deep reinforcement learning decision model is completed in an offline virtual environment constructed by a digital twin simulation module, and its training objective is driven by a multi-objective reward function. The multi-objective reward function is calculated based on the coating quality result and resource consumption result predicted by the digital twin simulation module after simulating the given decision state vector and the corresponding initial spraying parameter command vector; The multi-objective reward function is specifically composed of three sub-items: coating thickness uniformity reward, coating adhesion compliance reward, and spraying resource consumption penalty, which are linearly weighted and summed according to preset weight coefficients. The calculation logic of the multi-objective reward function is as follows: First, calculate the value of the coating thickness uniformity reward item, add the value of the coating adhesion achievement reward item, then subtract the value of the spraying resource consumption penalty item, and finally multiply the difference by a global scaling factor. The global scaling factor is used to balance the values of each reward sub-item with different physical dimensions in the early stage of training, so that the output value of the multi-objective reward function fluctuates within a reasonable range to facilitate deep neural network learning. It is usually set to a small positive number close to zero, such as 0.01. The preset weight coefficients are used to adjust the relative importance of different optimization objectives. For example, the thickness uniformity weight can be set to 0.5, the adhesion weight to 0.3, and the resource consumption weight to 0.2. The specific values can be adjusted according to the process priority. The coating thickness uniformity bonus is calculated based on the coating thickness distribution map obtained from simulation prediction. Uniformity is evaluated by quantifying the severity of spatial gradient changes in the thickness distribution; the smoother the change, the higher the bonus value. Specifically, the quantification method is to calculate the sum of the absolute values of the thickness differences between adjacent regions in the thickness distribution map, or to calculate the variance of the thickness distribution map. The smaller the variance, the smoother the change, and the higher the bonus value. One specific calculation method is to divide the simulated thickness distribution map into a grid, calculate the sum of the absolute values of the thickness differences between each grid cell and its adjacent cells, and then take the negative value as the uniformity bonus. The smaller the sum of the differences, the higher the bonus value (the less negative the sum). The coating adhesion compliance bonus is evaluated based on the simulated predicted coating film-forming kinetics. A positive bonus is given when the film-forming state meets the preset adhesion formation conditions. The preset adhesion formation conditions usually include a combination of indicators such as the coating curing degree reaching more than 50% and the interfacial shear stress being greater than the critical peel stress of the substrate material. For example, when the simulated curing degree is higher than the threshold of 60% and the interfacial stress is lower than the threshold of 0.5 MPa, the adhesion is judged to be compliant and a fixed positive bonus value (such as +1.0) is given; otherwise, zero bonus or a small negative bonus is given. The spraying resource consumption penalty term is proportional to the weighted sum of paint flow rate and spraying pressure in the initial spraying parameter command vector, which is used to encourage material and energy conservation; for example, the penalty term can be designed as: penalty value = 0.01 * paint flow rate + 0.05 * spraying pressure, where the coefficient converts the flow rate and pressure into an approximate cost metric; During the training of the deep reinforcement learning decision model, the output value of the calculated multi-objective reward function is used as a feedback signal and backpropagated to the policy network and value network. The weight parameters of the policy network are continuously adjusted so that the policy network learns to output the optimal initial spraying parameter instruction vector after comprehensively balancing coating quality and operation cost. The trained policy network is solidified and used directly for rapid inference in the online application stage to generate the initial spraying parameter instruction vector.
[0023] In this embodiment, it is specifically necessary to explain that in step S3, the process of loading the high-dimensional microenvironment state vector and the initial spraying parameter command vector into the digital twin simulation module for rapid simulation calculation and robustness evaluation is as follows: Using the high-dimensional microenvironment state vector as the environmental boundary condition and the initial spraying parameter command vector as the nominal action, a set of parallel disturbance actions is generated in the digital twin simulation module. The parallel disturbance action set is generated as follows: First, based on the historical control accuracy and sensor measurement error of the spraying equipment, a relative disturbance intensity vector is preset. The relative disturbance intensity vector defines the allowable small fluctuation percentage range of parameters such as spraying pressure, paint flow rate, atomization angle, and spray gun movement speed. Based on the parameter values of the nominal action, multiple parameter disturbance values within the set small fluctuation percentage range are randomly generated according to the small fluctuation percentage range defined by the relative disturbance intensity vector. Each parameter disturbance value is added to the nominal action parameter value item by item to obtain multiple disturbanced spraying parameter combinations. These disturbanced parameter combinations together with the original nominal action constitute the parallel disturbance action set. The relative disturbance intensity vector can be set according to the specific equipment performance. For example, the disturbance range of spraying pressure is set to ±5% of the nominal value, the disturbance range of paint flow rate is set to ±3%, the disturbance range of atomization angle is set to ±2%, and the disturbance range of spray gun movement speed is set to ±4%. Subsequently, the digital twin simulation module is driven to perform fast, coupled computational fluid dynamics and coating film-forming dynamics simulations in parallel for each spraying parameter combination in the parallel perturbation action set, using the high-dimensional microenvironment state vector as a unified initial field. Each simulation calculation outputs a corresponding subset of coating quality prediction data, which includes at least the predicted coating thickness spatial distribution matrix, the coating cure degree spatial distribution matrix, and the maximum stress value at the coating-substrate interface. The simulation calculations are accelerated in parallel on a graphics processor to meet millisecond-level response requirements. Each element of the coating thickness spatial distribution matrix represents the predicted thickness value of the coating on the corresponding micro-grid cell, and each element of the coating cure degree spatial distribution matrix represents the predicted degree of curing of the coating on the corresponding grid cell, with values between zero and one. The process of generating the parallel perturbation action set is further refined as follows: the nominal action is determined by the initial spraying parameter command vector and is denoted as the first action; then, based on the preset relative perturbation intensity vector, a set of random perturbation vectors is generated. Each random perturbation vector is subjected to a Hadamard product operation with the nominal action to obtain the corresponding perturbation increment. The perturbation increment is added to the nominal action item by item to generate the corresponding perturbation action; this process is repeated until a predetermined number of perturbation actions are generated. Finally, the first action and all perturbation actions are summarized to form a parallel perturbation action set; the predetermined number can be set according to the computing resources and accuracy requirements, for example, generating ten perturbation actions; the Hadamard product operation refers to multiplying the elements at corresponding positions of two vectors to obtain a new vector of the same dimension; The process of evaluating the robustness of the initial spraying parameter command vector under a preset perturbation and outputting the final spraying parameter command verified by simulation is as follows: Based on all the coating quality prediction data subsets obtained from parallel simulation, a robust loss value is calculated to quantify the degree of fluctuation in coating quality prediction data when the nominal action is subject to minor parameter perturbations. The robustness loss value is calculated as follows: First, the thickness field structural sensitivity index and the solidified state overall distribution offset index are calculated; the thickness field structural sensitivity index is multiplied by a preset first weighting coefficient, and the solidified state overall distribution offset index is multiplied by a preset second weighting coefficient; then the two products are added together, and finally the sum is used as the final robustness loss value; the sum of the first weighting coefficient and the second weighting coefficient is equal to one, which is used to assign importance weights to the two indices. For example, the first weighting coefficient can be set to 0.6 and the second weighting coefficient to 0.4. The robustness loss value is calculated by combining the thickness field structure sensitivity index and the overall distribution offset index of the cured state. The thickness field structure sensitivity index is calculated by comparing the difference between the thickness distribution matrix obtained from the nominal simulation and the thickness distribution matrix obtained from each perturbation simulation. Specifically, within the two-dimensional planar space representing the coating distribution, for each tiny spatial grid cell, the difference between the thickness value of the grid cell in the thickness distribution matrix corresponding to the nominal action and the thickness value of the corresponding grid cell in the thickness distribution matrix corresponding to the perturbation action is calculated. These differences are arranged according to spatial position to form a difference matrix. Then, the sum of squares of all elements in the difference matrix is calculated, and the arithmetic mean of the sum of squares of all spatial grid cells is obtained. The square root of the average value reflects the average change intensity of the thickness field spatial gradient structure caused by parameter perturbation. The arithmetic mean of the square roots of the average values corresponding to all perturbation actions is then obtained to obtain the thickness field structure sensitivity index. This calculation process essentially evaluates the local structural stability of the thickness distribution map. The smaller the thickness field structure sensitivity index value, the lower the sensitivity of the thickness distribution to parameter perturbation and the better the robustness. The overall distribution deviation index of the solidified state is calculated by comparing the overall probability distribution difference between the set of solidified state values obtained from the nominal action simulation and the set of solidified state values obtained from the simulation of various disturbance actions. Specifically, the Hellinger distance is used to measure the difference between the two probability distributions. The calculation method is as follows: First, the set of solidified state values obtained from the nominal action simulation is statistically represented as one probability density function, and the set of solidified state values obtained from the disturbance action simulation is statistically represented as another probability density function. Then, the integral of the square of the difference between the square roots of the two probability density functions is calculated over the entire numerical domain. The integral result is multiplied by one-half and the square root is taken. The calculation result can symmetrically and stably measure the overall change of the distribution pattern. The arithmetic mean of the calculation results corresponding to all disturbance actions is taken to obtain the overall distribution deviation index of the solidified state. The calculation result of the Hellinger distance ranges from zero to one, where zero indicates that the two probability distributions are exactly the same, and the larger the value, the greater the difference. The smaller the value of the overall distribution deviation index of the solidified state, the lower the sensitivity of the solidified state distribution to parameter disturbances. Among them, the first weighting coefficient and the second weighting coefficient are used to balance the contribution of the thickness field structure sensitivity index and the overall distribution deviation index of the curing state in the robustness loss value. They are usually set according to the emphasis of the process on the requirements of coating thickness uniformity and curing consistency. The calculated thickness field structural sensitivity index and the overall distribution offset index of the solidified state are weighted and summed according to preset weights to obtain the final robustness loss value; the preset weights are the first weight coefficient and the second weight coefficient. The robustness loss value is compared with a preset robustness threshold. If the robustness loss value is less than or equal to the robustness threshold, the initial spraying parameter command vector is directly used as the final spraying parameter command output verified by simulation. If the robustness loss value is greater than the robustness threshold, the initial spraying parameter command vector is determined to be insufficiently robust, triggering a local parameter optimization process running within the digital twin simulation module. This process aims to minimize the robustness loss value by fine-tuning the initial spraying parameter command vector, and the optimized parameter vector is used as the final spraying parameter command output verified by simulation. The robustness threshold can be set according to process tolerance requirements, for example, set to 0.1. The local parameter optimization process is triggered when the robustness loss value exceeds a preset robustness threshold. The specific iterative logic of the local parameter optimization process is as follows: using the initial spraying parameter command vector as the starting optimization variable, the robustness loss value corresponding to the current parameter vector is calculated in the digital twin simulation module, and the gradient of the loss value with respect to the parameter vector is calculated. The parameter vector is adjusted in the direction that reduces the robustness loss value until the robustness loss value corresponding to the adjusted parameter vector is less than or equal to the preset robustness threshold, or the preset maximum number of iterations is reached. The final parameter vector is then output as the final spraying parameter command verified by simulation. The local parameter optimization process uses the gradient descent method, with an initial learning rate set to 0.01 and a maximum number of iterations set to twenty. In each iteration, the gradient of the robustness loss value with respect to the spraying pressure, paint flow rate, atomization angle, and spray gun movement speed parameters is calculated through forward and backward propagation in the digital twin simulation, and the parameter values are updated accordingly.
[0024] In this embodiment, it is specifically necessary to explain that in step S4, the process of decoupling the final spraying parameter command verified by simulation into a sequence of coordinated control commands for the pressure regulating valve, flow pump, atomizer, and motion actuator, and then issuing and executing it, is as follows: Receive the final spraying parameter command verified by simulation. The final spraying parameter command includes the spraying pressure setting value, paint flow setting value, atomization angle setting value, spray gun moving speed setting value and spray gun spatial path planning sequence. The final spraying parameter commands verified by simulation are input into a pre-built collaborative command generator. The collaborative command generator performs a multi-objective optimization calculation based on the physical response dynamics model and electrical characteristic model of the pressure regulating valve, flow pump, atomizer, and each motion actuator. The physical response dynamics model and electrical characteristic model are collectively referred to as the actuator dynamic model, which is used to describe the mathematical relationship between the input commands and output responses of each actuator. The actuator dynamic model usually takes the form of a transfer function or state-space equation. For example, it can be a first-order inertia plus pure delay model, used to characterize the response characteristics from control voltage to output pressure, from control duty cycle to output flow, from control current to atomization angle, and from joint angular velocity commands to actual end pose. The objective function of the multi-objective optimization aims to simultaneously minimize the drastic changes in the control commands of each actuator in the generated sequence of cooperative control commands, and minimize the tracking error between the actual output pressure, actual output flow, actual output atomization angle and actual spray gun pose predicted based on the actuator dynamic model and the time-varying reference trajectory generated by the final spraying parameter commands. The degree of change in the control commands of each actuator is quantified by weighted sum of squares of the rate of change of the control command vector over time. The weights used in the weighted sum of squares calculation, namely the preset first weight matrix, have diagonal elements corresponding to the penalty intensity for the rate of change of pressure control command, flow control command, atomization control command, and motion control command, respectively, while the off-diagonal elements can be used to characterize the coupling penalty between the changes in different actuator commands. The tracking error of the actual output to the time-varying reference trajectory is quantified by weighted sum of squares of the differences between the actual output pressure, actual output flow, actual output atomization angle, and actual spray gun pose predicted based on the actuator dynamic model and their respective time-varying reference values. The other weight used in the weighted sum of squares calculation, namely the preset second weight matrix, has diagonal elements corresponding to the penalty intensity for pressure tracking error, flow tracking error, atomization angle tracking error, and pose tracking error, respectively. By solving a multi-objective optimization problem, specifically, within a preset time domain, a control command trajectory that minimizes the objective function is found; the preset time domain is the estimated total time required to complete the spraying at the current point; the time-varying reference trajectory is generated by smooth interpolation (such as fifth-order polynomial interpolation) from the set values in the final spraying parameter command to ensure that each physical quantity is continuous and differentiable in time, and to avoid step changes; The objective function is constructed as follows: First, the derivative of the control command vector corresponding to the coordinated control command sequence with respect to time is calculated. Then, the derivative is used to perform a quadratic operation with the preset first weight matrix to evaluate the degree of drastic change in the control command. Secondly, the difference between the actual output pressure, actual output flow, actual output atomization angle, and actual spray gun posture predicted by the actuator dynamic model and the corresponding time-varying reference trajectory is calculated. The difference is then used to perform a quadratic operation with the preset second weight matrix to evaluate the degree of deviation from the reference trajectory. Finally, the results of these two operations are integrated and summed in the time domain to obtain the final objective function value; the objective function value represents the sum of the smooth transition cost of the control command and the trajectory tracking error cost over the entire spraying time domain. The extreme value problem is solved by numerical iterative algorithm, and a set of coordinated control command sequences that are synchronized in time and coupled in control quantities are calculated. The coordinated control command sequence consists of the control voltage signal sequence of the pressure regulating valve, the drive duty cycle signal sequence of the flow pump, the control current signal sequence of the atomizer, and the angular velocity or position command sequence of each motion actuator. The numerical iterative algorithm can be a quadratic programming solver under the model predictive control framework. The coordinated control command sequence is synchronously sent to the corresponding pressure regulating valve, flow pump, atomizer and motion actuator in chronological order; The process of synchronously activating the multi-sensor array to perform high-frequency online detection on the coated area, obtaining real-time microscopic state data, and comparing the real-time microscopic state data with the coating quality prediction data to generate dynamic feedforward compensation is as follows: After the coordinated control command sequence begins execution, the multi-sensor array is immediately driven to continuously and synchronously scan the sprayed micro-area, which is in a wet film state, within a fixed lag distance behind the spray gun at a sampling rate higher than the scanning frequency in step S1. Infrared thermal radiation data, hyperspectral reflectance data, and macro image data of the sprayed micro-area are collected as real-time microscopic state data. The sampling rate of high-frequency online detection can reach more than 1,000 times per second to capture the rapid dynamic process of wet film formation. The fixed lag distance is determined according to the spray gun moving speed and the initial curing time of the coating to ensure that the coating in the detection area has been initially deposited but has not yet been fully cured. The collected real-time microscopic state data is input into a pre-defined lightweight online feature extraction network. The online feature extraction network shares some encoding layer parameters with the deep learning feature fusion network in step S1. This network is used to quickly extract a real-time microscopic state feature vector that characterizes the transient properties of the current wet film. The real-time microscopic state feature vector encodes the local film thickness variation trend, surface leveling texture features, and specified spectral features related to the solvent evaporation rate. The online feature extraction network typically contains only a few convolutional or fully connected layers to complete inference within milliseconds. The shared encoding layer parameters ensure that its feature representation has semantic consistency with the high-dimensional microenvironment state vector in step S1. Meanwhile, from the coating quality prediction data corresponding to the current spraying point output in step S3, the predicted micro-state feature vector on the sprayed micro-area is extracted as a reference benchmark for comparison; the predicted micro-state feature vector is output by the digital twin simulation module in step S3 in the corresponding area, or obtained by performing the same feature extraction operation on the thickness and curing degree distribution map output by the simulation. Calculate the deviation signal between the real-time microstate feature vector and the predicted microstate feature vector; The deviation signals from the current moment and multiple consecutive moments within a short time window in the past, along with the control commands from the cooperative control command sequence being issued at the current moment, are input into a pre-trained lightweight temporal attention compensation network. The deviation signals from the current moment and multiple consecutive moments within a short time window in the past constitute a deviation signal sequence. The short time window typically covers several to dozens of control cycles to provide sufficient temporal context information. The lightweight temporal attention compensation network analyzes the temporal pattern characteristics of the deviation signal sequence through its internal attention mechanism, identifies whether the deviation is continuous, oscillating, or divergent, and infers a dynamic feedforward compensation amount by combining the current control command state. The lightweight temporal attention compensation network can focus on the historical moment in the deviation sequence that is most relevant to the current compensation decision. The compensation network is trained on historical data through supervised learning or reinforcement learning, learning the mapping from "deviation pattern + current control command" to "optimal compensation action". The dynamic feedforward compensation is a vector whose dimension corresponds to the cooperative control command. It includes fine-tuning command values for pressure regulating valves, flow pumps, atomizers, and motion actuators. The fine-tuning command value is usually a small percentage adjustment of the original control command value, such as within ±5%. The process by which the dynamic feedforward compensation amount performs real-time fine-tuning compensation on the currently executing cooperative control command sequence is as follows: Within each control cycle, a dynamic feedforward compensation amount generated by a lightweight temporal attention compensation network is received; the control cycle is consistent with the discrete time interval of the cooperative control command sequence, typically on the order of milliseconds, for example, one to ten milliseconds; The dynamic feedforward compensation quantity is added to the original control command from the coordinated control command sequence that is planned to be issued in the current control cycle. That is, each parameter component of the original control command is added to the corresponding parameter component of the dynamic feedforward compensation quantity to generate a set of coordinated control commands that take effect immediately at the current moment and have been compensated and corrected. This operation is a vector addition and is completed before the next control command is issued. Next, the compensated and corrected collaborative control commands are sent in real time to the pressure regulating valve, flow pump, atomizer and motion actuator, directly covering and replacing the original control commands that were originally scheduled to be sent, thereby changing the instantaneous control quantity that will be applied to the physical spraying equipment; this means that the compensation is feedforward, aiming to correct the spraying effect that will be produced in the next control cycle, rather than correcting the deviation that has already occurred. Subsequently, in the next control cycle, the multi-sensor array again performs high-frequency detection on the latest state of the sprayed area to obtain new real-time micro-state data; and repeats the complete process: inputting the new real-time micro-state data into the online feature extraction network to extract the new real-time micro-state feature vector, calculating the new deviation signal between it and the predicted micro-state feature vector, adding the new deviation signal to the deviation signal sequence (removing the oldest signal), inputting it into the lightweight temporal attention compensation network to generate a new dynamic feedforward compensation amount, and again performing the correction of the command for the next control cycle based on this new compensation amount; The "execution-perception-compensation" loop runs continuously at millisecond intervals until the entire spraying process for the current target spraying point is completed. Through feedforward real-time fine-tuning, it offsets coating quality deviations caused by batch differences in paint, equipment dynamic response errors, and unmodeled environmental disturbances, ensuring that the actual film formation state closely tracks the coating quality prediction data output in step S3. This closed loop enables the spraying process to adaptively adjust even in the presence of model mismatch and unknown disturbances, ultimately achieving a high degree of consistency between key quality indicators such as coating thickness and curing state and simulation predictions.
[0025] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0026] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0027] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0028] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0029] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0030] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0031] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for targeted microenvironment control spraying for bridge bearing maintenance, characterized in that, Specifically, the following steps are included: Step S1: The target spraying point of the bridge support is synchronously scanned by a multi-sensor array deployed at the end of the spraying equipment to collect multimodal raw data, including infrared thermal radiation data, hyperspectral reflectance data, macro image data, and contact temperature and humidity data. The multimodal raw data is fused using a preset deep learning feature fusion network to generate a high-dimensional microenvironment state vector representing the comprehensive surface state of the target spraying point. Specifically, the multimodal raw data, namely infrared thermal radiation data, hyperspectral reflectance data, macro image data, and contact temperature and humidity data, are respectively input into the corresponding branch coding subnetworks in the deep learning feature fusion network for processing. Specifically, the branch coding subnetwork for processing infrared thermal radiation data and hyperspectral reflectance data is configured with a convolutional neural network structure, the branch coding subnetwork for processing macro image data is configured with a convolutional neural network structure, and the branch coding subnetwork for processing contact temperature and humidity data adopts a fully connected network structure; each branch coding subnetwork maps the input multimodal raw data into a high-dimensional feature vector of a unified dimension. Subsequently, the high-dimensional feature vectors output by each branch encoding subnetwork are input into the pre-built cross-modal attention module in the deep learning feature fusion network; the cross-modal attention module linearly transforms each high-dimensional feature vector to a common feature subspace, and calculates the original attention score of each high-dimensional feature vector in the common feature subspace by combining the learnable context vector; Based on the original attention score, the attention weight coefficients corresponding to each high-dimensional feature vector are calculated using a normalized exponential function; based on the calculated attention weight coefficients, each high-dimensional feature vector is weighted, summed, and finally transformed to generate a high-dimensional micro-environment state vector. Step S2: Input the high-dimensional microenvironment state vector and the preset target coating specification parameters into a pre-trained deep reinforcement learning decision model; The deep reinforcement learning decision model is trained based on a digital twin simulation module that integrates computational fluid dynamics and coating film formation dynamics. The deep reinforcement learning decision model outputs an initial spraying parameter command vector, which includes spraying pressure, coating flow rate, atomization angle and spray gun movement speed. Step S3: Load the high-dimensional microenvironment state vector and the initial spraying parameter command vector into the digital twin simulation module; drive the digital twin simulation module to perform rapid simulation calculations on the spraying process, predict the coating quality prediction data, and evaluate the robustness of the initial spraying parameter command vector under preset perturbations; output the final spraying parameter command verified by simulation. Step S4: Decouple the final spraying parameter command into a sequence of coordinated control commands for the pressure regulating valve, flow pump, atomizer, and motion actuator, and issue and execute them; simultaneously start a multi-sensor array to perform high-frequency online detection on the sprayed area to obtain real-time microscopic state data; compare the real-time microscopic state data with the coating quality prediction data, generate a dynamic feedforward compensation amount based on the deviation signal generated by the comparison, and use the dynamic feedforward compensation amount to perform real-time fine-tuning compensation on the currently executing coordinated control command sequence.
2. The fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 1, characterized in that: The high-dimensional microenvironment state vector is a structured multidimensional data object, which implicitly encodes the comprehensive state information of the surface of the target spraying point. The components of a high-dimensional microenvironment state vector include encoding the surface temperature field distribution pattern, encoding the surface humidity gradient variation trend, and encoding the types and concentration distribution of surface-attached pollutants. The surface temperature field distribution pattern is encoded from the features extracted from infrared thermal radiation data by a branch coding subnetwork and involved in attention-weighted fusion. The encoding of the surface humidity gradient trend comes from the specified spectral band features related to moisture absorption in hyperspectral reflectance data, as well as the features extracted from contact temperature and humidity data by a branch coding subnetwork and fused together. The encoding of surface-attached contaminant types and their concentration distributions is derived from the spectral features characterizing different chemical components in hyperspectral reflectance data, as well as the texture and morphological features extracted from macro image data, which are then combined into a joint feature representation after attention-weighted fusion.
3. The fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 2, characterized in that: In step S2, the process of inputting the high-dimensional microenvironment state vector and the preset target coating specification parameters into a pre-trained deep reinforcement learning decision model is as follows: A decision state vector is formed by concatenating a high-dimensional microenvironment state vector with preset target coating specification parameters; the decision state vector is then input into the policy network of a deep reinforcement learning decision model, which is a parameterized deep neural network trained using a proximal policy optimization algorithm. The objective function for near-end policy optimization is constrained by a pruning mechanism. The calculation process is as follows: First, the ratio of the action probability of the new and old policy networks under a given state is calculated during the training of the deep reinforcement learning decision model; then, the action probability ratio is constrained within a preset upper and lower limit range. The optimal network weight parameters are obtained by multiplying the ratio of the action probabilities after pruning with the advantage function, taking the expectation of the product and updating it using gradient ascent; whereby the advantage function is estimated by the value network trained in conjunction with the policy network. The policy network performs multi-layer nonlinear transformation and feature extraction on the input decision state vector, and finally generates the initial spraying parameter command vector directly through the output layer of the policy network. The spraying pressure, paint flow rate, atomization angle and spray gun movement speed in the initial spraying parameter command vector are all continuous real values. The range of their real values is predefined and constrained according to the physical execution capability of the spraying equipment during the training stage of the deep reinforcement learning decision model.
4. A fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 3, characterized in that: The pre-training process of the deep reinforcement learning decision model is completed in an offline virtual environment constructed by a digital twin simulation module, and its training objective is driven by a multi-objective reward function. The multi-objective reward function is calculated based on the coating quality result and resource consumption result predicted by the digital twin simulation module after simulating the given decision state vector and the corresponding initial spraying parameter command vector; The multi-objective reward function consists of three sub-items: coating thickness uniformity reward, coating adhesion compliance reward, and spraying resource consumption penalty, which are linearly weighted and summed according to preset weight coefficients. The coating thickness uniformity bonus is calculated based on the coating thickness distribution map obtained from simulation prediction, and the uniformity is evaluated by quantifying the severity of the spatial gradient change in the thickness distribution. The coating adhesion achievement bonus is evaluated based on the simulated prediction of the coating film formation dynamics. When the film formation state meets the preset adhesion formation conditions, a positive bonus is given. The penalty for spraying resource consumption is proportional to the weighted sum of paint flow rate and spraying pressure in the initial spraying parameter command vector; During the training of the deep reinforcement learning decision model, the output value of the calculated multi-objective reward function is used as a feedback signal and propagated back to the policy network and value network. The weight parameters of the policy network are continuously adjusted so that the policy network learns to output the optimal initial spraying parameter instruction vector after comprehensively balancing coating quality and operating cost.
5. A fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 4, characterized in that: In step S3, the process of loading the high-dimensional microenvironment state vector and the initial spraying parameter command vector into the digital twin simulation module for rapid simulation calculation and robustness evaluation is as follows: Using the high-dimensional microenvironment state vector as the environmental boundary condition and the initial spraying parameter command vector as the nominal action, a set of parallel disturbance actions is generated in the digital twin simulation module. The method for generating the parallel perturbation action set is as follows: based on the parameter value of the nominal action, according to a small fluctuation percentage range defined by a preset relative perturbation intensity vector, multiple parameter perturbation values are randomly generated. Each parameter perturbation value is added to the nominal action parameter value item by item to obtain multiple perturbed spraying parameter combinations. These perturbed parameter combinations together with the original nominal action constitute the parallel perturbation action set. The driving digital twin simulation module uses a high-dimensional microenvironment state vector as a unified initial field to perform fast simulation calculations of coupled computational fluid dynamics and coating film formation dynamics for each combination of spraying parameters in the set of parallel disturbance actions. Each simulation calculation outputs a corresponding subset of coating quality prediction data, which includes the predicted spatial distribution matrix of coating thickness, the spatial distribution matrix of coating curing degree, and the maximum stress value at the interface between the coating and the substrate.
6. A fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 5, characterized in that: The process of evaluating the robustness of the initial spraying parameter command vector under a preset perturbation and outputting the final spraying parameter command verified by simulation is as follows: A robust loss value is calculated based on a subset of all coating quality prediction data obtained from parallel simulation. The calculation logic for the robustness loss value is as follows: First, calculate the values of the thickness field structural sensitivity index and the solidification state overall distribution offset index; multiply the thickness field structural sensitivity index by a preset first weighting coefficient, and multiply the solidification state overall distribution offset index by a preset second weighting coefficient; then add the two product results together, and finally use the sum as the final robustness loss value. The thickness field structure sensitivity index is calculated by comparing the difference between the thickness distribution matrix obtained from the nominal simulation and the thickness distribution matrix obtained from the simulation of each perturbation action. Specifically, it is to calculate the average change intensity of the thickness field spatial gradient structure caused by parameter perturbation. The overall distribution offset index of the solidified state is calculated by comparing the overall probability distribution difference between the set of solidified state values obtained from the simulation of nominal actions and the set of solidified state values obtained from the simulation of each disturbance action. Specifically, the Hellinger distance is used to measure the difference between the two probability distributions. The robustness loss value is compared with a preset robustness threshold. If the robustness loss value is less than or equal to the robustness threshold, the initial spraying parameter command vector is directly used as the final spraying parameter command output verified by simulation. If the robustness loss value is greater than the robustness threshold, a local parameter optimization process running in the digital twin simulation module is triggered to minimize the robustness loss value. The initial spraying parameter command vector is fine-tuned, and the optimized parameter vector is used as the final spraying parameter command output verified by simulation.
7. A fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 6, characterized in that: In step S4, the process of decoupling the final spraying parameter commands verified by simulation into a sequence of collaborative control commands and issuing them for execution is as follows: Receive the final spraying parameter command verified by simulation; The final spraying parameter commands verified by simulation are input into a pre-built collaborative command generator. The collaborative command generator performs a multi-objective optimization calculation based on the physical response dynamics model and electrical characteristic model of the pressure regulating valve, flow pump, atomizer and each motion actuator. The physical response dynamics model and electrical characteristic model are collectively referred to as the actuator dynamic model. The objective function of the multi-objective optimization aims to simultaneously minimize the drastic changes in the control commands of each actuator in the generated sequence of cooperative control commands, and minimize the tracking error between the actual output pressure, actual output flow, actual output atomization angle and actual spray gun pose predicted based on the actuator dynamic model and the time-varying reference trajectory generated by the final spraying parameter commands. By solving a multi-objective optimization problem, a set of coordinated control command sequences that are synchronized in time and coupled in terms of control variables is calculated. The coordinated control command sequence is synchronously sent to the corresponding pressure regulating valve, flow pump, atomizer and motion actuator in chronological order.
8. A fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 7, characterized in that: The process of synchronously activating the multi-sensor array to perform high-frequency online detection on the coated area, obtaining real-time microscopic state data, and comparing the real-time microscopic state data with the coating quality prediction data to generate dynamic feedforward compensation is as follows: After the collaborative control command sequence begins to execute, the multi-sensor array is immediately driven to continuously and synchronously scan the sprayed micro area within a fixed lag distance behind the spray gun at a sampling rate higher than the scanning frequency in step S1, and collect real-time microscopic state data of the sprayed micro area. The collected real-time microstate data is input into a preset lightweight online feature extraction network to quickly extract a real-time microstate feature vector that characterizes the current transient properties of the wet film. Meanwhile, from the coating quality prediction data in step S3, the predicted micro-state feature vectors on the sprayed micro-areas are extracted as a reference benchmark for comparison. Calculate the deviation signal between the real-time microstate feature vector and the predicted microstate feature vector; The deviation signals from the current moment and multiple consecutive moments within a short time window in the past, along with the control commands in the current sequence of coordinated control commands, are input into a pre-trained lightweight temporal attention compensation network. The lightweight temporal attention compensation network analyzes the temporal pattern characteristics of the deviation signal sequence through its internal attention mechanism, and infers and generates a dynamic feedforward compensation amount by combining the current control command status.
9. A fixed-point microenvironment controlled spraying method for bridge bearing maintenance according to claim 8, characterized in that: The process by which the dynamic feedforward compensation amount performs real-time fine-tuning compensation on the currently executing cooperative control command sequence is as follows: Within each control cycle, receive dynamic feedforward compensation generated by a lightweight temporal attention compensation network; The dynamic feedforward compensation is added to the original control commands from the coordinated control command sequence that are planned to be issued in the current control cycle, and a set of coordinated control commands that have been compensated and corrected is generated. The compensated and corrected coordinated control commands are issued in real time to the pressure regulating valve, flow pump, atomizer and motion actuator, covering and replacing the original control commands that were originally planned to be issued. Subsequently, in the next control cycle, the steps of high-frequency online detection, feature extraction, bias calculation, compensation network inference, and instruction correction are repeated. The cycle of execution, perception, and compensation runs continuously at millisecond intervals until the entire spraying process on the current target spraying point is completed.
Citation Information
Patent Citations
Multi-mode environment cooperative regulation and control system for poultry egg preservation
CN121277057A
Spraying system based on multi-modal vision and artificial intelligence self-correction and control method thereof
CN121650004A