Multi-region in-situ working condition automatic monitoring system based on multi-modal fusion

CN122548629APending Publication Date: 2026-08-11INST OF CHEM CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本发明的一个目的在于提出基于多模态融合的多区域原位工况自动监测系统,利用三维精准位移台实现多区域原位工况监测,借助人工智能技术实现“监测-分析-调控”的闭环操作,解决现有方法中监测界面单一、物理误差难以纠偏与观测策略缺乏评估和自优化等问题

Benefits of technology

本发明通过执行多模态信号的光谱动态演变判定与图像形变校正,结合工况参数进行时空索引对应以构建六维标准化输入数据集;将输入数据集输入Uniformer模型,通过空间块与时间块交替堆叠以挖掘空间异质性并捕获动态演变趋势,经跨模态融合输出界面状态特征编码;将编码输入至改进的QPSO算法中,创新引入基于设备运动历史轨迹的概率分布非对称偏置机制,将对称指数分布分裂为具方向偏置的非对称概率波,在兼顾设备损耗与数据信息增益的适应度函数指导下输出最优观测坐标、频率与激励参数;基于最优参数下发多线程联动指令执行自动移位、信号触发、工况修改及区域增删,并通过闭环模块在无人值守下动态迭代观测策略生成时空演化数据集。针对化学体系在真实工作状态下呈现的动态非平衡特性,传统原位表征仅对单一位置监测,难以全面反映催化剂表面反应、电极腐蚀或电池充放电等过程在空间上的显著非均匀性与潜在异质性。人为捕获空间异质性需要长时间、高频次调节仪器,且反应持续数小时乃至数天,长时值守与重复操作对人员精力消耗极大,时序控制一致性差。本发明通过上述机制,将人工智能驱动的自主决策与实验参数实时调控深度融合于原位/工况研究策略中,以电化学体系为典型场景,最终实现了跨区域时空关联分析与系统自适应运行,有效解决了现有方法中空间异质性表征难、人机交互成本高与长时序控制一致性差等问题,突破现有研究瓶颈,推动化学原位表征向“无人化”智能实验范式演进。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548629A_ABST
    Figure CN122548629A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-region in-situ automatic monitoring system for operating conditions based on multimodal fusion, comprising the following modules: a multi-region monitoring and reaction state control module, used to acquire spatiotemporal sequences of active sites and multi-channel in-situ signals, and output a set of electrochemical-thermal-mechanical multi-field operating condition parameters; a data acquisition and preprocessing module, used for signal determination and image correction, and constructing a six-dimensional standardized dataset; a feature extraction module, which uses a Uniformer model to fuse spatiotemporal features and output them; a strategy optimization module, which introduces an asymmetric bias mechanism through an improved QPSO algorithm, and outputs the optimal observation position and excitation parameter set; and a linkage and closed-loop control module, which issues commands based on the optimal parameters to automatically control equipment movement and operating condition adjustment, and dynamically iterates strategies under unattended operation. This invention realizes unmanned intelligent closed-loop monitoring of dynamic interface processes in diverse chemical environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of in-situ and operational condition monitoring of chemical reaction processes, and in particular to a multi-regional in-situ automatic operational condition monitoring system based on multimodal fusion. Background Technology

[0002] Chemical systems (such as catalysis, material corrosion, and electrochemical energy) exhibit dynamic and non-equilibrium characteristics under real-world operating conditions. In-situ / operating-condition characterization allows for real-time, dynamic observation of reaction interfaces, revealing dynamic reaction processes and establishing a direct link from microscopic mechanisms to macroscopic performance. However, many chemical processes (such as catalyst surface reactions, electrode corrosion, and battery charging and discharging) exhibit significant spatial non-uniformity, making it difficult to comprehensively reflect the spatial heterogeneity of interfacial behavior by monitoring only a single location. Manually capturing spatial heterogeneity requires long-term, high-frequency adjustments and recordings, resulting in high manpower consumption and poor consistency in timing control. Furthermore, chemical reactions often last for hours or even days, requiring operators to maintain continuous monitoring and perform high-frequency repetitive operations to clearly capture the dynamic details of interfacial reactions. This poses a significant challenge to the physical and mental well-being of operators. Therefore, a multi-region in-situ automatic monitoring system for operating conditions, coupled with a three-dimensional precision displacement device and multimodal testing, has been developed to address this need. In practical applications, the deployment effect of multi-region in-situ automatic monitoring system based on multimodal fusion is still constrained by factors such as rapid evolution of interface state and complex multi-field coupling environment.

[0003] Currently, some systems only use preset procedures with fixed timing to evaluate monitoring paths, ignoring the combined effects of multiple factors such as data information gain, equipment wear and tear, and acquisition frequency adaptability, thus limiting the adaptive optimization capability of observation strategies. Furthermore, the black-box process of strategy generation lacks a physically meaningful interpretable path, making it difficult to provide researchers with clear guidelines for site selection and control, affecting the reliability and usability of monitoring results.

[0004] Furthermore, most existing automatic monitoring systems employ statically symmetrical probabilistic sampling mechanisms in translation stage control, failing to dynamically adjust their distribution patterns based on the equipment's historical motion trajectory. This makes them ill-suited for continuous tracking of the reaction interface at different spatial scales, severely impacting the system's positioning accuracy and data quality in long-term real-world experiments. Therefore, providing a multi-regional in-situ automatic monitoring system based on multimodal fusion is a problem urgently needing to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a multi-region in-situ automatic monitoring system based on multimodal fusion. This system utilizes a three-dimensional precision displacement stage to achieve multi-region in-situ monitoring of working conditions and leverages artificial intelligence technology to realize a closed-loop operation of "monitoring-analysis-control," addressing the problems of single monitoring interfaces, difficulty in correcting physical errors, and lack of evaluation and self-optimization of observation strategies in existing methods. This invention fully integrates key steps such as six-dimensional tensor construction, spatiotemporal feature extraction from the Uniformer model, improved QPSO algorithm optimization, and multi-threaded linkage closed-loop control, constructing a research paradigm with data standardization, cross-modal heterogeneity mining, and adaptive iteration of observation strategies. It achieves comprehensive perception, autonomous point selection, and closed-loop iteration of dynamic interfaces in complex chemical environments. By introducing an improved QPSO algorithm and innovatively incorporating an asymmetric bias mechanism based on the probability distribution of the equipment's historical motion trajectory, this invention effectively overcomes positioning deviations caused by mechanical clearances of the translation stage. Simultaneously, by combining the Uniformer model to deeply analyze interface heterogeneity and dynamic evolution trends, it achieves automatic capture of multi-regional reaction sites and intelligent optimization of observation strategies under long-term complex working conditions. This system possesses advantages such as deep multi-dimensional feature fusion, broad spatial exploration, and low equipment movement loss, significantly improving the comprehensiveness and accuracy of capturing long-cycle dynamic reaction processes. It effectively solves the problems of difficult spatial heterogeneity characterization and high human-computer interaction costs in existing monitoring, and promotes the evolution of in-situ chemical characterization towards an "unmanned" intelligent experimental paradigm.

[0006] The multi-region in-situ automatic monitoring system for working conditions based on multimodal fusion according to an embodiment of the present invention includes the following modules: The multi-region monitoring and reaction state control module is used to acquire the spatial coordinates and time series of active sites at the reaction interface, detect and acquire multi-channel in-situ characterization signals, read key state parameters corresponding to electric field, thermal field, force field or chemical field, and perform multi-modal information comprehensive analysis to output operating condition parameter dataset. The data acquisition and preprocessing module is used to read multi-channel in-situ characterization signals, determine the enhancement or reduction of spectral peaks and generate labels by comparing trends in multiple regions, perform background subtraction and correction on image signals, perform spatiotemporal index correspondence with the working condition parameter dataset, and construct and output a six-dimensional standardized input dataset. The feature extraction module is used to input the six-dimensional standardized input dataset into the Uniformer model, mine heterogeneous features in the spatial dimension, capture dynamic evolution trends in the temporal dimension, use working condition parameters as conditional embedding vectors to perform cross-modal interactive fusion with spatiotemporal features, and output interface state feature encoding through alternating stacking and fusion of spatial blocks and temporal blocks. The strategy optimization module is used to input the interface state feature encoding into the improved QPSO algorithm, construct the fitness function, introduce an asymmetric bias mechanism based on the probability distribution of the device's motion history trajectory, split the originally symmetric exponential distribution function into an asymmetric probability wave with directional bias, optimize in the multi-dimensional parameter space, and fuse and output the optimal observation position coordinates, acquisition frequency and chemical excitation parameter set. The linkage control module is used to issue multi-threaded linkage instructions based on the optimal observation position coordinates, acquisition frequency and chemical excitation parameter set, automatically control the high-precision three-dimensional translation stage to perform spatial movement, control the triggering of execution signals, and control the working station to dynamically modify the electric, thermal, mechanical or chemical field excitation conditions, and output multi-threaded linkage instructions and corresponding execution actions. The closed-loop control module is used to cyclically execute the steps from the multi-area monitoring and reaction control module to the linkage control module under unattended conditions. It dynamically iterates the observation strategy based on real-time data, executes closed-loop control, and automatically generates and outputs a multi-dimensional spatiotemporal evolution dataset.

[0007] Optionally, the multi-region monitoring and reaction state control module specifically includes: The high-precision three-dimensional translation stage is moved on the reaction interface by issuing control commands from the central control computer. The detector is positioned to multiple different spatial coordinate positions on the reaction interface in sequence. Each different spatial coordinate position is used as an active site on the reaction interface. The spatial coordinates of each active site on the reaction interface are read and recorded in real time. At the same time, the system clock is read to obtain the corresponding time series information. The control multi-channel in-situ characterization signal acquisition sensor detects the active sites of the reaction interface at the current location, acquires spectral data, and combines them to generate a multi-channel in-situ characterization signal. Extract the timestamps corresponding to the key state parameters, subtract the key state parameters of different reaction interface active sites at the same timestamp pairwise to calculate the difference, and obtain the reaction space difference value between different reaction interface active sites. Sum the reaction space difference values ​​of all reaction interface active sites and divide by the total number of reaction interface active sites to calculate the average value, and obtain the average space difference value. In the multi-channel in-situ characterization signal, the wavelength position of the spectral feature peak is found. The difference between the wavelength position of the spectral feature peak at the current time stamp and the wavelength position of the spectral feature peak at the previous time stamp is calculated to obtain the wavelength displacement of the spectral feature peak. The pixel coordinate position of the edge contour of the active site of the reaction interface is extracted from the image signal. The difference between the edge pixel coordinate of the current time stamp and the edge pixel coordinate of the previous time stamp is calculated to obtain the pixel displacement of the image edge contour. Extract the key state parameters of the current timestamp and the previous timestamp, subtract the previous key state parameter from the current key state parameter to obtain the voltage change or temperature change, divide the wavelength shift of the spectral feature peak by the voltage change or temperature change to calculate the first ratio, divide the pixel shift of the image edge contour by the voltage change or temperature change to calculate the second ratio, and add the first ratio and the second ratio to obtain the comprehensive analysis value of multimodal information. The calculated multimodal information is combined and packaged with the multi-regional spatial coordinates of the corresponding active sites on the reaction interface, the corresponding time series information, and the corresponding key state parameters. Spatiotemporal indexing is performed in the spatiotemporal grid composed of multi-regional spatial coordinates and time series information to output the working condition parameter dataset.

[0008] Optionally, the key state parameters specifically include: When performing electric field coupling, constant current and constant potential values ​​are sent down, and the electric field is output to the reaction interface. At the same time, the digital signal returned by the communication interface is read and analyzed into current, potential and capacity values ​​as key state parameters. When performing thermal field coupling, the preset range of variable temperature values ​​are sent to the working station, the thermal field is output to the reaction interface, and the digital signal returned by the communication interface is read and parsed into temperature values ​​as key status parameters. When performing force field coupling, the external pressure stacking value is sent to the working condition workstation, the force field is output to the reaction interface, and the digital signal returned by the communication interface is read and analyzed into pressure value as a key state parameter. When performing magnetic field or light field coupling, the magnetic field value or light field value is sent to the working condition workstation, the magnetic field or light field is output to the reaction interface, and the digital signal returned by the communication interface is read and analyzed into magnetic field strength value or light intensity value as a key state parameter. When performing chemical field coupling, the system sends different reaction electrode material values, electrolyte composition values, or conductivity values ​​to the working condition workstation. At the same time, it reads the digital signals returned by the communication interface and parses them into electrode material values, electrolyte composition ratio values, or conductivity values ​​as key state parameters. Optionally, the data acquisition and preprocessing module specifically includes: Read the multi-channel in-situ characterization signal. When the multi-channel in-situ characterization signal is a spectral signal, divide the absorption peak intensity of the spectral signal at the current time stamp by the absorption peak intensity at the previous time stamp to calculate the relative intensity ratio of the spectral peaks. When the relative intensity ratio of the spectral peaks is greater than the first set threshold, it is determined that the spectral peak is enhanced. When the relative intensity ratio of the spectral peaks is less than the second set threshold, it is determined that the spectral peak is weakened. Record the time stamp and wavelength position of the spectral peak enhancement or weakening as spectral dynamic evolution event data. The spectral dynamic evolution event data of all active sites at the same time point are compared one by one. When all active sites at the reaction interface increase the spectral peak or decrease the spectral peak at the same time point, it is determined that the trend change of the spectral peak in the multi-region is the same. When there is one or more active sites at the reaction interface that record the spectral peak increase and the other active sites at the reaction interface do not record the spectral increase, it is determined that the trend change of the spectral peak in the multi-region is different, and multi-region spectral trend comparison labels are generated. When the multi-channel in-situ characterization signal is an image signal, the average pixel gray value is calculated by cropping the blank area at the edge that does not contain the reaction interface. The average pixel gray value is subtracted from the gray value of each pixel in the image signal to complete the image background subtraction. The radial distortion coefficient and tangential distortion coefficient in the camera calibration file are read to calculate the offset of each pixel after distortion. The offset pixels are moved back to their original coordinate positions according to the opposite offset to complete the deformation correction. The preprocessed multi-channel in-situ characterization signal is then output. Read the working condition parameter dataset, extract the multimodal information comprehensive analysis values, multi-regional spatial coordinates of active sites on the reaction interface, time series information, key state parameters, and multi-regional spectral trend comparison labels from the working condition parameter dataset, and search for time series information that is exactly the same as the time series information of the multi-channel in-situ characterization signal and the corresponding multi-regional spatial coordinates as spatiotemporal index in the spatiotemporal grid composed of multi-regional spatial coordinates and time series information. The intensity of the characterization signal in the preprocessed multi-channel in-situ characterization signal is placed at the position corresponding to the spatiotemporal index. The multi-region spatial coordinates, time series information, key state parameters, multimodal information comprehensive analytical values, and multi-region spectral trend comparison labels with the characterization signal intensity are combined and arranged in a six-dimensional tensor combination according to the time dimension, spatial dimension, signal dimension, parameter dimension, analytical dimension, and label dimension to construct a six-dimensional standardized input dataset containing time, location, characterization signal intensity, operating condition parameters, comprehensive analytical values, and multi-region spectral trend comparison labels. The output is a six-dimensional standardized input dataset.

[0009] Optionally, the feature extraction module specifically includes: The signal intensity in the six-dimensional standardized input dataset is divided into non-overlapping image blocks according to the spatial coordinates of multiple regions. The values ​​of all pixels in each image block are flattened into a one-dimensional vector. The horizontal and vertical index values ​​of the current image block are extracted. The total dimension of the model embedding vector is set to d. The total dimension d is divided by 2 to obtain the maximum value of the dimension index k. Starting from 0 and incrementing by dimension index k, calculate 100002k / d as the division factor for the current dimension. Divide the horizontal and vertical index values ​​by the division factor respectively. Take the sine of the quotient of the horizontal coordinate to generate the first half-dimensional component, and take the cosine of the quotient of the vertical coordinate to generate the second half-dimensional component. Concatenate the first and second components into a preset position encoding vector, and add it to the one-dimensional vector one by one to generate a basic image patch sequence containing spatial position information. The basic image patch sequence is input into the spatial block of the Uniformer model. The dot product between any two image patch vectors in the sequence is calculated and divided by the square root of the vector dimension to obtain the attention score. The attention score is input into the Softmax function to transform the probability distribution weights, and the basic image patch sequence is weighted and summed to aggregate the spatial association information between different multi-region spatial coordinates and output the spatial feature map. The spatial feature map is input into the Uniformer Transformer block of the Uniformer model. A one-dimensional convolution operation is performed on all pixels in the spatial feature map along the channel dimension. The features of local neighboring pixels are extracted by sliding, and the intermediate spatial features are output. The intermediate spatial features corresponding to different time series information under the same multi-region spatial coordinates are arranged in chronological order to form a time feature sequence, which is input into the time block of the Uniformer model. The dot product of the feature vectors of any two time nodes is calculated within the time feature sequence and divided by the square root of the vector dimension. After transformation by the Softmax function, the probability distribution weight of the time dimension is obtained. The time feature sequence is weighted and summed to capture the dynamic evolution trend of the signal and output the time fusion feature sequence. The operating condition parameters in the six-dimensional standardized input dataset are extracted into numerical sequences, which are then transformed into vector dimensions with the same as the time fusion feature sequence through a linear mapping layer. At each time node, the transformed operating condition parameter vector is added element-wise to the corresponding vector in the time fusion feature sequence to generate a cross-modal interactive fusion feature sequence. The cross-modal interaction fusion feature sequence is input into the alternating stacked structure composed of spatial and temporal blocks in the next layer. The above operations of spatial aggregation, local convolution, temporal weighting and cross-modal addition are repeatedly performed to extract multi-layer deep features. The feature sequence output by the last layer is flattened in dimensions and mapped through a fully connected layer to output the interface state feature encoding.

[0010] Optionally, the strategy optimization module specifically includes: The interface state feature encoding is input into the improved QPSO algorithm. In the multi-dimensional parameter space composed of the observation position coordinates, acquisition frequency and chemical excitation parameter set, N initial particles are randomly generated. Each particle corresponds to a five-dimensional parameter vector containing the horizontal coordinate, vertical coordinate, height coordinate, acquisition frequency value and chemical excitation value. For each particle's x-coordinate, y-coordinate, and height coordinate, iterate through all the spatial coordinates that have been executed in the history record, calculate the three-dimensional Euclidean distance between the current particle's coordinates and each historical coordinate, add up all the three-dimensional Euclidean distances and take the average to obtain the average spatial distance, and input the five-dimensional parameter vector of the current particle into the fully connected mapping layer where the hidden layer neuron weights in the interface state feature encoding output by the Uniformer model are used as the connection weights of the corresponding input nodes. The linear transformation of the five-dimensional parameter vector is calculated and summed, then input into the ReLU activation function to calculate the feature activation response value. The average spatial distance is then multiplied by the feature activation response value to obtain the data information gain value. Extract the x-coordinate, y-coordinate, and height coordinate from the current five-dimensional parameter vector of each particle, and extract the x-coordinate, y-coordinate, and height coordinate from the optimal observation position coordinates output in the previous stage. Subtract the x-coordinate, y-coordinate, and height coordinates of the two and take the absolute value. Add the three absolute values ​​to obtain the device loss value. Subtract the device loss value from the data information gain value to obtain the fitness function value of each particle. Read the x-coordinate, y-coordinate, and height coordinates of the optimal observation position coordinates received by the high-precision three-dimensional translation stage in the previous stage, as well as the initial x-coordinate, initial y-coordinate, and initial height coordinates of the high-precision three-dimensional translation stage before it moved. Subtract the initial x-coordinate, initial y-coordinate, and initial height coordinates from the received x-coordinate, y-coordinate, and height coordinates to obtain the X-axis direction component, Y-axis direction component, and Z-axis direction component, and arrange them in order to generate the movement direction vector. An asymmetric bias mechanism based on the probability distribution of the device's motion history trajectory is introduced. A bias coefficient is preset, and the current value of the i-th parameter in the five-dimensional parameter vector of the current particle and the value of the global optimal position with the largest fitness function value among all current particles in the i-th dimension are extracted. The position deviation value is obtained by subtracting the value of the global optimal position in the i-th dimension from the current value. The previous stage movement direction vector component corresponding to the i-th parameter is extracted, and the movement direction vector component is multiplied by the preset bias coefficient to obtain the direction bias amount. Subtract the value of the global optimal position in the i-th dimension from the current value of the i-th parameter and subtract the direction bias to obtain the inner term of the exponential absolute value. Take the absolute value of the inner term of the exponential absolute value, multiply the result by -2 and divide it by the preset contraction and expansion coefficient that controls the convergence speed in the QPSO algorithm to obtain the exponential term. Calculate the result of the power operation with the natural constant e as the base and the exponential term as the exponent. Divide the result of the power operation by 2 times the preset contraction and expansion coefficient to obtain the asymmetric probability density function value of the i-th parameter. The updated five-dimensional parameter vectors for all particles are calculated using the five-dimensional parameter vectors and asymmetric probability density function values. The optimal observation position coordinates, acquisition frequency, and chemical excitation parameter set for the next stage are then output.

[0011] Optionally, the step of calculating the updated five-dimensional parameter vector for all particles using the five-dimensional parameter vector and asymmetric probability density function values, and outputting the optimal observation position coordinates, acquisition frequency, and chemical excitation parameter set for the next stage, specifically includes: For the acquisition frequency dimension and chemical excitation value dimension in the five-dimensional parameter vector, the corresponding movement direction vector component is directly set to 0, and multiplied by a preset bias coefficient to obtain a directional bias of 0. Then, the value of the global optimal position in the corresponding dimension is subtracted, and the value is subtracted to 0 to obtain the inner term of the exponential absolute value. The preset contraction and expansion coefficient calculation steps are repeated to maintain the original symmetrical shape of the probability density function of the acquisition frequency dimension and chemical excitation value dimension, which does not change with the movement direction of the device, to obtain an asymmetric probability wave. Perform asymmetric Monte Carlo random sampling. Starting from the current five-dimensional parameter vector of each particle, for each dimension, set a preset number of discrete calculation points between the value of the current dimension and the value of the global optimal position in the current dimension. Substitute each discrete calculation point into the step of calculating the asymmetric probability density function value to obtain the asymmetric probability density function value corresponding to each discrete calculation point. Add the asymmetric probability density function values ​​corresponding to all discrete calculation points in order to obtain the cumulative probability sum. Divide the asymmetric probability density function value corresponding to each discrete calculation point by the sum of cumulative probabilities to obtain the probability ratio of the interval occupied by each discrete calculation point. Accumulate the interval probability ratios in the order of the discrete calculation points to obtain the cumulative probability distribution value corresponding to each discrete calculation point. Randomly generate a decimal between 0 and 1 as the target sampling probability. Compare the target sampling probability with the cumulative probability distribution value corresponding to each discrete calculation point in turn. Find the first target discrete calculation point whose value is just greater than or equal to the target sampling probability. Use the value of the target discrete calculation point as the sampling base offset. Add the value of the global optimal position in the current dimension to the sampling base offset to obtain the updated value of the current dimension. Calculate the updated values ​​of the other four dimensions in the same way, and combine the updated values ​​of the five dimensions to generate the new five-dimensional parameter vector for each particle after the update. Substitute the updated five-dimensional parameter vectors of all particles back into the algorithm and calculate the new fitness function value. Compare the new fitness function value of each particle with the corresponding particle's historical maximum fitness function value. If the new fitness function value is larger, replace the particle's historical optimal position with the new five-dimensional parameter vector. Extract the particle with the largest new fitness function value among all particles. If the maximum value is greater than the historical global maximum fitness function value, replace the global optimal position with the particle's new five-dimensional parameter vector. Determine whether the current number of iterations has reached the preset iteration threshold. If not, return to the step of constructing the asymmetric probability wave and continue execution. If it has, stop the loop, extract the current global optimal position, and output the combination of the included x-coordinate, y-coordinate, and height coordinates as the optimal observation position coordinates for the next stage. Output the included acquisition frequency values ​​as the optimal acquisition frequency for the next stage, and output the included chemical excitation values ​​as the optimal chemical excitation parameter set for the next stage.

[0012] Optionally, the linkage control module specifically includes: The central control computer extracts the x-coordinate, y-coordinate, and height coordinates from the optimal observation position coordinates for the next stage, converts the x-coordinate, y-coordinate, and height coordinates into target pulse signal numbers, and sends the three sets of target pulse signal numbers to the stepper motor driver inside the high-precision three-dimensional translation stage to control the three-axis motor of the high-precision three-dimensional translation stage to rotate and move to the optimal observation position coordinates for the next stage. The optimal observation location coordinates for the next stage are compared with the spatial coordinate set of the monitoring area currently being monitored. If the optimal observation location coordinates for the next stage do not belong to the spatial coordinate set of the monitoring area currently being monitored, the optimal observation location coordinates for the next stage are added to the spatial coordinate set of the monitoring area to complete the automatic addition of the monitoring area. If there are historical spatial coordinates in the spatial coordinate set of the currently monitored area that have not appeared in the optimal observation position coordinates for a set number of consecutive times, the historical spatial coordinates that have not appeared consecutively will be removed from the spatial coordinate set of the monitored area to automatically reduce the monitored area. The central control computer extracts the optimal acquisition frequency for the next stage, converts the acquisition frequency value into a timing clock cycle parameter, sends the timing clock cycle parameter to the trigger controller, and automatically executes signal triggering and acquires multi-channel in-situ characterization signals according to the timing clock cycle parameter. The central control computer extracts the optimal chemical excitation parameter set for the next stage, reads the specific target values ​​of the electric field, thermal field, force field or chemical field corresponding to the chemical excitation parameter set, sends the specific target values ​​to the working condition workstation, controls the working condition workstation to adjust the output power of the electrochemical workstation or temperature controller, and modifies the voltage, current, and capacity parameters under electrochemical working conditions or the temperature and pressure parameters under thermal / mechanical working conditions to the specific target values. The central control computer packages the target pulse signal number sent to the motor driver, the timing clock cycle parameter sent to the trigger controller, the specific target value sent to the working condition workstation, and the update results of automatically adding or deleting monitoring areas, and generates and outputs multi-threaded linkage instructions and corresponding execution actions.

[0013] Optionally, the closed-loop control module specifically includes: The central control computer sets a preset total monitoring duration, reads the current system clock time, and determines whether the current clock time has reached the preset total monitoring duration. If it has not reached the preset total monitoring duration, the data acquisition and preprocessing module is automatically triggered to re-execute data acquisition. The multi-region monitoring and reaction state control module, data acquisition and preprocessing module, feature extraction module, strategy optimization module and linkage control module are treated as a complete working cycle. In each working cycle, the latest set of the next stage optimal observation position coordinates, acquisition frequency and chemical excitation parameter set output by the strategy optimization module is used to replace the historical parameters of the previous working cycle. The observation strategy is dynamically iterated based on the real-time acquired multi-channel in-situ characterization signals and key state parameters. Extract the six-dimensional standardized input dataset output by the preprocessing module in each work cycle, and concatenate the six-dimensional standardized input dataset of the current work cycle to the end of the dataset generated in the previous work cycle according to the time dimension. As the work cycles are executed cyclically, the amount of data is continuously expanded. When the current clock time reaches the preset total monitoring duration, the data acquisition and preprocessing module stops, and the datasets generated by splicing all work cycles are reorganized according to the dimensions of multi-regional spatial coordinates, time series information, characterizing signal strength and operating parameters, and a multi-dimensional spatiotemporal evolution dataset is automatically generated and output.

[0014] The beneficial effects of this invention are: This invention constructs a six-dimensional standardized input dataset by performing spectral dynamic evolution determination and image deformation correction of multimodal signals, combined with spatiotemporal index correspondence of operating parameters. The input dataset is then fed into a Uniformer model, which uses alternating stacking of spatial and temporal blocks to mine spatial heterogeneity and capture dynamic evolution trends. The resulting interface state feature encoding is then output through cross-modal fusion. This encoding is input into an improved QPSO algorithm, which innovatively introduces an asymmetric bias mechanism based on the historical trajectory of equipment motion. This splits the symmetric exponential distribution into asymmetric probability waves with directional bias. Under the guidance of a fitness function that balances equipment loss and data information gain, the optimal observation coordinates, frequency, and excitation parameters are output. Based on the optimal parameters, multi-threaded linkage commands are issued to execute automatic shifting, signal triggering, operating condition modification, and region addition / deletion. A spatiotemporal evolution dataset is generated through a closed-loop module using a dynamically iterative observation strategy under unattended operation. Traditional in-situ characterization, which only monitors a single location, fails to comprehensively reflect the significant spatial non-uniformity and potential heterogeneity of processes such as catalyst surface reactions, electrode corrosion, or battery charging / discharging, given the dynamic non-equilibrium characteristics of chemical systems under real-world operating conditions. Artificially capturing spatial heterogeneity requires long-term, high-frequency instrument adjustments, with reactions lasting hours or even days. This prolonged monitoring and repetitive operation is extremely demanding on personnel, and results in poor consistency in time-series control. This invention, through the aforementioned mechanism, deeply integrates AI-driven autonomous decision-making and real-time control of experimental parameters into in-situ / operational condition research strategies. Using electrochemical systems as a typical scenario, it ultimately achieves cross-regional spatiotemporal correlation analysis and system adaptive operation. This effectively solves the problems of difficult spatial heterogeneity characterization, high human-computer interaction costs, and poor consistency in long-term control found in existing methods, breaking through existing research bottlenecks and promoting the evolution of in-situ chemical characterization towards an "unmanned" intelligent experimental paradigm. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a structural diagram of the multi-region in-situ automatic monitoring system for working conditions based on multimodal fusion proposed in this invention. Figure 2 This is a flowchart of the spatiotemporal decoupling attention mechanism and cross-modal interaction fusion feature extraction based on the Uniformer model proposed in this invention; Figure 3 The flowchart shows the asymmetric probability wave construction and multi-parameter space adaptive optimization based on the improved QPSO algorithm proposed in this invention. Figure 4 This invention presents a data graph based on actual instrument measurements. Figure 5 This is a graph comparing the effects of the present invention with manual monitoring. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0017] refer to Figures 1-3 The multi-region in-situ automatic monitoring system for working conditions based on multimodal fusion includes the following modules: The multi-region monitoring and reaction state control module is used to acquire the spatial coordinates and time series of active sites at the reaction interface, detect and acquire multi-channel in-situ characterization signals, read key state parameters corresponding to electric field, thermal field, force field or chemical field, and perform multi-modal information comprehensive analysis to output operating condition parameter dataset. The data acquisition and preprocessing module is used to read multi-channel in-situ characterization signals, determine the enhancement or weakening of spectral peaks, and generate labels by comparing trends in multiple regions. It performs background subtraction and deformation correction on image signals, performs spatiotemporal index correspondence with the working condition parameter dataset, and constructs and outputs a six-dimensional standardized input dataset. The feature extraction module is used to input the six-dimensional standardized input dataset into the Uniformer model, mine heterogeneous features in the spatial dimension, capture dynamic evolution trends in the temporal dimension, use working condition parameters as conditional embedding vectors to perform cross-modal interactive fusion with spatiotemporal features, and output interface state feature encoding through alternating stacking and fusion of spatial blocks and temporal blocks. The strategy optimization module is used to input the interface state feature encoding into the improved QPSO algorithm, construct the fitness function, introduce an asymmetric bias mechanism based on the probability distribution of the device's motion history trajectory, split the originally symmetric exponential distribution function into an asymmetric probability wave with directional bias, optimize in the multi-dimensional parameter space, and fuse and output the optimal observation position coordinates, acquisition frequency and chemical excitation parameter set. The linkage control module is used to issue multi-threaded linkage instructions based on the optimal observation position coordinates, acquisition frequency and chemical excitation parameter set, automatically control the high-precision three-dimensional translation stage to perform spatial movement, control the triggering of execution signals, and control the working station to dynamically modify the electric, thermal, mechanical or chemical field excitation conditions, and output multi-threaded linkage instructions and corresponding execution actions. The closed-loop control module is used to cyclically execute the steps from the multi-area monitoring and reaction control module to the linkage control module under unattended conditions. It dynamically iterates the observation strategy based on real-time data, executes closed-loop control, and automatically generates and outputs a multi-dimensional spatiotemporal evolution dataset.

[0018] This invention significantly improves the spatiotemporal correlation and adaptability of observation strategies in multi-region in-situ monitoring of operating conditions. By constructing a six-dimensional standardized input dataset, it achieves unified modeling and accurate spatiotemporal indexing of multi-channel spectral and image signals, effectively addressing the challenge of characterizing spatial heterogeneity in complex chemical systems. During dynamic reactions, the Uniformer model is used to deeply explore spatiotemporal heterogeneity features, and operating condition parameters are used as conditional vectors for cross-modal interactive fusion, greatly enhancing the ability to express the dynamic evolution trend of interface states. Combined with an improved QPSO algorithm incorporating an asymmetric bias mechanism, it breaks through the limitations of traditional symmetric optimization, accurately outputting the optimal observation position, acquisition frequency, and excitation parameters in a multi-dimensional parameter space. After issuing multi-threaded linkage commands, it can automatically drive the three-dimensional translation stage and the operating condition workstation to execute collaboratively, achieving seamless integration of hardware actions and optimization strategies. By continuously performing dynamic closed-loop iteration of "perception-decision-control" under unattended conditions, it not only completely eliminates the problem of low efficiency due to long-term manual operation, but also demonstrates a strong level of automation and robustness, significantly improving the intelligence and precision of acquiring multi-dimensional spatiotemporal evolution data of complex interface processes.

[0019] In this embodiment, the multi-region monitoring and reaction state control module specifically includes: The high-precision three-dimensional translation stage is moved on the reaction interface by issuing control commands from the central control computer. The probe beam is positioned to multiple different spatial coordinate positions on the reaction interface in sequence. Each different spatial coordinate position is used as an active site on the reaction interface. The multi-region spatial coordinates of each active site on the reaction interface are read and recorded in real time. At the same time, the system clock is read to obtain the corresponding time series information. The control multi-channel in-situ characterization signal acquisition sensor detects the active sites of the reaction interface at the current location, acquires spectral wavelength and light intensity data or image pixel grayscale data, and combines them to generate multi-channel in-situ characterization signals. Extract the timestamps corresponding to the key state parameters, subtract the key state parameters of different reaction interface active sites at the same timestamp pairwise to calculate the difference, and obtain the reaction space difference value between different reaction interface active sites. Sum the reaction space difference values ​​of all reaction interface active sites and divide by the total number of reaction interface active sites to calculate the average value, thereby reducing the randomness of a single site. In the multi-channel in-situ characterization signal, the wavelength position of the spectral feature peak is found. The difference between the wavelength position of the spectral feature peak at the current time stamp and the wavelength position of the spectral feature peak at the previous time stamp is calculated to obtain the wavelength displacement of the spectral feature peak. The pixel coordinate position of the edge contour of the active site of the reaction interface is extracted from the image signal. The difference between the edge pixel coordinate of the current time stamp and the edge pixel coordinate of the previous time stamp is calculated to obtain the pixel displacement of the image edge contour. Extract the key state parameters of the current timestamp and the previous timestamp, subtract the previous key state parameter from the current key state parameter to obtain the voltage change or temperature change, divide the wavelength shift of the spectral feature peak by the voltage change or temperature change to calculate the first ratio, divide the pixel shift of the image edge contour by the voltage change or temperature change to calculate the second ratio, and add the first ratio and the second ratio to obtain the comprehensive analysis value of multimodal information. The calculated multimodal information, the multi-regional spatial coordinates of the corresponding active sites on the reaction interface, the corresponding time series information, and the corresponding key state parameters are packaged together. In the time-space grid composed of multi-regional spatial coordinates and time series information, time series information that is completely identical to the time series information in the packaged information and the multi-regional spatial coordinates of the corresponding active sites on the reaction interface are searched for spatiotemporal index correspondence. The output is a working condition parameter dataset containing multimodal information, multi-regional spatial coordinates of active sites on the reaction interface, time series information, and key state parameters.

[0020] In this embodiment, the key state parameters specifically include: When performing electric field coupling, constant current and constant potential values ​​are sent down, and the electric field is output to the reaction interface. At the same time, the digital signal returned by the communication interface is read and analyzed into current, potential and capacity values ​​as key state parameters. When performing thermal field coupling, the system sends variable temperature values ​​within a preset range to the working station, outputs the thermal field to the reaction interface, and reads the digital signal returned by the communication interface to parse it into temperature values ​​as key status parameters. The preset range is -20℃ to 60℃. When force field coupling is executed, the external pressure stacking value is sent to the working condition workstation, the force field is output to the reaction interface, and the digital signal returned by the communication interface is read and analyzed into pressure value as a key state parameter. When field coupling is executed, the magnetic field value or light field value is sent to the working condition workstation, the magnetic field or light field is output to the reaction interface, and the digital signal returned by the communication interface is read and analyzed into magnetic field strength value or light intensity value as a key state parameter. When performing chemical field coupling, the system sends different reaction electrode material values, electrolyte composition values, or conductivity values ​​to the working condition workstation. At the same time, it reads the digital signals returned by the communication interface and parses them into electrode material values, electrolyte composition ratio values, or conductivity values ​​as key state parameters. In this embodiment, the data acquisition and preprocessing module specifically includes: Read the multi-channel in-situ characterization signal. When the multi-channel in-situ characterization signal is a spectral signal, divide the absorption peak intensity of the spectral signal at the current time stamp by the absorption peak intensity at the previous time stamp to calculate the relative intensity ratio of the spectral peaks. When the relative intensity ratio of the spectral peaks is greater than a first set threshold, it is determined that the spectral peak has appeared. When the relative intensity ratio of the spectral peaks is less than a second set threshold, it is determined that the spectral peak has disappeared. Record the time stamp and wavelength position of the appearance or disappearance of the spectral peak as spectral dynamic evolution event data. The first set threshold is 1.2 and the second set threshold is 0.8. The spectral dynamic evolution event data of all active sites at the same time point are compared one by one. When all active sites at the same time point record the appearance or disappearance of spectral peaks, it is determined that the trend changes of spectral peaks in multiple regions are the same. When there is one or more active sites at the same time point that record the appearance of spectral peaks and the other active sites at the same time point do not record the appearance of spectral peaks, it is determined that the trend changes of spectral peaks in multiple regions are different, and multi-region spectral trend comparison labels are generated. When the multi-channel in-situ characterization signal is an image signal, the average pixel gray value is calculated by cropping the blank area at the edge that does not contain the reaction interface. The average pixel gray value is subtracted from the gray value of each pixel in the image signal to complete the image background subtraction. The radial distortion coefficient and tangential distortion coefficient in the camera calibration file are read to calculate the offset of each pixel after distortion. The offset pixels are moved back to their original coordinate positions according to the opposite offset to complete the deformation correction. The preprocessed multi-channel in-situ characterization signal is then output. Read the working condition parameter dataset, extract the multimodal information comprehensive analysis values, multi-regional spatial coordinates of active sites on the reaction interface, time series information, key state parameters, and multi-regional spectral trend comparison labels from the working condition parameter dataset, and search for time series information that is exactly the same as the time series information of the multi-channel in-situ characterization signal and the corresponding multi-regional spatial coordinates as spatiotemporal index in the spatiotemporal grid composed of multi-regional spatial coordinates and time series information. The intensity of the characterization signal in the preprocessed multi-channel in-situ characterization signal is placed at the position corresponding to the spatiotemporal index. The multi-region spatial coordinates, time series information, key state parameters, multi-modal information comprehensive analytical values, and multi-region spectral trend comparison labels with the characterization signal intensity are combined and arranged in a six-dimensional tensor combination according to the time dimension, spatial dimension, signal dimension, parameter dimension, analytical dimension, and label dimension to construct a six-dimensional standardized input dataset containing position, time, characterization signal intensity, operating condition parameters, comprehensive analytical values, and multi-region spectral trend comparison labels. The output is a six-dimensional standardized input dataset.

[0021] In this embodiment, the feature extraction module specifically includes: The signal intensity in the six-dimensional standardized input dataset is divided into non-overlapping image blocks according to the spatial coordinates of multiple regions. The values ​​of all pixels in each image block are flattened into a one-dimensional vector. The horizontal and vertical index values ​​of the current image block are extracted. The total dimension of the model embedding vector is set to d. The total dimension d is divided by 2 to obtain the maximum value of the dimension index k. Starting from 0 and incrementing by dimension index k, calculate 100002k / d as the division factor for the current dimension. Divide the horizontal and vertical index values ​​by the division factor respectively. Take the sine of the quotient of the horizontal coordinate to generate the first half-dimensional component, and take the cosine of the quotient of the vertical coordinate to generate the second half-dimensional component. Concatenate the first and second components into a preset position encoding vector, and add it to the one-dimensional vector one by one to generate a basic image patch sequence containing spatial position information. The basic image patch sequence is input into the spatial block of the Uniformer model. The dot product between any two image patch vectors in the sequence is calculated and divided by the square root of the vector dimension to obtain the attention score. The attention score is input into the Softmax function to transform the probability distribution weights, and the basic image patch sequence is weighted and summed to aggregate the spatial association information between different multi-region spatial coordinates and output the spatial feature map. The spatial feature map is input into the Uniformer Transformer block of the Uniformer model. A one-dimensional convolution operation is performed on all pixels in the spatial feature map along the channel dimension. The features of local neighboring pixels are extracted by sliding, and the intermediate spatial features are output. The intermediate spatial features corresponding to different time series information under the same multi-region spatial coordinates are arranged in chronological order to form a time feature sequence, which is input into the time block of the Uniformer model. The dot product of the feature vectors of any two time nodes is calculated within the time feature sequence and divided by the square root of the vector dimension. After transformation by the Softmax function, the probability distribution weight of the time dimension is obtained. The time feature sequence is weighted and summed to capture the dynamic evolution trend of the signal and output the time fusion feature sequence. The operating condition parameters in the six-dimensional standardized input dataset are extracted into numerical sequences, which are then transformed into vector dimensions with the same as the time fusion feature sequence through a linear mapping layer. At each time node, the transformed operating condition parameter vector is added element-wise to the corresponding vector in the time fusion feature sequence to generate a cross-modal interactive fusion feature sequence. The cross-modal interaction fusion feature sequence is input into the alternating stacked structure composed of spatial and temporal blocks in the next layer. The above operations of spatial aggregation, local convolution, temporal weighting and cross-modal addition are repeatedly performed to extract multi-layer deep features. The feature sequence output by the last layer is flattened in dimensions and mapped through a fully connected layer to output the interface state feature encoding.

[0022] This implementation introduces the Uniformer model as the core feature extraction architecture, which has significant advantages over traditional pure CNN, ViT, or pure RNN time series models. Traditional CNNs are limited by local receptive fields, making it difficult to establish global spatial dependencies between multiple regions; although ViT has a global perspective, its unrestricted self-attention mechanism introduces a large amount of irrelevant background noise and has extremely high computational complexity; while RNNs are prone to losing early key features when dealing with long-term evolution.

[0023] This invention effectively solves the aforementioned bottlenecks through the innovative spatiotemporal decoupling mechanism of Uniformer. In the spatial feature extraction stage, the self-attention mechanism of spatial blocks is used to accurately aggregate the heterogeneous correlations between coordinates of different regions. Subsequently, in the unified Transformer block, local one-dimensional convolution is used instead of global self-attention, effectively filtering instantaneous interference noise and fusing local texture details while significantly reducing computational cost. Finally, in the temporal dimension, the self-attention mechanism of temporal blocks is still used to capture the long-term dynamic evolution trend of the characterizing signal. Furthermore, by using dynamic operating parameters as conditional embedding vectors for cross-modal addition, the model can sensitively perceive changes in external stimuli. This alternating stacking and fusion strategy, while ensuring linear computational efficiency, achieves efficient and high-precision extraction of complex spatial heterogeneity and long-term dynamic evolution patterns of the interface, significantly enhancing the robustness of feature encoding.

[0024] In this embodiment, the strategy optimization module specifically includes: The interface state feature encoding is input into the improved QPSO algorithm. In the multi-dimensional parameter space composed of the observation position coordinates, acquisition frequency and chemical excitation parameter set, N initial particles are randomly generated. Each particle corresponds to a five-dimensional parameter vector containing the horizontal coordinate, vertical coordinate, height coordinate, acquisition frequency value and chemical excitation value. The specific value of N is 50. For each particle's x-coordinate, y-coordinate, and height coordinate, iterate through all the multi-region spatial coordinates that have been executed in the history record, calculate the three-dimensional Euclidean distance between the current particle's coordinates and each historical coordinate, add up all the three-dimensional Euclidean distances and take the average to obtain the average spatial distance, and input the current particle's five-dimensional parameter vector into the fully connected mapping layer where the hidden layer neuron weights in the interface state feature encoding output by the Uniformer model are used as the corresponding input node connection weights. The linear transformation of the five-dimensional parameter vector is calculated and summed, then input into the ReLU activation function. The positive value of the output is taken as the feature activation response value, which characterizes the dynamic intensity or information uncertainty of the reaction interface under the parameter node. The average spatial distance is multiplied by the feature activation response value to obtain the data information gain value that takes into account both the breadth of spatial exploration and the intensity of interface reaction. Extract the x-coordinate, y-coordinate, and height coordinate from the current five-dimensional parameter vector of each particle, and extract the x-coordinate, y-coordinate, and height coordinate from the optimal observation position coordinates output in the previous stage. Subtract the x-coordinate, y-coordinate, and height coordinates of the two and take the absolute value. Add the three absolute values ​​to obtain the device loss value. Subtract the device loss value from the data information gain value containing feature encoding weights to obtain the fitness function value of each particle. Read the x-coordinate, y-coordinate, and height coordinates of the optimal observation position coordinates received by the high-precision three-dimensional translation stage in the previous stage, as well as the initial x-coordinate, initial y-coordinate, and initial height coordinates of the high-precision three-dimensional translation stage before it moved. Subtract the initial x-coordinate, initial y-coordinate, and initial height coordinates from the received x-coordinate, y-coordinate, and height coordinates to obtain the X-axis direction component, Y-axis direction component, and Z-axis direction component, and arrange them in order to generate the movement direction vector. An asymmetric bias mechanism based on the probability distribution of the device's motion history trajectory is introduced. A bias coefficient is preset, and the current value of the i-th parameter in the five-dimensional parameter vector of the current particle and the value of the global optimal position with the largest fitness function value among all current particles in the i-th dimension are extracted. The position deviation value is obtained by subtracting the value of the global optimal position in the i-th dimension from the current value. The previous stage movement direction vector component corresponding to the i-th parameter is extracted, and the movement direction vector component is multiplied by the preset bias coefficient to obtain the direction bias amount. Subtract the value of the global optimal position in the i-th dimension from the current value of the i-th dimension parameter and subtract the direction bias to obtain the inner term of the exponential absolute value. Take the absolute value of the inner term of the exponential absolute value, multiply the result by -2 and divide it by the preset contraction and expansion coefficient that controls the convergence speed in the improved QPSO algorithm to obtain the exponential term. Calculate the result of the power operation with the natural constant e as the base and the exponential term as the exponent. Divide the result of the power operation by twice the preset contraction and expansion coefficient to obtain the asymmetric probability density function value of the i-th dimension parameter. The updated five-dimensional parameter vectors for all particles are calculated using the five-dimensional parameter vectors and asymmetric probability density function values. The optimal observation position coordinates, acquisition frequency, and chemical excitation parameter set for the next stage are then output.

[0025] In this embodiment, the updated five-dimensional parameter vectors for all particles are calculated using the five-dimensional parameter vector and the asymmetric probability density function value. The optimal observation position coordinates, acquisition frequency, and chemical excitation parameter set for the next stage are then output. Specifically, this includes: For the acquisition frequency dimension and chemical excitation value dimension in the five-dimensional parameter vector, the corresponding movement direction vector component is directly set to 0, and multiplied by a preset bias coefficient to obtain a directional bias of 0. Then, the value of the global optimal position in the corresponding dimension is subtracted, and the value is subtracted to 0 to obtain the inner term of the exponential absolute value. The preset contraction and expansion coefficient calculation steps are repeated to maintain the original symmetrical shape of the probability density function of the acquisition frequency dimension and chemical excitation value dimension, which does not change with the movement direction of the device, to obtain an asymmetric probability wave. Perform asymmetric Monte Carlo random sampling. Starting from the current five-dimensional parameter vector of each particle, for each dimension, set a preset number of discrete calculation points between the value of the current dimension and the value of the global optimal position in the current dimension. Substitute each discrete calculation point into the step of calculating the asymmetric probability density function value to obtain the asymmetric probability density function value corresponding to each discrete calculation point. Add the asymmetric probability density function values ​​corresponding to all discrete calculation points in order to obtain the cumulative probability sum. Divide the asymmetric probability density function value corresponding to each discrete calculation point by the sum of cumulative probabilities to obtain the probability ratio of the interval occupied by each discrete calculation point. Accumulate the interval probability ratios in the order of the discrete calculation points to obtain the cumulative probability distribution value corresponding to each discrete calculation point. Randomly generate a decimal between 0 and 1 as the target sampling probability. Compare the target sampling probability with the cumulative probability distribution value corresponding to each discrete calculation point in turn. Find the first target discrete calculation point whose value is just greater than or equal to the target sampling probability. Use the value of the target discrete calculation point as the sampling base offset. Add the value of the global optimal position in the current dimension to the sampling base offset to obtain the updated value of the current dimension. Calculate the updated values ​​of the other four dimensions in the same way, and combine the updated values ​​of the five dimensions to generate the new five-dimensional parameter vector for each particle after the update. Substitute the updated five-dimensional parameter vectors of all particles back into the algorithm and calculate the new fitness function value. Compare the new fitness function value of each particle with the corresponding particle's historical maximum fitness function value. If the new fitness function value is larger, replace the particle's historical optimal position with the new five-dimensional parameter vector. Extract the particle with the largest new fitness function value among all particles. If the maximum value is greater than the historical global maximum fitness function value, replace the global optimal position with the particle's new five-dimensional parameter vector. Determine whether the current number of iterations has reached the preset iteration threshold. If not, return to the step of constructing the asymmetric probability wave and continue execution. If it has, stop the loop, extract the current global optimal position, and output the combination of the included x-coordinate, y-coordinate, and height coordinates as the optimal observation position coordinates for the next stage. Output the included acquisition frequency values ​​as the optimal acquisition frequency for the next stage, and output the included chemical excitation values ​​as the optimal chemical excitation parameter set for the next stage.

[0026] This implementation achieves efficient adaptive optimization and mechanical error elimination for multi-dimensional observation strategies through an improved QPSO algorithm and an asymmetric bias mechanism based on historical trajectories. A fitness function is constructed by maximizing data information gain and minimizing equipment loss, precisely balancing high-value data acquisition with hardware protection. Addressing the inherent mechanical backlash problem of the 3D translation stage, the algorithm extracts historical movement direction vectors, splitting the symmetric exponential distribution of the spatial coordinate dimension into asymmetric probability waves with directional bias. Combined with asymmetric Monte Carlo sampling, more sampling points are allocated along the positive motion direction to actively offset the reverse backlash error. Simultaneously, a symmetric distribution is maintained for the frequency and excitation parameters, avoiding invalid interference from irrelevant dimensions. This strategy converges rapidly within 100 iterations, significantly improving the optimization accuracy of observation positions and chemical excitation parameters at complex interfaces, and achieving intelligent, accurate, and low-loss tracking of multi-region interfaces.

[0027] The improved QPSO algorithm of this invention is similar to the original QPSO algorithm in that both retain the core framework of quantum particle swarm optimization, namely, both randomly initialize the particle swarm in a multidimensional parameter space, update the individual state by tracking the historical best position of the current particle and the global best position of all particles, and both use an exponential probability density function with the natural constant e as the base to guide the position iteration of the particles, and finally find the parameter solution with the optimal global fitness function value through multiple iterations.

[0028] The difference lies in that this invention breaks the limitation of the original QPSO algorithm, which uses an absolutely symmetrical probability distribution for sampling in all dimensions, and introduces an asymmetric bias mechanism based on the device's historical motion trajectory. Instead of directly calculating the symmetric exponential function from the position deviation in the original algorithm, this invention adds a pre-processing step for extracting the movement direction vector. This is achieved by generating a spatial direction vector by reading the difference between the coordinates received by the translation stage in the previous stage and the initial coordinates. Next, when calculating the asymmetric probability density function, the deviation between the current particle position and the global optimal position is subtracted from the direction bias obtained by multiplying the direction vector component by the bias coefficient, splitting the originally symmetrical exponential distribution function into an asymmetric probability wave with a direction bias. Furthermore, for the acquisition frequency and chemical excitation dimensions in the five-dimensional parameter vector, the corresponding movement direction vector components are forcibly set to 0 to maintain their original symmetry, achieving differentiated processing of spatial and non-spatial dimensions. Finally, in the particle update stage, asymmetric Monte Carlo random sampling replaces traditional mean sampling.

[0029] Based on the aforementioned improvements, the beneficial effects of this invention are as follows: through direction vector offset and asymmetric probability wave design, the improved QPSO algorithm can address the inherent mechanical backlash problem of the 3D translation stage by allocating more probability sampling points in the positive motion direction. This allows it to actively cancel out the mechanical backlash that needs to be overcome during reverse motion, completely eliminating the systematic error caused by the mechanical structure and significantly improving the absolute accuracy of spatial positioning. Simultaneously, maintaining a symmetrical distribution of non-spatial parameters avoids invalid interference from irrelevant dimensions, ensuring the purity of the optimization of observation frequency and chemical excitation parameters. The asymmetric Monte Carlo sampling combined with a precise fitness function significantly improves the convergence accuracy and optimization efficiency of complex interface observation strategies within a finite number of iterations, enhancing the stability and robustness of the system under long-term continuous operation.

[0030] In this embodiment, the linkage control module specifically includes: The central control computer extracts the x-coordinate, y-coordinate, and height coordinates from the optimal observation position coordinates for the next stage, converts the x-coordinate, y-coordinate, and height coordinates into target pulse signal numbers, and sends the three sets of target pulse signal numbers to the stepper motor driver inside the high-precision three-dimensional translation stage to control the three-axis motor of the high-precision three-dimensional translation stage to rotate and move to the optimal observation position coordinates for the next stage. The optimal observation location coordinates for the next stage are compared with the spatial coordinate set of the monitoring area currently being monitored. If the optimal observation location coordinates for the next stage do not belong to the spatial coordinate set of the monitoring area currently being monitored, the optimal observation location coordinates for the next stage are added to the spatial coordinate set of the monitoring area to complete the automatic addition of the monitoring area. When there are historical spatial coordinates in the spatial coordinate set of the currently monitored area that have not appeared in the optimal observation position coordinates for a set number of consecutive times, the historical spatial coordinates that have not appeared for a set number of consecutive times will be removed from the spatial coordinate set of the monitored area to automatically reduce the monitored area. The set number of times is 10. The central control computer extracts the optimal acquisition frequency for the next stage, converts the acquisition frequency value into a timing clock cycle parameter, sends the timing clock cycle parameter to the trigger controller, and automatically executes signal triggering and acquires multi-channel in-situ characterization signals according to the timing clock cycle parameter. The central control computer extracts the optimal chemical excitation parameter set for the next stage, reads the specific target values ​​of the electric field, thermal field, force field or chemical field corresponding to the chemical excitation parameter set, sends the specific target values ​​to the working condition workstation, controls the working condition workstation to adjust the output power of the electrochemical workstation or temperature controller, and modifies the voltage, current, and capacity parameters under electrochemical working conditions or the temperature and pressure parameters under thermal / mechanical working conditions to the specific target values. The central control computer packages the target pulse signal number sent to the stepper motor driver, the timing clock cycle parameter sent to the trigger controller, the specific target value sent to the working condition workstation, and the update results of automatically adding or deleting monitoring areas, and generates and outputs multi-threaded linkage instructions and corresponding execution actions.

[0031] In this embodiment, the closed-loop control module specifically includes: The central control computer sets a preset total monitoring duration, reads the current system clock time, and determines whether the current clock time has reached the preset total monitoring duration. If it has not reached the preset total monitoring duration, the data acquisition and preprocessing module is automatically triggered to re-execute data acquisition. The multi-region monitoring and reaction state control module, data acquisition and preprocessing module, feature extraction module, strategy optimization module and linkage control module are treated as a complete working cycle. In each working cycle, the latest set of the next stage optimal observation position coordinates, acquisition frequency and chemical excitation parameter set output by the strategy optimization module is used to replace the historical parameters of the previous working cycle. The observation strategy is dynamically iterated based on the real-time acquired multi-channel in-situ characterization signals and key state parameters. Extract the six-dimensional standardized input dataset output by the preprocessing module in each work cycle, and concatenate the six-dimensional standardized input dataset of the current work cycle to the end of the dataset generated in the previous work cycle according to the time dimension. As the work cycles are executed cyclically, the amount of data is continuously expanded. When the current clock time reaches the preset total monitoring duration, the data acquisition and preprocessing module stops, and the datasets generated by splicing all work cycles are reorganized according to the dimensions of multi-regional spatial coordinates, time series information, characterizing signal strength and operating parameters, and a multi-dimensional spatiotemporal evolution dataset is automatically generated and output.

[0032] Example 1: To verify the feasibility of this invention in practice, it was applied to in-situ intelligent monitoring of high-energy-density lithium-sulfur batteries in the laboratory. Due to the spatial heterogeneity of the liquid-phase reaction components within lithium-sulfur batteries during the electrode / electrolyte interface transformation process, monitoring only a single location is insufficient to reflect the true evolution of the interface across multiple regions. Traditional in-situ characterization methods primarily rely on researchers manually operating a three-dimensional electric translation stage and spectrometer, frequently clicking computer software buttons to switch observation points. This manual method of capturing spatial heterogeneity requires operators to perform long-term, high-frequency manual adjustments and recordings, resulting in high manpower consumption and poor consistency in timing control. Furthermore, battery charge-discharge cycles often last for tens of hours. To clearly capture the details of the dynamic evolution of the polysulfide interface, researchers need to maintain continuous monitoring for extended periods and perform tedious, repetitive operations. This poses a significant challenge to their physical and mental well-being, easily leading to missed data collection at key reaction inflection points due to fatigue, severely hindering in-depth analysis of the reaction process and failure mechanisms.

[0033] In practical deployment, the method of this invention integrates a high-precision, fully adaptable three-dimensional electric translation stage, an in-situ UV-Vis spectrometer, and an electrochemical workstation, all connected to a central control computer for unified scheduling. During sample preparation, a drilling machine is used to drill two holes, each six millimeters in diameter, at the center and edge of the aluminum-plastic film of the lithium-sulfur soft-pack battery. These holes are then sealed with quartz plates to serve as optical windows, and the three-dimensional coordinates of the two holes are recorded. Operators only need to set a few parameters, such as the battery's operating voltage, current range, and wavelength range, and the system can automatically execute cyclic tasks unattended. In each scan, the translation stage precisely positions the center of each sub-region under visible light according to an optimized path, automatically acquiring UV-Vis spectra and simultaneously reading and recording precise time points, spatial coordinates, and corresponding electrochemical parameters such as voltage, current, and capacity. After completing a scan of all sites, the system automatically repeats the entire spatial scan sequence according to a set time interval, thereby obtaining the absorption spectra at a series of time points at each spatial location, constructing a three-dimensional dataset containing location, time, and reflectance. To verify the reliability of the automated system, the laboratory simultaneously conducted a comparative experiment with traditional manual multi-region translation stage monitoring. After starting the electrochemical program, researchers manually clicked the start button to test the first coordinate point. Immediately after the test, they manually controlled the translation stage to align the second hole with the incident window and test again. This process was repeated after collecting data from all regions and waiting several minutes until the battery charging and discharging were complete. During the testing and analysis phase, the electrochemical curves showed similarities between single-hole and multi-hole batteries, and the charge-discharge curves from the automated and manual tests of this invention were also highly similar, indicating that the number of holes and the AI-assisted automated testing did not significantly affect the battery's electrochemical process. Further UV-Vis spectroscopy analysis revealed significant differences in peak position and intensity of polysulfide spectra at different locations (center and edge) of the same battery, confirming the spatial heterogeneity of the reaction kinetics. The results from the same location in different batteries also differed, indicating that differences in battery assembly interfere with heterogeneity studies, necessitating in-situ monitoring at different locations within the same battery. Simultaneously, the automated and manual test results for multiple locations of the same pouch battery showed a strong agreement, fully demonstrating that the automated testing of this invention can completely replace manual experimentation. Figure 4 The data graphs and data based on actual instrument measurements proposed in this invention are Figure 5 The table below shows the comparison between the intelligent monitoring method of this invention and the traditional manual method in the multi-region in-situ operating condition tracking task of lithium-sulfur soft-pack batteries during a two-week continuous testing period: Table 1. Comprehensive Comparison of Multi-Dimensional Performance between Intelligent Monitoring System and Traditional Manual Methods

[0034] Based on the comparative data shown in Table 1, it can be seen that the multi-region in-situ automatic monitoring system based on multimodal fusion proposed in this invention exhibits significant performance advantages over traditional manual methods in the in-situ characterization of lithium-sulfur batteries, achieving comprehensive improvements in key indicators such as scanning efficiency, data throughput, positioning accuracy, and temporal consistency. Regarding the efficiency of single-round multi-region scanning, traditional manual methods require researchers to closely monitor the screen, manually click the software start button, and manually control the translation stage to move and align with the next optical window after completing one site test. This process is prone to interruptions, with a single round taking up to four and a half minutes. In contrast, this invention uses a central control computer to issue multi-threaded linkage commands, seamlessly connecting the translation stage and the spectrometer, drastically reducing the single-round time to 3.5 minutes, an efficiency improvement of over 70%. Regarding the amount of effective data acquired throughout the day, manual operation cannot operate continuously day and night, and manually switching sites is too time-consuming, resulting in less than 300 sets of effective data acquired within 24 hours. This invention can continuously cycle with extremely high consistency, acquiring more than 1,100 sets of data per day, providing an extremely rich and high-quality data foundation for subsequent use of deep learning to explore battery failure mechanisms.

[0035] In terms of spatial positioning accuracy and timing control, traditional methods rely on manual operation of a joystick. Due to mechanical inertia, this often leads to overshoot or undershoot, resulting in an average positioning error exceeding eleven micrometers. Furthermore, human fatigue causes significant fluctuations in the time interval between adjacent data acquisitions. This invention controls a stepper motor with precise pulse signal counts, suppressing the positioning error to the two-micrometer level. Simultaneously, the time interval variance is drastically reduced from 15 seconds squared to 0.3 seconds squared, ensuring absolute rigor in the time series. Regarding the key scientific research indicator of inflection point capture rate, since polysulfide phase transitions often occur within extremely short time windows, the lag in manual operation causes nearly one-third of inflection points to be missed. This system's high-frequency, no-miss scan improves the capture rate to over 96%, truly achieving in-situ dynamic quantitative tracking of the spatial non-uniformity of liquid-phase reaction components. Overall, this invention significantly extends continuous fault-free operation time from six hours to seventy-two hours, filling the technical shortcomings of traditional methods in long-cycle, multi-region dynamic interface control. It provides a reliable, universal, and intelligent methodological support for a deeper understanding of the failure mechanisms and performance evolution laws of electrochemical systems.

[0036] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multi-region in-situ working condition automatic monitoring system based on multi-modal fusion, characterized in that, Includes the following modules: The multi-region monitoring and reaction state control module is used to acquire the spatial coordinates and time series of active sites at the reaction interface, detect and acquire multi-channel in-situ characterization signals, read key state parameters corresponding to electric field, thermal field, force field or chemical field, and perform multi-modal information comprehensive analysis to output operating condition parameter dataset. The data acquisition and preprocessing module is used to read multi-channel in-situ characterization signals, determine the enhancement or reduction of spectral peaks and generate labels by comparing trends in multiple regions, perform background subtraction and correction on image signals, perform spatiotemporal index correspondence with the working condition parameter dataset, and construct and output a six-dimensional standardized input dataset. The feature extraction module is used to input the six-dimensional standardized input dataset into the Uniformer model, mine heterogeneous features in the spatial dimension, capture dynamic evolution trends in the temporal dimension, use working condition parameters as conditional embedding vectors to perform cross-modal interactive fusion with spatiotemporal features, and output interface state feature encoding through alternating stacking and fusion of spatial blocks and temporal blocks. The strategy optimization module is used to input the interface state feature encoding into the improved QPSO algorithm, construct the fitness function, introduce an asymmetric bias mechanism based on the probability distribution of the device's motion history trajectory, split the originally symmetric exponential distribution function into an asymmetric probability wave with directional bias, optimize in the multi-dimensional parameter space, and fuse and output the optimal observation position coordinates, acquisition frequency and chemical excitation parameter set. The linkage control module is used to issue multi-threaded linkage instructions based on the optimal observation position coordinates, acquisition frequency and chemical excitation parameter set, automatically control the high-precision three-dimensional translation stage to perform spatial movement, control the triggering of execution signals, and control the working station to dynamically modify the electric, thermal, mechanical or chemical field excitation conditions, and output multi-threaded linkage instructions and corresponding execution actions. The closed-loop control module is used to cyclically execute the steps from the multi-area monitoring and reaction control module to the linkage control module under unattended conditions. It dynamically iterates the observation strategy based on real-time data, executes closed-loop control, and automatically generates and outputs a multi-dimensional spatiotemporal evolution dataset.

2. The multi-region in-situ automatic monitoring system for working conditions based on multimodal fusion according to claim 1, characterized in that, The multi-region monitoring and reaction state control module specifically includes: The high-precision three-dimensional translation stage is moved on the reaction interface by issuing control commands from the central control computer. The detector is positioned to multiple different spatial coordinate positions on the reaction interface in sequence. Each different spatial coordinate position is used as an active site on the reaction interface. The spatial coordinates of each active site on the reaction interface are read and recorded in real time. At the same time, the system clock is read to obtain the corresponding time series information. The control multi-channel in-situ characterization signal acquisition sensor detects the active sites of the reaction interface at the current location, acquires spectral data, and combines them to generate a multi-channel in-situ characterization signal. Extract the timestamps corresponding to the key state parameters, subtract the key state parameters of different reaction interface active sites at the same timestamp pairwise to calculate the difference, and obtain the reaction space difference value between different reaction interface active sites. Sum the reaction space difference values ​​of all reaction interface active sites and divide by the total number of reaction interface active sites to calculate the average value, and obtain the average space difference value. In the multi-channel in-situ characterization signal, the wavelength position of the spectral feature peak is found. The difference between the wavelength position of the spectral feature peak at the current time stamp and the wavelength position of the spectral feature peak at the previous time stamp is calculated to obtain the wavelength displacement of the spectral feature peak. The pixel coordinate position of the edge contour of the active site of the reaction interface is extracted from the image signal. The difference between the edge pixel coordinate of the current time stamp and the edge pixel coordinate of the previous time stamp is calculated to obtain the pixel displacement of the image edge contour. Extract the key state parameters of the current timestamp and the previous timestamp, subtract the previous key state parameter from the current key state parameter to obtain the voltage change or temperature change, divide the wavelength shift of the spectral feature peak by the voltage change or temperature change to calculate the first ratio, divide the pixel shift of the image edge contour by the voltage change or temperature change to calculate the second ratio, and add the first ratio and the second ratio to obtain the comprehensive analysis value of multimodal information. The calculated multimodal information is combined and packaged with the multi-regional spatial coordinates of the corresponding active sites on the reaction interface, the corresponding time series information, and the corresponding key state parameters. Spatiotemporal indexing is performed in the spatiotemporal grid composed of multi-regional spatial coordinates and time series information to output the working condition parameter dataset.

3. The multi-modal fusion based multi-region in-situ working condition automatic monitoring system according to claim 2, characterized in that, The key state parameters specifically include: When performing electric field coupling, constant current and constant potential values ​​are sent down, and the electric field is output to the reaction interface. At the same time, the digital signal returned by the communication interface is read and analyzed into current, potential and capacity values ​​as key state parameters. When performing thermal field coupling, the preset range of variable temperature values ​​are sent to the working station, the thermal field is output to the reaction interface, and the digital signal returned by the communication interface is read and parsed into temperature values ​​as key status parameters. When performing force field coupling, the external pressure stacking value is sent to the working condition workstation, the force field is output to the reaction interface, and the digital signal returned by the communication interface is read and analyzed into pressure value as a key state parameter. When performing magnetic field or light field coupling, the magnetic field value or light field value is sent to the working condition workstation, the magnetic field or light field is output to the reaction interface, and the digital signal returned by the communication interface is read and analyzed into magnetic field strength value or light intensity value as a key state parameter. When performing chemical field coupling, the system sends different reaction electrode material values, electrolyte composition values, or conductivity values ​​to the working station. At the same time, it reads the digital signals returned by the communication interface and parses them into electrode material values, electrolyte composition ratio values, or conductivity values ​​as key state parameters.

4. The multi-modal fusion based multi-region in-situ working condition automatic monitoring system according to claim 1, characterized in that, The data acquisition and preprocessing module specifically includes: Read the multi-channel in-situ characterization signal. When the multi-channel in-situ characterization signal is a spectral signal, divide the absorption peak intensity of the spectral signal at the current time stamp by the absorption peak intensity at the previous time stamp to calculate the relative intensity ratio of the spectral peaks. When the relative intensity ratio of the spectral peaks is greater than the first set threshold, it is determined that the spectral peak is enhanced. When the relative intensity ratio of the spectral peaks is less than the second set threshold, it is determined that the spectral peak is weakened. Record the time stamp and wavelength position of the spectral peak enhancement or weakening as spectral dynamic evolution event data. The spectral dynamic evolution event data of all active sites at the same time point are compared one by one. When all active sites at the reaction interface increase the spectral peak or decrease the spectral peak at the same time point, it is determined that the trend change of the spectral peak in the multi-region is the same. When there is one or more active sites at the reaction interface that record the spectral peak increase and the other active sites at the reaction interface do not record the spectral increase, it is determined that the trend change of the spectral peak in the multi-region is different, and multi-region spectral trend comparison labels are generated. When the multi-channel in-situ characterization signal is an image signal, the average pixel gray value is calculated by cropping the blank area at the edge that does not contain the reaction interface. The average pixel gray value is subtracted from the gray value of each pixel in the image signal to complete the image background subtraction. The radial distortion coefficient and tangential distortion coefficient in the camera calibration file are read to calculate the offset of each pixel after distortion. The offset pixels are moved back to their original coordinate positions according to the opposite offset to complete the deformation correction. The preprocessed multi-channel in-situ characterization signal is then output. Read the working condition parameter dataset, extract the multimodal information comprehensive analysis values, multi-regional spatial coordinates of active sites on the reaction interface, time series information, key state parameters, and multi-regional spectral trend comparison labels from the working condition parameter dataset, and search for time series information that is exactly the same as the time series information of the multi-channel in-situ characterization signal and the corresponding multi-regional spatial coordinates as spatiotemporal index in the spatiotemporal grid composed of multi-regional spatial coordinates and time series information. The intensity of the characterization signal in the preprocessed multi-channel in-situ characterization signal is placed at the position corresponding to the spatiotemporal index. The multi-region spatial coordinates, time series information, key state parameters, multimodal information comprehensive analytical values, and multi-region spectral trend comparison labels with the characterization signal intensity are combined and arranged in a six-dimensional tensor combination according to the time dimension, spatial dimension, signal dimension, parameter dimension, analytical dimension, and label dimension to construct a six-dimensional standardized input dataset containing time, location, characterization signal intensity, operating condition parameters, comprehensive analytical values, and multi-region spectral trend comparison labels. The output is a six-dimensional standardized input dataset.

5. The multi-modal fusion based multi-zone in-situ working condition automatic monitoring system according to claim 1, characterized in that, The feature extraction module specifically includes: The signal intensity in the six-dimensional standardized input dataset is divided into non-overlapping image blocks according to the spatial coordinates of multiple regions. The values ​​of all pixels in each image block are flattened into a one-dimensional vector. The horizontal and vertical index values ​​of the current image block are extracted. The total dimension of the model embedding vector is set to d. The total dimension d is divided by 2 to obtain the maximum value of the dimension index k. Calculate 10000 using dimension index k that increments from 0. 2k / d As the division factor for the current dimension, the horizontal and vertical index values ​​are divided by the division factor respectively. The sine value of the quotient of the horizontal coordinate is used to generate the first half-dimensional component, and the cosine value of the quotient of the vertical coordinate is used to generate the second half-dimensional component. The two components are concatenated into a preset position encoding vector, and then added to the one-dimensional vector one by one to generate a basic image block sequence containing spatial position information. The basic image patch sequence is input into the spatial block of the Uniformer model. The dot product between any two image patch vectors in the sequence is calculated and divided by the square root of the vector dimension to obtain the attention score. The attention score is input into the Softmax function to transform the probability distribution weights, and the basic image patch sequence is weighted and summed to aggregate the spatial association information between different multi-region spatial coordinates and output the spatial feature map. The spatial feature map is input into the Uniformer Transformer block of the Uniformer model. A one-dimensional convolution operation is performed on all pixels in the spatial feature map along the channel dimension. The features of local neighboring pixels are extracted by sliding, and the intermediate spatial features are output. The intermediate spatial features corresponding to different time series information under the same multi-region spatial coordinates are arranged in chronological order to form a time feature sequence, which is input into the time block of the Uniformer model. The dot product of the feature vectors of any two time nodes is calculated within the time feature sequence and divided by the square root of the vector dimension. After transformation by the Softmax function, the probability distribution weight of the time dimension is obtained. The time feature sequence is weighted and summed to capture the dynamic evolution trend of the signal and output the time fusion feature sequence. The operating condition parameters in the six-dimensional standardized input dataset are extracted into numerical sequences, which are then transformed into vector dimensions with the same as the time fusion feature sequence through a linear mapping layer. At each time node, the transformed operating condition parameter vector is added element-wise to the corresponding vector in the time fusion feature sequence to generate a cross-modal interactive fusion feature sequence. The cross-modal interaction fusion feature sequence is input into the alternating stacked structure composed of spatial and temporal blocks in the next layer. The above operations of spatial aggregation, local convolution, temporal weighting and cross-modal addition are repeatedly performed to extract multi-layer deep features. The feature sequence output by the last layer is flattened in dimensions and mapped through a fully connected layer to output the interface state feature encoding.

6. The multi-modal fusion based multi-zone in-situ working condition automatic monitoring system according to claim 1, characterized in that, The strategy optimization module specifically includes: The interface state feature encoding is input into the improved QPSO algorithm. In the multi-dimensional parameter space composed of the observation position coordinates, acquisition frequency and chemical excitation parameter set, N initial particles are randomly generated. Each particle corresponds to a five-dimensional parameter vector containing the horizontal coordinate, vertical coordinate, height coordinate, acquisition frequency value and chemical excitation value. For each particle's x-coordinate, y-coordinate, and height coordinate, iterate through all the spatial coordinates that have been executed in the history record, calculate the three-dimensional Euclidean distance between the current particle's coordinates and each historical coordinate, add up all the three-dimensional Euclidean distances and take the average to obtain the average spatial distance, and input the five-dimensional parameter vector of the current particle into the fully connected mapping layer where the hidden layer neuron weights in the interface state feature encoding output by the Uniformer model are used as the connection weights of the corresponding input nodes. The linear transformation of the five-dimensional parameter vector is calculated and summed, then input into the ReLU activation function to calculate the feature activation response value. The average spatial distance is then multiplied by the feature activation response value to obtain the data information gain value. Extract the x-coordinate, y-coordinate, and height coordinate from the current five-dimensional parameter vector of each particle, and extract the x-coordinate, y-coordinate, and height coordinate from the optimal observation position coordinates output in the previous stage. Subtract the x-coordinate, y-coordinate, and height coordinates of the two and take the absolute value. Add the three absolute values ​​to obtain the device loss value. Subtract the device loss value from the data information gain value to obtain the fitness function value of each particle. Read the x-coordinate, y-coordinate, and height coordinates of the optimal observation position coordinates received by the high-precision three-dimensional translation stage in the previous stage, as well as the initial x-coordinate, initial y-coordinate, and initial height coordinates of the high-precision three-dimensional translation stage before it moved. Subtract the initial x-coordinate, initial y-coordinate, and initial height coordinates from the received x-coordinate, y-coordinate, and height coordinates to obtain the X-axis direction component, Y-axis direction component, and Z-axis direction component, and arrange them in order to generate the movement direction vector. An asymmetric bias mechanism based on the probability distribution of the device's motion history trajectory is introduced. A bias coefficient is preset, and the current value of the i-th parameter in the five-dimensional parameter vector of the current particle and the value of the global optimal position with the largest fitness function value among all current particles in the i-th dimension are extracted. The position deviation value is obtained by subtracting the value of the global optimal position in the i-th dimension from the current value. The previous stage movement direction vector component corresponding to the i-th parameter is extracted, and the movement direction vector component is multiplied by the preset bias coefficient to obtain the direction bias amount. Subtract the value of the global optimal position in the i-th dimension from the current value of the i-th parameter and subtract the direction bias to obtain the inner term of the exponential absolute value. Take the absolute value of the inner term of the exponential absolute value, multiply the result by -2 and divide it by the preset contraction and expansion coefficient that controls the convergence speed in the QPSO algorithm to obtain the exponential term. Calculate the result of the power operation with the natural constant e as the base and the exponential term as the exponent. Divide the result of the power operation by 2 times the preset contraction and expansion coefficient to obtain the asymmetric probability density function value of the i-th parameter. The updated five-dimensional parameter vectors for all particles are calculated using the five-dimensional parameter vectors and asymmetric probability density function values. The optimal observation position coordinates, acquisition frequency, and chemical excitation parameter set for the next stage are then output.

7. The multi-region in-situ automatic monitoring system for working conditions based on multimodal fusion according to claim 6, characterized in that, The process of calculating the updated five-dimensional parameter vectors for all particles using the five-dimensional parameter vector and asymmetric probability density function values, and outputting the optimal observation position coordinates, acquisition frequency, and chemical excitation parameter set for the next stage, specifically includes: For the acquisition frequency dimension and chemical excitation value dimension in the five-dimensional parameter vector, the corresponding movement direction vector component is directly set to 0, and multiplied by a preset bias coefficient to obtain a directional bias of 0. Then, the value of the global optimal position in the corresponding dimension is subtracted, and the value is subtracted to 0 to obtain the inner term of the exponential absolute value. The preset contraction and expansion coefficient calculation steps are repeated to maintain the original symmetrical shape of the probability density function of the acquisition frequency dimension and chemical excitation value dimension, which does not change with the movement direction of the device, to obtain an asymmetric probability wave. Perform asymmetric Monte Carlo random sampling. Starting from the current five-dimensional parameter vector of each particle, for each dimension, set a preset number of discrete calculation points between the value of the current dimension and the value of the global optimal position in the current dimension. Substitute each discrete calculation point into the step of calculating the asymmetric probability density function value to obtain the asymmetric probability density function value corresponding to each discrete calculation point. Add the asymmetric probability density function values ​​corresponding to all discrete calculation points in order to obtain the cumulative probability sum. Divide the asymmetric probability density function value corresponding to each discrete calculation point by the sum of cumulative probabilities to obtain the probability ratio of the interval occupied by each discrete calculation point. Accumulate the interval probability ratios in the order of the discrete calculation points to obtain the cumulative probability distribution value corresponding to each discrete calculation point. Randomly generate a decimal between 0 and 1 as the target sampling probability. Compare the target sampling probability with the cumulative probability distribution value corresponding to each discrete calculation point in turn. Find the first target discrete calculation point whose value is just greater than or equal to the target sampling probability. Use the value of the target discrete calculation point as the sampling base offset. Add the value of the global optimal position in the current dimension to the sampling base offset to obtain the updated value of the current dimension. Calculate the updated values ​​of the other four dimensions in the same way, and combine the updated values ​​of the five dimensions to generate the new five-dimensional parameter vector for each particle after the update. Substitute the updated five-dimensional parameter vectors of all particles back into the algorithm and calculate the new fitness function value. Compare the new fitness function value of each particle with the corresponding particle's historical maximum fitness function value. If the new fitness function value is larger, replace the particle's historical optimal position with the new five-dimensional parameter vector. Extract the particle with the largest new fitness function value among all particles. If the maximum value is greater than the historical global maximum fitness function value, replace the global optimal position with the particle's new five-dimensional parameter vector. Determine whether the current number of iterations has reached the preset iteration threshold. If not, return to the step of constructing the asymmetric probability wave and continue execution. If it has, stop the loop, extract the current global optimal position, and output the combination of the included x-coordinate, y-coordinate, and height coordinates as the optimal observation position coordinates for the next stage. Output the included acquisition frequency values ​​as the optimal acquisition frequency for the next stage, and output the included chemical excitation values ​​as the optimal chemical excitation parameter set for the next stage.

8. The multi-modal fusion based multi-zone in-situ working condition automatic monitoring system according to claim 1, characterized in that, The linkage control module specifically includes: The central control computer extracts the x-coordinate, y-coordinate, and height coordinates from the optimal observation position coordinates for the next stage, converts the x-coordinate, y-coordinate, and height coordinates into target pulse signal numbers, and sends the three sets of target pulse signal numbers to the stepper motor driver inside the high-precision three-dimensional translation stage to control the three-axis motor of the high-precision three-dimensional translation stage to rotate and move to the optimal observation position coordinates for the next stage. The optimal observation location coordinates for the next stage are compared with the spatial coordinate set of the monitoring area currently being monitored. If the optimal observation location coordinates for the next stage do not belong to the spatial coordinate set of the monitoring area currently being monitored, the optimal observation location coordinates for the next stage are added to the spatial coordinate set of the monitoring area to complete the automatic addition of the monitoring area. If there are historical spatial coordinates in the spatial coordinate set of the currently monitored area that have not appeared in the optimal observation position coordinates for a set number of consecutive times, the historical spatial coordinates that have not appeared consecutively will be removed from the spatial coordinate set of the monitored area to automatically reduce the monitored area. The central control computer extracts the optimal acquisition frequency for the next stage, converts the acquisition frequency value into a timing clock cycle parameter, sends the timing clock cycle parameter to the trigger controller, and automatically executes signal triggering and acquires multi-channel in-situ characterization signals according to the timing clock cycle parameter. The central control computer extracts the optimal chemical excitation parameter set for the next stage, reads the specific target values ​​of the electric field, thermal field, force field or chemical field corresponding to the chemical excitation parameter set, sends the specific target values ​​to the working condition workstation, controls the working condition workstation to adjust the output power of the electrochemical workstation or temperature controller, and modifies the voltage, current, and capacity parameters under electrochemical working conditions or the temperature and pressure parameters under thermal / mechanical working conditions to the specific target values. The central control computer packages the target pulse signal number sent to the motor driver, the timing clock cycle parameter sent to the trigger controller, the specific target value sent to the working condition workstation, and the update results of automatically adding or deleting monitoring areas, and generates and outputs multi-threaded linkage instructions and corresponding execution actions.

9. The multi-modal fusion based multi-zone in-situ working condition automatic monitoring system according to claim 1, characterized in that, The closed-loop control module specifically includes: The central control computer sets a preset total monitoring duration, reads the current system clock time, and determines whether the current clock time has reached the preset total monitoring duration. If it has not reached the preset total monitoring duration, the data acquisition and preprocessing module is automatically triggered to re-execute data acquisition. The multi-region monitoring and reaction state control module, data acquisition and preprocessing module, feature extraction module, strategy optimization module and linkage control module are treated as a complete working cycle. In each working cycle, the latest set of the next stage optimal observation position coordinates, acquisition frequency and chemical excitation parameter set output by the strategy optimization module is used to replace the historical parameters of the previous working cycle. The observation strategy is dynamically iterated based on the real-time acquired multi-channel in-situ characterization signals and key state parameters. Extract the six-dimensional standardized input dataset output by the preprocessing module in each work cycle, and concatenate the six-dimensional standardized input dataset of the current work cycle to the end of the dataset generated in the previous work cycle according to the time dimension. As the work cycles are executed cyclically, the amount of data is continuously expanded. When the current clock time reaches the preset total monitoring duration, the data acquisition and preprocessing module stops, and the datasets generated by splicing all work cycles are reorganized according to the dimensions of multi-regional spatial coordinates, time series information, characterizing signal strength and operating parameters, and a multi-dimensional spatiotemporal evolution dataset is automatically generated and output.