Standardized detection result calibration method based on multi-modal fusion
By using multimodal fusion methods for time synchronization and spatial alignment, combined with local anomaly factors and a Bayesian online learning framework, the time synchronization and calibration accuracy issues of heterogeneous multi-source data are solved, achieving high-accuracy and adaptive calibration of the system in a dynamic environment.
Patent Information
- Application Number
- CN202511106339.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
When processing heterogeneous multi-source data, existing technologies face difficulties in time synchronization and standardization processing, insufficient stability in anomaly detection, difficulty for traditional models to adapt to dynamic changes in the system, and insufficient calibration accuracy and automation level.
Through the multimodal fusion method, time synchronization and spatial alignment are performed, the local outlier factor (LOF) is used to identify and eliminate outliers, a high-confidence system state representation model is constructed, and the Bayesian online learning framework is combined for parameter identification and calibration. Closed-loop recalibration is performed, and predictive calibration is achieved using a data-driven anomaly detection algorithm.
It achieves precise alignment and adaptive calibration of multimodal data, ensuring that the model maintains high diagnostic accuracy in dynamic environments, providing reliable probabilistic output, and improving the system's calibration accuracy and adaptability through comprehensive closed-loop and predictive calibration.
Smart Images

Figure CN120597221A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a standardized detection result calibration method based on multimodal fusion. Background Art
[0002] With the acceleration of digitalization and intelligentization in various fields, real-time and accurate monitoring of the operating status of complex systems and calibration of results have become crucial. Traditional monitoring methods usually rely on a single type of data source or sensor modality. Although these single-modality data can provide basic information in specific applications, they are difficult to fully and accurately reflect the true status of the system when faced with a complex, changeable and uncertain actual environment.
[0003] In existing technologies, to improve monitoring accuracy and analysis depth, some methods attempt to introduce advanced data processing techniques, such as signal enhancement, feature extraction, and machine learning-based models to identify anomalies. However, the core challenge lies in the processing of heterogeneous multi-source data. Data sources collected by different sensors have different structures, frequencies, and formats, making time synchronization and standardization processing fundamental challenges in achieving data fusion. Outliers and noise are prevalent in the data collection process and need to be eliminated. However, existing anomaly detection methods mostly target a single data modality, resulting in insufficient stability in anomaly identification. The dynamic evolution of system operating states makes it difficult for traditional fixed models to adapt, and diagnostic accuracy will inevitably decline over time. Failure to compensate for offsets in a timely manner will gradually lead to large errors from small offsets. Even corrections rely on time-consuming offline retraining. The inherent bias and drift of measurement equipment affect the credibility of the data. Existing calibration methods lack in-depth utilization of the synergistic effects of multi-source data, resulting in insufficient calibration accuracy and automation.
[0004] To this end, a standardized detection result calibration method based on multimodal fusion is proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a standardized test result calibration method based on multimodal fusion to address the deficiencies of the prior art in heterogeneous multi-source data processing, anomaly detection, model adaptability, and calibration accuracy. To achieve the above purpose, the present invention provides the following technical solutions: A standardized detection result calibration method based on multimodal fusion, comprising: Perform temporal synchronization and spatial alignment on the collected multimodal data of different modalities, different sampling frequencies, and different formats; perform standardized preprocessing on the multimodal data, use local outlier factors (LOFs) to perform real-time quality assessment on the multimodal data, identify and remove outliers and noise; and use the inherent correlation between different modal data for cross-validation during the outlier detection process. Build and continuously optimize a high-confidence system state representation model that characterizes the performance of system components, integrate the pre-processed and quality-assessed multimodal data in real time, and apply a Bayesian online learning framework to perform online parameter identification and adaptive calibration of the internal parameters of the system state representation model; For system state deviations, closed-loop recalibration is performed by combining the cross-validation and residual analysis, and predictive calibration is achieved using a data-driven anomaly detection algorithm. Based on the calibrated system state representation model, degradation diagnosis and prediction of the performance of the monitored systems and equipment are performed.
[0006] Preferably, the steps of time synchronization and space alignment include: A high-precision dual calibration strategy is used to record the exact time information for each data point as a timestamp. Accurate matching is performed based on the timestamp to link the consistency of different modal data points on the time axis. The cross-correlation function between different modal data sequences is then calculated, and the cross-correlation peak is found to determine the optimal time deviation to align asynchronously collected multimodal data. Based on the spatiotemporal alignment results, the non-uniformly sampled data is resampled and corrected using the spline interpolation method. The position and posture relationship of each sensor is calibrated, the data is unified into a global coordinate system through coordinate transformation, an inertial measurement unit is introduced, and small displacements and attitude drifts of the sensor during use are dynamically compensated; feature points are extracted from the multimodal data respectively; the matching relationship between the feature points is determined; a transformation matrix is calculated based on the matching feature points; and the multimodal data is spatially aligned according to the transformation matrix.
[0007] Preferably, the parameters of the local anomaly factor algorithm include a neighborhood size k value and a threshold value, which are dynamically adjusted according to historical data and / or expert experience. The dynamic adjustment of the neighborhood size k value is adaptively adjusted according to the density distribution or cluster characteristics of historical detection data, and the dynamic adjustment of the threshold value is adaptively adjusted according to the abnormal proportion and severity of historical detection results.
[0008] Preferably, the cross-validation using the intrinsic correlation between different modal data includes: Construct multimodal data association rules and use the potential correlations learned by the machine learning model. When the sensor data of one modality shows abnormal characteristics of the device, automatically trigger the synchronous analysis of at least one modal sensor that has physical and logical associations with the single modal sensor, cross-check the sensor data from the associated modalities, and use statistical methods, pattern matching algorithms and / or pre-trained anomaly classifiers to determine whether the data presents abnormal characteristics.
[0009] Preferably, the constructing and continuously optimizing the high-confidence system state representation model specifically includes: Based on historical data, expert knowledge and physical principles, deep neural networks and hybrid modeling architecture are used to construct an initial system state representation model to map the relationship between multimodal input data and the performance indicators of each system component. Through online learning, transfer learning and reinforcement learning, continuous optimization is carried out, taking into account environmental changes, new faults and performance degradation, multimodal data is calibrated in real time, and model parameters are dynamically adjusted.
[0010] Preferably, the online parameter identification and calibration specifically includes: A variational Bayesian neural network architecture is constructed, in which the network weights and biases are modeled as probability distributions, and the variational parameters of these distributions are initialized. When new multimodal data is input, samples are taken from the variational distributions of each weight and bias, and forward propagation is performed to obtain model predictions. Based on the difference between the model predictions and the actual observations, the gradient of the loss function with respect to the variational parameters is calculated through a backpropagation algorithm. Using the gradient, an optimization algorithm is used to update the variational parameters to optimize the lower bound of the evidence. The construction, sampling, forward propagation, gradient calculation and variational parameter update processes are continuously and cyclically executed to perform online parameter identification and adaptive calibration of the system state representation model parameters.
[0011] Preferably, the method further comprises applying a reinforcement learning agent to generate an optimal action sequence: The system state representation calibrated with multimodal data is used as the core input to drive an intelligent multi-objective reinforcement learning agent, which adopts a hierarchical reinforcement learning architecture and integrates model predictive control technology. By using the calibrated system state representation model as a predictor, the agent simulates the potential impact of different control strategies over a period of time in the future, selects the optimal action sequence that can optimize multiple conflicting objectives, and dynamically adjusts the multi-objective reward weights based on system state changes and priority adjustments. At the same time, combined with adversarial training, it accurately predicts the key indicators of the monitored object.
[0012] Preferably, calibration for specific events is also included, specifically including: For specific events, which are non-periodic, unconventional, and special events that have a significant impact on the accuracy of the detection results, a high-frequency data collection and analysis mode is immediately activated, and preset calibration rules and model fine-tuning strategies for specific events are called according to the event type; when the specific event occurs, the new data generated is used to quickly update and calibrate the internal parameters of the system state representation model and the reinforcement learning agent, and based on the context of the specific event, the impact of the specific event on the performance indicators is predicted, and the optimal action sequence is generated after adjusting the optimization target.
[0013] Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the data that has been initially aligned by timestamp, cross-correlation technology can further fine-tune the alignment effect and eliminate residual time deviations caused by sensor characteristics, transmission delays or subtle system desynchronization. For example, if two sensors record the same event slightly earlier or later than the other, the cross-correlation function can help identify this lag and accurately align the data.
[0014] 2. The present invention constructs and optimizes a high-confidence system state representation model. By utilizing deep learning and a Bayesian online learning framework, the model can learn and adapt to environmental changes and system degradation in real time. This continuous learning capability ensures that the model maintains high diagnostic accuracy during dynamic operation and provides reliable probabilistic output.
[0015] 3. This invention achieves comprehensive closed-loop and predictive calibration. It ensures sustained accuracy through multi-sensor cross-validation and model residual analysis. Data-driven algorithms predict and correct potential biases. Reinforcement learning agents further enable multi-objective optimization and enable instantaneous calibration and contextualized predictions for specific events, supporting intelligent maintenance and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flowchart of the steps of a standardized detection result calibration method based on multimodal fusion proposed by the present invention; Figure 2 The spatiotemporal synchronization and standardization process of multimodal data proposed in this invention; Figure 3 This is a flow chart of online calibration of the system state representation model of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] See also Figures 1 to 3 The present invention provides a standardized detection result calibration method based on multimodal fusion. The technical solution is as follows: Figure 1 , the specific steps include: Perform temporal synchronization and spatial alignment on collected multimodal data of different modalities, different sampling frequencies, and different formats; standardize the multimodal data and use local outlier factors (LOFs) to perform real-time quality assessment on the multimodal data to identify and remove outliers and noise; during the outlier detection process, use the inherent correlation between different modal data for cross-validation; Build and continuously optimize a high-confidence system state representation model that characterizes the performance of system components, combine the real-time fusion of the pre-processed and quality-assessed multimodal data, and apply a Bayesian online learning framework to perform online parameter identification and calibration of the internal parameters of the system state representation model; In response to system state deviations, closed-loop recalibration is performed by combining multi-sensor cross-validation and residual analysis, and predictive calibration is achieved using a data-driven anomaly detection algorithm. Based on the calibrated system state representation model, degradation diagnosis and prediction of the performance of the monitored systems and equipment are performed. Example 1:
[0019] The scenario of this embodiment is a relatively independent and static smart agriculture subsystem - an intelligent irrigation water pump system, which realizes accurate monitoring of the health status of the equipment, fault prediction and operation optimization.
[0020] First, the processing requires that all data acquisition devices accurately record the current time information when acquiring data. This is called a timestamp. The timestamp accuracy is better than 50 microseconds. Nanosecond time synchronization is achieved through the IEEE 1588 (PTP) protocol. A very small time window is set, and data points with timestamps falling within the same window in different modal data are considered to be initially synchronized. For example, if the temperature sensor and humidity sensor record data at the same millisecond, then a direct comparison of the timestamps can closely link these two data points. The specific multimodal data collected is: Modality A is environmental data, including air temperature, relative humidity, light intensity, and carbon dioxide concentration. The acquisition device is synchronized with a standard time server either built-in or through the network time protocol to ensure clock accuracy. Modality B is crop canopy imagery, and the shooting time is recorded in the metadata of the image file. Modality C is soil data, including soil moisture content, electrical conductivity (EC), and soil temperature. For example, the ambient temperature (22.5°C) recorded at 2023-11-15T08:00:00.125Z, the crop canopy image captured at 2023-11-15T08:01:04.500Z, and the soil moisture content (35.2%) obtained at 2023-11-15T08:00:01.850Z, each data point is accompanied by a UTC timestamp accurate to milliseconds.
[0021] A dual calibration strategy is used to achieve high-precision time synchronization, referring to Figure 2 : It is a spatiotemporal synchronization and standardization processing process for multimodal data; specifically, first record the precise time information for each data point as a timestamp. For example, a high-precision timestamp is attached to the ambient temperature data of 22.5°C at 2023-11-15T08:00:00.125Z and the crop image data at 2023-11-15T08:01:04.500Z, and then accurately match them based on the timestamps to link the consistency of different modal data points on the time axis; this is the first-level calibration, which is specifically implemented as follows: based on the image timestamp, find the environmental data point closest in time within the preset window, for example, the data with a timestamp of 2023-11-15T08:01:00.128Z, to complete the preliminary association.
[0022] On this basis, by calculating the cross-correlation function between different modal data sequences, using the normalized cross-correlation function, with a calculation window size of 10 seconds, the time lag is determined by finding the cross-correlation peak, and the optimal time deviation is determined by finding the cross-correlation peak to calibrate and align the asynchronously collected multimodal data; this serves as the second-level calibration, which is used to compensate for slight errors in timestamps or process data collected completely asynchronously; for example, through calculation, it is found that the image acquisition sequence has a systematic time deviation of -1.8 seconds relative to the environmental monitoring sequence. This deviation is the time lag corresponding to the cross-correlation peak; accordingly, the original timestamp of the aforementioned image will be corrected to 2023-11-15T08:01:02.700Z.
[0023] Furthermore, based on the spatiotemporal alignment results, cubic spline interpolation is used to resample the data to a unified sampling frequency of 100 Hz. This ensures that all data are processed at a uniform temporal resolution and the position relationship of each sensor is calibrated. For example, the relative position of the high-definition camera installed on the agricultural robot relative to the lidar sensor is determined to be x: 0.5m, y: 0.0m, z: 0.2m, and the attitude rotation is roll: 0°, pitch: -15°, yaw: 0°, where roll refers to the rotation of an object around its own front-to-back axis (usually the x-axis). ) rotation, Pitch refers to the rotation of an object around its own left-right axis (usually the y-axis), and Yaw refers to the rotation of an object around its own vertical axis (usually the z-axis). Attitude rotation roll: 0°, pitch: -15°, yaw: 0° means that the object has no roll (no tilt left or right) and no yaw (no left or right rotation) in space, but it is diving down or tilting 15 degrees; then coordinate transformation is used to unify the data into a global coordinate system with a corner of the field as the origin, ensuring that the spatial data from different sensors are in the same reference frame.
[0024] Furthermore, this method introduces an inertial measurement unit to dynamically compensate for the sensor's slight displacement and attitude drift during use. This is achieved by fusing IMU and GNSS data through an extended Kalman filter. For example, when the vehicle is driving, the IMU detects a heading angle drift of +0.1 degrees. The system will use this data in real time to compensate for the attitude of all subsequently collected spatial data. The inertial measurement unit will continuously observe whether there is drift, realize long-term correction and timely compensation of offset, and fuse the positions of different data sensors to obtain a more accurate position estimate.
[0025] Specifically, for modal data containing spatial features, such as images and LiDAR point clouds, this method extracts feature points from each modality and determines their matching relationships. For example, if the coordinates of a crop leaf tip identified in an image are px:450,py:620, the corresponding 3D coordinates of x:5.21m, y:1.55m, z:-0.10m are found in the LiDAR point cloud. Based on these matching feature points, a homogeneous transformation matrix is calculated from the camera pixel coordinate system to the LiDAR 3D coordinate system.
[0026] By combining high-precision timestamp alignment and cross-correlation function calibration, accurate time synchronization of data in different modalities is achieved, effectively addressing the challenges of asynchronous acquisition; at the same time, combined with sensor pose calibration, IMU dynamic compensation and feature point matching, accurate spatial correspondence of multimodal data is ensured, greatly improving the reliability of data fusion.
[0027] By constructing a semantic representation of multimodal data and integrating ontological knowledge of the external environment, and adopting empathic cognitive learning, the calibration system can go beyond preset rules and generate a set of nonlinear calibration strategies for specific complex and unseen scenarios. It can also dynamically identify and quantify the impact of scenario changes on the calibration target weights, ensuring soft adaptive calibration in highly uncertain environments, rather than just hard parameter adjustments.
[0028] The parameters of the local anomaly factor algorithm include the neighborhood size k value and the threshold. These parameters will be dynamically adjusted based on historical data and / or expert experience. For example, the initial setting k value range is 5 to 30, and the anomaly score threshold is set to 1.5. Specifically, for vibration data with high noise levels, for example, for equipment vibration acceleration data, whose normal fluctuation range is between 0.2-0.5m / s² but accompanied by high-frequency noise, the system will use a larger neighborhood size, such as k=25, and set a higher threshold LOF>2.0 to smooth the noise impact and accurately identify true anomalies caused by bearing wear that continuously exceed 1.2m / s². For temperature data with relatively stable changes, for example, when monitoring greenhouse ambient temperature, its changes are usually gentle, a smaller neighborhood size, such as k=5, is used, and a more sensitive threshold LOF>1.6 is set to effectively capture the small but continuous abnormal temperature rise of 0.2°C per minute caused by ventilation system failure.
[0029] By dynamically adjusting the parameters of the local anomaly factor algorithm, intelligent adaptive processing of data with different noise levels is achieved. This enables the system to more effectively smooth high-noise vibration data and sensitively capture tiny anomalies in smooth temperature data, thereby improving the accuracy of data quality assessment.
[0030] Furthermore, the cross-validation using the intrinsic correlation between different modal data specifically includes: first, constructing multimodal data association rules and utilizing the potential correlation learned by the machine learning model, the association rules are mined from historical data based on the Apriori algorithm, and the machine learning model is a cross-modal anomaly detection model based on the Transformer architecture; for example, the system learns from historical data the strong positive correlation between the motor current, water flow rate and pipeline pressure of the irrigation water pump, and constructs an association rule: when the motor current rises from the normal 4.5A to 6.0A, the water flow rate should synchronously increase from 50L / min to 65L / min.
[0031] When sensor data of one modality indicates abnormal characteristics of the equipment, for example, the vibration sensor installed on the water pump detects an instantaneous surge in vibration amplitude to 3.5g at 10:32:15Z, far exceeding the normal threshold of 1.5g and is initially marked as abnormal, the system will automatically trigger a synchronous analysis of the data of at least one modal sensor that has physical and logical associations with the single modal sensor, namely the motor current, water flow rate and pipeline pressure sensor within a few seconds before and after 10:32:15Z.
[0032] Construct and maintain in real time a topologically heterogeneous mapping diagram of the multimodal sensor network to characterize the logical and physical multiple connections and data flow paths of different types of sensors. When any sensor or its data flow fails, the system intelligently reconstructs the transmission path based on the mapping diagram, and dynamically selects and fuses redundant data from a heterogeneous sensor combination with topological proximity and information complementarity to perform fault-tolerant calibration, thereby enhancing the reliability of the system in the event of local sensor failure or data anomalies and achieving a higher level of system adaptability.
[0033] Furthermore, the system cross-checks sensor data from associated modalities and applies statistical methods, pattern matching algorithms, and / or pre-trained anomaly classifiers. In this case, the anomaly classifier used is a one-dimensional convolutional neural network trained on a historical fault dataset to determine whether the data exhibits abnormal characteristics. In this case, the system cross-check found that at the same time, the motor current also soared to 7.2A (normal value is approximately 4.5A), while the water flow rate dropped sharply to 10L / min (normal value >50L / min). This "high vibration, high current, low flow" data pattern highly matched the pre-trained "water pump impeller blockage" fault classifier with a matching confidence level of 98.5%, confirming that this was a real equipment anomaly event.
[0034] By utilizing the inherent correlation between multimodal data for cross-validation, when an anomaly occurs in one modal data, it can automatically trigger the synchronous analysis of the associated modal data. This significantly enhances the robustness and accuracy of anomaly identification, effectively reduces false alarms and missed alarms, and improves the reliability of the system's judgment of abnormal events.
[0035] In the smart agriculture implementation scenario, taking a smart irrigation pump system as an example, a model that can assess the health of the pump system in real time is constructed as follows: A real-time data stream from the pump system is aligned in time and space, including motor current (A), outlet pipe pressure (kPa), frequency domain data collected by the vibration sensor, and motor housing temperature (°C). The model outputs a state vector containing performance indicators such as energy efficiency (L / kWh), bearing wear index (0-1), and estimated remaining useful life (RUL) in hours. The model architecture adopts hybrid modeling, which embeds the basic operating concept of water pump efficiency into the physical principle layer based on expert knowledge and experience, and forms the benchmark of the model according to the theoretical relationship between flow, pressure and power consumption.
[0036] The deep neural network layer specifically uses a long short-term memory network, which is particularly suitable for deep neural networks that process time series data. The LSTM network specializes in learning the complex patterns of vibration spectra and motor temperature changes over time in historical data. These patterns are difficult to describe with physical formulas, but are highly correlated with progressive degradation processes such as bearing wear and material fatigue. The hybrid model is trained using several months of historical operating data, including normal operation, minor faults such as minor blockages, and a complete data set before and after a planned maintenance bearing replacement. After training, the model can output accurate status assessments such as {"energy efficiency": 92.5L / kWh, "bearing wear index": 0.45, "RUL": 2500 hours} based on the input real-time multimodal data. Figure 3 , which is a flowchart of online calibration of the system state representation model of this application; Furthermore, to optimize the irrigation strategy, this method introduces a reinforcement learning architecture, in which the state representation model provides the state, that is, the real-time health status of the water pump. The reinforcement learning agent is the irrigation controller, whose action is to adjust the water pump speed or switch. The reward function is designed to maximize the satisfaction of crop water needs while minimizing the electricity cost and the exponential growth of water pump wear. Through continuous trial and error, the agent can learn an optimal strategy. For example, it will be found that when water is not needed urgently, running at 80% power for a long time can meet the irrigation needs while reducing the RUL consumption of the water pump by 15%, compared with intermittent operation at 100% power. This not only realizes the active management of equipment performance degradation, but also maximizes the overall benefit of the system.
[0037] By building a system state representation model based on deep neural networks and hybrid modeling architecture, and using online learning, transfer learning, and reinforcement learning for continuous optimization, the model can adapt to environmental changes, new faults, and performance degradation in real time, ensuring high confidence and dynamic adaptability in the performance characterization of each component of the system.
[0038] The online parameter identification and calibration specifically includes: first constructing a variational Bayesian neural network architecture, wherein the network weights and biases of the architecture are modeled as probability distributions. For example, a key weight is no longer a single value, but is initialized to a Gaussian distribution with variational parameters of mean μ=0.8 and variance σ²=0.01.
[0039] When new multimodal data is input, the system samples from the variational distribution of each weight and bias and performs forward propagation to obtain the model prediction; this process generates a prediction result with uncertainty. For example, the model predicts that the system energy efficiency at that moment is 90.5L / kWh and gives a 95% confidence interval of [89.9,91.1]L / kWh. This confidence interval quantifies the model's confidence in its own prediction.
[0040] Then, based on the difference between the model prediction and the actual observation, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the variational parameter. The loss function is the negative of the evidence lower bound. In this scenario, the actual energy efficiency observation calculated by the real-time flow meter and power meter is 92.0 L / kWh. There is a significant difference between the model-predicted mean of 90.5 and the actual value of 92.0.
[0041] Finally, the gradient is used to update the variational parameters using the Adam optimization algorithm to optimize the lower bound of evidence. Specifically, because the model underestimates efficiency, the optimization process will adjust the distribution of relevant weights. For example, the mean μ of the weights will be updated to a higher value, such as 0.82, in order to produce predictions closer to the actual value when encountering similar inputs in the future. This method will continuously execute the construction, sampling, forward propagation, gradient calculation, and variational parameter update processes in a loop to perform online identification of the parameters of the system state representation model. The standardized test result calibration method uses an endogenous evolutionary algorithm to drive the calibration model architecture and parameter adaptive growth. Periodic structural mutation, cross-breeding, and selection are performed to dynamically evolve the optimal calibration model architecture that best adapts to the current operating environment and data characteristics. At the same time, combined with the simulated annealing optimization strategy, a certain degree of non-optimal changes are allowed during the model structure evolution process to escape the local optimal solution, achieving global convergence and adaptation of the calibration performance.
[0042] By constructing a variational Bayesian neural network architecture, using the probability results output by the architecture as the basis for parameter updating, and continuously and periodically executing the cyclic parameter update process, online parameter identification and adaptive calibration of the system state representation model are achieved, overcoming the defect of the system's real-time dynamic response in a dynamic environment that requires maintaining high precision for a long time.
[0043] The system state representation, calibrated with multimodal data, serves as the core input to drive an intelligent multi-objective reinforcement learning agent. This agent uses a hierarchical reinforcement learning architecture and integrates model predictive control technology. In the smart greenhouse implementation scenario, the agent's goal is to achieve refined autonomous control of greenhouse environments such as temperature, humidity, light, and CO2 concentration.
[0044] The core input is the calibrated system state. The agent's decision-making is based not on raw sensor readings but on high-level semantic information output by the aforementioned "system state representation model." For example, at a given moment, the agent receives the following state input: {"crop photosynthesis efficiency": 85%, "downy mildew risk index": 0.12, "current energy consumption level": 'high', "remaining life of wind turbine equipment": 4500 hours}.
[0045] The hierarchical reinforcement learning architecture is specifically designed with a high-level agent serving as the policy layer, responsible for formulating long-term macro-strategies, such as those for the next four hours. This agent determines the greenhouse's overall operating mode based on the current crop growth stage and external weather forecast. For example, it might issue a high-level goal: enter energy-saving and growth-promoting mode. The lower-level agents, serving as the execution layer, receive the high-level goal and break it down into specific, short-term, control instructions for equipment, such as those for the next five minutes. In this "energy-saving and growth-promoting" mode, the lighting control agent, ventilation system control agent, and fertigation agent work together to calculate the optimal lighting duration, fan speed, and irrigation volume.
[0046] Specifically, by utilizing a calibrated system state representation model as a predictor, the agent is able to simulate the potential impact of different control strategies over a period of time in the future.
[0047] Model predictive control fusion uses the state representation model to perform forward-looking simulations when the high-level agent considers whether to execute "Strategy A: Increase lighting by 10%" or "Strategy B: Maintain current lighting and reduce ventilation energy consumption by 5%" in the next hour. The specific simulation states are as follows: The simulation model of state A predicts that implementing strategy A will increase the "photosynthesis efficiency" to 92%, but the "energy consumption level" will become extremely high, and the "downy mildew risk index" will rise slightly to 0.13; the simulation model of state B predicts that implementing strategy B will maintain the "photosynthesis efficiency" at 85%, but the "energy consumption level" will drop to medium, and the "downy mildew risk index" will remain unchanged.
[0048] Based on the simulation results, the agent is able to select the optimal action sequence that optimizes multiple conflicting objectives; for example, maximizing production efficiency, minimizing energy consumption, and improving resource utilization efficiency, and can dynamically adjust the multi-objective reward weights based on system state changes and priority adjustments.
[0049] The reward function of the agent in the dynamic multi-objective optimization setting is designed as Crop yield gain)- Energy costs)- disease risk), of which is the weight corresponding to the crop yield gain, is the weight corresponding to the energy cost, is the weight corresponding to the disease risk.
[0050] During critical periods of crop growth, such as the fruiting period, the system automatically increases the crop yield weight gain, making the agent more inclined to choose strategies that maximize yield, such as Strategy A mentioned above. When receiving a peak electricity price warning from the power grid, the system dynamically increases the energy cost weight, guiding the agent to choose a more energy-saving strategy, such as Strategy B, which achieves intelligent peak shaving and valley filling.
[0051] At the same time, adversarial training enhances model stability by injecting simulated noise and faults, ultimately enabling accurate predictions of key indicators of the monitored object. During the training phase, the system introduces an adversarial module that simulates unexpected sensor failures, such as a momentary 5°C drift in the thermometer reading or a partial actuator failure, such as a fan stall. By training in this "adversity," the agent learns to make safe, suboptimal decisions even under nonideal conditions. Ultimately, through repeated simulations, the agent not only outputs optimal control command sequences—for example, "For the next hour, maintain the fill light power at 80%, the skylight opening angle at 30%, and the circulation fan speed at 65%—but also accurately predicts key indicators, such as "It is expected that the tomato biomass in the greenhouse will increase by 3.2% over the next 24 hours, and the total power consumption will be 210 kWh."
[0052] The fully calibrated system state representation model is used as the core input of the multi-objective reinforcement learning agent, reflecting the operating status of the real scenario in real time. It integrates a hierarchical architecture to simulate and predict the impact of different control strategies. Under the constraint of optimizing multiple conflicting objectives, a long-term and optimal action sequence can be determined, and key indicators in the scenario can be accurately predicted and provided decision support.
[0053] This method enables instantaneous response and calibration to specific events, such as equipment failures or sudden changes in the external environment. In smart agriculture scenarios, when the system identifies a localized blockage in a drip irrigation line through multimodal data, such as an abnormally high pressure increase or a sudden drop in flow, it immediately activates high-frequency data acquisition and performs event-driven parameter updates. This update is rapid and localized. For example, the "irrigation efficiency" parameter for the corresponding area in the system state representation model would be immediately adjusted down from 0.95 to 0.20 without the time-consuming retraining of the entire model, ensuring that the model can quickly adapt to the sudden change in system state caused by the event.
[0054] The system then makes contextualized predictions based on this event. It can predict both short-term and long-term impacts, such as "Without intervention, crop yields in this area will drop by 20% due to water shortages within six hours." Based on this precise prediction, the reinforcement learning agent dynamically adjusts its optimization objective, switching from the conventional "minimize cost" to "save crops," and generates an optimal sequence of actions. For example, it can send a precise alert to administrators that "pipeline blockage in Area 3 poses a risk of water shortages," and propose an intelligent intervention recommendation to "temporarily increase irrigation in adjacent areas to replenish water through lateral infiltration."
[0055] By activating high-frequency acquisition for specific events, calling preset rules, and fine-tuning models, instantaneous response and precise calibration are achieved; this enables the system to use event data to quickly update parameters and make contextual predictions, greatly improving the predictive calibration and optimized control capabilities under critical incidents.
[0056] This embodiment targets an intelligent irrigation water pump system. It uses timestamps and cross-correlation functions to perform high-precision spatiotemporal synchronization of multimodal data such as the environment, crop images, and soil. It uses a dynamic local anomaly factor algorithm to evaluate data quality and uses multimodal data correlation cross-validation to identify anomalies. It then constructs and optimizes a system state model, combines it with a Bayesian online learning framework to achieve adaptive parameter calibration, and implements accurate prediction and intelligent irrigation optimization through a reinforcement learning agent. Example 2
[0057] This embodiment is applied to a 100-hectare large-scale smart farm. The core operations are completed by a collaborative heterogeneous robot cluster, including a high-altitude reconnaissance drone and a ground-based precision spraying vehicle. The system's goal is to achieve early detection of pests and diseases and precise, on-demand removal of variable quantities.
[0058] In this embodiment, the UAV, as the first modality, draws a "pest and disease heat map" at an altitude of 50 meters, and the ground spraying vehicle, as the second modality, needs to perform precise operations based on this map. The accuracy of spatiotemporal alignment directly determines whether the pesticide can be sprayed on the target plants. In the time synchronization step, in addition to adding a high-precision timestamp to each data point, such as each hyperspectral image taken by the UAV and each LiDAR scan frame of the ground vehicle, the second-level cross-correlation calibration is used here to correct the variable delay caused by the wireless communication link between the UAV-ground station-ground vehicle. For example, by comparing the RTK-GPS trajectory data recorded by the UAV and the ground vehicle respectively, the system can calculate the average time deviation between the two, such as the data of the ground vehicle has a 250ms lag compared to the UAV, and compensate for it.
[0059] First, each intelligent agent's own sensors are statically calibrated. The key lies in dynamic alignment: the drone identifies the diseased area in the image it captures and records the global geographic coordinates of its center point. These coordinates are sent to the ground vehicle, which uses its own RTK-GPS positioning to convert these global coordinates into its own local coordinate system in real time. For example, it can convert them into a navigation target of "15.3 meters ahead, 2.8 meters to the right", thus achieving precise mapping from aerial reconnaissance to ground execution.
[0060] The model no longer describes the physical wear and tear of a single piece of equipment. Instead, it dynamically constructs a digital twin map reflecting the crop's growth status while simultaneously monitoring the status of the robot cluster. State representation model construction: The system constructs a coupled state model. One component is a "crop health model" based on a convolutional neural network. This model uses drone hyperspectral imagery as input and outputs a high-resolution grid map. The value of each grid represents the probability of disease occurrence in that area. The other component is a "cluster status model" that monitors each piece of equipment's remaining battery / fuel level, pesticide level, current location, and operating efficiency.
[0061] The spectral characteristics of crops used in online parameter identification vary with the growth cycle and environmental changes. For example, in the early stages of crop growth, mild nitrogen stress can be very similar to the spectral characteristics of certain leaf spot diseases. Once a drone initially identifies a suspected diseased area, a ground vehicle approaches and uses its high-definition camera for secondary confirmation. This manual or semi-manual confirmation, the "true value label," is immediately used as a new training sample to update the network weight distribution of the crop health model online using a variational Bayesian method. This enables the model to continuously learn, dynamically distinguishing similar but different visual features at different growth stages, and maintain high accuracy.
[0062] The reinforcement learning agent uses a hierarchical reinforcement learning architecture. The high-level agent is the task planning layer, responsible for receiving crop health maps and weather forecasts. The decision space is "which area to prioritize," "which UGV to dispatch," and "which spraying strategy to adopt, such as low-dose broad-spectrum or high-dose spot spraying." The goal is to maximize the value of crops cured per unit time while minimizing pesticide costs and environmental risks. The low-level agent is the path planning layer, installed on each UGV. It receives specific tasks from the high-level agent, such as "spot spray area A." Using model predictive control, it uses its own LiDAR data to plan an optimal driving path in real time, avoiding field obstacles such as rocks and fallen crops, and precisely controls the timing of the spray nozzles.
[0063] The specific event calibration is set to occur if communication between the drone and ground vehicle is suddenly interrupted during an operation. The ground vehicle immediately switches from "drone-guided mode" to "autonomous emergency mode." Based on the last received damage map and its own sensor data, it will continue the mission within the known operation area or navigate to a safe waiting point.
[0064] The system immediately predicted: "Communications are expected to be interrupted for 10 minutes. If not restored, 15% of the diseased points in the target area will be missed. The wind speed is increasing, and the operation window will close in 10 minutes." At the same time, the system issued an alarm to the main control center and gave advice: "It is recommended that the operator manually guide the UGV to area B with a stronger signal and stand by, and prepare a backup drone to take off." This realizes an intelligent and predictive response to emergencies in complex collaborative tasks.
[0065] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A standardized detection result calibration method based on multimodal fusion, characterized in that: include: Perform time synchronization and spatial alignment on multimodal data collected in different modalities, different sampling frequencies and different formats; Perform standardized preprocessing on multimodal data, use local outlier factors (LOFs) to perform real-time quality assessment on the multimodal data, identify and remove outliers and noise; during the outlier detection process, use the inherent correlation between different modal data for cross-validation; Build and continuously optimize a high-confidence system state representation model that characterizes the performance of system components, integrate the pre-processed and quality-assessed multimodal data in real time, and apply a Bayesian online learning framework to perform online parameter identification and adaptive calibration of the internal parameters of the system state representation model; For system state deviations, closed-loop recalibration is performed by combining the cross-validation and residual analysis, and predictive calibration is achieved using a data-driven anomaly detection algorithm. Based on the calibrated system state representation model, degradation diagnosis and prediction of the performance of the monitored systems and equipment are performed.
2. The method for calibrating standardized detection results based on multimodal fusion according to claim 1, wherein: The steps of time synchronization and space alignment include: A high-precision dual calibration strategy is used to record the exact time information for each data point as a timestamp. Accurate matching is performed based on the timestamp to link the consistency of different modal data points on the time axis. The cross-correlation function between different modal data sequences is then calculated, and the cross-correlation peak is found to determine the optimal time deviation to align asynchronously collected multimodal data. Based on the spatiotemporal alignment results, the non-uniformly sampled data is resampled and corrected using the spline interpolation method. The position and posture relationship of each sensor is calibrated, the data is unified into a global coordinate system through coordinate transformation, an inertial measurement unit is introduced, and small displacements and attitude drifts of the sensor during use are dynamically compensated; feature points are extracted from the multimodal data respectively; the matching relationship between the feature points is determined; a transformation matrix is calculated based on the matching feature points; and the multimodal data is spatially aligned according to the transformation matrix.
3. The method for calibrating standardized detection results based on multimodal fusion according to claim 1, wherein: The parameters of the local anomaly factor algorithm include the neighborhood size k value and the threshold, which are dynamically adjusted based on historical data and / or expert experience. The dynamic adjustment of the neighborhood size k value is adaptively adjusted based on the density distribution or cluster characteristics of historical detection data, and the dynamic adjustment of the threshold is adaptively adjusted based on the anomaly proportion and severity of historical detection results.
4. The method for calibrating standardized detection results based on multimodal fusion according to claim 1, wherein: The cross-validation using the inherent correlation between different modal data includes: Construct multimodal data association rules and use the potential correlations learned by the machine learning model. When the sensor data of one modality shows abnormal characteristics of the device, automatically trigger the synchronous analysis of at least one modal sensor that has physical and logical associations with the single modal sensor, cross-check the sensor data from the associated modalities, and use statistical methods, pattern matching algorithms and / or pre-trained anomaly classifiers to determine whether the data presents abnormal characteristics.
5. The method for calibrating standardized detection results based on multimodal fusion according to claim 1, wherein: The construction and continuous optimization of a high-confidence system state representation model specifically includes: Based on historical data, expert knowledge and physical principles, deep neural networks and hybrid modeling architecture are used to construct an initial system state representation model to map the relationship between multimodal input data and the performance indicators of each system component. Through online learning, transfer learning and reinforcement learning, continuous optimization is carried out, taking into account environmental changes, new faults and performance degradation, multimodal data is calibrated in real time, and model parameters are dynamically adjusted.
6. The method for calibrating standardized detection results based on multimodal fusion according to claim 1, wherein: The online parameter identification and calibration specifically includes: A variational Bayesian neural network architecture is constructed, in which the network weights and biases are modeled as probability distributions, and the variational parameters of these distributions are initialized. When new multimodal data is input, samples are taken from the variational distributions of each weight and bias, and forward propagation is performed to obtain model predictions. Based on the difference between the model predictions and the actual observations, the gradient of the loss function with respect to the variational parameters is calculated through a backpropagation algorithm. Using the gradient, an optimization algorithm is used to update the variational parameters to optimize the lower bound of the evidence. The construction, sampling, forward propagation, gradient calculation and variational parameter update processes are continuously and cyclically executed to perform online parameter identification and adaptive calibration of the system state representation model parameters.
7. The method for calibrating standardized detection results based on multimodal fusion according to claim 1, characterized in that: It also includes applying reinforcement learning agents to generate optimal action sequences: The system state representation calibrated with multimodal data is used as the core input to drive an intelligent multi-objective reinforcement learning agent, which adopts a hierarchical reinforcement learning architecture and integrates model predictive control technology. By using the calibrated system state representation model as a predictor, the agent simulates the potential impact of different control strategies over a period of time in the future, selects the optimal action sequence that can optimize multiple conflicting objectives, and dynamically adjusts the multi-objective reward weights based on system state changes and priority adjustments. At the same time, combined with adversarial training, it accurately predicts the key indicators of the monitored object.
8. The method for calibrating standardized detection results based on multimodal fusion according to claim 7, characterized in that: Also included are calibrations for specific events, including: For specific events, which are non-periodic, unconventional, and special events that have a significant impact on the accuracy of the detection results, a high-frequency data collection and analysis mode is immediately activated, and preset calibration rules and model fine-tuning strategies for specific events are called according to the event type; when the specific event occurs, the new data generated is used to quickly update and calibrate the internal parameters of the system state representation model and the reinforcement learning agent, and based on the context of the specific event, the impact of the specific event on the performance indicators is predicted, and the optimal action sequence is generated after adjusting the optimization target.
Citation Information
Patent Citations
Industrial intelligent detection method and system based on multi-modal large model
CN118503832A
Multi-modal data driven power distribution network equipment state evaluation and optimization method
CN119809439A
Dynamic multi-modal data fusion and real-time analysis method
CN120046119A
Non-corresponding auditing system for data sharing of Internet of Things
CN120145024A
Hybrid energy storage system optimization scheduling method based on AI intelligent regulation and control
CN120150194A
Cited By
Broaching tool machining data acquisition method and system
CN121131866A
A method and system for collecting machining data of a broaching tool
CN121131866B
Beidou passive indoor distribution dynamic closed-loop calibration method based on reinforcement learning
CN121165697A
Intelligent voltage monitoring system with adaptive calibration function
CN121186434A
Three-eccentric center butterfly valve pressure distribution real-time monitoring system based on multi-mode sensing network
CN121253150A