Multi-modal data fusion method and device, equipment, storage medium and computer program product

By mapping and time-aligning multimodal data in the three-dimensional twin model of the power plant, the fuzzy area problem caused by overlapping sensor detection ranges is solved, the accuracy and reliability of data fusion are improved, and the accuracy of equipment status evaluation and fault prediction is improved.

CN120429809APending Publication Date: 2025-08-05济南作为科技有限公司

Patent Information

Application Number
CN202510401652.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The data caused by different sampling frequency and installation locations of existing sensors have differences in time and space, and cannot effectively process fuzzy area data due to overlapping detection ranges, affecting the accuracy and reliability of data fusion results.

Method used

By mapping the multimodal data in the target area into the three-dimensional twin model of the power plant, the fuzzy area is determined, and the target mode data is projected under a unified time ruler for time alignment and fusion, the spatial misalignment of fuzzy area caused by overlapping detection ranges is eliminated.

Benefits of technology

It significantly improves the accuracy and timeliness of equipment status evaluation, fault prediction and optimization decision-making in complex industrial scenarios, and provides key technical support for the digital transformation of smart power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429809A_ABST
    Figure CN120429809A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a multi-modal data fusion method, device and equipment, a storage medium and a computer program product, and the method comprises the steps: after receiving a data fusion request, mapping multi-modal data collected by each sensor in a target region into a three-dimensional twin model of a power plant, and obtaining a plurality of fuzzy regions; for each fuzzy region, determining target modal data of each fuzzy region in a preset time period from the multi-modal data; projecting the target modal data to the corresponding time nodes to obtain a data sequence under a unified time scale; and selecting a target time node corresponding to the data fusion request from the data sequence according to the data fusion request, and fusing target modal data in the target time node to obtain a data fusion result. Target data of sensors at different installation positions are projected to a unified time scale, so that time alignment of multi-sampling-frequency data is realized, and the precision of equipment state evaluation in an industrial scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a multi-modal data fusion method, apparatus, device, storage medium, and computer program product. Background Art

[0002] Due to different sampling frequencies and installation positions of existing sensors, the data collected by different sensors have differences in time and space, and it is impossible to effectively process the data in the fuzzy area that appears due to the overlapping detection ranges, which affects the accuracy and reliability of the data fusion result. Summary of the Invention

[0003] The main purpose of this application is to provide a multi-modal data fusion method, apparatus, device, storage medium, and computer program product, aiming to solve the technical problem that existing sensors cannot effectively process the data in the fuzzy area that appears due to the overlapping detection ranges.

[0004] To achieve the above purpose, this application proposes a multi-modal data fusion method, and the multi-modal data fusion method includes:

[0005] After receiving a data fusion request, map the multi-modal data collected by each sensor in the target area to the three-dimensional twin model of the power plant to obtain multiple fuzzy areas;

[0006] For each of the fuzzy areas, determine the target modal data of each of the fuzzy areas within a preset time period from the multi-modal data;

[0007] Project the target modal data onto the corresponding time nodes to obtain a data sequence under a unified time scale;

[0008] According to the data fusion request, select the target time node corresponding to the data fusion request from the data sequence, and fuse the target modal data in the target time node to obtain a data fusion result.

[0009] Optionally, the step of projecting the target modal data onto the corresponding time nodes to obtain a data sequence under a unified time scale includes:

[0010] Determine the initial timestamp and the maximum acquisition period of the target modal data;

[0011] Based on the maximum acquisition period and the time resolution requirement of the data fusion request, divide the continuous time axis into fixed time windows, and assign corresponding initial time nodes to each of the time windows;

[0012] Perform time alignment on the target modal data according to the initial time node and the initial timestamp to obtain a data sequence.

[0013] Optionally, the step of performing time alignment on the target modal data according to the initial time node and the initial timestamp to obtain a data sequence includes:

[0014] Analyze the initial timestamp to obtain the sampling strategy for each target modal data;

[0015] If the sampling strategy is uniform sampling, allocate the target modal data to the initial time node adjacent to and before the initial timestamp to obtain the first alignment information;

[0016] If the sampling strategy is non-uniform sampling, allocate the target modal data to the initial time node closest to the initial timestamp to obtain the second alignment information;

[0017] Generate a data sequence according to the first alignment information and / or the second alignment information.

[0018] Optionally, the step of selecting a target time node corresponding to the data fusion request from the data sequence and fusing the target modal data in the target time node to obtain a data fusion result includes:

[0019] Parse the data fusion request to obtain the time range constraint of the target area;

[0020] Select the target time nodes that meet the conditions from the data sequence according to the time range constraint;

[0021] After preprocessing the target modal data in the target time node based on a preset fusion algorithm, perform multi-level fusion to generate a data fusion result.

[0022] Optionally, after the step of performing multi-level fusion to generate a data fusion result after preprocessing the target modal data in the target time node based on a preset fusion algorithm, it further includes:

[0023] Calculate the deviation value and deviation direction between the data fusion result and the physical constraints of each sensor operating parameter;

[0024] If the deviation value exceeds the dynamic adjustment threshold, adjust the parameters of the preset fusion algorithm according to the deviation direction and re-execute the fusion process.

[0025] Optionally, the step of mapping the multi-modal data collected by each sensor in the target area to the power plant three-dimensional twin model after receiving the data fusion request to obtain multiple fuzzy areas includes:

[0026] After receiving a data fusion request, according to the position parameters and acquisition ranges of each sensor, the collected multimodal data is uniformly mapped into the three-dimensional twin model of the power plant through a coordinate system conversion algorithm;

[0027] Based on the mapped three-dimensional twin model of the power plant, calculate the detection probability density distribution of each sensor in three-dimensional space, and superimpose the probability density distributions of the sensors to obtain a probability distribution model;

[0028] Based on the probability distribution model, evaluate the confidence level of the spatial overlapping regions of the three-dimensional twin model of the power plant, and mark the regions with a confidence level lower than the preset detection threshold as fuzzy regions.

[0029] In addition, to achieve the above object, the present application also proposes a multimodal data fusion device, which includes:

[0030] A region determination module, configured to map the multimodal data collected by each sensor in the target region into the three-dimensional twin model of the power plant after receiving a data fusion request, to obtain a plurality of fuzzy regions;

[0031] A data determination module, configured to, for each of the fuzzy regions, determine the target modal data of each of the fuzzy regions within a preset time period from the multimodal data;

[0032] A dependency determination module, configured to project the target modal data onto the corresponding time nodes to obtain a data sequence under a unified time scale;

[0033] A data fusion module, configured to select the target time nodes corresponding to the data fusion request from the data sequence according to the data fusion request, and fuse the target modal data in the target time nodes to obtain a data fusion result.

[0034] In addition, to achieve the above object, the present application also proposes a multimodal data fusion device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the multimodal data fusion method as described above.

[0035] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the multimodal data fusion method as described above.

[0036] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the multi-modal data fusion method described above.

[0037] In the present application, after receiving a data fusion request, multi-modal data collected by various sensors in a target area is mapped into a three-dimensional twin model of a power plant to obtain multiple fuzzy areas; for each of the fuzzy areas, target modal data of each of the fuzzy areas within a preset time period is determined from the multi-modal data; the target modal data is projected onto corresponding time nodes to obtain a data sequence under a unified time scale; according to the data fusion request, a target time node corresponding to the data fusion request is selected from the data sequence, and the target modal data in the target time node is fused to obtain a data fusion result. Through the space mapping mechanism of the three-dimensional twin model, the data collected by sensors at different installation positions is used to extract target modal data and projected onto a unified time scale, eliminating the problem of spatial misalignment of fuzzy areas caused by overlapping detection ranges, achieving time alignment of multi-sampling frequency data, significantly improving the accuracy and timeliness of equipment status evaluation, fault prediction, and optimization decision-making in complex industrial scenarios, and providing key technical support for the digital transformation of smart power plants. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0039] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a schematic flowchart of the first embodiment of the multi-modal data fusion method of the present application;

[0041] Figure 2 It is a schematic flowchart of the second embodiment of the multi-modal data fusion method of the present application;

[0042] Figure 3 It is a schematic flowchart of the third embodiment of the multi-modal data fusion method of the present application;

[0043] Figure 4 It is a schematic module structure diagram of the multi-modal data fusion device in the embodiment of the present application;

[0044] Figure 5This is a schematic diagram of the device structure of the hardware operating environment involved in the multi-modal data fusion method in the embodiments of this application.

[0045] The implementation, functional features, and advantages of this application will be further described in conjunction with the embodiments and the accompanying drawings. Specific implementation manners

[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0047] To better understand the technical solutions of this application, the following will be described in detail in conjunction with the drawings of the specification and specific implementation manners.

[0048] The main solution of the embodiments of this application is: after receiving a data fusion request, map the multi-modal data collected by each sensor in the target area to the three-dimensional twin model of the power plant to obtain multiple fuzzy areas; for each of the fuzzy areas, determine the target modal data of each of the fuzzy areas within a preset time period from the multi-modal data; project the target modal data onto the corresponding time nodes to obtain a data sequence under a unified time scale; select the target time node corresponding to the data fusion request from the data sequence according to the data fusion request, and fuse the target modal data in the target time node to obtain a data fusion result.

[0049] As the core equipment of industrial production and energy systems, the operation stability of boilers directly affects production efficiency and safety. During the boiler startup, pressure rise, and constant-pressure operation stages, the dynamic matching of parameters such as main steam pressure, bypass valve control, and combustion efficiency is the key to avoiding faults. However, the boiler system has strong coupling and non-linear characteristics. Traditional detection methods often lead to untimely fault identification or even chain reactions due to isolated parameter analysis or lag in dynamic response. Most existing technologies judge faults by simply comparing the main steam pressure with a fixed threshold, without considering associated parameters such as the tube bundle heating time and dynamic regulation of the bypass valve, and are prone to false alarms due to working condition fluctuations, such as misjudgment triggered by short-term pressure fluctuations; only static statistics are performed on parameters such as main steam flow and valve position deviation, lacking modeling of change trends, such as deviation change rate and load dynamic fluctuations, and unable to capture early fault signs; at the same time, fault mode recognition depends on offline clustering of historical data, without calculating the matching degree between the dynamic feature vector and the fault mode in real time, resulting in insufficient response speed for fault classification.

[0050] Therefore, in view of the problems in the multi-modal data fusion of power plants, such as low spatio-temporal alignment accuracy, insufficient detection of sensor coverage blind spots, and poor adaptability to dynamic scenarios, this application proposes a fusion method based on a three-dimensional twin model and a unified time scale. Traditional methods result in data fragmentation due to differences in the spatio-temporal benchmarks of sensors, and are unable to effectively identify low-confidence regions such as sensor blind spots. At the same time, they lack adaptability to the dynamic changes in the operating conditions of equipment, and are prone to generating fusion results that conflict with physical constraints. This solution solves the above technical bottlenecks through the spatial mapping, time-axis alignment, and fuzzy region positioning of multi-modal data, improving the integrity and reliability of the fusion results.

[0051] It should be noted that the execution entity of this embodiment can be a computing service device with data processing and program running functions, such as a fusion decision-making system, or an electronic device capable of implementing the above functions. The following takes the multi-modal data processing system of a power plant as an example to illustrate this embodiment and the following embodiments.

[0052] Based on this, the embodiment of this application provides a multi-modal data fusion method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the multi-modal data fusion method of this application.

[0053] In this embodiment, the multi-modal data fusion method includes:

[0054] Step S10, after receiving a data fusion request, map the multi-modal data collected by each sensor in the target area to the three-dimensional twin model of the power plant to obtain multiple fuzzy regions.

[0055] It should be noted that the data fusion request refers to an instruction initiated by an external system (such as a power plant monitoring platform) to perform fusion analysis on the multi-source sensor data of a specified area (such as a steam turbine, boiler, etc.) to support fault diagnosis or status assessment. The target area is the physical space range that needs to be monitored in the power plant. The multi-modal data is heterogeneous sensor data of different physical quantities (such as vibration, temperature, pressure), including temporal, spatial, and semantic differences. The three-dimensional twin model of the power plant is a digital mirror model of the physical entity of the power plant constructed based on the actual structure of the power plant. The fuzzy region is an area in this application where data conflicts and spatial misalignments occur due to the overlapping detection ranges of sensors.

[0056] It should be understood that after receiving the data fusion request, relevant sensors in the target area need to be screened, and the data needs to be normalized and outliers removed.

[0057] Furthermore, in order to quantify the spatial coverage ability through the superposition of the detection probability densities of sensors, accurately locate the low-confidence regions, and provide a priority guide for subsequent fusion. The step S10 may include:

[0058] After receiving a data fusion request, according to the position parameters and acquisition ranges of each sensor, the multi-modal data collected is uniformly mapped into the three-dimensional twin model of the power plant through a coordinate system transformation algorithm; based on the mapped three-dimensional twin model of the power plant, the detection probability density distribution of each sensor in three-dimensional space is calculated, and the probability density distributions of the sensors are superimposed to obtain a probability distribution model; based on the probability distribution model, the confidence level of the spatial overlapping area of the three-dimensional twin model of the power plant is evaluated, and the area with a confidence level lower than the preset detection threshold is marked as a fuzzy area.

[0059] It should be noted that the position parameters and acquisition ranges of each sensor may include the spatial coordinates, installation angles, and detection ranges of the sensors in the physical power plant. These data are input parameters of the coordinate system transformation algorithm and are used to accurately map the sensor data to the corresponding nodes of the three-dimensional model. The coordinate system transformation algorithm is an algorithm that transforms sensor data from a local coordinate system to a global three-dimensional model based on the spatial relationship between the sensor installation position and the three-dimensional model through a rotation translation matrix or an affine transformation. The detection probability density distribution can be calculated based on sensor accuracy, coverage range, and environmental interference and is used to quantify data confidence.

[0060] It should be understood that each sensor in the target area is deployed at different positions, and the position information of the target detected by the current sensor is usually a polar coordinate point centered on its own sensor and cannot be directly fused. Therefore, to implement the fusion function of multi-source heterogeneous data, it is necessary to perform coordinate transformation on the target points of each sensor to convert them into polar coordinate data relative to a unified reference point, and then perform data fusion processing. Therefore, the coordinate system transformation algorithm can unify the text information of each sensor.

[0061] It can be understood that based on the sensor position parameters (such as GPS coordinates, installation inclination angles), the sensor data can be mapped to the corresponding grid nodes of the three-dimensional model through an affine transformation matrix. For example, the two-dimensional temperature field of an infrared thermal imager needs to be converted into a three-dimensional thermal map on the surface of the device. For the probability density distribution of a single sensor, a Gaussian distribution or a uniform distribution model can be constructed according to the sensor detection range and accuracy to characterize its detection ability in three-dimensional space; for the overlapping area of multiple sensors, its probability distribution is obtained by superimposing the probability density distributions of each sensor through weighted average or probability product rules. For example, the joint detection area of a vibration sensor and a temperature sensor needs to consider weight distribution.

[0062] In one example, the sensor local coordinate system data is mapped to the global coordinate system of the three-dimensional twin model of the power plant through an affine transformation. The calculation formula for the unified mapping of each sensor is as follows:

[0063]

[0064] Among them, x global , y global , z global are the mapped coordinates in the global coordinate system respectively; x local , y local , z local are the measured coordinates in the local coordinate system of the sensor respectively; R is the rotation matrix, determined by the installation angle of the sensor; T is the translation vector, used to characterize the position of the sensor in the global coordinate system.

[0065] Based on the mapped coordinates in the global coordinate system, the detection probability density distribution can be modeled through the three-dimensional Gaussian distribution formula, and the formula is as follows:

[0066]

[0067] Among them, μ = (x global , y global , z global ), representing the position of the sensor detection center; v = (x, y, z), indicating the position of the target evaluation point in the target area. Σ is the covariance matrix, characterizing the detection range and accuracy of the sensor.

[0068] Next, through the detection probability density distribution of each target point, probability density superposition and confidence evaluation are carried out. Among them, the probability density of each point can be superposed based on the weighted superposition formula, as follows:

[0069]

[0070] In the above formula, w i represents the sensor weight, which is inversely proportional to the accuracy of the corresponding sensor; p total represents the total probability density after fusion. By comparing p total with the confidence threshold α, it can be determined whether the area pointed to by the target point is a fuzzy area. The confidence threshold needs to be adjusted by the historical data weight and taking the data within the nearest preset time period to improve the detection accuracy.

[0071] Step S20, for each of the fuzzy areas, determine the target modal data of each of the fuzzy areas within a preset time period from the multimodal data.

[0072] It can be understood that in the specific implementation process, a unique identifier can be assigned to each fuzzy area so that different fuzzy areas can be accurately distinguished in subsequent processing, and this identifier can be associated with the area information in the three-dimensional twin model. The specific range of each fuzzy area in three-dimensional space can be represented by the coordinate range.

[0073] It should be understood that the preset time period needs to be determined according to actual requirements, such as the urgency of decision-making requirements. For the data filtered out within the preset time period, the identifier of the corresponding sensor of the data can be parsed first, and then the identifier is corresponded with the identifier of the fuzzy area to filter out the data whose spatial position is within the range of the fuzzy area.

[0074] Step S30: Project the target modal data onto the corresponding time node to obtain a data sequence under a unified time scale.

[0075] It can be understood that to solve the problem of inconsistent data time caused by different sampling frequencies of different sensors and align multi-modal data in the time dimension, it is necessary to project the target modal data onto the corresponding time node to obtain a data sequence under a unified time scale.

[0076] Step S40: Select the target time node corresponding to the data fusion request from the data sequence according to the data fusion request, and fuse the target modal data in the target time node to obtain a data fusion result.

[0077] It should be understood that for different decision requirements, the corresponding generated data fusion requests are different, and the required data is also different. If the data fusion request is a historical request, the time nodes of the start time and the end time need to be extracted; if the data fusion request is a real-time request, the time node of the current time or the historical target time period needs to be obtained.

[0078] In this embodiment, after receiving a data fusion request, the multi-modal data collected by each sensor in the target area is mapped into the three-dimensional twin model of the power plant to obtain a plurality of fuzzy areas; for each of the fuzzy areas, the target modal data of each of the fuzzy areas within a preset time period is determined from the multi-modal data; the target modal data is projected onto the corresponding time node to obtain a data sequence under a unified time scale; the target time node corresponding to the data fusion request is selected from the data sequence according to the data fusion request, and the target modal data in the target time node is fused to obtain a data fusion result. Through the spatial mapping mechanism of the three-dimensional twin model, the target modal data of the data collected by sensors at different installation positions is extracted and projected onto a unified time scale, eliminating the problem of spatial misalignment of the fuzzy area caused by overlapping detection ranges, realizing the time alignment of multi-sampling frequency data, significantly improving the accuracy and timeliness of equipment status evaluation, fault prediction and optimization decision-making in complex industrial scenarios, and providing key technical support for the digital transformation of smart power plants.

[0079] Refer to Figure 2 , Figure 2This is a flowchart of the second embodiment of the multi-modal data fusion method of the present application. Based on the above first embodiment, the second embodiment of the multi-modal data fusion method of the present application is proposed.

[0080] In the second embodiment, step S30 includes:

[0081] Step S301, determining the initial timestamp and the maximum acquisition period of the target modal data.

[0082] It should be noted that the initial timestamp is the starting time mark of each data point in the target modal data, used to determine the absolute position of the data on the time axis. The maximum acquisition period is mainly to find the sensor with the maximum time period. Among them, if a newly connected sensor is available, it is necessary to judge the acquisition time period of the newly connected sensor. If it exceeds the current maximum acquisition period, the maximum acquisition period will be replaced; if the sensor with the maximum time period has been disconnected, it is necessary to re-find and calculate the maximum acquisition period of the currently connected sensors.

[0083] Step S302, based on the maximum acquisition period and the time resolution requirement of the data fusion request, dividing the continuous time axis into fixed time windows, and assigning corresponding initial time nodes to each of the time windows.

[0084] It should be noted that the time resolution represents the time accuracy required in the data fusion request, which determines the division accuracy of the time windows.

[0085] It can be understood that dividing the continuous time axis into fixed time windows can align the time nodes of different modal data. The initial time node can represent the identification information of each time window, and the starting time of each window can be set as the initial time node.

[0086] Step S303, performing time alignment on the target modal data according to the initial time node and the initial timestamp to obtain a data sequence.

[0087] It can be understood that for each initial timestamp, there are usually 2 adjacent initial time nodes. The distances between the initial timestamps obtained by different sampling methods and their adjacent initial time nodes are different. When performing time alignment, it is necessary to separately determine the initial time node corresponding to the initial timestamp according to different situations.

[0088] Furthermore, in order to allocate data to adjacent or the nearest time nodes according to the sampling strategy, by combining uniform sampling and non-uniform sampling, the continuity of the time series is improved, the interpolation error is reduced, and data distortion is avoided. Step S303 may include:

[0089] Analyze the initial timestamp to obtain the sampling strategy for each target modal data; if the sampling strategy is uniform sampling, allocate the target modal data to the initial time node adjacent to and before the initial timestamp to obtain the first alignment information; if the sampling strategy is non-uniform sampling, allocate the target modal data to the initial time node closest to the initial timestamp to obtain the second alignment information; generate a data sequence according to the first alignment information and / or the second alignment information.

[0090] It should be noted that sensors with uniform sampling collect data at fixed time intervals (such as once per second), and the data points are evenly distributed on the time axis; sensors with non-uniform sampling do not collect data at fixed time intervals (such as event-triggered), and the data points are unevenly distributed on the time axis.

[0091] For ease of understanding, taking the actual situation as an example without limiting the present application, in one example, the earliest timestamp is extracted from the target modal data as the alignment starting point t init , as follows:

[0092] t init = min(t1, t2, …, t n )

[0093] where t1, t2, …, t n represent the earliest timestamps of each target modal data. At the same time, the maximum acquisition period T max can be set according to system requirements or device characteristics.

[0094] When dividing the time window and allocating the initial time node, first set the window interval Δt according to the accuracy of the data fusion request, as follows:

[0095]

[0096] Then, when allocating the initial time node t k , take the start time of each window as the alignment benchmark, and its calculation formula is as follows:

[0097] t k = t init + k·Δt (k = 0, 1, …, N - 1)

[0098] If the maximum acquisition period is 12 hours and the time resolution requirement is 5 minutes, then the window interval Δt = 5 minutes and the total number of windows N = 144. The initial time nodes are "08:00, 08:05, 08:10, …, 19:55". After obtaining the initial time nodes on the time axis, the data can be allocated to the time nodes according to the sampling strategy to reduce the interpolation error.

[0099] If the data is uniformly sampled, the data is allocated to adjacent nodes earlier than the initial timestamp as follows:

[0100] t align = max{t k |t k ≤ t data}

[0101] where t data represents the initial timestamp of the data. t align represents the target data node to which the data is allocated. For example, if the initial timestamp of a certain temperature data is "08:03" and it is collected every 5 seconds at a fixed interval, then this temperature data needs to be allocated to the "08:00" node.

[0102] If the data is non-uniformly sampled, the data is allocated to the closest initial time node, and the corresponding calculation formula is as follows:

[0103]

[0104] For example, if the initial timestamp of a sudden vibration data is "08:08" and the closest nodes are "08:05" and "08:10", then this vibration data needs to be allocated to "08:10". In particular, for non-uniformly sampled data, when the distances to the front and back initial time nodes are equal, the former is preferred.

[0105] Finally, taking each initial time node as a set, the multi-modal data is respectively identified and stored.

[0106] In this embodiment, the initial timestamp and the maximum acquisition period of the target modal data are determined; based on the maximum acquisition period and the time resolution requirement of the data fusion request, the continuous time axis is divided into fixed time windows, and a corresponding initial time node is allocated to each time window; the target modal data is time-aligned according to the initial time node and the initial timestamp to obtain a data sequence. By increasing the maximum acquisition period and the time resolution requirement to divide fixed time windows and allocate initial time nodes, the problem of time axis fragmentation caused by different sampling periods of multiple sensors is solved, providing a unified benchmark for time alignment.

[0107] Refer to Figure 3 , Figure 3 which is a schematic flowchart of the third embodiment of the multi-modal data fusion method of the present application. Based on the above second embodiment, the third embodiment of the multi-modal data fusion method of the present application is proposed.

[0108] In the third embodiment, the step S40 includes:

[0109] Step S401: Analyze the data fusion request to obtain the time range constraint of the target area.

[0110] It should be noted that the time range constraint can be the time boundary conditions defined in the data fusion request, including the start time, end time, and time resolution, which are used to filter eligible data; it can also be the target time node inferred according to the user's decision-making requirements. Of course, it can also be based on the device operation log or operation and maintenance rules to derive the target time node that meets the data fusion request, especially in the case of device fault analysis.

[0111] Step S402: Select the target time nodes that meet the conditions from the data sequence according to the time range constraint.

[0112] It can be understood that in the specific implementation process, selecting the target time nodes that meet the conditions from the data sequence according to the time range constraint can be to select the time nodes within the time range constraint, or to select the time nodes within the time range constraint and some time nodes before and after the time range.

[0113] Step S403: After preprocessing the target modal data in the target time nodes based on a preset fusion algorithm, perform multi-level fusion to generate a data fusion result.

[0114] It should be understood that before data fusion, data preprocessing such as noise reduction and normalization needs to be performed to eliminate the dimensional difference.

[0115] It can be understood that multi-level fusion can include data-level fusion, feature-level fusion, and decision-level fusion. The fusion strategy can be a predefined fusion strategy, such as neural network, Bayesian network, D-S evidence theory, etc.

[0116] Furthermore, in order to dynamically adjust the fusion parameters by calculating the deviation value between the fusion result and the physical constraint. Ensure that the fusion result meets the actual operation constraints of the device and avoid misjudgment caused by data anomalies.

[0117] After the step S403, it may include:

[0118] Calculate the deviation value and deviation direction between the data fusion result and the physical constraints of each sensor operation parameter; if the deviation value exceeds the dynamic adjustment threshold, adjust the parameters of the preset fusion algorithm according to the deviation direction, and re-execute the fusion process.

[0119] It should be understood that physical constraints represent the reasonable value ranges and variation rules of the operating parameters of each sensor in the actual physical environment. For example, the temperature value measured by a temperature sensor should be within the temperature range that the device can withstand during normal operation, and the pressure value of a pressure sensor also has its corresponding safety range. Deviation value: The numerical difference between the data fusion result and the physical constraints of the sensor operating parameters, which is used to measure the degree to which the fusion result deviates from the normal physical range. The deviation direction indicates whether the data fusion result is higher or lower than the physical constraint range. The dynamic adjustment threshold is a preset deviation value limit. When the calculated deviation value exceeds this threshold, it indicates that there may be a large error in the fusion result, and the parameters of the preset fusion algorithm need to be adjusted.

[0120] It can be understood that for the positive deviation where the data fusion result is higher than the physical constraint, it is necessary to reduce the weight of the high-weight sensor or increase its noise covariance to weaken its influence on the fusion result; for the negative deviation where the data fusion result is lower than the physical constraint, it is necessary to increase the weight of the low-weight sensor or reduce its noise covariance to enhance its contribution.

[0121] It should be understood that after re-executing the fusion process with the adjusted parameters and obtaining the fusion result again, the above steps of calculating the deviation value and adjusting the parameters also need to be repeated until the deviation value is within the threshold range.

[0122] In this embodiment, the data fusion request is parsed to obtain the time range constraint of the target area; the target time nodes that meet the conditions are selected from the data sequence according to the time range constraint; after preprocessing the target modal data in the target time nodes based on a preset fusion algorithm, multi-level fusion is performed to generate a data fusion result. By improving the multi-level fusion after parsing the time range constraint to generate a preliminary result, the solution supports task-driven dynamic fusion and improves the flexibility of the fusion strategy.

[0123] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the multi-modal data fusion method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0124] This application also provides a multi-modal data fusion device. Please refer to Figure 4 , the multi-modal data fusion device includes:

[0125] An area determination module 10, which is used to map the multi-modal data collected by each sensor in the target area to the three-dimensional twin model of the power plant after receiving a data fusion request, so as to obtain multiple fuzzy areas;

[0126] The data determination module 20 is configured to determine, for each of the fuzzy regions, the target modal data of each of the fuzzy regions within a preset time period from the multimodal data;

[0127] The dependency determination module 30 is configured to project the target modal data onto corresponding time nodes to obtain a data sequence under a unified time scale;

[0128] The data fusion module 40 is configured to select, according to the data fusion request, the target time nodes corresponding to the data fusion request from the data sequence, and fuse the target modal data in the target time nodes to obtain a data fusion result.

[0129] The multimodal data fusion device provided by the present application adopts the multimodal data fusion method in the above embodiment, and can solve the technical problem that existing sensors cannot effectively process the data of fuzzy regions caused by overlapping detection ranges. Compared with the prior art, the beneficial effects of the multimodal data fusion device provided by the present application are the same as those of the multimodal data fusion method provided by the above embodiment, and other technical features in the multimodal data fusion device are the same as the features disclosed in the method of the above embodiment, and will not be elaborated herein.

[0130] The present application provides a multimodal data fusion device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the multimodal data fusion method in the first embodiment above.

[0131] Next, refer to Figure 5 , which shows a schematic structural diagram of a multimodal data fusion device suitable for implementing the embodiments of the present application. The multimodal data fusion device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The shown multimodal data fusion device is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0132] As Figure 5As shown, the multimodal data fusion device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the multimodal data fusion device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the multimodal data fusion device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a multimodal data fusion device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.

[0133] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0134] The multimodal data fusion device provided by the present application adopts the multimodal data fusion method in the above-mentioned embodiments, and can solve the technical problem that existing sensors cannot effectively process the data in the fuzzy area due to the overlapping detection ranges. Compared with the prior art, the beneficial effects of the multimodal data fusion device provided by the present application are the same as those of the multimodal data fusion method provided by the above-mentioned embodiments, and the other technical features in the multimodal data fusion device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0135] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0136] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

[0137] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the multi-modal data fusion method in the above embodiments.

[0138] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0139] The above computer-readable storage medium can be included in the multi-modal data fusion device; or it can exist independently and not be assembled into the multi-modal data fusion device.

[0140] The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the multi-modal data fusion device, the multi-modal data fusion device is caused to execute the multi-modal data fusion method described above.

[0141] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0143] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0144] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned multi-modal data fusion method, and can solve the technical problem that existing sensors cannot effectively process data in the fuzzy area due to overlapping detection ranges. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the multi-modal data fusion method provided by the above embodiments, and will not be elaborated here.

[0145] The present application also provides a computer program product, including a computer program, which implements the steps of the multimodal data fusion method as described above when executed by a processor.

[0146] The computer program product provided by the present application can solve the technical problem that existing sensors cannot effectively process data in fuzzy areas caused by overlapping detection ranges. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the multimodal data fusion method provided by the above embodiments, and will not be elaborated here.

[0147] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A multimodal data fusion method, characterized in that: The multimodal data fusion method includes: After receiving the data fusion request, the multimodal data collected by each sensor in the target area is mapped to the 3D twin model of the power plant to obtain multiple fuzzy areas; For each of the fuzzy areas, determining target modal data of each of the fuzzy areas within a preset time period from the multimodal data; Projecting the target modal data onto corresponding time nodes to obtain a data sequence under a unified time scale; According to the data fusion request, a target time node corresponding to the data fusion request is selected from the data sequence, and the target modality data in the target time node is fused to obtain a data fusion result.

2. The multimodal data fusion method according to claim 1, wherein: The step of projecting the target modal data onto corresponding time nodes to obtain a data sequence under a unified time scale includes: Determining an initial timestamp and a maximum acquisition period of the target modal data; Based on the maximum acquisition period and the time resolution requirement of the data fusion request, the continuous time axis is divided into fixed time windows, and a corresponding initial time node is assigned to each of the time windows; The target modality data is time-aligned according to the initial time node and the initial timestamp to obtain a data sequence.

3. The multimodal data fusion method according to claim 2, wherein: The step of performing time alignment on the target modality data according to the initial time node and the initial timestamp to obtain a data sequence includes: Analyzing the initial timestamp to obtain a sampling strategy for each target modal data; If the sampling strategy is uniform sampling, the target modal data is allocated to an initial time node adjacent to the initial timestamp and before the initial timestamp to obtain first alignment information; If the sampling strategy is non-uniform sampling, the target modal data is allocated to the initial time node closest to the initial timestamp to obtain second alignment information; A data sequence is generated according to the first alignment information and / or the second alignment information.

4. The multimodal data fusion method according to claim 1, wherein: The step of selecting a target time node corresponding to the data fusion request from the data sequence according to the data fusion request, and fusing the target modality data in the target time node to obtain a data fusion result includes: Parsing the data fusion request to obtain a time range constraint for a target area; Selecting a target time node that meets the conditions from the data sequence according to the time range constraint; After preprocessing the target modality data in the target time node based on a preset fusion algorithm, multi-level fusion is performed to generate a data fusion result.

5. The multimodal data fusion method according to claim 4, wherein: After pre-processing the target modality data at the target time node based on a preset fusion algorithm and performing multi-level fusion to generate a data fusion result, the method further includes: Calculating a deviation value and a deviation direction between the data fusion result and the physical constraints of the operating parameters of each sensor; If the deviation value exceeds the dynamic adjustment threshold, the parameters of the preset fusion algorithm are adjusted according to the deviation direction, and the fusion process is re-executed.

6. The multimodal data fusion method according to any one of claims 1 to 5, characterized in that: After receiving the data fusion request, the multimodal data collected by each sensor in the target area is mapped to the three-dimensional twin model of the power plant to obtain multiple fuzzy areas, including: After receiving the data fusion request, the collected multimodal data is uniformly mapped to the power plant 3D twin model through the coordinate system conversion algorithm according to the location parameters and acquisition range of each sensor; Calculating the detection probability density distribution of each sensor in the three-dimensional space based on the mapped three-dimensional twin model of the power plant, and superimposing the probability density distributions of the sensors to obtain a probability distribution model; A confidence evaluation is performed on the spatially overlapping areas of the three-dimensional twin model of the power plant based on the probability distribution model, and areas where the confidence is lower than a preset detection threshold are marked as fuzzy areas.

7. A multimodal data fusion device, characterized in that: The device comprises: The region determination module is used to map the multimodal data collected by each sensor in the target area to the three-dimensional twin model of the power plant after receiving the data fusion request, and obtain multiple fuzzy regions; a data determination module, configured to determine, for each of the fuzzy areas, target modal data of the fuzzy area within a preset time period from the multimodal data; A dependency determination module, configured to project the target modal data onto corresponding time nodes to obtain a data sequence under a unified time scale; The data fusion module is used to select a target time node corresponding to the data fusion request from the data sequence according to the data fusion request, and fuse the target modality data in the target time node to obtain a data fusion result.

8. A multimodal data fusion device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the multimodal data fusion method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the multimodal data fusion method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the multimodal data fusion method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Microseismic positioning and sensor layout optimization method and device based on probability density distribution field

    CN119355804A

  • Sensor integration system and sensor integration method

    JP2012163495A

Cited By

  • Children psychological monitoring system based on multi-modal space-time alignment

    CN120636823A

  • Multi-source data fusion method and device for gas turbine test

    CN121959468A

  • Method and apparatus for multi-source data fusion in gas turbine testing

    CN121959468B