Multi-modal data fusion, fault diagnosis and alarm method, medium and equipment

By generating a unified structured fusion record at the end of the time window and using missing state markers in multimodal data fusion, the problem of inconsistent fusion results caused by missing modal data is solved, and the stability and simplification of the system are achieved.

CN121808701APending Publication Date: 2026-04-07WISDRI WUHAN AUTOMATION
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing multimodal data fusion technologies struggle to generate structurally complete and consistent fusion results when modal acquisition frequencies are inconsistent or some modal data is missing, leading to increased system complexity and decreased processing stability.

Method used

By using the end of the time window as the sole trigger condition for generating fusion records, a fusion record with a uniform structure is generated, and missing data is filled with missing status markers, ensuring that a fusion record with a complete structure is generated for each time window.

Benefits of technology

It achieves absolute continuity and structural consistency of fused records on the timeline, reduces system development complexity and maintenance costs, and improves processing stability and data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808701A_ABST
    Figure CN121808701A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal data fusion method, a multi-modal data fault diagnosis method, a multi-modal data alarm method, a medium and equipment, and relates to the technical field of multi-modal data processing and fusion. The method comprises the following steps: acquiring original data acquired by collecting a monitored object, and processing the original data to generate corresponding modal features; dividing a monitoring time axis of the original data into a plurality of continuous time windows; for each time window, sequentially judging whether each data mode of the mode set has corresponding feature data or not, and if not, generating a missing state mark corresponding to the data mode; automatically triggering to generate fusion records in one-to-one correspondence with the time windows according to the ending time of the time windows, wherein the fusion records have a unified data structure; and outputting the generated fusion record for subsequent data processing, storage or calling. According to the method, the window-level output continuity can still be ensured under the condition of mode missing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data processing and fusion technology, and in particular to a multimodal data fusion, fault diagnosis, and alarm method, medium, and device. Background Technology

[0002] In existing multimodal data processing systems, common data sources include image data, sensor time-series data, and structured record data generated by information systems. These different modalities typically have different acquisition frequencies, time granularities, and generation methods. Therefore, when fusing multimodal data, it is usually necessary to perform time correlation or alignment processing on data from different modalities to reflect the state information of the same monitored object at different times.

[0003] In existing technologies, the following processing methods are typically used to achieve the fusion of multimodal data:

[0004] 1. Fusion method based on strict time alignment:

[0005] In this implementation, the system uses a unified time point or timestamp as the fusion benchmark, and generates the corresponding fusion result only when multiple modalities have data at the same or similar time points.

[0006] This method can achieve synchronization of multimodal data in the time dimension under ideal conditions. However, when there are large differences in the acquisition frequency of different modalities, or when a certain modality does not generate effective data in some time periods, the following problems may occur: strict alignment conditions cannot be met in some time periods, resulting in missing fusion results; data already acquired in other modalities is discarded to ensure time alignment, reducing data utilization; and it is difficult to form a continuous and stable sequence of fusion results.

[0007] 2. Methods for generating fusion results based on existing data:

[0008] In this type of existing technology, the system does not require that data exist for different modalities at the same point in time. Instead, it dynamically generates fusion results based on the modal data that actually exist in each time period.

[0009] When a certain modality is missing within a corresponding time period, the system only generates fusion results for modalities with available data, or skips that time period directly.

[0010] While this approach improves data utilization to some extent, it still has significant shortcomings in practical applications: the number and types of fields generated at different time periods may differ; the data structure of the fusion results changes depending on the presence of modal data; and subsequent processing modules need to perform additional judgments and adaptations for fusion results with different structures, increasing the complexity of system implementation.

[0011] 3. Data aggregation method based on time windows:

[0012] This type of existing technology uses time-window-based data aggregation or statistical methods to process multi-source data within a certain time range.

[0013] In this type of approach, time windows are typically used as tools to limit the time range of data or for statistical aggregation. Their main function is to accommodate data sampling results from different frequencies, rather than to control the generation of the fusion result. When data for a certain modality is missing within a given time window, existing technologies usually adopt one of the following approaches: not generating a fusion result for the corresponding time window; or generating a fusion result only for the modality where data exists. Therefore, this type of approach also struggles to ensure that the data generated under different time windows maintains structural consistency.

[0014] Based on the above-mentioned existing technical solutions, when the multimodal data acquisition frequency is inconsistent and some modal data is missing within certain time ranges, the existing technologies have at least the following technical defects:

[0015] 1) The generation of fusion results depends on the existence of modal data. When some modalities are missing, the time window may be skipped or the fusion results may be missing, making it difficult to form a continuous time series output.

[0016] 2) In the presence of modal missing data, the data structure of the fusion results changes over time, resulting in inconsistencies in the field level of the fusion results output at different time periods;

[0017] 3) Due to the inconsistent structure of the fusion results, the subsequent processing module needs to perform additional structure judgment and adaptation logic for the data structure of different time periods, which increases the system complexity and affects the processing stability.

[0018] 4) The time window in the existing technology is mainly used for time alignment or data aggregation, but it fails to constrain the generation behavior of the fusion result at the system mechanism level, making it difficult to fundamentally solve the problem of structural inconsistency. Summary of the Invention

[0019] The purpose of this invention is to provide a method, medium, and device for multimodal data fusion, fault diagnosis, and alarm, aiming to generate a structurally complete and consistent fusion record at the end of each time window even when the multimodal data acquisition frequency is inconsistent and some modalities lack valid data within certain time ranges. The specific technical solution is as follows:

[0020] A multimodal data fusion method, the method comprising the following steps:

[0021] S10. Determine the modality set corresponding to the monitored object, wherein the modality set includes at least two different data modalities;

[0022] S20. Obtain raw data collected from the monitored object, wherein the raw data includes monitoring data corresponding to at least a portion of the data modes in the modality set, and process the raw data to generate corresponding modal features;

[0023] S30. Divide the monitoring timeline of the raw data into multiple consecutive time windows;

[0024] S40. For each time window, sequentially determine whether each data modality in the modality set has corresponding feature data. If it does, generate an associated feature corresponding to the data modality based on the feature data. If it does not, generate a missing status marker corresponding to the data modality. The fusion record corresponding to each time window is automatically triggered by the end time of the time window. The fusion record has a unified data structure, which includes an identification field of the monitored object, an identification field of the time window, and a modality data field corresponding to each modality in the modality set. The modality data field is the associated feature or the missing status marker. Output the generated fusion record for subsequent data processing, storage, or retrieval.

[0025] Furthermore, in step S10, the mode set is represented as follows:

[0026]

[0027] in, It is a pre-determined set of all data modalities corresponding to the same monitoring object. Each modal data is associated with the same monitoring object and carries a corresponding collection timestamp. Let i be the i-th data mode, 1≤i≤n, where n is the preset total number of modes.

[0028] Further, in step S20, each modal data is processed separately to generate corresponding modal features, wherein the modal feature generated for data modality Mᵢ at timestamp t is represented as follows: The modal features It includes the following information: timestamp information t used to identify the collection time, object identification information obj_id used to identify the monitored object, and feature information used to characterize the state or content of the modal data.

[0029] Furthermore, in step S30, the method for dividing the monitoring time axis of the original data into multiple consecutive time window sets is either a fixed step size division rule or a clock alignment division rule; the fixed step size division rule is: a constant time step size ΔT is preset, and each time window... The length of each is equal to ΔT; the clock alignment division rule is: alignment division is based on the natural clock.

[0030] Furthermore, in step S40, for each time window W... k From the modal feature set corresponding to each modality, select modal features whose timestamps satisfy the following condition: t∈W k For each mode M in the mode set M i In time window W k The following processing logic is executed within the time window W: If the modality is within the time window W k The memory contains at least one candidate modal feature F i (t), then based on at least one candidate modality feature, generate the associated feature corresponding to the data modality; if the modality is within the time window W k If no modal feature exists within the data modality, then a corresponding missing state label δ is generated for that data modality. i (W k Thus forming a time window W; k Corresponding cross-modal correlation unit:

[0031]

[0032] in, In time window W k Within, a set of structured data is formed by processing and associating the features of each data mode in the modality set M; For the i-th data mode in the modality set within the time window W k The corresponding associated features or placeholder status within.

[0033] Furthermore, if data mode M i In time window W k The memory contains multiple candidate modal features F i (t), then a modal feature is determined according to any of the following methods to generate the associated feature corresponding to the data modality: last valid value rule, statistical aggregation rule, time-weighted average rule, feature extreme value rule; the last valid value rule is: filtering within the time window W k The time stamp t collected within the window is closest to the end time of the window. Modal characteristics; statistical aggregation rule is: for time window W k All modal features F i (t) Perform mathematical statistical calculations, taking the arithmetic mean, median, or mode; the time-weighted average rule is: based on the modal characteristics F... i (t) Distance of sampling point t from the end of the window The duration is assigned different weight coefficients wt, and the weighted average is calculated using the following formula: The eigenvalue extremum rule is to select the maximum or minimum value from the candidate set.

[0034] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the multimodal data fusion method as described above.

[0035] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the multimodal data fusion method as described above.

[0036] The present invention also provides a method for diagnosing equipment faults, the method comprising the following steps:

[0037] S100. Acquisition and parsing of fusion records: Acquire the fusion records output by the multimodal data fusion method described above from the message queue or database;

[0038] S200, Tensor Transformation: Traverse the modal data fields in the fusion record. For each data modality, if the modal data field is an associated feature, write the associated feature value into the corresponding index position of the input tensor; if the modal data field is a missing state marker, fill the missing state marker with a preset placeholder.

[0039] S300, Mask Generation: Synchronously generate a mask vector to mark the positions in the modal data fields that are marked as missing states as invalid;

[0040] S400, Model Inference and Output: Input the aligned tensor into the diagnostic neural network model to obtain the diagnostic results.

[0041] The present invention also provides an alarm method, the method comprising the following steps:

[0042] S1000, Rule Parsing and Mapping: The alarm engine loads preset rules and maps the physical quantities involved in the preset rules to fixed field indices in the fusion record output by the multimodal data fusion method as described above;

[0043] S2000, Fast Status Retrieval and Deterministic Logic Execution: The engine reads the modal data field. If the modal data field is a missing status marker, the incomplete data warning logic is triggered directly without entering the numerical comparison process. If the modal data field is a related feature, the corresponding related feature is extracted and compared with a preset threshold.

[0044] S3000, Logical Combination Judgment: Based on the judgment result of the data mode, output the final alarm signal.

[0045] The present invention provides a multimodal data fusion, fault diagnosis, and alarm method, medium, and device, which have the following beneficial effects:

[0046] This invention uses the end time of the time window as the sole trigger condition for generating fused records, unlike traditional logic that triggers after data alignment. This ensures that the system forces output at the time window boundary regardless of whether modal data arrives within the window, guaranteeing absolute continuity of the fused records on the timeline and eliminating window skipping or sequence interruptions caused by missing data. By pre-defining an immutable modal set, the system iterates through all data modalities in the set to fill in gaps during fused record generation. Even if a modal data is missing, the system will still generate a fixed-position field for it according to a unified structure, ensuring structural consistency of records output from different time windows. This allows downstream modules, such as AI inference engines, to directly read data with a fixed offset without complex dynamic schema parsing. Furthermore, by introducing a dedicated missing status marker, data existence is stored as an explicit modal attribute, enabling semantic traceability. Downstream logic can intuitively distinguish between device inactivity and sensor failure / network packet loss, and the processing logic does not require writing complex exception handling code, achieving deep decoupling between business logic and the existence of underlying data. In summary, this invention is not a simple data alignment, but an engineering solution that achieves system determinism through mechanism innovation. It greatly reduces the development complexity and maintenance cost of multimodal monitoring systems by sacrificing minimal storage bandwidth. Attached Figure Description

[0047] Figure 1 A flowchart illustrating a multimodal data fusion method provided in an embodiment of the present invention;

[0048] Figure 2 This is a structural block diagram of a computer device according to an embodiment of the present invention;

[0049] Figure 3 This is a flowchart illustrating a device fault diagnosis method provided in an embodiment of the present invention.

[0050] Figure 4 This is a flowchart illustrating an alarm method provided in an embodiment of the present invention. Detailed Implementation

[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. The advantages and features of the present invention will become clearer from the following description. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clearly illustrate the purpose of the embodiments of the present invention.

[0052] Terminology Explanation:

[0053] 1. Multimodal Data

[0054] This refers to data types that differ in origin, acquisition method, or data representation, including but not limited to image data, sensor time-series data, log data, or structured record data. In this invention, multimodal data corresponds to different data modes, and each modality of data is associated with the same monitoring object.

[0055] 2. Modality

[0056] This refers to a single data type with independent data sources and acquisition characteristics in a multimodal data processing system. In this invention, multiple modalities are determined and constitute a modality set before the system runs, and each modality occupies a fixed field position in the fused record.

[0057] 3. Modal set (M)

[0058] This refers to the set of all modalities corresponding to the same monitored object, which is predetermined before the system starts operating. The modality set is used to constrain the field structure of the fused records and remains unchanged during system operation to ensure consistency in the number and position of fields in the fused records generated in different time windows.

[0059] 4. Modal characteristics

[0060] Modal features refer to the feature representations generated from raw modal data after processing, used for participating in cross-modal association. Modal features include at least acquisition timestamp information and monitoring object identification information, and may include feature values ​​used to characterize the state or content of the modal data.

[0061] 5. Time window

[0062] This refers to a continuous time range divided on the time axis according to preset rules. In this invention, the time window is not only used to limit the associated range of modal features, but also serves as a control unit for the generation of fused records; the generation of the corresponding fused record is triggered when each time window ends.

[0063] 6. Triggered when window ends

[0064] This refers to a mechanism that automatically initiates the fusion record generation process corresponding to a certain time window after the system detects that the time range corresponding to that time window has ended. This triggering mechanism does not depend on the existence of modal features.

[0065] 7. Cross-modal correlation unit

[0066] This refers to the data set formed after filtering and associating the modal features of each modality in the modality set within the same time window. The associated unit corresponds one-to-one with the modality set in structure, and even if a certain modality does not have modal features within the window, the corresponding position is still reserved for that modality.

[0067] 8. Missing status marker

[0068] This refers to a status flag used to explicitly indicate whether a certain modality has valid modal features within a corresponding time window. When a modality does not have modal features within a time window, a corresponding missing status flag is generated to indicate that the modality is in a missing state.

[0069] 9. Unified data structure

[0070] This refers to the pre-determined field structure of the fusion record before system operation, used to constrain the generation process of the fusion record. The unified data structure includes at least an object identifier field, a time window identifier field, a modal feature field group corresponding one-to-one with the modality set, and a corresponding missing state field group.

[0071] 10. Fusion Records

[0072] This refers to the data record generated by the system at the end of each time window, based on the cross-modal association units and missing state markers within that time window. The fused records generated across different time windows maintain a consistent data structure.

[0073] 11. Placeholder generation

[0074] This refers to a generation method in which, even if a certain modality does not have valid modal features within the corresponding time window, the system still generates corresponding fields for that modality according to a unified data structure and represents them through missing status markers.

[0075] 12. Structural consistency

[0076] This refers to the consistency of the number of fields, field order, and field semantics in fused records generated in different time windows, which does not change due to missing modal data or differences in collection frequency.

[0077] 13. Post-processing module

[0078] This refers to a system module that further processes, analyzes, stores, or retrieves fused records. Its processing logic is based on the unified data structure design of fused records, eliminating the need for structure adaptation or branch judgment for different time windows.

[0079] Example 1

[0080] This embodiment provides a multimodal data fusion method, see reference. Figure 1 As shown, the method includes the following steps:

[0081] S10. Determine the modality set corresponding to the monitored object, wherein the modality set includes at least two different data modalities.

[0082] Specifically, the modality set is represented as follows:

[0083]

[0084] in, It is a pre-determined set of all data modalities corresponding to the same monitoring object. Each modal data is associated with the same monitoring object and carries a corresponding collection timestamp. Let i be the i-th data mode, 1≤i≤n, where n is the preset total number of modes.

[0085] S20. Obtain the raw data collected from the monitored object, wherein the raw data includes monitoring data corresponding to at least some data modes in the modality set, and process the raw data to generate corresponding modal features.

[0086] Specifically, each modal data is processed separately to generate corresponding modal features, among which data modalities... The modal feature generated at timestamp t is represented as: The modal features It includes the following information: timestamp information t used to identify the collection time, object identification information obj_id used to identify the monitored object, and feature information used to characterize the state or content of the modal data.

[0087] Through the above processing, data from different modalities are all associated with time window level in a unified form of modal features in subsequent processing.

[0088] S30. Divide the monitoring timeline of the raw data into multiple consecutive time windows.

[0089] Specifically, the monitoring timeline of the raw data is divided into multiple consecutive time window sets, represented as follows: k is the sequence index number of the time window, k = 1, 2, ...; each time window Corresponding to a continuous time range: , For time window The start time of the corresponding continuous time range For time window The end time of the corresponding continuous time range.

[0090] In an optional embodiment, the method of dividing the monitoring time axis of the raw data into multiple consecutive time window sets is either a fixed step size division rule or a clock alignment division rule.

[0091] Specifically, the fixed step size division rule is as follows: a constant time step ΔT (e.g., 100ms, 1s, or 1min) is preset, and each time window... The length of each is equal to ΔT.

[0092] Start time calculation:

[0093] End time calculation:

[0094] Where T_base is the base time after system startup, and k is a sequence of positive integers.

[0095] Technical effect: This rule ensures that the timeline is cut evenly and that the windows are connected end to end without gaps.

[0096] Specifically, the clock alignment division rule is: alignment division is based on the natural clock.

[0097] For example, if ΔT is set to 1 minute, the window's switching point will be forced to align to 00 seconds per minute in natural time.

[0098] Technical effect: This rule can ensure that the generated fusion record R k It has a natural periodicity in terms of time semantics, which makes it easy for downstream modules to perform "hourly" or "daily" big data statistics and facilitates subsequent cross-system data comparison.

[0099] S40. For each time window, sequentially determine whether each data modality in the modality set has corresponding feature data. If it does, generate an associated feature corresponding to the data modality based on the feature data. If it does not, generate a missing status marker corresponding to the data modality. The fusion record corresponding to each time window is automatically triggered by the end time of the time window. The fusion record has a unified data structure, which includes an identification field of the monitored object, an identification field of the time window, and a modality data field corresponding to each modality in the modality set. The modality data field is the associated feature or the missing status marker. Output the generated fusion record for subsequent data processing, storage, or retrieval.

[0100] In one embodiment, for each time window W k From the modal feature set corresponding to each modality, select modal features whose timestamps satisfy the following condition: t∈W k For each mode M in the mode set Mi In time window W k The following processing logic is executed within the time window W: If the modality is within the time window W k The memory contains at least one candidate modal feature F i (t), then based on at least one candidate modality feature, generate the associated feature corresponding to the data modality; if the modality is within the time window W k If no modal feature exists within the data modality, then a corresponding missing state label δ is generated for that data modality. i (W k ), where δ i (W k )=1 indicates mode M i In time window W k The internal state is missing.

[0101] After the above processing, a time window W is formed. k Corresponding cross-modal correlation unit:

[0102]

[0103] in, In time window W k Within, a set of structured data is formed by processing and associating the features of each data mode in the modality set M; For the i-th data mode in the modality set within the time window W k The corresponding associated features or placeholder status within.

[0104] The specific value depends on the data modality M i The presence of original features and their corresponding missing status markers within this time window:

[0105] When valid data exists: if mode M i In time window W k The memory contains at least one candidate modal feature F i (t) (that is, satisfying t∈W) k If the missing state marker δ is missing, then... i (W k F = 0. i (W k ) is from one or more F i The associated feature values ​​generated by filtering or aggregating in (t) are used.

[0106] When data is missing: If data mode M i In time window W k There is no modal feature F that satisfies the condition. i If (t), the system triggers the missing status handling mechanism and generates a missing status marker δ.i (W k F = 1. At this time, F i (W k This is represented by placeholders generated according to a unified data structure, used to explicitly indicate that the modality is missing within the window, thereby ensuring that U... k The length and structure of the vector are constant.

[0107] In one embodiment, if data mode M i In time window W k The memory contains multiple candidate modal features F i If (t), then a modal feature is determined according to any of the following methods to generate the associated feature corresponding to the data modality: last valid value rule, statistical aggregation rule, time weighted average rule, feature extreme value rule.

[0108] Specifically, the rule for the last valid value is: filter within the time window W. k The time stamp t collected within the window is closest to the end time of the window. Modal characteristics.

[0109] Applicable scenarios: Suitable for industrial condition monitoring with high real-time requirements, to reflect the latest status of the monitored object at the end of the window.

[0110] Specifically, the statistical aggregation rule is as follows: for time window W k All modal features F i (t) Perform mathematical statistical calculations, and calculate the arithmetic mean, median, or mode.

[0111] Applicable scenarios: Suitable for high-frequency sampled sensor data, reducing the impact of random noise on the fused record through aggregation processing.

[0112] Specifically, the time-weighted average rule is as follows: based on the modal features F i (t) Distance of sampling point t from the end of the window The duration is assigned different weight coefficients wt, and the weighted average is calculated using the following formula: .

[0113] Specifically, the eigenvalue extremum rule is to select the maximum or minimum value from the candidate set.

[0114] Applicable scenarios: Suitable for safety warning scenarios, such as recording the peak values ​​of pressure or temperature sensors within a window to avoid masking abnormal peaks due to averaging.

[0115] Regardless of which selection rule is used, the system will trigger the fusion logic once and only once at the end of each time window. Even if the rule logic is complex, its output F i (Wk In fusion records The position, type, and field width of each field are all subject to the mandatory constraints of a unified data structure, thereby ensuring structural consistency across windows.

[0116] In this embodiment of the invention, the generation of fused records is not contingent on the existence of modal features. Even if a certain modality does not have valid modal features within the current time window, the system still generates corresponding fields for that modality and uses missing status fields as placeholders. This method ensures that each time window generates a fused record with complete fields, thereby maintaining the consistency of the fused record structure across different time windows.

[0117] In this embodiment of the invention, since all fused records are generated based on the same unified data structure, subsequent modules do not need to perform structure adaptation or branch judgment for different time windows or different modal missing cases when processing fused records, thereby improving the stability and consistency of the system processing flow.

[0118] This invention uses the end time of the time window as the sole trigger condition for generating fused records, unlike the traditional logic that triggers after data alignment. This ensures that the system forces output at the time window boundary regardless of whether modal data arrives within the window, guaranteeing the absolute continuity of the fused records on the timeline and eliminating window skipping or sequence interruptions caused by missing data. By pre-defining an immutable modal set, the system iterates through all data modalities in the set to fill in gaps when generating fused records. Even if a modal data is missing, the system will still generate a fixed-position field for it according to a unified structure. This ensures structural consistency of records output from different time windows, allowing downstream modules such as AI inference engines to directly read data with a fixed offset without complex dynamic schema parsing. By introducing a dedicated missing status marker, data existence is stored as an explicit modal attribute, enabling semantic traceability. Downstream logic can intuitively distinguish between device inactivity and sensor failure / network packet loss, and the processing logic does not require writing complex exception handling code, achieving deep decoupling between business logic and the existence of underlying data.

[0119] Example 2

[0120] This embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the multimodal data fusion method described above.

[0121] The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium may also include combinations of the above types of memory.

[0122] Example 3

[0123] This embodiment provides a computer device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the multimodal data fusion method described above.

[0124] like Figure 2 As shown, the computer device 70 may include: at least one processor 71, such as a CPU (Central Processing Unit), at least one communication interface 73, a memory 74, and at least one communication bus 72. The communication bus 72 is used to enable communication between these components. The communication interface 73 may include a display screen and a keyboard; optionally, the communication interface 73 may also include a standard wired interface or a wireless interface. The memory 74 may be high-speed RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory 74 may also be at least one storage device located remotely from the aforementioned processor 71. The memory 74 stores application programs, and the processor 71 calls the program code stored in the memory 74 to execute any of the above-described method steps.

[0125] The communication bus 72 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 72 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0126] The memory 74 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 74 may also include a combination of the above types of memory.

[0127] The processor 71 can be a central processing unit (CPU), a network processor (NP), or a combination of CPU and NP.

[0128] The processor 71 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0129] Optionally, the memory 74 is also used to store program instructions. The processor 71 can invoke the program instructions to implement the multimodal data fusion method of the present invention.

[0130] Example 4

[0131] This embodiment provides a method for diagnosing equipment faults. (See attached document.) Figure 3 As shown, the steps are as follows:

[0132] S100. Acquisition and parsing of fusion records: Acquire the fusion records as described above from the message queue or database; since the number and position of fields are fixed, the module directly extracts the modal features through the preset fixed-length offset.

[0133] S200, Tensor Transformation: Traverse the modal data fields in the fusion record. For each data modality, if the modal data field is an associated feature, write the associated feature value into the corresponding index position of the input tensor; if the modal data field is a missing state marker, fill the missing state marker with a preset placeholder (such as 0 or mean).

[0134] S300, Mask Generation: Synchronously generate a mask vector to mark the positions where the modality data field is marked as missing as invalid, so as to inform the diagnostic neural network model to ignore the noise effect of the modality during calculation.

[0135] S400, Model Inference and Output: Input the aligned tensor into the diagnostic neural network model to obtain the diagnostic results.

[0136] The embodiments of the present invention completely solve the common problem of dynamic tensor alignment in multimodal deep learning, and ensure the computational stability of the inference process.

[0137] Example 5

[0138] This embodiment provides an alarm method, see below. Figure 4 As shown, it includes the following steps:

[0139] S1000, Rule Parsing and Mapping: The alarm engine loads preset rules and maps the physical quantities (such as temperature and vibration) involved in the preset rules to the fixed field indexes in the fusion record as described above.

[0140] S2000, Fast Status Retrieval and Deterministic Logic Execution: The engine reads the modal data field. If the modal data field is a missing status marker, the incomplete data warning logic is triggered directly without entering the numerical comparison process. If the modal data field is a related feature, the corresponding related feature is extracted and compared with a preset threshold.

[0141] S3000, Logical Combination Judgment: Based on the judgment result of the data mode, output the final alarm signal.

[0142] In large-scale concurrent monitoring scenarios, this invention significantly reduces CPU branch prediction overhead and improves the response speed of high-frequency alarm requests by using logical debranching.

[0143] In summary, this invention is not a simple data alignment, but an engineering solution that achieves system determinism through mechanism innovation. It greatly reduces the development complexity and maintenance cost of multimodal monitoring systems by sacrificing minimal storage bandwidth (placeholders).

[0144] Those skilled in the art should understand that the present invention can be implemented in many other specific forms without departing from the spirit and scope of the invention. Any changes or modifications made by those skilled in the art based on the embodiments of the present invention and the above disclosure shall fall within the protection scope of the claims.

Claims

1. A multimodal data fusion method, characterized in that, The method includes the following steps: S10. Determine the modality set corresponding to the monitored object, wherein the modality set includes at least two different data modalities; S20. Obtain raw data collected from the monitored object, wherein the raw data includes monitoring data corresponding to at least a portion of the data modes in the modality set, and process the raw data to generate corresponding modal features; S30. Divide the monitoring timeline of the raw data into multiple consecutive time windows; S40. For each time window, sequentially determine whether each data modality in the modality set has corresponding feature data. If it does, generate an associated feature corresponding to the data modality based on the feature data. If it does not, generate a missing status marker corresponding to the data modality. The fusion record corresponding to each time window is automatically triggered by the end time of the time window. The fusion record has a unified data structure, which includes an identification field of the monitored object, an identification field of the time window, and a modality data field corresponding to each modality in the modality set. The modality data field is the associated feature or the missing status marker. Output the generated fusion record for subsequent data processing, storage, or retrieval.

2. The multimodal data fusion method according to claim 1, characterized in that, In step S10, the mode set is represented as follows: ; in, It is a pre-determined set of all data modalities corresponding to the same monitoring object. Each modal data is associated with the same monitoring object and carries a corresponding collection timestamp. Let i be the i-th data mode, 1≤i≤n, where n is the preset total number of modes.

3. The multimodal data fusion method according to claim 2, characterized in that, In step S20, each modal data is processed to generate corresponding modal features, wherein the data modes... The modal feature generated at timestamp t is represented as The modal features It includes the following information: timestamp information t used to identify the collection time, object identification information obj_id used to identify the monitored object, and feature information used to characterize the state or content of the modal data.

4. The multimodal data fusion method according to claim 3, characterized in that, In step S30, the method for dividing the monitoring time axis of the original data into multiple consecutive time window sets can be either a fixed step size division rule or a clock alignment division rule. The fixed step size division rule is as follows: a constant time step size ΔT is preset, and each time window... The length of each is equal to ΔT; the clock alignment division rule is: alignment division is based on the natural clock.

5. The multimodal data fusion method according to claim 4, characterized in that, In step S40, for each time window W k From the modal feature set corresponding to each modality, select modal features whose timestamps satisfy the following condition: t∈W k For each mode M in the mode set M i In time window W k The following processing logic is executed within the time window W: If the modality is within the time window W k The memory contains at least one candidate modal feature F i (t), then based on at least one candidate modality feature, generate the associated feature corresponding to the data modality; if the modality is within the time window W k If no modal feature exists within the data modality, then a corresponding missing state label δ is generated for that data modality. i (W k Thus forming a time window W; k Corresponding cross-modal correlation unit: ; in, In time window W k Within, a set of structured data is formed by processing and associating the features of each data mode in the modality set M; For the i-th data mode in the modality set within the time window W k The corresponding associated features or placeholder status within.

6. The multimodal data fusion method according to claim 5, characterized in that, If the data mode M i In time window W k The memory contains multiple candidate modal features F i (t), then a modal feature is determined according to any of the following methods to generate the associated feature corresponding to the data modality: last valid value rule, statistical aggregation rule, time-weighted average rule, feature extreme value rule; the last valid value rule is: filtering within the time window W k The time stamp t collected within the window is closest to the end time of the window. Modal characteristics; The statistical aggregation rule is: for time window W k All modal features F i (t) Perform mathematical statistical calculations, taking the arithmetic mean, median, or mode; the time-weighted average rule is: based on the modal characteristics F... i (t) Distance of sampling point t from the end of the window The duration is assigned different weight coefficients wt, and the weighted average is calculated using the following formula: The eigenvalue extremum rule is to select the maximum or minimum value from the candidate set.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the multimodal data fusion method as described in any one of claims 1-6.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multimodal data fusion method as described in any one of claims 1-6.

9. A method for diagnosing equipment faults, characterized in that, The method includes the following steps: S100, Fusion Record Acquisition and Parsing: Acquire the fusion record output by the multimodal data fusion method as described in any one of claims 1-6 from the message queue or database; S200, Tensor Transformation: Traverse the modal data fields in the fusion record. For each data modality, if the modal data field is an associated feature, write the associated feature value into the corresponding index position of the input tensor; if the modal data field is a missing state marker, fill the missing state marker with a preset placeholder. S300, Mask Generation: Synchronously generate a mask vector to mark the positions in the modal data fields that are marked as missing states as invalid; S400, Model Inference and Output: Input the aligned tensor into the diagnostic neural network model to obtain the diagnostic results.

10. An alarm method, characterized in that, The method includes the following steps: S1000, Rule parsing and mapping: The alarm engine loads preset rules and maps the physical quantities involved in the preset rules to the fixed field indexes in the fusion record output by the multimodal data fusion method as described in any one of claims 1-6; S2000, Fast Status Retrieval and Deterministic Logic Execution: The engine reads the modal data field. If the modal data field is a missing status marker, the incomplete data warning logic is triggered directly without entering the numerical comparison process. If the modal data field is a related feature, the corresponding related feature is extracted and compared with a preset threshold. S3000, Logical Combination Judgment: Based on the judgment result of the data mode, output the final alarm signal.

Citation Information

Patent Citations

  • Multi-modal data processing method and device, equipment and medium

    CN121071487A

  • Health management method and device based on multi-modal data, equipment and storage medium

    CN121122752A

  • Multi-source data fusion pipeline monitoring method and system

    CN121352512A

  • Fault diagnosis method and device of permanent magnet motor system and readable storage medium

    CN121456612A

  • Numerical control machine tool fault prediction system

    CN121523226A