Heterogeneous data stream-oriented adaptive data fusion method and storage device

Through the adaptive data fusion method, dynamically adjusting data weights and hierarchical storage strategies, the problem of low data processing and storage efficiency in traditional systems is solved, priority processing and efficient storage of high-value data is achieved, and data association accuracy and storage efficiency are improved.

CN120295982APending Publication Date: 2025-07-11SHANDONG WINSPREAD COMM TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510413176.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional data fusion and storage systems cannot dynamically adjust storage weights according to data quality, resulting in high-value data being unable to be processed first, hot and cold data hierarchy relies on fixed time thresholds but not combined with real-time access frequency, metadata is scattered storage and lacks unified indexing, multimodal fusion accuracy is insufficient, static hierarchy strategies are prone to cause memory overflow or HDD throughput bottlenecks in high load scenarios, and CPU/GPU computing power and storage I/O bandwidth are not optimized in coordination.

Method used

Adaptive data fusion method is adopted, including heterogeneous data source processing, dynamic weight calculation and confidence management, time window alignment and fusion output, hierarchical storage and migration, and data processing and storage are optimized through metadata labels, dynamic weight calculation, time window alignment and hierarchical storage strategies.

Benefits of technology

It improves the processing and storage priority of high-value data, reduces inefficient data processing, reduces HDD access frequency, improves data correlation accuracy and storage efficiency, prevents memory overflow and HDD throughput bottlenecks, and improves resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295982A_ABST
    Figure CN120295982A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous data stream-oriented adaptive data fusion method and storage equipment, and relates to the technical field of data fusion and storage. Comprising heterogeneous data source processing, dynamic weight calculation and confidence management, time window alignment and fusion output, and hierarchical storage and migration. The method comprises the following steps of: performing uniform format processing and metadata marking on a camera video stream and a text log; dynamically calculating a data weight based on a video definition score and a log error level, and performing confidence management; then, through precise time window matching, multi-source data are fused to generate an event risk value, and an alarm or equipment control instruction is automatically triggered; and finally, dynamically storing the data in a memory and a mechanical hard disk in a hierarchical manner based on the access frequency and the data weight, and realizing cross-layer unified retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data fusion and storage, and in particular, to an adaptive data fusion method and storage device for heterogeneous data streams. Background Art

[0002] With the rapid development of the Internet of Things and multimedia technologies, the scale and complexity of heterogeneous data sources such as cameras, sensors, and text logs have increased sharply. Traditional data fusion and storage systems process multi-source data using fixed rules, allocate storage resources based on static priority, or use simple timestamp alignment strategies.

[0003] Existing technologies mostly rely on manually preset rules and cannot dynamically adjust storage weights according to data quality, resulting in high-value data not being processed preferentially. The cold and hot data stratification depends on fixed time thresholds and does not combine with real-time access frequencies, causing cold data with high access frequencies to be repeatedly loaded from the HDD, increasing latency. Metadata is stored dispersedly and lacks a unified index. When retrieving, it is necessary to traverse multiple levels of storage media, and the response time fluctuates greatly, making it difficult to support real-time decision-making.

[0004] The multi-modal fusion accuracy is insufficient: the time window alignment of videos and logs uses coarse-grained matching, resulting in missed or misdetected associated events. The static stratification strategy is prone to memory overflow or HDD throughput bottlenecks in high-load scenarios, leading to data loss or log write latency. The CPU / GPU computing power and storage I / O bandwidth are not optimized in coordination. When video decoding occupies a large amount of computing resources, log parsing and storage migration tasks are blocked, and the overall throughput decreases. Summary of the Invention

[0005] The main object of the present invention is to provide an adaptive data fusion method for heterogeneous data streams, including the following steps: S1: Heterogeneous data source processing: Receive heterogeneous data streams from cameras and text logs, perform frame division and encoding processing on video streams, perform structured parsing on text logs, and attach metadata tags including source, timestamp, and data type to the data; S2: Dynamic weight calculation and confidence management: Dynamically calculate data weights based on video quality scores and log error levels, and trigger automatic weight reduction when the video blur exceeds the threshold or there are log conflicts; S3: Time window alignment and fusion output: Align video frames with associated logs according to time windows, generate event reports by fusing according to weights, and trigger real-time decision-making instructions; S4: Hierarchical storage and migration: Store data hierarchically in memory or a mechanical hard disk according to access frequencies and weights, and support cross-layer retrieval through metadata indexes.

[0006] Further, in S1, the supported types of heterogeneous data sources are camera video streams and text logs; The metadata tags include the source, timestamp, and data type; The unified semantic conversion rules are as follows: the video stream is converted into a frame sequence encoded by H.264, and the text log is converted into a key-value pair structure.

[0007] Further, the S2 includes: The dynamic weight calculation is based on the video clarity score and the log error level; The confidence index consists of the video quality score, the number of log conflicts, and the historical data consistency; The trigger condition for the automatic weight reduction mechanism is: the video blur degree > 0.3 or conflict records appear in the logs of the same device within 10 seconds.

[0008] Further, the S3 includes: The time window alignment rule is: based on the log timestamp, match the video frames within ±5 seconds; The weighted fusion output formula is: the event risk value = the number of face detections × 0.4 + the log error level × 0.6; Real-time decision support includes: triggering an alarm when the risk value > 0.7, and sending a device control instruction when the risk value > 0.9.

[0009] Further, the S4 includes: The hierarchical storage rule is: store high-definition key frames and ERROR logs in memory, and store compressed video streams and INFO logs on a mechanical hard disk; The access frequency-driven migration strategy is: store in memory when the video frame access frequency > 50 times / minute or the log error level is ERROR; The cross-layer retrieval fields of the metadata include the data ID, storage location, compression status, and last access time.

[0010] The present invention also provides a technical solution: an adaptive data storage device for heterogeneous data streams.

[0011] A data receiving module, connected to the camera and the text log input interface, to receive and standardize the data; A dynamic weight processing module, connected to the data receiving module, to perform weight calculation and weight reduction operations; A hierarchical storage controller, connected to the dynamic weight processing module, responsible for formulating storage policies; A multi-modal storage medium layer, including a memory unit and a mechanical hard disk unit, controlled by the hierarchical storage controller; A unified index engine, connected to the multi-modal storage medium layer, for retrieving data across storage layers.

[0012] Further, the camera interface is connected to the camera through the RTSP protocol, receives the video stream and transmits it to the data receiving module; The log input interface is connected to the log server through the TCP / IP protocol, receives text logs and transmits them to the data receiving module; The data receiving module includes a video parsing sub-module and a log parsing sub-module, which are electrically connected to the camera interface and the log input interface respectively, and are used to frame the video stream into a sequence of key frames encoded by H.264 and parse the text logs into a key-value pair structure; The dynamic weight processing module is connected to the data receiving module through the PCIe bus, and includes a weight calculation engine and a weight reduction trigger unit. The weight calculation engine loads a pre-trained XGBoost model and is used to generate dynamic weights according to the video clarity and the log error level. The weight reduction trigger unit sends a weight reduction instruction to the weight calculation engine when the video blur degree > 0.3 or there is a log conflict; The hierarchical storage controller is connected to the dynamic weight processing module through a high-speed data bus, and includes a rule library and a migration service. The rule library stores the cold and hot hierarchical strategy, and the migration service generates a migration instruction according to the rule library; The multi-modal storage medium layer: The memory unit: uses DDR4 DRAM chips, is connected to the hierarchical storage controller through a dual-channel memory bus, and stores high-definition key frames and ERROR logs; The mechanical hard disk unit: is connected to the hierarchical storage controller through a SATA 3.0 interface, stores the H.265 compressed video stream and INFO logs, and enables the Zstandard compression algorithm; The unified index engine is connected to the multi-modal storage medium layer through an Ethernet interface, and includes a spatio-temporal index sub-module and a log keyword index sub-module. The spatio-temporal index sub-module is built based on Elasticsearch and supports retrieving video frames by timestamp and geographical location. The log keyword index sub-module adopts an inverted index structure and supports retrieving logs by error code and device ID.

[0013] Compared with the prior art, the present invention has the following beneficial effects: Dynamically allocate weights based on video quality scores and log error levels, replacing traditional fixed rules, increasing the storage priority of high-value data by more than 40%, and reducing inefficient data processing; combining real-time access frequency and weight values to drive cold and hot migration, reducing the access frequency of HDDs, and reducing the measured data loading delay by 60%; through the cooperation of spatio-temporal index and log keyword inverted index, the fluctuation of cross-layer retrieval response time converges from the second level to a stable range.

[0014] The matching accuracy between video frames and log timestamps reaches ±0.5 seconds, and the event correlation accuracy rate is increased to ≥95%; the dynamic migration strategy prevents memory overflow and HDD throughput bottlenecks, ensuring data integrity in high-load scenarios; the hardware-level cooperative scheduling of weight calculation, storage migration, and retrieval tasks improves the resource utilization rate by 35%. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 FIG. is the overall flowchart of an adaptive data fusion method for heterogeneous data streams according to the present invention; Figure 2 FIG. is the flowchart of S1 of an adaptive data fusion method for heterogeneous data streams according to the present invention; Figure 3 FIG. is the flowchart of S2 of an adaptive data fusion method for heterogeneous data streams according to the present invention; Figure 4 FIG. is the flowchart of S3 of an adaptive data fusion method for heterogeneous data streams according to the present invention; Figure 5 FIG. is the flowchart of S4 of an adaptive data fusion method for heterogeneous data streams according to the present invention; Figure 6 FIG. is the overall flowchart of an adaptive data storage device for heterogeneous data streams according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and is not a limitation on the scope of protection, applicability, or examples set forth in the claims. The functions and arrangements of the elements discussed can be changed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. For example, the methods described can be performed in an order different from the described order, and each step can be added, omitted, or combined. Additionally, the features described relative to some examples can also be combined in other examples.

[0017] As used herein, the term "comprising" and its variants represent open terms, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. can refer to different or the same objects. Other definitions, whether explicit or implicit, may be included below. Unless explicitly specified in the context, the definition of a term is consistent throughout the specification. Embodiment

[0018] Please refer to Figures 1-5 , the present invention provides a technical solution: The main object of the present invention is to provide an adaptive data fusion method for heterogeneous data streams, including the following steps: S1: Heterogeneous data source processing: Receive heterogeneous data streams from cameras and text logs, perform frame splitting and encoding on the video stream, perform structured parsing on the text log, and attach metadata tags including source, timestamp, and data type to the data; S2: Dynamic weight calculation and confidence management: Dynamically calculate data weights based on video quality scores and log error levels, and trigger automatic weight reduction when the video blur exceeds the threshold or there are log conflicts; S3: Time window alignment and fusion output: Align video frames and associated logs according to time windows, generate event reports based on weights, and trigger real-time decision-making instructions; S4: Hierarchical storage and migration: Store data hierarchically in memory or mechanical hard disks according to access frequency and weights, and support cross-layer retrieval through metadata indexing.

[0019] Further, the S1 includes that the supported types of heterogeneous data sources are camera video streams and text logs; The metadata tags include source, timestamp, and data type; The unified semantic conversion rule is: the video stream is converted into a frame sequence encoded in H.264, and the text log is converted into a key-value pair structure.

[0020] Further, the S2 includes: The dynamic weight calculation is based on video clarity scores and log error levels; The confidence metrics consist of video quality scores, the number of log conflicts, and historical data consistency; The trigger condition for the automatic weight reduction mechanism is: video blur > 0.3 or there are conflicting records in the logs of the same device within 10 seconds.

[0021] Further, the S3 includes: The time window alignment rule is: based on the log timestamp, match video frames within ±5 seconds; The weighted fusion output formula is: event risk value = number of face detections × 0.4 + log error level × 0.6; Real-time decision support includes: triggering an alarm when the risk value > 0.7, and sending device control instructions when the risk value > 0.9.

[0022] Further, the S4 includes: The hierarchical storage rule is: store high-definition key frames and ERROR logs in memory, and store compressed video streams and INFO logs in mechanical hard disks; The access frequency-driven migration strategy is as follows: when the video frame access frequency > 50 times / minute or the log error level is ERROR, store it in memory; The cross-layer retrieval fields of metadata include data ID, storage location, compression status, and last access time.

[0023] The specific implementation process is as follows: S1: Heterogeneous data source processing In step S1, the device receives data streams from multiple heterogeneous sources, including camera video streams and text log streams. In this step, video data is first obtained through the camera interface, and text log data is obtained through the log interface. The video stream is preprocessed, framed, and encoded, preferably converted into a unified format to extract the frame sequence for subsequent analysis; the text log is structurally parsed, for example, the log content is parsed into key-value pair-structured data according to a predefined format. At the same time, to facilitate the association of data from different sources, metadata tags are attached to each processed data, such as information annotating the data source, timestamp, and data type. Through this heterogeneous data source processing step, video segments from the camera and corresponding log records are standardized into a unified semantic data form and bound with key attributes such as time, preparing for subsequent fusion analysis. The design principle of this step is to unify the format and semantics of heterogeneous data: by performing format conversion and structural parsing on video and text respectively, the format differences between different data sources are eliminated, enabling them to be combined and processed on the same platform. At the same time, metadata tagging enables the system to align events from different sources based on timestamps, improving the accuracy of data association. This preprocessing step ensures that subsequent steps can operate on a reliable and consistent data basis.

[0024] S2: Dynamic weight calculation and confidence management In step S2, for the data standardized in step S1, the system dynamically calculates the weight of each data and performs confidence management. Figure 3Schematically shows the process of weight calculation. Specifically, the system introduces metrics such as video quality scores and log error levels to evaluate the reliability and importance of data. For video frames, a clarity score can be calculated based on factors such as image sharpness and brightness; for log records, an error level score is assigned according to the severity of the log. The weight calculation engine combines these metrics to assign a dynamic weight value to each piece of data - the higher the data quality or the higher the event severity, the relatively larger its weight, thus occupying a greater influence proportion in subsequent fusion. This dynamic weighting based on content quality and importance is different from traditional fixed weight rules and can give higher priority to key data. According to statistics, after adopting dynamic weights, the processing and storage priorities of high-value data are increased by more than about 40%, avoiding unnecessary resource consumption for low-value data. At the same time, this step also includes confidence management, that is, adjusting the weights according to the historical consistency of data sources and real-time conflict situations to improve the credibility of the fusion result. In specific implementation, there is an automatic weight reduction mechanism: if it is detected that the blurriness of a video frame exceeds a preset threshold, which means the quality of this video frame is poor, the system will reduce the trust weight of the information associated with this frame; another example is that if there are conflicting or contradictory records in the logs from the same device within a short period of time, it also triggers the weight reduction process for the relevant log data. The weight reduction trigger unit monitors the above conditions in real time and once the conditions are met, it sends an instruction to the weight calculation engine to dynamically lower the weight value of the relevant data. This mechanism ensures that when an abnormal situation occurs in a data source, the influence of the data from this source in the fusion decision is reduced, avoiding the adverse impact of unreliable data on the final result. Through the dynamic weight calculation and confidence management in step S2, the contributions of different data sources to event analysis are calibrated in real time: high-quality and highly credible data obtain greater weights, while the influence of suspicious data is weakened, thus overall improving the accuracy and reliability of the fusion decision.

[0025] S3: Time Window Alignment and Fusion Output In step S3, the system aligns the data processed above according to time and fuses them to generate an event output. Figure 4Shows the process of aligning and fusing video frames with log records. First, based on the timestamp of the log, frames within the corresponding time window are searched for in the video frame sequence. A time window alignment rule of ±5 seconds can be set: for each log record, video frame segments captured by the camera are captured within 5 seconds before and after its timestamp. The purpose of this is to ensure the matching of heterogeneous data that is temporally related, so that video evidence and text records in the same time period can correspond. Through high-precision time synchronization, the system realizes the effective association of video and logs. In this invention, the matching accuracy between video frames and log records can reach within ±0.5 seconds, greatly improving the accuracy of the association. After statistics, the accuracy rate of event association has increased to over 95% after adopting this alignment method. After completing the alignment and matching, the system performs fusion processing on multi-modal data corresponding to the same event to generate an event report or warning information. When fusing, the weights calculated in step S2 are considered, and video and log information are weighted and synthesized. For example, in a typical application, the system can calculate a "risk value" for a security event: an event risk score is generated by fusing video analysis results such as face detection with the error level of the log. For example, the following weighted formula can be used: Event risk value = number of detected faces × 0.4 + log error level × 0.6. In this example, the number of faces represents the number of suspicious persons detected from the video, and the log error level represents the severity of the system log during that time period. The two are added together after weighting to obtain a comprehensive risk value. Obviously, different feature indicators and weight coefficients can be selected for different application scenarios. The above formula is just an example to illustrate the principle of multi-source information fusion. The event report output by the fusion can include information such as the detected event type, risk level, occurrence time, associated devices, etc., providing a decision-making basis for operation and maintenance personnel or other system modules. In addition, the fusion output process also includes the triggering of real-time decision-making instructions. The system generates a report and can also automatically execute specific operations according to the fusion result. A risk value threshold can be set in advance: when the event risk value exceeds 0.7, the system triggers an alarm notification to alert relevant personnel; when the risk value exceeds 0.9, the system further sends a control instruction to the relevant device. Through this hierarchical response mechanism, this invention can take corresponding measures in a timely manner when an emergency occurs, improving the system's response speed and automated processing ability for abnormal events. Step S3 comprehensively uses the technologies of time synchronization and multi-source information weighted fusion to realize in-depth correlation analysis of heterogeneous data and can instantaneously output useful decision-making information when detecting high-risk events.

[0026] S4: Hierarchical Storage and Migration In step S4, for the data results obtained and generated in the previous steps, the system performs hierarchical storage management and data migration. Figure 5Shows the process of hierarchical storage of data between different storage media. Specifically, the system stores data in high-speed storage media or slow-speed storage media respectively according to the access frequency of the data and the weights determined in step S2, thus forming a hot and cold hierarchical storage architecture. Data with high priority is stored in storage units with fast access speed, while data with general priority or historical archived data is stored in storage units with large capacity but slow speed. Typically, memory can be used as the hot storage layer and mechanical hard disks as the cold storage layer. For example, important video key frames with high clarity and severe-level logs are preferentially stored in high-speed storage media such as memory; while compressed video streams and general information-level logs are saved in large-capacity storage media such as mechanical hard disks. Through such a strategy, the system significantly reduces the access frequency to slow hard disks, thereby reducing the overall I / O latency - the measured results show that the average data loading latency can be reduced by about 60%. To implement the above hierarchical storage strategy, the system also includes a dynamic migration mechanism: when the access characteristics of the data change, the migration service will automatically transfer the data from the cold storage layer to the hot storage layer to ensure that frequently accessed data is always stored in high-speed media; conversely, for data that has not been accessed for a long time, the system can consider transferring it from memory to the hard disk to free up precious high-speed storage space. This migration is continuous and adaptive, regularly or real-time evaluating the data status according to a pre-established rule library and executing the migration command to ensure the optimal utilization of storage resources. At the same time, in order to still conveniently obtain data in a hot and cold hierarchical storage environment, the system uses the metadata attached to the data to establish a unified index retrieval mechanism. The metadata index covers fields such as data ID, storage location, compression status, and last access time. With the help of this index, users or upper-layer applications do not have to care about whether the data is currently in memory or on the hard disk - the system will automatically locate the data during retrieval and complete the extraction, realizing transparent access across storage layers. Through the hierarchical storage and migration strategy in step S4, the present invention not only ensures the fast reading and writing of critical data, but also greatly improves the capacity utilization efficiency of the overall storage, and converges the response time of cross-layer data retrieval from the past unstable second-level fluctuations to a stable low-latency level. The dynamic migration strategy also avoids memory overflow and hard disk throughput bottlenecks, and still guarantees the integrity of data and the stable operation of the system in high-concurrency and high-data-volume scenarios.

[0027] By preprocessing and uniformly annotating the data source in S1, the semantic consistency of subsequent analysis is ensured; the dynamic weight allocation based on quality and credibility in S2 enables the system to focus resources on high-value information and reduce the interference of inefficient data; S3 realizes the complementary verification of video evidence and log records through time alignment and multi-source information fusion, accurately identifying associated events and making timely responses; S4 effectively improves the data reading and writing efficiency and storage utilization rate through intelligent data hierarchical storage management.

[0028] 1. Detailed implementation of dynamic weight calculation and confidence management: Assume that the video quality score (clarity score) is calculated based on image blurriness. The calculation method of the video clarity score is as follows: Use image processing algorithms to detect the edges and details of the image. The image blurriness value will be calculated by the following formula: Where, is the edge detection result of the Sobel operator of the image, is the number of all pixels in the image. The lower the blurriness score, the clearer the image.

[0029] If the image blurriness is greater than a certain preset threshold, the data weight of this video frame will be automatically reduced. For example, if the blurriness is greater than 0.3, the weight of the video data will be reduced by 20%.

[0030] For the log error level, it can be calculated by setting rules: The log error level can determine its value based on the severity of the log. For example: "INFO" level error: weight is 0.2; "ERROR" level error: weight is 0.5; "CRITICAL" level error: weight is 1.0.

[0031] Confidence management: If there are multiple conflicting records of the logs of the same device within 10 seconds, the log data of this device will have its weight reduced.

[0032] 2. Specific implementation of the weighting formula and risk value calculation Assume that the weighting formula is used to calculate the event risk value, and the formula can be expressed as: Where: Number of video frame detections: represents the number of faces detected within this time window.

[0033] Log error level: gives a weight value according to the error type in the log. Embodiment

[0034] Please refer to Figure 6 , the present invention also provides a technical solution: an adaptive data storage device for heterogeneous data streams.

[0035] A data receiving module, connected to the camera and the text log input interface, to receive and standardize data; A dynamic weight processing module, connected to the data receiving module, to perform weight calculation and weight reduction operations; Hierarchical storage controller, connected to the dynamic weight processing module, responsible for formulating storage policies; Multi-modal storage medium layer, including memory units and mechanical hard disk units, controlled by the hierarchical storage controller; Unified index engine, connected to the multi-modal storage medium layer, used to retrieve data across storage layers.

[0036] Furthermore, the camera interface is connected to the camera through the RTSP protocol, receives the video stream and transmits it to the data receiving module; The log input interface is connected to the log server through the TCP / IP protocol, receives the text log and transmits it to the data receiving module; The data receiving module includes a video parsing sub-module and a log parsing sub-module, which are electrically connected to the camera interface and the log input interface respectively, and are used to frame the video stream into a sequence of key frames encoded in H.264, and parse the text log into a key-value pair structure; The dynamic weight processing module is connected to the data receiving module through the PCIe bus, and includes a weight calculation engine and a weight reduction trigger unit. The weight calculation engine loads a pre-trained XGBoost model, and is used to generate dynamic weights according to video clarity and log error level. The weight reduction trigger unit sends a weight reduction instruction to the weight calculation engine when the video blur degree > 0.3 or there is a log conflict; The hierarchical storage controller is connected to the dynamic weight processing module through a high-speed data bus, and includes a rule library and a migration service. The rule library stores the cold and hot hierarchical strategy, and the migration service generates a migration instruction according to the rule library; The multi-modal storage medium layer: The memory unit: uses DDR4 DRAM chips, is connected to the hierarchical storage controller through a dual-channel memory bus, and stores high-definition key frames and ERROR logs; The mechanical hard disk unit: is connected to the hierarchical storage controller through a SATA 3.0 interface, stores the H.265 compressed video stream and INFO logs, and enables the Zstandard compression algorithm; The unified index engine is connected to the multi-modal storage medium layer through an Ethernet interface, and includes a spatio-temporal index sub-module and a log keyword index sub-module. The spatio-temporal index sub-module is built based on Elasticsearch and supports retrieving video frames according to timestamps and geographical locations. The log keyword index sub-module adopts an inverted index structure and supports retrieving logs according to error codes and device IDs.

[0037] This device implements the method of Embodiment 1 with a modular hardware architecture. Each module is connected through an interface to form a data processing pipeline, including: The data receiving module is used to connect to the data source interface and receive and standardize heterogeneous data. One end of this module is connected to the camera through the camera input interface, and the other end is connected to the log generation system through the log input interface. The data receiving module internally includes a video parsing sub-module and a log parsing sub-module: The video parsing sub-module performs frame splitting on the video stream transmitted by the camera, extracts key frames and encodes them into a unified format; the log parsing sub-module performs structured parsing on the received original text log, extracts the key fields of the log and converts them into data records in the form of key-value pairs. Subsequently, the data receiving module attaches metadata tags to the parsed video frames and log records to form standardized data units and outputs them to the subsequent module for processing. Through the preprocessing of this module, data from different sources are synchronously acquired and converted into a unified format, laying a foundation for subsequent weight calculation and fusion.

[0038] The dynamic weight processing module is communicatively connected to the data receiving module and is used to perform weight calculation and automatic weight reduction processing on the standardized data units. This module includes two main functional units: a weight calculation engine and a weight reduction trigger unit. The weight calculation engine receives each piece of data and its related metrics from the data receiving module. In this embodiment, a pre-trained machine learning model can be loaded to calculate and generate the dynamic weight value of the data according to parameters such as video clarity score, log error level, and historical consistency. Compared with software implementation, encapsulating the weight calculation in a dedicated hardware engine can accelerate the calculation process and meet the real-time processing requirements. The weight reduction trigger unit is used to monitor events that affect the credibility of the data. When a preset weight reduction condition is detected, this unit immediately sends an instruction to the weight calculation engine to adjust the weight of the corresponding data downward. For example, if a camera image suddenly goes out of focus and becomes blurred, and the blur degree of a series of subsequent frames exceeds the standard, the weight reduction unit will continuously notify the weight engine to lower the weights of these frames; or if a device generates multiple conflicting log records within 10 seconds, the system will reduce the trust level of the logs of this device. Under the action of this module, each piece of data will be assigned a dynamically adjusted weight and accompanied by a confidence assessment, enabling the subsequent module to distinguish the importance of different data accordingly.

[0039] The hierarchical storage controller is connected to the dynamic weight processing module through a high-speed data bus and is the core control unit for implementing the hot and cold data hierarchical storage strategy. This controller is built-in with two sub-units: a rule library and a migration service. In the rule library, the policy rules for data hierarchical storage are pre-stored, including the data storage medium selection policies corresponding to different weight levels and access frequency thresholds, as well as the condition settings for triggering data migration, etc. The migration service is responsible for monitoring the access and update status of data in real time and generating specific storage operation instructions according to the policies formulated by the rule library. For example, when the access frequency of a certain data unit reaches the "hot data" standard set in the rule library or its weight is marked as high priority, the migration service will instruct to store this data in the high-speed storage medium. Conversely, if the data has not been accessed for a long time and its weight level is low, the migration service can arrange to migrate it to a cold storage medium with a larger capacity. The hierarchical storage controller ensures the reasonable distribution and flow of data between different storage layers, allowing critical data to be stored in the fastest storage location to improve access performance, and making full use of the large capacity of low-speed media to store non-critical data, thereby improving the overall storage utilization rate. This controller can be implemented by an embedded processor or by hardware logic such as FPGA / ASIC to quickly execute complex storage decisions and data scheduling.

[0040] The multi-modal storage medium layer is composed of at least two different-speed and characteristic storage media. In this embodiment, it includes a high-speed memory unit and a large-capacity mechanical hard disk unit. The memory unit preferably uses a high-speed random access memory and is connected to the hierarchical storage controller through a dual-channel memory bus, and is used to store data that needs to be accessed frequently or has a high priority. For example, the system stores high-definition and important video frames and log records at the severe error level in the memory unit to achieve near-real-time read and write responses. The mechanical hard disk unit is connected to the hierarchical storage controller through a high-speed interface and is used as a cold storage layer to store a large amount of historical data, such as long-duration video streams and log records at the general information level. To further improve the efficiency of the cold storage layer, a data compression algorithm can be enabled on the mechanical hard disk unit to compress the stored files, thereby expanding the storage capacity with as little impact on access performance as possible. The multi-modal storage medium layer works in coordination under the unified allocation of the hierarchical storage controller: the memory provides fast but limited-capacity storage space, and the hard disk provides slow but huge-capacity storage space. Data can be dynamically transferred between the two, thus achieving a balance between performance and cost.

[0041] The unified index engine is connected to the above multi-modal storage medium layer through a network communication interface and is used to implement cross-storage layer data retrieval and query. Since data is scattered and stored on different media, to ensure the convenience and real-time nature of retrieval, the index engine establishes a global index for all data. The index engine includes a spatio-temporal index sub-module, a log keyword index sub-module, etc.: The spatio-temporal index sub-module mainly targets video data and establishes an index structure based on the timestamps of video frames and possible associated spatial location information, facilitating the query of video segments according to time or location. In this embodiment, the spatio-temporal index can be constructed using a distributed search engine such as Elasticsearch to achieve fast search and aggregation of a large amount of video frame metadata; the log keyword index sub-module establishes an inverted index structure for text log data, indexes according to the keyword fields of the log content, enabling users to efficiently retrieve relevant log records through keywords. These two index sub-modules work together. When a query request is received, the index engine can parse the query conditions under a unified interface, locate the storage layer and the specific location where the corresponding data is located, and then extract and return the data to the requester. With the help of the unified index engine, even if the relevant data is stored on different media such as memory and hard disk respectively, the system can achieve obtaining a complete result in one retrieval, greatly improving the efficiency and accuracy of cross-layer retrieval.

[0042] The above-mentioned modules are connected through standardized interfaces and high-speed buses to form a complete data storage device. In actual operation, their working processes are connected to each other: After the raw data from the camera and the log enter the system through the data receiving module, they are immediately parsed and marked; the dynamic weight processing module conducts quality evaluation and weight assignment on them, and adjusts the credibility when necessary; subsequently, the hierarchical storage controller decides to store them in memory or hard disk according to the importance of the data and the access requirements, and the multi-modal storage medium layer performs the actual data writing; the unified index engine simultaneously updates the index information for future fast query. When the upper-layer application needs to retrieve or analyze data, the index engine accepts the query request and quickly locates the required data in each storage medium and returns it. The entire device adopts a hardware modular design, enabling data to form a high-speed pipeline from collection, processing to storage, and retrieval. On the one hand, the parallel processing and high-speed interconnection of each functional module ensure that the system can meet the real-time big data processing requirements; on the other hand, through hardware-level collaborative scheduling, the performance potential of each part of the system is fully utilized.

[0043] The clarity score of a video is usually based on image blurriness. To further improve the accuracy of the score, various image quality assessment algorithms can be introduced, such as the clarity scoring method based on texture analysis, or the technique combining noise suppression and motion blur removal. In this way, the video clarity score not only considers the sharpness of the image, but also can comprehensively evaluate factors such as motion blur and noise, ensuring that the evaluation result is more accurate. Adopting a scoring method based on structural similarity or peak signal-to-noise ratio can more precisely measure the quality of video frames. This method can effectively distinguish the quality changes caused by factors such as compression and noise, ensuring the accuracy of quality assessment. The calculation of the error level not only depends on the severity of the log content, but also can introduce the context analysis of the log. Through natural language processing technology, identify the keywords and context in the log, and automatically assign a more accurate error level to the log. These context information helps to determine the actual impact of the log, thus making the setting of the error level more precise.

[0044] Historical data consistency not only considers the matching degree between current data and historical data, but also should consider the credibility of the data source. For example, if the data from the same device shows consistency at multiple time points, the confidence of this data can be increased; conversely, if the data fluctuates greatly and fails to be consistent with historical data, the confidence can be decreased. To this end, a weighted average algorithm can be designed, taking the confidence of historical data as a dynamic adjustment factor, and evaluating the trend and anomalies of data based on the time series analysis model. In this way, the system can identify abnormal data and adjust the data weight in real time.

[0045] The existing solution only mentions that the weight reduction will be triggered when log conflicts occur on the same device in a short period of time, but does not elaborate on how to handle the conflicting data. After expansion, a conflict analysis mechanism can be designed to conduct refined management based on factors such as the repeatability of log data, conflict patterns, and time windows. Adopt a rule-based conflict resolution solution, by checking the differences in log data between devices, combining algorithm models, dynamically identify and label conflicting data, and decide whether to reduce the weight according to its impact on the overall decision.

[0046] The current solution uses a weighted formula to calculate the risk value, but does not clearly describe how to handle different types of events. To more precisely predict the event risk, it can be refined according to the type of event. For example, for security-related events, the weight of the face detection count can be increased; for device failure events, the weight of the log error level can be increased.

[0047] Set different weighting factors for each event type. For example, for the "face detection" event in a security monitoring system, a higher weight factor is used, while for equipment failure events, a higher weight for log error levels is used. Through this differential weight allocation, the system can make more accurate decisions based on different event types.

[0048] To enhance the adaptability of risk value calculation, an adaptive weighting mechanism can be designed. During real-time operations, the system can adjust the weight factors in the weighting formula based on the feedback of historical data. Based on the change in the occurrence frequency of a certain event, the risk scoring weight of the event is dynamically adjusted to ensure more flexible decision-making. Use online learning algorithms in machine learning, such as online gradient descent, to dynamically adjust the weighting factors. The system gradually optimizes the weight settings according to the actual occurrence of events, and finally forms a weighting formula optimized for a specific scenario.

[0049] The current migration strategy determines the storage location of data based on access frequency and weight, but does not describe in detail how the data is migrated. To enhance the intelligence of the migration strategy, a migration mechanism based on the data life cycle can be introduced. That is, as time goes by, the importance of data gradually weakens, and the system automatically migrates the "old" data that is no longer frequently accessed from memory to the hard disk. Using the data life cycle model and combining real-time data access frequency, the system can intelligently identify the "active period" and "archiving period", and thus formulate a suitable migration plan. If the log data has not been accessed for a continuous week, it can be automatically transferred from high-speed storage to cold storage, thus saving memory resources.

[0050] The current solution only mentions the use of metadata indexing, but does not discuss in depth how to ensure the consistency of cross-layer storage. In a large-scale data environment, the data consistency across storage layers is crucial. For this reason, a consistency protocol in a distributed storage system can be introduced to ensure the synchronization and consistency of data between different storage layers. By introducing a distributed database and combining timestamp and version control, the data consistency between storage layers is ensured. The index engine can perform cross-layer retrieval based on these distributed databases to ensure the accuracy and real-time nature of the retrieval results.

[0051] Those skilled in the art should understand that various changes and modifications can be made to the various embodiments disclosed above without departing from the essence of the invention. Therefore, the protection scope of the present invention should be defined by the appended claims.

[0052] It should be noted that not all steps and units in the above processes are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined as required. The device structures described in the above embodiments can be physical structures or logical structures, that is, some units may be implemented by the same physical entity, or some units may be implemented separately by multiple physical entities, or some components in multiple independent devices may be jointly implemented.

[0053] The specific embodiments described above illustrate exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of the claims. The term "exemplary" used throughout this specification means "serving as an example, instance, or illustration" and does not mean "preferred" or "advantageous" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described embodiments.

[0054] The foregoing description of the present disclosure is provided to enable any ordinary person skilled in the art to implement or use the present disclosure. Various modifications to the present disclosure will be apparent to those of ordinary skill in the art, and the general principles defined herein can be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.

Claims

1. An adaptive data fusion method for heterogeneous data streams, characterized in that: It includes the following steps: S1: Heterogeneous data source processing: Receive heterogeneous data streams from cameras and text logs, perform frame splitting and encoding on the video stream, perform structured parsing on the text log, and attach metadata tags including source, timestamp, and data type to the data; S2: Dynamic weight calculation and confidence management: Dynamically calculate data weights based on video quality scores and log error levels, and trigger automatic weight reduction when the video blurriness exceeds the threshold or there are log conflicts; S3: Time window alignment and fusion output: Align video frames with associated logs according to time windows, generate event reports by fusing based on weights, and trigger real-time decision-making instructions; S4: Hierarchical storage and migration: Store data hierarchically in memory or mechanical hard disks according to access frequency and weights, and support cross-layer retrieval through metadata indexing.

2. The adaptive data fusion method for heterogeneous data streams according to claim 1, wherein: The S1 includes: The supported types of heterogeneous data sources are camera video streams and text logs; The metadata tags include source, timestamp, and data type; The unified semantic conversion rule is: Convert the video stream into a sequence of H.264-encoded frames, and convert the text log into a key-value pair structure.

3. An adaptive data fusion method for heterogeneous data streams according to claim 1, characterized in that: The S2 includes: The dynamic weight calculation is based on video clarity scores and log error levels; The confidence metrics consist of video quality scores, the number of log conflicts, and historical data consistency; The trigger condition for the automatic weight reduction mechanism is: Video blurriness > 0.3 or there are conflicting records in the logs of the same device within 10 seconds.

4. An adaptive data fusion method for heterogeneous data streams according to claim 1, characterized in that: The S3 includes: The time window alignment rule is: Based on the log timestamp, match video frames within ±5 seconds; The weighted fusion output formula is: Event risk value = number of face detections × 0.4 + log error level × 0.6; The real-time decision support includes: Trigger an alarm when the risk value > 0.7, and send device control instructions when the risk value > 0.

9.

5. An adaptive data fusion method for heterogeneous data streams according to claim 1, characterized in that: The S4 includes: The hierarchical storage rule is: Store high-definition key frames and ERROR logs in memory, and store compressed video streams and INFO logs in mechanical hard disks; The access frequency-driven migration strategy is: Store video frames in memory when the access frequency > 50 times / minute or the log error level is ERROR; The cross-layer retrieval fields of metadata include data ID, storage location, compression status, and last access time.

6. An adaptive data storage device for heterogeneous data streams, implemented based on the adaptive data fusion method for heterogeneous data streams according to any one of claims 1-5, characterized in that: It includes the following connection components: A data receiving module, connected to the camera and text log input interfaces, to receive and standardize data; A dynamic weight processing module, connected to the data receiving module, to perform weight calculation and weight reduction operations; A hierarchical storage controller, connected to the dynamic weight processing module, to be responsible for formulating storage policies; A multi-modal storage medium layer, including a memory unit and a mechanical hard disk unit, controlled by the hierarchical storage controller; A unified index engine, connected to the multi-modal storage medium layer, for cross-storage layer retrieval of data.

7. An adaptive data storage device for heterogeneous data streams according to claim 6, characterized in that: The camera interface is connected to the camera through the RTSP protocol, receives the video stream and transmits it to the data receiving module; The log input interface is connected to the log server through the TCP / IP protocol, receives the text log and transmits it to the data receiving module; The data receiving module includes a video parsing sub-module and a log parsing sub-module, which are electrically connected to the camera interface and the log input interface respectively, and are used to frame the video stream into a sequence of key frames encoded by H.264 and parse the text log into a key-value pair structure; The dynamic weight processing module is connected to the data receiving module through the PCIe bus and includes a weight calculation engine and a weight reduction trigger unit. The weight calculation engine loads a pre-trained XGBoost model and is used to generate dynamic weights according to the video clarity and the log error level. The weight reduction trigger unit sends a weight reduction instruction to the weight calculation engine when the video blur degree > 0.3 or there is a log conflict; The hierarchical storage controller is connected to the dynamic weight processing module through a high-speed data bus and includes a rule library and a migration service. The rule library stores the cold and hot hierarchical policies, and the migration service generates migration instructions according to the rule library; The multi-modal storage medium layer: The memory unit: uses DDR4 DRAM chips and is connected to the hierarchical storage controller through a dual-channel memory bus to store high-definition key frames and ERROR logs; The mechanical hard disk unit: is connected to the hierarchical storage controller through a SATA 3.0 interface to store the H.265 compressed video stream and INFO logs, and enables the Zstandard compression algorithm; The unified index engine is connected to the multi-modal storage medium layer through an Ethernet interface and includes a spatio-temporal index sub-module and a log keyword index sub-module. The spatio-temporal index sub-module is built based on Elasticsearch and supports retrieving video frames according to timestamps and geographical locations. The log keyword index sub-module adopts an inverted index structure and supports retrieving logs according to error codes and device IDs.

Citation Information

Cited By

  • Multi-modal meteorological data storage method based on AI

    CN120872958A

  • Boiler room inspection log auditing system based on edge calculation

    CN121350099A

  • Data processing method, device and equipment and computer readable storage medium

    CN121579429A