Multi-dimensional data fusion real-time filtering system and method based on internet of things
By using the data stream interruption monitoring and filtering accuracy assessment module, the problem of data stream interruption caused by network instability during multi-dimensional data collection by IoT terminals was solved, realizing accurate fusion and efficient storage of multi-dimensional data, and improving the accuracy and reliability of data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 北京中电华安科技股份有限公司
- Filing Date
- 2025-11-05
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, unstable networks can cause data flow interruptions during multi-dimensional data collection by IoT terminals, leading to differences in cloud data quality and insufficient spatiotemporal alignment accuracy, which affects the accuracy of multi-dimensional data fusion and real-time filtering.
By using the data stream interruption monitoring module, the filtering accuracy monitoring module, and the filtering accuracy optimization judgment module, the impact of data stream interruption is judged and the filtering accuracy is evaluated to ensure that the spatiotemporal alignment is qualified, and to achieve accurate fusion of multi-dimensional data and storage quality analysis.
It improves the accuracy of multi-dimensional data fusion and real-time filtering in the Internet of Things, ensures data processing efficiency and quality, avoids resource waste, and ensures quality control throughout the entire process from data collection to fusion.
Smart Images

Figure CN121456814B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real-time data filtering technology, and in particular to a real-time filtering system and method based on multi-dimensional data fusion of the Internet of Things. Background Technology
[0002] To ensure the credibility of multi-dimensional data sources from IoT terminals, cloud security capabilities can be integrated into the terminals to filter data and prevent unauthorized device access and data theft. Existing methods first collect multi-dimensional data from IoT terminals (including but not limited to traffic flow and average vehicle speed detected by video radar during traffic flow monitoring). Before the multi-dimensional data enters the cloud, edge computing nodes perform a first round of real-time filtering and lightweight fusion, such as filtering out meaningless steady-state data or making preliminary judgments using local rule engines. This greatly reduces the bandwidth pressure on the core links of the cloud network. After the multi-dimensional data enters the cloud, the core process unfolds: the multi-dimensional data is stored in different cloud storage layers (such as real-time layer Kafka, near-real-time layer HBase, historical layer InfluxDB, etc.). For example, real-time streaming data enters a message queue (such as Kafka) for real-time processing, while historical data is stored in a time-series database or data lake. Next, real-time fusion engines (such as Flink and Spark Streaming) concurrently read data from multiple data sources and perform spatiotemporal alignment, correlation analysis, and quality assessment based on preset fusion models (such as the ImageBind model), eliminating outliers, redundant information, and untrusted data caused by network jitter. During this process, cloud security capabilities intervene again, ensuring the confidentiality and integrity of the data processing chain through two-way authentication between microservices and encrypted processing of streaming data. The fused and filtered multi-dimensional data is then output, driving real-time monitoring dashboards and decision-making systems in the cloud. Furthermore, through the tight integration of elastic cloud network connectivity, layered cloud storage processing, and end-to-end cloud security protection, accurate, reliable, and real-time information is efficiently obtained, providing strong data support for applications such as smart cities and the industrial internet.
[0003] For example, the Chinese invention patent with announcement number CN120449108B discloses a real-time data filtering method and system based on multi-dimensional feature fusion, which includes: acquiring data transmission logs and extracting features of the original transmission data, and parsing the data frame field structure information; identifying multi-dimensional abnormal coupling of data transmission based on the field time gradient change rate and dynamic offset status; further detecting abnormal dilution of transmission data, predicting the data transmission complexity exceeding the limit, and detecting the overload status of the data filtering chip structure; subsequently evaluating the decay trend of real-time data filtering efficiency, optimizing the real-time data filtering path accordingly, completing the real-time filtering processing of the data transmission complexity exceeding the limit, and outputting the filtering result.
[0004] For example, Chinese invention patent CN110674125B discloses a method, device, and readable storage medium for filtering data to be fused, which includes: acquiring multiple data to be fused from a cached database, as well as the data identifier and time information of each data to be fused; dividing the multiple data to be fused into multiple datasets to be filtered based on the acquisition time of each data to be fused and multiple preset time intervals; determining the target data and its time information from the multiple data to be fused that belong to the same data identifier in each dataset to be filtered; and filtering the other data to be fused in each dataset to be filtered except for the target data to obtain the dataset to be fused.
[0005] The above-mentioned technology has at least the following technical problems:
[0006] Because IoT terminals may experience data flow interruptions during multi-dimensional data collection due to unstable networks (manifested as multiple tasks running simultaneously at the IoT terminal edge, storage interruption resume, network jitter, or communication interruptions), such as traffic flow monitoring in weak signal areas in smart cities, existing methods may fail to obtain continuous and complete multi-dimensional data sources. This can lead to data quality differences in the layered storage stage after the filtered multi-dimensional data arrives at the cloud (e.g., real-time streaming data stored in Kafka contains incomplete data with many missing fields, and historical data stored in time-series databases has data gaps at key time nodes). Existing completion methods often use offline cloud-based completion or simple linear prediction, which lacks accuracy and cannot guarantee spatiotemporal alignment precision, resulting in spatiotemporal alignment relationship failure or offset, and abrupt changes in alignment results. Lightweight fusion of multi-dimensional data may not be able to accurately achieve basic spatiotemporal alignment and correlation verification of multi-dimensional data. Consequently, when the cloud-based real-time fusion engine concurrently reads multi-dimensional data, due to incomplete data sources and basic spatiotemporal alignment deviations, it cannot complete accurate correlation analysis and quality assessment through the preset fusion model, leading to low accuracy in multi-dimensional data fusion and real-time data filtering based on IoT. Summary of the Invention
[0007] To address the low accuracy of existing technologies for multi-dimensional data fusion and real-time data filtering based on the Internet of Things (IoT), this invention provides a multi-dimensional data fusion and real-time filtering system and method based on IoT. The technical solution is as follows:
[0008] On one hand, a real-time filtering system for multi-dimensional data fusion based on the Internet of Things (IoT) is provided, including: a data flow interruption monitoring module, a filtering accuracy monitoring module, and a filtering accuracy optimization judgment module. The data flow interruption monitoring module is used to determine the impact of data flow interruptions during the fusion and filtering of multi-dimensional data collected by IoT terminals in a specified scenario, generating a judgment result reflecting the data flow interruption situation during multi-dimensional data collection and transmitting it to the corresponding module. The filtering accuracy monitoring module is used to perform spatiotemporal alignment and correlation verification to evaluate the spatiotemporal alignment of multi-dimensional data if the received judgment result is not to perform filtering accuracy evaluation. Based on the output verification result, it determines whether to perform multi-dimensional data fusion to output the fused multi-dimensional data. The filtering accuracy evaluation is used to evaluate the filtering situation of multi-dimensional data at edge computing nodes and the cloud. Multi-dimensional data represents a collection of structured and unstructured data collected by IoT terminals. The filtering accuracy optimization judgment module is used to determine whether to perform hierarchical storage quality analysis to reflect the quality qualification of multi-dimensional data in the hierarchical storage stage if the received judgment result is to perform filtering accuracy evaluation. If hierarchical storage quality analysis is performed, spatiotemporal alignment and correlation verification are performed after the analysis is completed; otherwise, spatiotemporal alignment and correlation verification are performed directly.
[0009] On the other hand, a real-time filtering method for multi-dimensional data fusion based on the Internet of Things (IoT) is provided, which is applied to a real-time filtering system for multi-dimensional data fusion based on IoT. This method includes: during the fusion and filtering of multi-dimensional data from a specified scenario collected by an IoT terminal, determining the impact of data flow interruption to generate a judgment result reflecting the data flow interruption during multi-dimensional data collection and transmitting it to the corresponding module; if the received judgment result indicates no filtering accuracy assessment, then performing spatiotemporal alignment and correlation verification to assess the spatiotemporal alignment of the multi-dimensional data, and determining whether to perform multi-dimensional data fusion based on the output verification result to output the fused multi-dimensional data; if the received judgment result indicates filtering accuracy assessment, then determining whether to perform hierarchical storage quality analysis based on the output result to reflect the quality compliance of the multi-dimensional data in the hierarchical storage stage, and if hierarchical storage quality analysis is performed, then spatiotemporal alignment and correlation verification are performed after the analysis; otherwise, spatiotemporal alignment and correlation verification are performed directly.
[0010] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0011] 1. By determining the impact of data flow interruptions and deciding whether to conduct a filtering accuracy assessment based on the output results, it helps to accurately locate data quality problems and avoid the transmission of quality risks throughout the entire chain. If a filtering accuracy assessment is not conducted, spatiotemporal alignment and correlation verification are performed. Based on the output verification results, it is determined whether to perform multi-dimensional data fusion, which helps to achieve "resource allocation on demand" and balance data processing efficiency. If a filtering accuracy assessment is conducted, it is determined whether to perform hierarchical storage quality analysis based on the output results. If hierarchical storage quality analysis is performed, spatiotemporal alignment and correlation verification are performed after the analysis is completed; otherwise, spatiotemporal alignment and correlation verification are performed directly. This helps to remove quality obstacles in the fusion process, ensure the accuracy of dimensional data fusion, and thus improve the accuracy of multi-dimensional data fusion and real-time filtering based on the Internet of Things, solving the problem of low accuracy in existing technologies for multi-dimensional data fusion and real-time filtering based on the Internet of Things.
[0012] 2. By inputting the data stream interruption impact judgment value into the preset data stream interruption related set, the interruption interference correlation value is output. When the interruption interference correlation value is greater than the preset interruption interference correlation value, the filtering accuracy is evaluated; otherwise, spatiotemporal alignment and correlation verification are performed. This helps to achieve refined hierarchical control of the impact of data stream interruption. For example, continuous disconnection of terminals in weak signal areas of smart cities leads to a large number of time gaps and missing fields in multi-dimensional data. The filtering accuracy evaluation can accurately identify the low-quality data that remains after edge node filtering (such as incomplete traffic flow data that has not been removed), thereby helping to avoid the ineffective consumption of edge and cloud computing resources.
[0013] 3. When multiple tasks need to be processed simultaneously on an edge computing node (such as processing multiple tasks simultaneously in traffic flow monitoring in a smart city), some data streams may be interrupted due to limited processing resources or excessive load. At this time, data streams of different dimensions may not be able to be transmitted to the cloud simultaneously due to packet loss or interruption, resulting in data loss or missing data during the data storage stage. By obtaining the average duration of the interruption when the load on the edge computing node exceeds the preset edge computing node load, and by evaluating the accuracy of filtering when the average duration of the interruption exceeds the preset average duration of the interruption, and performing spatiotemporal alignment and correlation verification, it is helpful to achieve differentiated management of data stream interruption risks in high-load edge node scenarios, accurately intercept low-quality data caused by overload, and thus ensure the integrity of multi-dimensional data transmission and storage quality from the edge to the cloud, laying a solid data foundation for subsequent cloud-based integrated analysis.
[0014] 4. By acquiring multi-dimensional data, and when the spatial alignment qualification value is greater than the preset spatial qualification value and the time alignment qualification value is greater than the preset time qualification value, multi-dimensional data fusion is performed; otherwise, a fusion anomaly prompt is sent. Existing technologies often judge whether data can be fused based on a single dimension (such as data format or numerical range), ignoring the data's "spatial location matching degree" (such as whether environmental data collected by IoT devices in the same area correspond to the same geographical coordinates) and "time synchronization". Compared with existing technologies, this helps to achieve accurate access control before multi-dimensional data fusion, blocking unqualified data from entering the fusion process from the spatiotemporal dimensions, thereby ensuring the accuracy and reliability of multi-dimensional data fusion in the cloud and providing high-quality data support for IoT business decisions.
[0015] 5. By conducting edge-side filtering accuracy assessment to obtain edge-side filtering timeliness deviation results, and when the edge-side filtering timeliness deviation results do not meet the edge-side filtering timeliness qualification conditions, edge computing node filtering performance assessment is conducted to obtain an average filtering performance assessment index. Edge computing node filtering optimization can only be performed when the filtering performance is qualified, i.e., when the average filtering performance assessment index is greater than 0. Compared with existing technologies that mostly rely on "fixed periodic triggering" or "single deviation index triggering," such as directly starting optimization as soon as filtering accuracy is detected as substandard, ignoring the filtering performance foundation of the edge node itself, this helps to achieve precise triggering and risk control of edge-side filtering optimization, avoiding ineffective optimization or over-optimization that damages the stability of edge nodes; by conducting cloud real-time filtering accuracy assessment to obtain an average cloud real-time filtering accuracy deviation index, and when the average cloud real-time filtering accuracy deviation index is greater than 0, cloud real-time filtering optimization is performed, which helps to achieve dynamic iteration and precise improvement of cloud filtering quality, ensuring the final quality before multi-dimensional data fusion. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the structure of the real-time filtering system for multi-dimensional data fusion based on the Internet of Things provided in this embodiment of the invention;
[0018] Figure 2 This is an overview diagram of the real-time filtering system for multi-dimensional data fusion based on the Internet of Things provided in this embodiment of the invention;
[0019] Figure 3 This is a schematic diagram of the hierarchical storage quality analysis of the multi-dimensional data fusion real-time filtering system based on the Internet of Things provided in this embodiment of the invention;
[0020] Figure 4 This is a flowchart of a real-time filtering method for multi-dimensional data fusion based on the Internet of Things provided in an embodiment of the present invention. Detailed Implementation
[0021] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0022] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0023] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0024] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0025] This invention provides a real-time filtering system based on multi-dimensional data fusion using the Internet of Things (IoT). For example... Figure 1 The schematic diagram of the IoT-based multi-dimensional data fusion real-time filtering system shown includes: a data flow interruption monitoring module, a filtering accuracy monitoring module, and a filtering accuracy optimization judgment module. The data flow interruption monitoring module is used to determine the impact of data flow interruptions during the fusion and filtering of multi-dimensional data collected from a specified scenario based on IoT terminals. This results in a judgment that reflects the data flow interruption situation during the multi-dimensional data collection process and is transmitted to the corresponding module. By monitoring data flow interruptions, it is helpful to achieve risk prevention and control at the data collection source and accurate diversion of processing paths, preventing interrupted data from flowing into subsequent fusion stages and reducing "invalid data processing" (such as using interrupted data for fusion, leading to distorted results) from the source.
[0026] The filtering accuracy monitoring module is used to perform spatiotemporal alignment and correlation verification to assess the spatiotemporal alignment of multi-dimensional data if the received judgment result is no filtering accuracy assessment. Based on the output verification result, it determines whether to perform multi-dimensional data fusion to output fused multi-dimensional data. The filtering accuracy assessment is used to evaluate the filtering of multi-dimensional data at edge computing nodes and the cloud. Multi-dimensional data represents a collection of structured and unstructured data collected by IoT terminals (such as traffic flow and average vehicle speed detected by video radar when monitoring traffic flow). By monitoring filtering accuracy, it helps to achieve spatiotemporal consistency verification and fusion feasibility assurance, ensuring that only "spatiotemporally qualified" data enters the fusion process and avoiding the failure of fusion results due to data misalignment.
[0027] The filter accuracy optimization judgment module is used to determine whether to perform tiered storage quality analysis based on the output result if the received judgment result is to perform a filter accuracy assessment. This is to reflect the quality compliance of multi-dimensional data in the tiered storage stage. If tiered storage quality analysis is performed, spatiotemporal alignment and correlation verification are performed after the analysis is completed. Otherwise, spatiotemporal alignment and correlation verification are performed directly. By performing filter accuracy optimization judgment, it is helpful to achieve in-depth quality optimization and compliance control of the storage integration link throughout the entire process.
[0028] like Figure 2 The diagram shown is an overview of the real-time filtering system for multi-dimensional data fusion based on the Internet of Things provided in this embodiment of the invention. Figure 2 It can be seen that: by determining the impact of data flow interruption, a judgment result is generated. If no filtering accuracy assessment is performed, spatiotemporal alignment and correlation verification are directly performed to obtain spatiotemporal alignment data. If a filtering accuracy assessment is performed, the edge-side filtering accuracy assessment is performed first, followed by obtaining the edge-side filtering timeliness deviation result. If the edge-side filtering timeliness deviation result meets the edge-side filtering timeliness qualification conditions, the cloud real-time filtering accuracy assessment is performed; otherwise, the edge computing node filtering performance assessment is performed to verify the qualification of the edge computing node filtering performance. When the average filtering performance assessment index is greater than 0, the edge computing node filtering is optimized. After the edge-side filtering accuracy assessment is completed, the cloud real-time filtering accuracy assessment is performed to obtain the average cloud real-time filtering accuracy deviation index. When the monitored average cloud real-time filtering accuracy deviation index is greater than 0, cloud real-time filtering optimization is performed; otherwise, spatiotemporal alignment and correlation verification are performed after the tiered storage quality analysis is qualified.
[0029] Prior to the design of the IoT-based multi-dimensional data fusion real-time filtering system provided in this application, a database is established to store various types of preset data. The database includes, but is not limited to, the request sending duration of preset data streams, preset interruption interference correlation values, etc. The various values are directly set by technical personnel. The database adopts an efficient and reliable storage architecture to ensure the integrity and accessibility of the data. The database uses the mature storage architecture of MySQL or PostgreSQL, which can organize and manage data in the form of tables. Through this structured storage method, various types of data can be accurately classified and quickly retrieved.
[0030] In this embodiment, the data flow interruption monitoring module, the filtering accuracy monitoring module, and the filtering accuracy optimization judgment module influence and transmit to each other, which helps to realize a closed loop of quality control for the entire process of IoT multi-dimensional data from collection to fusion. The modules are linked through "judgment result transmission" and "processing result feedback" to ensure that there is a quality control node at each step of data processing, thereby improving the availability and reliability of IoT multi-dimensional data.
[0031] Furthermore, the specific process for determining the impact of data stream interruptions to generate a judgment result reflecting the data stream interruption situation during multi-dimensional data acquisition is as follows: An assessment of the data stream interruption situation during multi-dimensional data acquisition based on IoT terminals is performed based on the node-request sending duration to obtain a data stream interruption impact judgment value; this value quantifies the interference of data stream interruptions on the accuracy of multi-dimensional data filtering; the data stream interruption impact judgment value is represented by the difference between the node-request sending duration and the preset data stream request sending duration; the IoT terminal monitored by the timer sends a time synchronization request to the edge computing node. The duration corresponding to the time of the request is used as the node-request sending duration; by inputting the data stream interruption impact judgment value into the preset data stream interruption related set, the interruption interference correlation value is output; a judgment is made based on the interruption interference correlation value; if the interruption interference correlation value is greater than the preset interruption interference correlation value, the corresponding judgment result is marked as undergoing filtering accuracy evaluation, otherwise, the corresponding judgment result is marked as not undergoing filtering accuracy evaluation, and the corresponding multi-dimensional data is marked as qualified multi-dimensional data for spatiotemporal alignment and correlation verification. The preset interruption interference correlation value is represented by the average interruption interference correlation value of historical time periods.
[0032] It should be added that, in the embodiments of this application, preset data flow interruption related sets and preset storage quality related sets with mapping relationships are presented, which are retrieved from the database. These mapping relationships have dynamic characteristics, and can achieve both one-to-one mapping between single parameters and many-to-one mapping between multiple parameters and single parameters. Specifically, information such as the combination of data flow interruption impact judgment values, qualified cloud real-time filtering accuracy deviation index, and tiered storage quality analysis values collected by preset personnel within a historical time period are first input into a machine learning model (e.g., a decision tree model) used to reflect the importance of features. Through the feature splitting operation of the model, the corresponding weights or data can be obtained, namely, interruption interference correlation values, tiered storage quality qualification assessment scores, etc. Next, the historical time period data is associated and matched with the corresponding weights or data to obtain preset data flow interruption related sets, preset storage quality related sets, etc. By inputting the combination of real-time collected data flow interruption impact judgment value, qualified cloud real-time filtering accuracy deviation index and tiered storage quality analysis value into the corresponding preset data flow interruption related sets and preset storage quality related sets, according to the preset mapping relationship, the corresponding interruption interference relatedness value and tiered storage quality qualified assessment score, whose value range is limited to the 0-1 interval, are output.
[0033] In this embodiment, the degree of interference related to the interruption is obtained by determining the impact of data flow interruption. When the degree of interference related to the interruption is greater than the preset degree of interference related to the interruption, the accuracy of filtering is evaluated; otherwise, spatiotemporal alignment and correlation verification are performed. This helps to achieve risk classification and control of multi-dimensional data processing in the Internet of Things, as well as a balance between efficiency and quality. By prioritizing the investigation of the potential impact of interruptions on the filtering process (such as the distortion of the input data of the filtering logic due to the interruption), the "filtering results with risks" are prevented from entering the fusion, laying the foundation for outputting reliable fused data.
[0034] Furthermore, the specific process for determining the impact of data stream interruptions to generate a judgment result reflecting the data stream interruption situation during multi-dimensional data acquisition is as follows: Obtain the edge computing node load to reflect the load status of the edge computing nodes; if the detected edge computing node load is not greater than the preset edge computing node load, directly perform spatiotemporal alignment and correlation verification; if the detected edge computing node load is greater than the preset edge computing node load, make a judgment based on the obtained average interruption duration: if the average interruption duration is greater than the preset average interruption duration, mark the corresponding judgment result as requiring filtering accuracy evaluation; otherwise, mark the corresponding judgment result as not requiring filtering accuracy evaluation, and mark the corresponding multi-dimensional data as qualified multi-dimensional data for spatiotemporal alignment and correlation verification; the average interruption duration is used to quantify the interference of data stream interruptions on the multi-dimensional data filtering accuracy, wherein the preset average interruption duration is represented by the average of the average interruption duration over historical time periods.
[0035] Specifically, the average duration of the interruption can be obtained through the following methods:
[0036] ;
[0037] In the formula, T i d K represents the average duration of interruptions in the data stream corresponding to the i-th multi-dimensional data within the window. i d T represents the number of interruptions in the data stream corresponding to the i-th multi-dimensional data within the window. i,K d This represents the duration of the kth interruption of the data stream corresponding to the i-th multi-dimensional data, where i represents the number of the data stream corresponding to the multi-dimensional data, i=1,2,3,,,m, m represents the total number of data streams corresponding to the multi-dimensional data, and d represents the statistics of the interrupted dimension, used to mark the interrupted dimension.
[0038] The interruption count of the data stream corresponding to the i-th multi-dimensional data within the window is represented by the total number of times when the arrival interval of consecutive data packets within the window (the interval from the data packet corresponding to the multi-dimensional data to the arrival window) is greater than the preset arrival interval set by the preset personnel, which is monitored by a timer and a network traffic monitor. The duration of the interruption is represented by the difference between the timestamp of the first data packet recorded when the interruption ends and the timestamp of the last data packet when the interruption occurs. If no interruption occurs within the window, the average duration of the interruption is defined as 0.
[0039] In this embodiment, when multiple tasks need to be processed simultaneously on an edge computing node (such as real-time traffic flow statistics, violation capture, and traffic incident recognition in smart city traffic flow monitoring), the average duration of interruptions is used to determine when the load on the edge computing node exceeds the preset edge computing node load. This helps to achieve efficient resource allocation for multi-task processing on edge nodes, ensure the quality of core tasks, and maintain stable system operation. It avoids waste of computing resources, achieves "on-demand allocation," thereby alleviating node load pressure and maintaining stable system operation.
[0040] Furthermore, the filtering accuracy assessment involves sequentially performing edge-side filtering accuracy assessment and cloud-based real-time filtering accuracy assessment. Edge-side filtering accuracy assessment evaluates the filtering status of multi-dimensional data at edge computing nodes before transmission to the cloud. Cloud-based real-time filtering accuracy assessment reflects the filtering status of multi-dimensional data in the cloud. The specific process of edge-side filtering accuracy assessment is as follows: The filtering accuracy of multi-dimensional data at edge computing nodes is quantified based on the total amount of multi-dimensional data processed by the edge computing nodes, and an edge filtering value is output. The total amount of multi-dimensional data filtered by the edge computing nodes within a specified edge-side assessment time period is monitored using a network traffic monitor and used as the edge filtering value. The specified edge-side assessment time period refers to the preset time period corresponding to the edge-side filtering accuracy assessment. A judgment is made based on the average edge filtering value and the preset edge filtering value. If the average edge filtering value is greater than the preset edge filtering value, a qualified edge filtering data volume prompt is sent, and filtering time assessment is performed; otherwise, edge computing node filtering optimization is performed. The preset edge filtering value is represented by the average of the average edge filtering values over historical time periods. The average edge filtering value represents the average edge filtering value obtained after performing a preset number of edge-side filtering accuracy assessments.
[0041] Specifically, the filtering time assessment reflects the filtering time of edge computing nodes under the condition that the amount of filtered data is acceptable. The specific process is as follows: Based on the node-to-data filtering time and the preset data filtering time, the filtering timeliness of multi-dimensional data at the edge computing node is quantified, and the edge-side filtering timeliness assessment result is output. The edge-side filtering timeliness assessment result is represented by the ratio of the preset data filtering time and the node-to-data filtering time. The edge-side filtering timeliness assessment result reflects the acceptable level of filtering timeliness of multi-dimensional data at the edge computing node. The average time from receiving multi-dimensional data to completing filtering at the edge computing node is monitored by a timer and used as the node-to-data filtering time. The preset data filtering time is represented by the average value of the node-to-data filtering time over a historical time period. The degree of deviation is determined based on the edge-side filtering timeliness assessment result and the preset edge-side filtering timeliness assessment result. The process quantifies and outputs the edge-side filtering timeliness deviation result. This deviation is represented by the difference between the edge-side filtering timeliness assessment result and the preset edge-side filtering timeliness assessment result, reflecting the deviation in the filtering timeliness of multi-dimensional data at the edge computing node. The process then determines whether the output edge-side filtering timeliness deviation result meets the edge-side filtering timeliness qualification criteria: if it does, a real-time cloud-based filtering accuracy assessment is performed; otherwise, an edge computing node filtering performance assessment is conducted to verify the qualification of the edge computing node's filtering performance. The edge-side filtering timeliness qualification criteria indicate that the edge-side filtering timeliness deviation result is greater than 0. The preset edge-side filtering timeliness assessment result is represented by the average value of edge-side filtering timeliness assessment results over a historical time period.
[0042] In this embodiment, an average edge filtering value is obtained by evaluating the accuracy of edge-side filtering. When the average edge filtering value is greater than a preset edge filtering value, the filtering time consumption is evaluated to obtain the edge-side filtering timeliness deviation result. If the edge-side filtering timeliness deviation result does not meet the edge-side filtering timeliness qualification condition, the edge computing node filtering performance is evaluated to verify the qualification of the edge computing node filtering performance. Through multi-level verification from "filtering accuracy" to "filtering timeliness" and then to "node performance adaptability", it helps to achieve closed-loop control of the quality, efficiency and performance of edge-side data filtering, thereby ensuring the real-time performance and availability of edge-side output data.
[0043] Furthermore, the specific process for evaluating the filtering performance of edge computing nodes is as follows: A filtering performance evaluation index is obtained by quantifying the deviation between the filtering efficiency value and the preset filtering efficiency value; this deviation quantification represents performing a difference calculation. The filtering performance evaluation index is used to assess the pass / fail status of the filtering performance of the edge computing node itself. The amount of multi-dimensional data filtered by the edge computing node during a specified edge-side evaluation period is monitored by a network traffic monitor and used as the filtering efficiency value. The preset filtering efficiency value is represented by the average value of filtering efficiency values over historical periods. A judgment is made based on the average filtering performance evaluation index. If the average filtering performance evaluation index is greater than 0, a filtering performance pass / fail prompt is sent, and edge computing node filtering optimization is performed in the next adjacent specified edge-side evaluation period; otherwise, a filtering performance abnormality alarm is sent. The average filtering performance evaluation index represents the average value of the filtering performance evaluation index obtained after performing a preset number of edge computing node filtering performance evaluations.
[0044] It should be added that edge computing node filtering optimization is used to buffer data flow interruptions caused by network jitter or outages, providing a complete data window for lightweight edge preprocessing and improving filtering accuracy. The specific process is as follows: The edge-side filtering timeliness deviation results and the bandwidth occupied by multi-dimensional data monitored through the Linux platform are input into a preset filtering correlation model, which outputs the buffer queue length of the edge computing node; within the preset range of the buffer queue length, the amplitude corresponding to the output value of the buffer queue length of the edge computing node is used as the adjustment amount, and the buffer queue length of the edge computing node is incremented based on the initial buffer queue length of the edge computing node; the edge node buffer queue is used to temporarily store multi-dimensional data to be processed (such as vehicle data collected by traffic radar). Traffic data (camera video frames) need to be matched with the real-time data input rate and node processing capabilities. Increasing the cache queue length helps prevent excessive cache resource consumption and enables dynamic load adaptation of edge computing node cache resources and smooth data flow. Each increment operation obtains the corresponding edge-side filtering timeliness deviation result. When the monitored edge-side filtering timeliness deviation result meets the edge-side filtering timeliness qualification conditions, edge computing node filtering optimization is stopped, and a real-time cloud filtering accuracy assessment is performed. If the edge computing node cache queue length reaches the maximum increment number preset by the pre-defined personnel, and the edge-side filtering timeliness deviation result still does not meet the edge-side filtering timeliness qualification conditions, an edge-side filtering anomaly prompt is sent.
[0045] It should be added that, specifically, the embodiments of this application provide a preset correlation analysis model (such as a random forest model) retrieved from the database, including a preset filtering correlation model, a preset cloud filtering correlation model, a preset hierarchical storage correlation model, etc. By inputting the combination of edge-side filtering timeliness deviation results and multi-dimensional data bandwidth occupied by preset personnel within a historical time period, the combination of cloud real-time filtering accuracy deviation index and the number of multi-dimensional data recorded in the cloud, the combination of qualified cloud real-time filtering accuracy deviation index, and the combination of hierarchical storage quality analysis value and the amount of multi-dimensional data into the corresponding preset filtering correlation model, preset cloud filtering correlation model, preset hierarchical storage correlation model, etc., which are used to reflect the importance of features, feature splitting can be performed by the model to obtain the corresponding cache queue length output value, cloud parallel unit number output value, partition replica number output value, etc.
[0046] It should be added that the training data (a combination of historical time-period edge filtering timeliness deviation results and multi-dimensional data bandwidth usage, a combination of cloud real-time filtering accuracy deviation index and the number of multi-dimensional data recorded in the cloud, a qualified cloud real-time filtering accuracy deviation index, a combination of hierarchical storage quality analysis values and the amount of multi-dimensional data, cache queue length output value, cloud parallel unit number output value, partition replica number output value, etc.) can be randomly divided into training and validation sets. The training set is input into a preset correlation analysis model (such as a random forest model) to obtain the preset correlation analysis model for training. The re-acquired combination of edge filtering timeliness deviation results and multi-dimensional data bandwidth usage, cloud real-time filtering accuracy deviation index and the number of multi-dimensional data recorded in the cloud, a qualified cloud real-time filtering accuracy deviation index, hierarchical storage quality analysis values and the amount of multi-dimensional data, etc., are input into the preset correlation analysis model for training to obtain the corresponding cache queue length output value, cloud parallel unit number output value, partition replica number output value, etc.
[0047] In this embodiment, an average filtering performance evaluation index is obtained by evaluating the filtering performance of edge computing nodes. Edge computing node filtering optimization is only allowed when the average filtering performance evaluation index is greater than 0, which helps to achieve accurate and necessary control over edge computing node filtering optimization. By linking edge computing node filtering performance evaluation and edge computing node filtering optimization, it helps to achieve closed-loop iteration and continuous evolution of edge computing node filtering capabilities, thereby improving the adaptability and stability of edge computing node filtering performance. This enables edge node filtering performance to dynamically adapt to changes in business needs, achieving a virtuous cycle of "problems can be located, optimizations can be verified, and capabilities can be evolved".
[0048] Furthermore, the specific process of cloud-based real-time filtering accuracy assessment is as follows: The degree of deviation is quantified based on residual outliers in the cloud and preset residual outliers, outputting a cloud-based real-time filtering accuracy deviation index; the cloud-based real-time filtering accuracy deviation index is represented by the difference between residual outliers in the cloud and preset residual outliers; the data volume of repeated multi-dimensional data and the total amount of cloud-filtered data within a specified cloud-based filtering assessment period are monitored using a network traffic monitor, and their ratio is used as the residual outlier in the cloud, which is represented by the average value of residual outliers in historical time periods; the specified cloud-based filtering assessment period refers to the preset time period corresponding to the cloud-based real-time filtering accuracy assessment; a judgment is made based on the average cloud-based real-time filtering accuracy deviation index: if the average cloud-based real-time filtering accuracy deviation index is greater than 0, cloud-based real-time filtering optimization is performed in the next adjacent specified cloud-based filtering assessment period; otherwise, tiered storage quality analysis is performed; the cloud-based real-time filtering accuracy deviation index is used to quantify the unqualified filtering of multi-dimensional data in the cloud; the average cloud-based real-time filtering accuracy deviation index represents the average value of the cloud-based real-time filtering accuracy deviation index obtained after performing a preset number of cloud-based real-time filtering accuracy assessments.
[0049] It should be added that cloud-based real-time filtering optimization is used to improve the throughput, alignment accuracy, and anti-jump capability of cloud-based real-time filtering. The specific process is as follows: The cloud-based real-time filtering accuracy deviation index and the number of multi-dimensional data recorded in the cloud are input into a preset cloud-based filtering correlation model, and the output value of the number of cloud parallel units is output. Within the preset range of the number of cloud parallel units, the amplitude corresponding to the output value of the number of cloud parallel units is used as an adjustment amount. Based on the initial number of cloud parallel units, the number of cloud parallel units is incremented. This helps to achieve elastic adaptation of cloud computing resources and dynamic matching of the needs of real-time processing of massive data, thereby improving the throughput and response timeliness of cloud-based real-time data processing. Each increment operation obtains the corresponding cloud-based real-time filtering accuracy deviation index. When the monitored cloud-based real-time filtering accuracy deviation index is not greater than 0, cloud-based real-time filtering optimization is stopped, and hierarchical storage quality analysis is performed. If the number of cloud parallel units reaches the maximum number of increments set in advance by preset personnel, and the cloud-based real-time filtering accuracy deviation index is still greater than 0, a cloud-based filtering optimization anomaly prompt is sent.
[0050] In this embodiment, an average real-time cloud filtering accuracy deviation index is obtained by performing a real-time cloud filtering accuracy assessment. When the average real-time cloud filtering accuracy deviation index is greater than 0, real-time cloud filtering optimization is performed; otherwise, tiered storage quality analysis is performed. This helps to achieve accurate quality distribution and efficient resource allocation of cloud data throughout the entire link from filtering to storage. The interaction and interconnection between the real-time cloud filtering accuracy assessment and real-time cloud filtering optimization help to achieve "dynamic closed-loop iteration" and long-term stability assurance of the real-time cloud filtering capability. This interaction allows the cloud filtering capability to dynamically adapt to changes in data characteristics (such as a surge in traffic data during morning and evening rush hours or sudden data format adjustments), ensuring that the cloud outputs high-quality filtered data over the long term.
[0051] like Figure 3 The diagram shown is a schematic representation of the hierarchical storage quality analysis of a multi-dimensional data fusion real-time filtering system based on the Internet of Things provided in an embodiment of the present invention. Figure 3 It can be seen that: when the average tiered storage quality qualification assessment score meets the storage quality qualification conditions, spatiotemporal alignment and correlation verification are performed; otherwise, tiered storage quality optimization is performed, that is, the number of partition replicas in the real-time layer is increased by using the magnitude corresponding to the output value of the number of partition replicas in the real-time layer as the adjustment amount.
[0052] Furthermore, the specific process of tiered storage quality analysis is as follows: Obtain tiered storage quality analysis values reflecting the quality compliance status of multi-dimensional data during the tiered storage stage; monitor the amount of multi-dimensional data successfully written and stored from edge computing nodes to the real-time layer using a network traffic monitor, and express the ratio as the tiered storage quality analysis value; input the qualified cloud real-time filtering accuracy deviation index and the tiered storage quality analysis value into a preset storage quality correlation set, and output the tiered storage quality compliance assessment score; determine whether the obtained average tiered storage quality compliance assessment score meets the storage quality compliance conditions; if it does, perform spatiotemporal alignment and correlation. Verification is performed; otherwise, tiered storage quality optimization is performed in the next adjacent specified tiered storage time period. The average tiered storage quality pass assessment score represents the average of the tiered storage quality pass assessment scores obtained after performing a preset number of tiered storage quality analyses. The specified tiered storage time period represents the preset time period corresponding to the tiered storage quality analysis. The qualified cloud real-time filtering accuracy deviation index is represented by the cloud real-time filtering accuracy deviation index that is not greater than the preset cloud real-time filtering accuracy deviation. The storage quality pass condition indicates that the average tiered storage quality pass assessment score is within the preset tiered storage pass range. The preset tiered storage pass range is set in advance by preset personnel and includes the endpoints of the upper and lower limits of the range.
[0053] Specifically, tiered storage quality optimization is used to improve the fault tolerance of real-time layer data and ensure the reliability of tiered storage data. The specific process is as follows: Within a preset range corresponding to the number of partition replicas in the real-time layer, the number of partition replicas in the real-time layer is increased based on the initial number of partition replicas. This helps improve the access efficiency, availability, and fault tolerance of real-time layer data. Real-time layer data needs to support high-concurrency access (such as multiple edge nodes simultaneously querying real-time traffic flow, and multiple cloud service modules synchronously calling device status). Increasing the number of partition replicas helps reduce cross-node transmission latency, thereby helping to achieve continuous and stable operation and timely response of IoT real-time services. This represents the number of replicas in each partition after partitioning the multi-dimensional data of the real-time layer. Each increment operation yields a corresponding tiered storage quality pass / fail score. When the score meets the storage quality pass / fail conditions, tiered storage quality optimization stops, and spatiotemporal alignment and correlation verification are performed. If the number of increments corresponding to the number of partition replicas reaches the preset maximum number of increments set by pre-defined personnel, and the tiered storage quality pass / fail score still does not meet the storage quality pass / fail conditions, a tiered storage quality warning is sent. The partition replica count output value is obtained by inputting the qualified cloud real-time filtering accuracy deviation index, the tiered storage quality analysis value, and the data volume of multi-dimensional data monitored by the network traffic monitor into a preset tiered storage correlation model.
[0054] In this embodiment, an average hierarchical storage quality pass assessment score is obtained by performing hierarchical storage quality analysis. If the average hierarchical storage quality pass assessment score meets the storage quality pass conditions, spatiotemporal alignment and correlation verification are performed; otherwise, hierarchical storage quality optimization is performed. This helps to achieve comprehensive linkage between hierarchical storage quality analysis, spatiotemporal alignment and correlation verification, and hierarchical storage quality optimization, thereby improving the reliability of hierarchical storage data and the effectiveness of spatiotemporal alignment verification, ensuring the availability of hierarchical storage data, and enhancing the accuracy of spatiotemporal alignment verification results. This helps to provide high-quality data assurance for multi-dimensional data fusion of the Internet of Things and downstream business decisions.
[0055] Furthermore, the specific process of spatiotemporal alignment and correlation verification is as follows: Spatiotemporal alignment data is acquired, including spatial alignment pass values representing the spatial alignment compliance of multi-dimensional data and temporal alignment pass values representing the temporal alignment compliance of multi-dimensional data; the ratio of the number of successful matches between the spatial identifiers corresponding to the multi-dimensional data and preset spatial identifiers to the total number of spatial alignments is used as the spatial alignment pass value; the ratio of the number of successful matches between the temporal identifiers corresponding to the multi-dimensional data and preset temporal identifiers to the total number of temporal alignments is used as the temporal alignment pass value; it is determined whether the spatiotemporal alignment data meets the spatiotemporal alignment pass conditions; if the spatiotemporal alignment data meets the spatiotemporal alignment pass conditions, multi-dimensional data fusion is performed; otherwise, a fusion anomaly prompt is sent; the spatiotemporal alignment pass conditions indicate that the spatial alignment pass value is greater than the preset spatial pass value, and the temporal alignment pass value is greater than the preset temporal pass value; multi-dimensional data fusion means inputting the corresponding multi-dimensional data into a preset fusion model (such as the ImageBind model) for data fusion output, where the preset spatial pass value is represented by the average of the spatial alignment pass values over a historical time period, and the preset temporal pass value is represented by the average of the temporal alignment pass values over a historical time period.
[0056] like Figure 4 The diagram shows a flowchart of a real-time filtering method for multi-dimensional data fusion based on the Internet of Things (IoT) provided in this application embodiment. This method is applied to a real-time filtering system for multi-dimensional data fusion based on the IoT and includes:
[0057] Data stream interruption monitoring: During the fusion and filtering of multi-dimensional data collected from a specified scenario based on IoT terminals, the impact of data stream interruption is determined to generate a judgment result reflecting the data stream interruption situation during the multi-dimensional data collection process and transmit it to the corresponding module.
[0058] Filtering accuracy monitoring: If the received judgment result is that no filtering accuracy assessment is performed, then spatiotemporal alignment and correlation verification is performed to evaluate the spatiotemporal alignment of multi-dimensional data. Based on the output verification result, it is determined whether to perform multi-dimensional data fusion to output the fused multi-dimensional data.
[0059] Filtering accuracy optimization judgment: If the received judgment result is to perform a filtering accuracy assessment, then based on the output result, it is determined whether to perform tiered storage quality analysis to reflect the quality compliance of multi-dimensional data in the tiered storage stage. If tiered storage quality analysis is performed, then spatiotemporal alignment and correlation verification are performed after the analysis is completed; otherwise, spatiotemporal alignment and correlation verification are performed directly.
[0060] In this embodiment, spatiotemporal aligned data is obtained by performing spatiotemporal alignment and correlation verification. When the spatial alignment qualification value is greater than the preset spatial qualification value and the temporal alignment qualification value is greater than the preset temporal qualification value, the corresponding multi-dimensional data is input into the preset fusion model for data fusion output. This helps to ensure the controllability of input quality and the effectiveness of fusion results in the fusion of multi-dimensional data in the Internet of Things, thereby helping to improve the reliability of fused data and ensure the accuracy and stability of downstream applications of the Internet of Things. For example, in smart transportation, the fused vehicle trajectory data can ensure the accuracy of traffic signal scheduling.
[0061] In summary, the embodiments of this application, by determining the impact of data flow interruptions and deciding whether to perform a filtering accuracy assessment based on the output judgment results, help to accurately locate data quality problems and avoid the transmission of quality risks throughout the entire chain. If a filtering accuracy assessment is not performed, spatiotemporal alignment and correlation verification are performed, and the output verification results determine whether to perform multi-dimensional data fusion, which helps to achieve "resource allocation on demand" and balance data processing efficiency. If a filtering accuracy assessment is performed, the output results determine whether to perform hierarchical storage quality analysis. If hierarchical storage quality analysis is performed, spatiotemporal alignment and correlation verification are performed after the analysis is completed; otherwise, spatiotemporal alignment and correlation verification are performed directly, which helps to clear quality obstacles in the fusion process, ensure the accuracy of dimensional data fusion, and thus improve the accuracy of multi-dimensional data fusion and real-time filtering based on the Internet of Things, solving the problem of low accuracy in multi-dimensional data fusion and real-time filtering based on the Internet of Things in the prior art.
[0062] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A multi-dimensional data fusion real-time filtering system based on Internet of Things, characterized in that, include: Data stream interruption monitoring module, filtering accuracy monitoring module, and filtering accuracy optimization judgment module; The data stream interruption monitoring module is used to determine the impact of data stream interruption during the fusion and filtering of multi-dimensional data collected from a specified scenario based on IoT terminals, so as to generate a determination result reflecting the data stream interruption situation during the multi-dimensional data collection process and transmit it to the corresponding module. The filtering accuracy monitoring module is used to perform spatiotemporal alignment and correlation verification to evaluate the spatiotemporal alignment of multi-dimensional data if the received judgment result is that no filtering accuracy assessment is performed. Based on the output verification result, it determines whether to perform multi-dimensional data fusion to output fused multi-dimensional data. The filtering accuracy assessment is used to evaluate the filtering status of multi-dimensional data at edge computing nodes and the cloud. The multi-dimensional data represents the collection of structured and unstructured data collected by IoT terminals. The filtering accuracy optimization judgment module is used to determine whether to perform hierarchical storage quality analysis based on the output result if the received judgment result is to perform filtering accuracy evaluation, so as to reflect the quality qualification of multi-dimensional data in the hierarchical storage stage. If hierarchical storage quality analysis is performed, spatiotemporal alignment and correlation verification are performed after the analysis is completed; otherwise, spatiotemporal alignment and correlation verification are performed directly. The filtering accuracy assessment refers to performing edge-side filtering accuracy assessment and cloud-based real-time filtering accuracy assessment sequentially. The edge-side filtering accuracy assessment is used to evaluate the filtering of multi-dimensional data at edge computing nodes before it is transmitted to the cloud. The cloud-based real-time filtering accuracy assessment is used to reflect the filtering status of multi-dimensional data in the cloud. The specific process for evaluating the accuracy of edge-side filtering is as follows: The accuracy of filtering multi-dimensional data at edge computing nodes is quantified based on the total amount of multi-dimensional data processed by edge computing nodes, and the edge filtering value is output. The judgment is based on the average edge filtering value and the preset edge filtering value. If the average edge filtering value is greater than the preset edge filtering value, an edge filtering data volume qualified prompt is sent and the filtering time is evaluated. Otherwise, the edge computing node filtering is optimized. The filtering time evaluation is used to reflect the filtering time of edge computing nodes under the condition that the amount of data filtered is sufficient. The specific process is as follows: Based on the node-data filtering time and the preset data filtering time, the timeliness of filtering data at the edge computing node is quantified in multiple dimensions, and the edge-side filtering timeliness evaluation results are output. The edge-side filtering timeliness evaluation results are used to reflect the passability of the filtering timeliness of multi-dimensional data at the edge computing node; The node-data filtering time refers to the average time from when the edge computing node receives multi-dimensional data to when it completes the filtering process. The deviation is quantified based on the edge-side filtering timeliness assessment results and the preset edge-side filtering timeliness assessment results, and the edge-side filtering timeliness deviation results are output. The edge-side filtering timeliness deviation results are used to reflect the deviation of the filtering timeliness of multi-dimensional data at the edge computing node; Determine whether the output edge-side filtering aging deviation result meets the edge-side filtering aging qualification conditions: If the edge-side filtering timeliness deviation result meets the edge-side filtering timeliness qualification condition, then perform a cloud-based real-time filtering accuracy assessment. If the edge-side filtering timeliness deviation result does not meet the edge-side filtering timeliness qualification condition, then the edge computing node filtering performance evaluation is performed to verify the qualification of the edge computing node filtering performance. The edge-side filtration aging qualification condition indicates that the edge-side filtration aging deviation result is greater than 0.
2. The real-time filtering system for multi-dimensional data fusion based on the Internet of Things according to claim 1, characterized in that, The specific process for determining the impact of data stream interruption to generate a determination result reflecting the data stream interruption situation during multi-dimensional data acquisition is as follows: An assessment of data stream interruption during multi-dimensional data collection based on IoT terminals is conducted based on node-request sending time, and the impact value of data stream interruption is obtained. The data stream interruption impact judgment value is used to quantify the interference of data stream interruption on the accuracy of multi-dimensional data filtering; The node-request sending duration refers to the duration corresponding to when the IoT terminal sends a time synchronization request to the edge computing node; By inputting the data stream interruption impact judgment value into a preset data stream interruption related set, the interruption interference correlation value is output. If the interruption interference correlation value is greater than the preset interruption interference correlation value, the corresponding judgment result will be marked as a filter accuracy assessment; otherwise, the corresponding judgment result will be marked as no filter accuracy assessment, and the corresponding multi-dimensional data will be marked as qualified multi-dimensional data for spatiotemporal alignment and correlation verification. 3.The Internet of Things based multi-dimension data fusion real-time filtering system according to claim 1, characterized in that, The specific process for determining the impact of data stream interruption to generate a determination result reflecting the data stream interruption situation during multi-dimensional data acquisition is as follows: Obtain the edge computing node load, which reflects the load status of the edge computing nodes; If the load of the edge computing node is not greater than the preset edge computing node load, spatiotemporal alignment and correlation verification are performed directly. If the load on an edge computing node is detected to be greater than the preset edge computing node load, the determination is made based on the obtained average interruption duration: If the average duration of the interruption is greater than the preset average duration of the interruption, the corresponding judgment result will be marked as a filter accuracy assessment; otherwise, the corresponding judgment result will be marked as no filter accuracy assessment will be performed, and the corresponding multi-dimensional data will be marked as qualified multi-dimensional data for spatiotemporal alignment and correlation verification. The average duration of the interruption is used to quantify the impact of data stream interruptions on the accuracy of multi-dimensional data filtering. 4.The Internet of Things based multi-dimension data fusion real-time filtering system according to claim 1, wherein, The specific process for evaluating the filtering performance of the edge computing node is as follows; A filtration performance evaluation index is obtained by quantifying the degree of deviation between the filtration efficiency value and the preset filtration efficiency value. The filtering performance evaluation index is used to assess the passability of the filtering performance of the edge computing node itself; The judgment is based on the average filtering performance evaluation index. If the average filtering performance evaluation index is greater than 0, a filtering performance qualified prompt is sent, and edge computing node filtering optimization is performed in the next adjacent specified edge side evaluation time period. Otherwise, an abnormal filtering performance alarm is sent. The edge computing node filtering optimization is used to buffer data flow breaks caused by network jitter or interruptions, providing a complete data window for lightweight edge preprocessing and improving filtering accuracy. The specific process is as follows: Input the edge-side filtering timeliness deviation results and the bandwidth occupied by multi-dimensional data into the preset filtering correlation model, and output the cache queue length of the edge computing node; Within the preset range of the cache queue length, the magnitude corresponding to the output value of the cache queue length of the edge computing node is used as the adjustment amount, and the cache queue length of the edge computing node is incremented based on the initial cache queue length of the edge computing node. Each increment operation is performed, and the corresponding edge-side filtering timeliness deviation result is obtained. When the monitored edge-side filtering timeliness deviation result meets the edge-side filtering timeliness qualification condition, the edge computing node filtering optimization is stopped, and the cloud real-time filtering accuracy assessment is performed. If the cache queue length of the edge computing node reaches the corresponding maximum increment count, and the edge-side filtering timeliness deviation result still does not meet the edge-side filtering timeliness qualification condition, an edge-side filtering abnormality prompt will be sent. 5.The Internet of Things based multi-dimension data fusion real-time filtering system according to claim 4, characterized in that, The specific process for evaluating the accuracy of real-time cloud-based filtering is as follows: The degree of deviation is quantified based on residual outliers in the cloud and preset residual outliers in the cloud, and the real-time cloud filtering accuracy deviation index is output. The residual outlier in the cloud is represented by the ratio of the amount of repeated multi-dimensional data within a specified cloud filtering evaluation period to the total amount of cloud filtered data. The judgment is based on the average cloud real-time filtering accuracy deviation index: if the average cloud real-time filtering accuracy deviation index is greater than 0, cloud real-time filtering optimization is performed in the next adjacent specified cloud filtering evaluation time period; otherwise, tiered storage quality analysis is performed. The cloud-based real-time filtering accuracy deviation index is used to quantify the failure of multi-dimensional data filtering in the cloud. The cloud-based real-time filtering optimization is used to improve the cloud-based real-time filtering throughput, alignment accuracy, and anti-jump capability. The specific process is as follows: Input the real-time cloud filtering accuracy deviation index and the number of multi-dimensional data recorded in the cloud into the preset cloud filtering correlation model, and output the number of cloud parallel units. Within the preset range of the number of parallel units in the cloud, the magnitude corresponding to the output value of the number of parallel units in the cloud is used as the adjustment amount, and the number of parallel units in the cloud is incremented based on the initial number of parallel units in the cloud. Each increment operation is performed, and the corresponding cloud real-time filtering accuracy deviation index is obtained. When the monitored cloud real-time filtering accuracy deviation index is not greater than 0, the cloud real-time filtering optimization is stopped, and the hierarchical storage quality analysis is performed. If the number of parallel units in the cloud reaches the corresponding maximum increment number, and the real-time filtering accuracy deviation index in the cloud is still greater than 0, a cloud filtering optimization anomaly prompt will be sent. 6.The Internet of Things based multi-dimension data fusion real-time filtering system according to claim 5, characterized in that, The specific process of the tiered storage quality analysis is as follows: Obtain tiered storage quality analysis values that reflect the quality compliance status of multi-dimensional data during the tiered storage phase; The hierarchical storage quality analysis value is represented by the ratio of the amount of multi-dimensional data sent from the edge computing node to the real-time layer that was successfully written and stored to the total amount of data sent by the edge computing node. Input the qualified cloud real-time filtering accuracy deviation index and the tiered storage quality analysis value into the preset storage quality correlation set, and output the tiered storage quality qualified assessment score; Determine whether the obtained average tiered storage quality qualification assessment score meets the storage quality qualification conditions. If it does, perform spatiotemporal alignment and correlation verification; otherwise, perform tiered storage quality optimization in the next adjacent specified tiered storage time period. The qualified cloud real-time filtering accuracy deviation index is represented by a cloud real-time filtering accuracy deviation index that is not greater than the preset cloud real-time filtering accuracy deviation. The storage quality qualification condition means that the average tiered storage quality qualification assessment score is within the preset tiered storage qualification range. 7.The Internet of Things based multi-dimension data fusion real-time filtering system according to claim 6, characterized in that, The tiered storage quality optimization is used to improve the fault tolerance of real-time layer data and ensure the reliability of tiered storage data. The specific process is as follows: Within a preset range corresponding to the number of partition replicas in the real-time layer, the number of partition replicas in the real-time layer is incremented based on the magnitude corresponding to the output value of the number of partition replicas in the real-time layer. The number of partition replicas represents the number of replicas for each partition after the multi-dimensional data of the real-time layer is partitioned. Each increment operation obtains the corresponding tiered storage quality qualification assessment score. When the tiered storage quality qualification assessment score meets the storage quality qualification conditions, the tiered storage quality optimization is stopped, and spatiotemporal alignment and correlation verification are performed. If the number of increments corresponding to the number of partition replicas reaches the preset maximum number of replica increments, and the tiered storage quality qualification assessment score still does not meet the storage quality qualification conditions, a tiered storage quality warning is sent. The output value of the number of partition replicas is obtained by inputting the data volume of qualified cloud real-time filtering accuracy deviation index, hierarchical storage quality analysis value and multi-dimensional data into a preset hierarchical storage correlation model.
8. The real-time filtering system for multi-dimensional data fusion based on the Internet of Things according to claim 7, characterized in that, The specific process of spatiotemporal alignment and correlation verification is as follows: Acquire spatiotemporal alignment data, which includes spatial alignment qualification values for characterizing the spatial alignment qualification of multi-dimensional data, and time alignment qualification values for characterizing the temporal alignment qualification of multi-dimensional data. If the spatiotemporal aligned data meets the spatiotemporal alignment qualification conditions, multi-dimensional data fusion will be performed; otherwise, a fusion anomaly prompt will be sent. The spatiotemporal alignment qualification condition means that the spatial alignment qualification value is greater than the preset spatial qualification value, and the temporal alignment qualification value is greater than the preset temporal qualification value; The multi-dimensional data fusion refers to inputting the corresponding multi-dimensional data into a preset fusion model to perform data fusion and output fused multi-dimensional data.
9. A real-time filtering method for multi-dimensional data fusion based on the Internet of Things, applied to the real-time filtering system for multi-dimensional data fusion based on the Internet of Things as described in any one of claims 1-8, characterized in that, include: During the fusion and filtering of multi-dimensional data collected from a specified scenario based on IoT terminals, the impact of data stream interruption is determined to generate a determination result reflecting the data stream interruption during the multi-dimensional data collection process and transmit it to the corresponding module. If the received judgment result is that no filtering accuracy assessment is performed, then spatiotemporal alignment and correlation verification is performed to evaluate the spatiotemporal alignment of multi-dimensional data. Based on the output verification result, it is determined whether to perform multi-dimensional data fusion to output the fused multi-dimensional data. If the received judgment result is to perform a filtering accuracy assessment, then based on the output result, it is determined whether to perform tiered storage quality analysis to reflect the quality compliance of multi-dimensional data in the tiered storage stage. If tiered storage quality analysis is performed, then spatiotemporal alignment and correlation verification are performed after the analysis is completed; otherwise, spatiotemporal alignment and correlation verification are performed directly.
Citation Information
Patent Citations
A method, device, and readable storage medium for filtering data to be fused.
CN110674125B
Real-time data filtering method and system based on multi-dimensional feature fusion
CN120449108B
Railway passenger train operation state monitoring method and system based on artificial intelligence
CN120534406A
Heterogeneous data processing optimization system based on intelligent edge computing
CN120872595A