Real-time processing method of multi-modal data for internet of things gateway

By monitoring data latency and diagnostic model version consistency in real time at the IoT gateway, eliminating redundant information, and performing link compensation and feedback adjustments, the transmission latency and link instability issues of IoT gateways in multimodal and high-bandwidth scenarios are resolved, thereby improving the real-time performance and reliability of the system.

CN121691294BActive Publication Date: 2026-06-30BEIJING KINGDOES RFID TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING KINGDOES RFID TECH
Filing Date
2025-12-16
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

IoT gateways suffer from transmission delays and link instability in multimodal, high-bandwidth scenarios, resulting in insufficient real-time decision-making capabilities. Existing technologies struggle to cope with dynamic network environments and multi-source data redundancy, leading to low processing efficiency and insufficient reliability.

Method used

By monitoring data latency in real time at the gateway, diagnosing model version consistency, eliminating redundant information, and performing link compensation and feedback adjustments, including model version diagnosis, information redundancy calculation, and decision cycle adjustment, an optimization closed loop is formed to improve system real-time performance and reliability.

Benefits of technology

It enables refined assessment of latency, improves system real-time performance and reliability, prevents misjudgments due to a single abnormal indicator, and ensures the effectiveness of decision-making and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121691294B_ABST
    Figure CN121691294B_ABST
Patent Text Reader

Abstract

This invention relates to the field of IoT data processing technology, and more particularly to a real-time multimodal data processing method for IoT gateways. The method includes: acquiring multi-source operational data and encapsulating it into a standard operational dataset; acquiring high-bandwidth modal data from the standard operational dataset; when the maximum link latency time exceeds a preset latency anomaly threshold, acquiring the version effective fingerprint of the target model and the fingerprint of the local model; calculating information redundancy when obtaining a model version anomaly diagnosis result; acquiring a corrected latency time and comparing it with a preset normal baseline latency interval; and adjusting the current decision cycle when the corrected latency time is greater than or equal to the upper limit of the normal baseline latency interval. This invention achieves accurate processing of high-bandwidth modal data by real-time monitoring of latency, diagnosing model version consistency, eliminating redundant information, and performing link adjustments, thereby improving the real-time performance and reliability of IoT gateway data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) data processing technology, and in particular to a method for real-time processing of multimodal data for IoT gateways. Background Technology

[0002] In the field of IoT gateways and multimodal data processing, with the increase in the number of sensors and the frequency of data acquisition, the amount of video, audio and sensor operation data that the gateway needs to process is growing rapidly, resulting in transmission delay and link instability problems for high-bandwidth modal data, which in turn affects the system's real-time decision-making capabilities.

[0003] Existing technologies typically rely on fixed acquisition frequencies or single latency judgment methods, which are difficult to cope with dynamic network environments and multi-source data redundancy. They cannot accurately identify redundant information or detect model version anomalies in a timely manner, resulting in low processing efficiency, large latency fluctuations, and insufficient system reliability.

[0004] Therefore, a method is needed that can monitor data latency in real time at the gateway, diagnose model version consistency, remove redundant information, and perform link compensation and feedback adjustment to improve the real-time performance, stability, and reliability of multimodal data processing in IoT gateways.

[0005] Chinese Patent Publication No. CN114374603A discloses a data processing method and an API gateway based on an API gateway. When a first message is received, the API gateway queries configuration information and obtains a conversion rule from the configuration information for converting the first message into a second message. The API gateway converts the first message into the second message according to the conversion rule. The API gateway then sends the second message to the recipient of the first message.

[0006] However, existing IoT data processing technologies have poor adaptability, real-time performance, and reliability in multimodal, high-bandwidth IoT scenarios. Summary of the Invention

[0007] To address this issue, the present invention provides a method for real-time processing of multimodal data for IoT gateways. By introducing model version diagnosis and redundant information removal, the method achieves dynamic latency correction and decision cycle feedback adjustment for high-bandwidth data, thereby solving the problem of low data processing efficiency in gateways.

[0008] To achieve the above objectives, the present invention provides a method for real-time processing of multimodal data for an Internet of Things (IoT) gateway, comprising:

[0009] Within the current acquisition cycle, multi-source operational data is acquired synchronously at a set initial acquisition frequency, and the multi-source operational data is encapsulated into a standard operational dataset;

[0010] High-bandwidth modal data from the standard running dataset is obtained to perform latency determination;

[0011] When the maximum link latency in high-bandwidth modal data exceeds a preset latency anomaly threshold, obtain the target model's version effective fingerprint and the local model fingerprint to diagnose the model version.

[0012] When the model version anomaly diagnosis result is obtained, the information redundancy is calculated based on the total information content of the high-bandwidth modal data and the preset redundancy information judgment standard.

[0013] Based on the information redundancy, the corrected delay time of the high-bandwidth modal data after link delay compensation is obtained, and the corrected delay time is compared with the preset normal reference delay range;

[0014] When the corrected delay time is greater than or equal to the upper limit of the normal baseline delay interval, the current decision cycle is adjusted based on the comparison result between the corrected delay time and the normal baseline delay interval.

[0015] Furthermore, the diagnostic model version process includes:

[0016] Obtain the version effective fingerprint of the target model and the fingerprint of the local model, and determine whether the version effective fingerprint and the local model fingerprint are consistent;

[0017] If the effective fingerprint of the version is inconsistent with the fingerprint of the local model, obtain the maximum cache queue length of high-bandwidth modal data in the current collection period and compare it with the corresponding threshold.

[0018] If the first cache comparison result is obtained, threshold verification is performed on the maximum time drift and network link status of high-bandwidth modal data;

[0019] If the first threshold verification result is obtained, the execution time of this fingerprint comparison is obtained, and the deviation is calculated based on the execution time and the average value of the historical normal comparison cycle to obtain the deviation magnitude.

[0020] When the deviation exceeds the set anomaly identification threshold, an anomaly diagnosis result for the model version is obtained.

[0021] Furthermore, the process for threshold verification of the maximum time drift and network link status of high-bandwidth modal data includes:

[0022] The maximum time drift is compared with the preset time drift threshold;

[0023] If the maximum time drift is less than the preset time drift threshold, a normal comparison result of the drift time is obtained, the average packet loss rate in the current collection period is obtained, and the average packet loss rate is compared with the preset packet loss rate threshold.

[0024] If the average packet loss rate is less than the preset packet loss rate threshold, a link stability comparison result is obtained, corresponding to the first threshold verification result.

[0025] Furthermore, the process for calculating information redundancy includes:

[0026] Obtain the total number of bytes in the standard running dataset to determine the total amount of information;

[0027] Based on the initial sampling frequency and the corresponding sampling points, the standard operating dataset is divided into several standard operating data subsets corresponding to the sampling points.

[0028] Adjacent standard running data subsets are combined into a continuous running dataset, the structural similarity value of each continuous running dataset is calculated, and compared with a preset similarity threshold;

[0029] If all structural similarity values ​​are greater than the preset similarity threshold, the structural similarity results are obtained, and the non-redundant information content is calculated.

[0030] The information redundancy is calculated based on the amount of non-redundant information and the total amount of information.

[0031] Furthermore, the process for calculating non-redundant information includes:

[0032] Preset decision weight coefficients for different feature parameters of the continuously running dataset;

[0033] The weighted calculation result is calculated based on the continuously running dataset and the decision weight coefficients;

[0034] The weighted calculation result is compared with a preset weighted threshold to obtain all continuously running datasets that are greater than the preset weighted threshold, and then updated to the decision key region set.

[0035] Furthermore, the process for calculating non-redundant information also includes:

[0036] The first running dataset in each of the aforementioned key decision-making regions is selected as the reference baseline dataset;

[0037] The reference benchmark dataset within each of the aforementioned decision key region sets is compared with each subsequent target continuous running dataset for each pair of execution decision items to obtain the decision variation set of each target continuous running dataset relative to the reference benchmark dataset.

[0038] Based on the decision change set and the preset local motion weight coefficient, the relative change intensity corresponding to each target continuous running dataset is calculated;

[0039] The relative change intensity is compared with a preset change intensity threshold. If the relative change intensity is lower than the preset change intensity threshold, the change portion of the target continuously running dataset relative to the reference baseline dataset is defined as redundant information and removed to obtain the non-redundant information amount.

[0040] Furthermore, the set of decision changes includes:

[0041] The series variation term obtained based on numerical differences;

[0042] The directional variation term obtained based on the directional difference;

[0043] The decision variation set is composed of the decision parameters corresponding to the aforementioned number-level variation term and direction variation term.

[0044] Furthermore, the process of comparing the corrected delay time with a preset normal reference delay interval includes:

[0045] The corrected delay time is compared with the preset normal baseline delay range;

[0046] If the corrected delay time is greater than or equal to the upper limit of the normal baseline delay range, it is determined that there is still an abnormal delay after compensation;

[0047] If the corrected delay time is within the normal baseline delay range, there is no abnormal delay after compensation;

[0048] If the corrected delay time is less than or equal to the lower limit of the normal baseline delay interval, it is determined that there is overcompensation, and step rollback is performed.

[0049] Furthermore, the process for adjusting the current decision-making cycle includes:

[0050] The delay deviation step size is calculated based on the difference between the corrected delay time and the normal reference delay interval.

[0051] Based on the aforementioned delay deviation step size, the duration of the current decision cycle is adjusted to obtain the corrected decision cycle;

[0052] The revised decision cycle is used as the duration of the next standard running dataset decision cycle.

[0053] Furthermore, the process for calculating the delay deviation step size includes:

[0054] Calculate the ratio of the difference between the corrected delay time and the normal baseline delay interval to the current decision cycle to obtain the intermediate ratio;

[0055] The intermediate ratio is calculated using a preset compensation correction function to obtain the delay deviation step size.

[0056] Compared with existing technologies, the advantages of this invention lie in that it encapsulates multi-source operational data into a standard operational dataset and extracts high-bandwidth modal data from it for latency assessment. When a latency exceeding a preset latency anomaly threshold is detected, instead of simple compensation, it first performs model version diagnosis. By comparing the effective fingerprint of the version with the local model fingerprint and combining multiple checks such as cache queue length, maximum time drift, and network link status, it accurately locates whether the latency originates from a local model version anomaly. After confirming the model version anomaly, it further introduces information redundancy calculation. By analyzing the structural similarity of a subset of standard operational data and based on preset decision weight coefficients... The method screens a set of key decision-making regions, then uses the decision change set and relative change intensity to identify and eliminate redundant information, thereby quantifying the proportion of effective information in high-bandwidth data. Based on this information redundancy, the method calculates the corrected delay time and compares it with the normal baseline delay interval, achieving a refined assessment of the delay status. Finally, if it is determined that there is still an abnormal delay after compensation, the current decision cycle is dynamically adjusted by calculating the delay deviation step size based on the difference between the corrected delay time and the baseline interval, forming an optimized closed loop that adjusts based on data value density and the actual processing capacity of the system, thereby improving the real-time performance of the system while ensuring the effectiveness of decision-making.

[0057] Furthermore, from the initial judgment of version fingerprint consistency to the check of data backlog, to the verification of external link quality, and finally to the verification of system load by comparing the time consumption of the operation itself, a tight logical chain is formed, which improves the reliability of the diagnostic conclusion and prevents misjudgment due to a single abnormal indicator. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating the real-time multimodal data processing method for an IoT gateway according to an embodiment of the present invention.

[0059] Figure 2 This is a logic decision diagram for a diagnostic model version of an embodiment of the present invention;

[0060] Figure 3 This is a schematic diagram illustrating the process of calculating information redundancy in an embodiment of the present invention;

[0061] Figure 4 This is a flowchart illustrating the process of adjusting the current decision-making cycle in an embodiment of the present invention. Detailed Implementation

[0062] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0063] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0064] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0065] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0066] Please see Figure 1 The diagram shown is a flowchart illustrating a real-time multimodal data processing method for an IoT gateway according to an embodiment of the present invention. The present invention provides a real-time multimodal data processing method for an IoT gateway, comprising:

[0067] Step S1: Within the current acquisition cycle, acquire multi-source running data synchronously at the set initial acquisition frequency, and encapsulate the multi-source running data into a standard running dataset;

[0068] Step S2: Obtain high-bandwidth modal data from the standard running dataset to perform latency determination;

[0069] Step S3: When the maximum link latency in the high-bandwidth modal data exceeds the preset latency anomaly threshold, obtain the target model's version activation fingerprint and the local model fingerprint to diagnose the model version.

[0070] Step S4: When the model version anomaly diagnosis result is obtained, calculate the information redundancy based on the total information content of the high-bandwidth modal data and the preset redundancy information judgment standard;

[0071] Step S5: Based on the information redundancy, obtain the corrected delay time of the high-bandwidth modal data after link delay compensation processing, and compare the corrected delay time with the preset normal reference delay range;

[0072] Step S6: When the corrected delay time is greater than or equal to the upper limit of the normal reference delay interval, adjust the current decision cycle according to the comparison result between the corrected delay time and the normal reference delay interval.

[0073] In this embodiment, the current acquisition cycle is the basic time unit for completing one round of data acquisition, processing, and decision-making; the initial acquisition frequency is the set optimal data acquisition rate for various data sources.

[0074] Multi-source operational data consists of raw data from different devices and of different natures;

[0075] The standard running dataset is a data package with uniform specifications and spatiotemporal alignment obtained by packaging raw data of different formats and sources according to a unified timestamp, data structure and format.

[0076] High-bandwidth modal data is a type of data with huge data volume and high transmission and processing overhead among multi-source data. It is usually a continuous streaming media, such as high-definition video, audio, and high-precision radar point cloud.

[0077] Delay determination involves calculating the maximum time difference between data generation and reception by the gateway, i.e., the maximum link delay time.

[0078] The preset delay anomaly threshold is a preset threshold for determining whether the system has abnormal delays;

[0079] The target model is the latest version of the algorithm model that is expected to be deployed on the gateway for data analysis;

[0080] The local model is a version of the algorithm model deployed on the gateway for data analysis. It may or may not be the same as the target model.

[0081] The version activation fingerprint and the local model fingerprint are identifiers that represent the unique identity of the model. They are usually a string of cryptographic hash values ​​generated by the model version number, structure hash value, training data batch, etc. The version activation fingerprint is the model version identifier issued by the cloud management platform and specified by the gateway to be used. The local model fingerprint is the model version identifier that the gateway actually loads and runs.

[0082] Total information volume refers to the volume of raw high-bandwidth modal data within the current acquisition cycle, such as the number of bytes.

[0083] Information redundancy is a quantitative indicator that represents the proportion of high-bandwidth data that can be removed without affecting the accuracy of the decision within the current decision objective and context.

[0084] The corrected delay time is a theoretical calculation value; its core logic is: since there is a considerable amount of data redundancy, the delay caused by processing this redundant data can be compressed or canceled out.

[0085] The normal baseline latency range refers to the reasonable range of latency that high-bandwidth data should have under ideal and healthy conditions.

[0086] The current decision-making cycle is the time allowed from receiving data to making a control decision;

[0087] Adjusting the current decision-making cycle is the feedback control action of the entire closed loop.

[0088] By encapsulating multi-source operational data into a standard operational dataset and extracting high-bandwidth modal data from it for latency assessment, when a latency exceeding a preset latency anomaly threshold is detected, instead of simple compensation, a model version diagnosis is first performed. This involves comparing the effective fingerprint of the current version with the fingerprint of the local model, combined with multiple checks such as cache queue length, maximum time drift, and network link status, to accurately pinpoint whether the latency originates from an anomaly in the local model version. After confirming the model version anomaly, information redundancy calculation is further introduced. By analyzing the structural similarity of a subset of the standard operational data and filtering a set of key decision regions based on preset decision weight coefficients, redundant information is identified and eliminated using the decision change set and relative change intensity, thereby quantifying the proportion of effective information in the high-bandwidth data. Based on this information redundancy, the method calculates the corrected latency time and compares it with the normal baseline latency interval, achieving a refined assessment of latency status. Finally, if an abnormal latency is still found after compensation, the current decision cycle is dynamically adjusted based on the difference between the corrected latency time and the baseline interval by calculating the latency deviation step size, forming an optimization closed loop that adjusts based on data value density and the actual processing capacity of the system, thereby improving the system's real-time performance while ensuring the effectiveness of decision-making.

[0089] See Figure 2 As shown, it is a logic decision diagram for model version diagnosis in an embodiment of the present invention;

[0090] Specifically, the process for diagnosing the model version includes:

[0091] Obtain the version effective fingerprint of the target model and the fingerprint of the local model, and determine whether the version effective fingerprint and the local model fingerprint are consistent;

[0092] If the effective fingerprint of the version is inconsistent with the fingerprint of the local model, obtain the maximum cache queue length of high-bandwidth modal data in the current collection period and compare it with the corresponding threshold.

[0093] If the first cache comparison result is obtained, threshold verification is performed on the maximum time drift and network link status of high-bandwidth modal data;

[0094] If the first threshold verification result is obtained, the execution time of this fingerprint comparison is obtained, and the deviation is calculated based on the execution time and the average value of the historical normal comparison cycle to obtain the deviation magnitude.

[0095] When the deviation exceeds the set anomaly identification threshold, an anomaly diagnosis result for the model version is obtained.

[0096] In this embodiment, the version effective fingerprint of the target model and the local model fingerprint are obtained, and the version effective fingerprint and the local model fingerprint are compared.

[0097] When the effective fingerprint of the version and the local model fingerprint are consistent, the model version is consistent, the diagnostic process is terminated, and the delay problem is attributed to non-model factors.

[0098] If the version activation fingerprint and the local model fingerprint are inconsistent, proceed to the next step of cache queue check;

[0099] Obtain the maximum buffer queue length of high-bandwidth modal data within the current acquisition period, and compare it with the corresponding threshold, namely the queue backlog anomaly threshold. The queue backlog anomaly threshold is an empirically selected threshold, and in this embodiment, the queue backlog anomaly threshold is set to 5MB.

[0100] When the maximum cache queue length exceeds the queue backlog anomaly threshold, the second cache comparison result is obtained, the diagnostic process terminates, and it is determined that the model version anomaly has caused serious data backlog, triggering an alarm.

[0101] When the maximum cache queue length is less than or equal to the queue backlog exception threshold, the first cache comparison result is obtained, and the next step of threshold verification is performed.

[0102] Threshold verification is performed on the maximum time drift and network link status of high-bandwidth modal data:

[0103] The maximum time drift is compared with a preset time drift tolerance threshold, wherein the time drift tolerance threshold is an empirically selected threshold, and in this embodiment, the time drift tolerance threshold is 35 milliseconds.

[0104] When the maximum time drift exceeds the time drift tolerance threshold, an abnormal drift result is obtained, the diagnostic process terminates, and the data source timing abnormality is determined to be the main problem.

[0105] When the maximum time drift is less than or equal to the time drift tolerance threshold, a normal time drift comparison result is obtained, and the network link is further verified.

[0106] The average packet loss rate within the current collection period is obtained, and the average packet loss rate is compared with a preset packet loss rate anomaly threshold, wherein the packet loss rate anomaly threshold is an empirically selected threshold, and in this embodiment, the packet loss rate anomaly threshold is 0.2%;

[0107] When the average packet loss rate exceeds the abnormal packet loss rate threshold, the link instability result is obtained, the diagnostic process is terminated, and the poor network link quality is determined to be the main problem.

[0108] When the average packet loss rate is less than or equal to the abnormal packet loss rate threshold, the link stability comparison results are obtained.

[0109] When the drift time normal comparison result and the link stability comparison result are obtained, the first threshold verification result is summarized, and the next step of deviation calculation is performed.

[0110] The execution time of this fingerprint comparison is obtained, and the deviation is calculated based on the execution time and the average value of historical normal comparison cycles to obtain the deviation range. The deviation range is then compared with a preset anomaly identification threshold, wherein the anomaly identification threshold is an empirically selected threshold, and in this embodiment, the anomaly identification threshold is set to 150%.

[0111] When the deviation is less than or equal to the anomaly identification threshold, the comparison time is normal, the diagnostic process is terminated, and the model version is determined to be inconsistent but has not caused significant performance degradation.

[0112] When the deviation exceeds the anomaly identification threshold, an abnormal comparison time is obtained;

[0113] When the version fingerprint inconsistency, the first cache comparison result, the first threshold verification result, and the comparison time abnormal result are obtained in sequence, the model version abnormal diagnosis result is finally obtained.

[0114] From initial judgment of version fingerprint consistency to checking the degree of data backlog, to verifying the quality of external links, and finally verifying the system load by comparing the time consumption of the operation itself, a tight logical chain is formed, which improves the reliability of the diagnostic conclusion and prevents misjudgment due to the abnormality of a single indicator.

[0115] Specifically, the process for threshold verification of the maximum time drift and network link status of high-bandwidth modal data includes:

[0116] The maximum time drift is compared with the preset time drift threshold;

[0117] If the maximum time drift is less than the preset time drift threshold, a normal comparison result of the drift time is obtained, the average packet loss rate in the current collection period is obtained, and the average packet loss rate is compared with the preset packet loss rate threshold.

[0118] If the average packet loss rate is less than the preset packet loss rate threshold, a link stability comparison result is obtained, corresponding to the first threshold verification result.

[0119] In this embodiment, the maximum time drift refers to the maximum deviation between the actual arrival time interval of consecutive data units and the theoretical interval in a high-bandwidth data stream, reflecting the timing stability of the data stream itself.

[0120] The preset time drift threshold is a critical value for judging whether the time drift is abnormal; in this embodiment, it is 35 milliseconds.

[0121] The average packet loss rate is the proportion of high-bandwidth data packets lost during network transmission within the current collection period.

[0122] The preset packet loss rate threshold is a critical value for judging whether the network link is stable; in this embodiment, it is 0.2%.

[0123] By sequentially verifying the time drift of the data stream itself and the packet loss rate of network transmission, it is possible to clearly distinguish the delay caused by external problems such as unstable data source acquisition and jitter of the upstream network from the internal processing capability problems of the gateway; ensuring that complex redundancy analysis and compensation mechanisms are only initiated for delays caused by internal reasons.

[0124] See Figure 3 As shown, it is a flowchart illustrating the calculation of information redundancy in an embodiment of the present invention;

[0125] Specifically, the process for calculating information redundancy includes:

[0126] Obtain the total number of bytes in the standard running dataset to determine the total amount of information;

[0127] Based on the initial sampling frequency and the corresponding sampling points, the standard operating dataset is divided into several standard operating data subsets corresponding to the sampling points.

[0128] Adjacent standard running data subsets are combined into a continuous running dataset, the structural similarity value of each continuous running dataset is calculated, and compared with a preset similarity threshold;

[0129] If all structural similarity values ​​are greater than the preset similarity threshold, the structural similarity results are obtained, and the non-redundant information content is calculated.

[0130] The information redundancy is calculated based on the amount of non-redundant information and the total amount of information.

[0131] In this embodiment, the total number of bytes in the standard running dataset is obtained to determine the total amount of information; in this embodiment, the total number of bytes is 50 megabytes.

[0132] Based on the initial sampling frequency and the corresponding sampling points, the standard operating dataset is divided into several standard operating dataset subsets corresponding to the sampling points; for example, it is divided into 3 standard operating dataset subsets within a decision cycle.

[0133] Adjacent standard running data subsets are combined into a continuous running dataset, the structural similarity value of each continuous running dataset is calculated, and compared with a preset similarity threshold;

[0134] For example, the structural similarity values ​​of two consecutively running datasets are calculated to be 0.92 and 0.89, respectively. The preset similarity threshold is an empirically selected threshold with a value of 0.85.

[0135] If all structural similarity values ​​are greater than the preset similarity threshold, a structural similarity result is obtained; this embodiment meets this condition.

[0136] After obtaining the structural similarity results, redundant information within the current decision-making cycle is extracted and removed to obtain non-redundant information.

[0137] The information redundancy is calculated based on the amount of non-redundant information and the total amount of information. In this embodiment, the amount of non-redundant information after the elimination process is 28 megabytes. According to the formula Information Redundancy = 1 - (Non-redundant information / Total information), the result is 0.44.

[0138] After calculating the corresponding information redundancy, the method compresses the theoretical processing time occupied by high-bandwidth modal data proportionally based on this value.

[0139] By dividing the data into subsets and calculating the structural similarity of continuous sets, we first determine from a macroscopic perspective whether there are repetitive patterns in the data. Based on the confirmation of similarity, we then perform fine-grained extraction of redundant information, which reduces the computational cost of performing full-scale detailed analysis from the outset while ensuring the accuracy of the evaluation.

[0140] Specifically, the process for calculating non-redundant information includes:

[0141] Preset decision weight coefficients for different feature parameters of the continuously running dataset;

[0142] The weighted calculation result is calculated based on the continuously running dataset and the decision weight coefficients;

[0143] The weighted calculation result is compared with a preset weighted threshold to obtain all continuously running datasets that are greater than the preset weighted threshold, and then updated to the decision key region set.

[0144] In this embodiment, the continuously running dataset is divided into several parallel and non-overlapping feature parameter categories based on the correlation between its data content and the determination of abnormal equipment status, and a decision weight coefficient is preset for each category.

[0145] In this embodiment, the features of high-bandwidth modal data within a time window are divided into three categories:

[0146] Characteristic parameters that are highly correlated with historical failure modes, such as the energy amplitude of a specific frequency band, are preset to have a decision weight coefficient of 0.6 in this embodiment;

[0147] For characteristic parameters that fluctuate under normal operating conditions but have a certain potential correlation with anomalies, such as time-domain peak values, the decision weight coefficient is preset to 0.3 in this embodiment.

[0148] In this embodiment, the decision weight coefficient of a characteristic parameter that remains relatively stable under both normal and abnormal conditions, such as the signal baseline, is preset to 0.1.

[0149] The sum of the above weighting coefficients is 1.0;

[0150] For each continuously running dataset, extract its corresponding significant abnormal feature values, general fluctuation feature values ​​and background steady-state feature values, and calculate its weighted result according to the preset weight coefficients;

[0151] The calculation formula is: Weighted result = Significant abnormal characteristic value × 0.6 + General fluctuation characteristic value × 0.3 + Background steady-state characteristic value × 0.1;

[0152] For example, the three feature values ​​of the continuously running dataset C1 are 5.2, 8.1, and 1.0, respectively, and its weighted result is 5.2×0.6+8.1×0.3+1.0×0.1=3.12+2.43+0.1=5.65;

[0153] The three feature values ​​of the continuously running dataset C2 are 1.1, 7.9, and 1.0, respectively. The weighted result is 1.1×0.6+7.9×0.3+1.0×0.1=0.66+2.37+0.1=3.13;

[0154] The weighted calculation result of each continuously running dataset is compared with a preset weighting threshold; the weighting threshold is an empirically selected threshold used to identify data that has sufficient influence on decision-making, and in this embodiment, the value is 4.0.

[0155] If the weighted result of a dataset is greater than the weighted threshold, the dataset is determined to contain information that significantly contributes to the final anomaly detection, and it is filtered out.

[0156] If the weighted result of a dataset is less than or equal to the weighted threshold, then the information contained in the dataset is deemed to contribute little to the current decision and will not be filtered.

[0157] In this embodiment, the weighted result of dataset C1 is greater than 4.0, so it is acquired and updated to the decision critical region set; the weighted result of dataset C2 is less than 4.0, so it is not included in the critical region set; the subsequent redundancy removal operation will only be performed within the decision critical region set.

[0158] By pre-setting weights related to the final decision objective for different feature parameters and calculating weighted values, the system can automatically identify regions that contribute highly to the current judgment from massive amounts of data, thus establishing an importance index for the data.

[0159] Specifically, the process for calculating non-redundant information also includes:

[0160] The first running dataset in each of the aforementioned key decision-making regions is selected as the reference baseline dataset;

[0161] The reference benchmark dataset within each of the aforementioned decision key region sets is compared with each subsequent target continuous running dataset for each pair of execution decision items to obtain the decision variation set of each target continuous running dataset relative to the reference benchmark dataset.

[0162] Based on the decision change set and the preset local motion weight coefficient, the relative change intensity corresponding to each target continuous running dataset is calculated;

[0163] The relative change intensity is compared with a preset change intensity threshold. If the relative change intensity is lower than the preset change intensity threshold, the change portion of the target continuously running dataset relative to the reference baseline dataset is defined as redundant information and removed to obtain the non-redundant information amount.

[0164] In this embodiment, the decision-making key region set includes continuously running datasets C1 and C3, so C1 is selected as the reference benchmark dataset;

[0165] By comparing each pair of decision items in the reference baseline dataset C1 with the subsequent target continuously running dataset C3, the set of decision changes in C3 relative to C1 is obtained.

[0166] In this embodiment, three key characteristic parameters are selected for comparison: peak value, amplitude of main frequency components, and waveform factor.

[0167] The peak value is 10.0 in C1 and 10.1 in C3, with a magnitude variation of 0.1 and a direction variation of +1.

[0168] The amplitude of the main frequency component is 0.5 in C1 and 0.49 in C3, with a magnitude variation of 0.01 and a direction variation of -1.

[0169] The waveform factor has a value of 1.5 in C1 and 1.52 in C3, with a magnitude variation of 0.02 and a direction variation of +1.

[0170] Thus, the decision change set of C3 relative to C1 contains the order and direction change terms of the above three characteristic parameters;

[0171] Based on the aforementioned decision change set and the preset local motion weight coefficients, the relative change intensity of the target dataset C3 is calculated; in this embodiment, the preset local motion weight coefficients for the changes in peak value, main frequency component amplitude, and waveform factor are 0.6, 0.3, and 0.1, respectively.

[0172] The formula for calculating the relative change intensity S is: S equals the sum of the products of each level of change term and its corresponding weight coefficient;

[0173] Substituting the decision change set numerical calculation into C3: S equals 0.1 multiplied by 0.6, plus 0.01 multiplied by 0.3, plus 0.02 multiplied by 0.1, the calculation result is 0.065;

[0174] The calculated relative change intensity of 0.065 is compared with a preset change intensity threshold; the change intensity threshold is an empirically selected threshold, and in this embodiment, the value is 0.1.

[0175] If the relative change intensity is greater than or equal to the preset change intensity threshold of 0.1, the change in the target dataset is determined to be significant and it is retained.

[0176] If the relative change intensity is less than the preset change intensity threshold of 0.1, the change in the target dataset relative to the reference dataset is determined to be redundant information and is removed.

[0177] In this embodiment, the relative change intensity of C3 is 0.065, which is less than the change intensity threshold of 0.1. Therefore, the C3 dataset is determined to be redundant information and is removed.

[0178] The initial decision-making key region set includes C1 and C3, each with a data volume of approximately 16.7MB. After comparison and removal of C3, only the reference baseline dataset C1 is retained. Therefore, the amount of non-redundant information obtained after removing redundant information is 16.7MB.

[0179] By establishing a reference baseline and calculating the relative change intensity of subsequent data, it is possible to accurately identify data segments that, although within the critical region, have not undergone significant and effective changes relative to the baseline state. This allows for the removal of static or slightly variable redundant parts while retaining key decision-making characteristics.

[0180] Specifically, the set of decision changes includes:

[0181] The series variation term obtained based on numerical differences;

[0182] The directional variation term obtained based on the directional difference;

[0183] The decision variation set is composed of the decision parameters corresponding to the aforementioned number-level variation term and direction variation term.

[0184] In this embodiment, the decision variation set is a structured dataset used to accurately characterize the differences between the target dataset and the reference dataset across all key decision dimensions, and its composition is as follows:

[0185] The numerical variation term quantifies the absolute magnitude of the change in the numerical value of each decision parameter;

[0186] The direction of change term qualitatively describes the trend direction of the changes in the values ​​of each decision parameter, such as increasing, decreasing, or remaining unchanged;

[0187] The decision change set uses all the decision parameters involved in the comparison as an index, and each parameter corresponds to the sum set of its order of change and direction of change; this set completely defines all the quantitative change information from one state to another.

[0188] The decision change set decomposes data changes into two orthogonal dimensions: order and direction. This provides a structured input for calculating the relative intensity of change, enabling parameter changes of different features and dimensions to be uniformly measured and compared according to weights. This is the foundation for achieving automated, universal, and redundant judgment.

[0189] Specifically, the process of comparing the corrected delay time with a preset normal baseline delay range includes:

[0190] The corrected delay time is compared with the preset normal baseline delay range;

[0191] If the corrected delay time is greater than or equal to the upper limit of the normal baseline delay range, it is determined that there is still an abnormal delay after compensation;

[0192] If the corrected delay time is within the normal baseline delay range, there is no abnormal delay after compensation;

[0193] If the corrected delay time is less than or equal to the lower limit of the normal baseline delay interval, it is determined that there is overcompensation, and step rollback is performed.

[0194] In this embodiment, the process of comparing the corrected delay time with the preset normal baseline delay interval is as follows:

[0195] The preset normal baseline latency range is [90 milliseconds, 120 milliseconds], with a lower limit of 90 milliseconds and an upper limit of 120 milliseconds;

[0196] The calculated corrected delay time is compared with the normal reference delay range in this embodiment:

[0197] If the corrected delay time is greater than or equal to the upper limit of the normal baseline delay range, it is determined that there is still an abnormal delay after compensation, and a warning is issued.

[0198] If the corrected delay time is within the normal baseline delay range, i.e. greater than 90 milliseconds and less than 120 milliseconds, it is determined that there is no abnormal delay after compensation, and the compensation result is retained.

[0199] If the corrected delay time is less than or equal to the lower limit of the normal reference delay interval, it is determined that there is overcompensation, and step rollback is performed.

[0200] Specifically, the execution step rollback process involves reverting the standard running dataset to its uncompensated state to avoid compromising data validity.

[0201] By comparing the correction delay with an interval, three states are distinguished, effectively preventing the system from falling into an oscillation or unstable state due to an overly aggressive compensation strategy, thus enhancing the robustness of the control loop.

[0202] See Figure 4 As shown, it is a flowchart illustrating the process of adjusting the current decision-making cycle according to an embodiment of the present invention;

[0203] Specifically, the process for adjusting the current decision-making cycle includes:

[0204] The delay deviation step size is calculated based on the difference between the corrected delay time and the normal reference delay interval.

[0205] Based on the aforementioned delay deviation step size, the duration of the current decision cycle is adjusted to obtain the corrected decision cycle;

[0206] The revised decision cycle is used as the duration of the next standard running dataset decision cycle.

[0207] In this embodiment, the preset normal baseline delay interval is [lower limit, upper limit]; the difference between the corrected delay time and the upper limit is calculated; the ratio of this difference to the current decision cycle duration is calculated to obtain an intermediate ratio; this intermediate ratio is input into the preset linear compensation correction function f(x)=k×x for calculation, where k is a preset empirical coefficient, with an optimal value of 0.5; x is the intermediate ratio; the output value of the function is the delay deviation step size;

[0208] Based on the aforementioned delay deviation step size, the correction decision period is calculated using the following formula:

[0209] Corrected decision cycle = Current decision cycle × (1 + Delay deviation step size).

[0210] Based on the degree of deviation from the health baseline, the decision cycle step size that needs to be adjusted is quantitatively calculated, and the processing window is dynamically extended accordingly.

[0211] Specifically, the process for calculating the delay deviation step size includes:

[0212] Calculate the ratio of the difference between the corrected delay time and the normal baseline delay interval to the current decision cycle to obtain the intermediate ratio;

[0213] The intermediate ratio is calculated using a preset compensation correction function to obtain the delay deviation step size.

[0214] By normalizing the absolute time difference to a ratio relative to the current decision-making cycle, and then smoothing it using a conservative linear function, the resulting delay deviation step size reflects both the severity of the problem and avoids excessively large single adjustment magnitudes. This design ensures that system adjustments are gradual and progressive, which helps maintain the overall stability of the system and prevents drastic fluctuations.

[0215] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0216] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for real-time processing of multimodal data for an Internet of Things (IoT) gateway, characterized in that, include: Within the current acquisition cycle, multi-source operational data is acquired synchronously at a set initial acquisition frequency, and the multi-source operational data is encapsulated into a standard operational dataset; High-bandwidth modal data from the standard running dataset is obtained to perform latency determination; When the maximum link latency in high-bandwidth modal data exceeds a preset latency anomaly threshold, obtain the target model's version effective fingerprint and the local model fingerprint to diagnose the model version. When the model version anomaly diagnosis result is obtained, the information redundancy is calculated based on the total information content of the high-bandwidth modal data and the preset redundancy information judgment standard. Based on the information redundancy, the corrected delay time of the high-bandwidth modal data after link delay compensation is obtained, and the corrected delay time is compared with the preset normal reference delay range; When the corrected delay time is greater than or equal to the upper limit of the normal baseline delay interval, the current decision cycle is adjusted according to the comparison result between the corrected delay time and the normal baseline delay interval. The diagnostic model version process includes: Obtain the version effective fingerprint of the target model and the fingerprint of the local model, and determine whether the version effective fingerprint and the local model fingerprint are consistent; If the effective fingerprint of the version is inconsistent with the fingerprint of the local model, obtain the maximum cache queue length of high-bandwidth modal data in the current collection period and compare it with the corresponding threshold. If the first cache comparison result is obtained, threshold verification is performed on the maximum time drift and network link status of high-bandwidth modal data; If the first threshold verification result is obtained, the execution time of this fingerprint comparison is obtained, and the deviation is calculated based on the execution time and the average value of the historical normal comparison cycle to obtain the deviation magnitude. When the deviation exceeds the set anomaly identification threshold, an anomaly diagnosis result for the model version is obtained; The process for calculating information redundancy includes: Obtain the total number of bytes in the standard running dataset to determine the total amount of information; Based on the initial sampling frequency and the corresponding sampling points, the standard operating dataset is divided into several standard operating data subsets corresponding to the sampling points. Adjacent standard running data subsets are combined into a continuous running dataset, the structural similarity value of each continuous running dataset is calculated, and compared with a preset similarity threshold; If all structural similarity values ​​are greater than the preset similarity threshold, structural similarity results are obtained, and non-redundant information is calculated. Calculate the information redundancy based on the non-redundant information amount and the total information amount; The process for calculating non-redundant information includes: Preset decision weight coefficients for different feature parameters of the continuously running dataset; The weighted calculation result is calculated based on the continuously running dataset and the decision weight coefficients; The weighted calculation result is compared with a preset weighted threshold to obtain all continuously running datasets that are greater than the preset weighted threshold, and then updated to the decision key region set. The process for calculating non-redundant information also includes: The first running dataset in each of the aforementioned key decision-making regions is selected as the reference baseline dataset; The reference benchmark dataset within each of the aforementioned decision key region sets is compared with each subsequent target continuous running dataset for each pair of execution decision items to obtain the decision variation set of each target continuous running dataset relative to the reference benchmark dataset; Based on the decision change set and the preset local motion weight coefficient, the relative change intensity corresponding to each target continuous running dataset is calculated; The relative change intensity is compared with a preset change intensity threshold. If the relative change intensity is lower than the preset change intensity threshold, the change portion of the target continuous running dataset relative to the reference baseline dataset is defined as redundant information and removed to obtain the non-redundant information amount.

2. The method for real-time processing of multimodal data for an IoT gateway according to claim 1, characterized in that, The process for threshold verification of maximum time drift and network link status for high-bandwidth modal data includes: The maximum time drift is compared with the preset time drift threshold; If the maximum time drift is less than the preset time drift threshold, a normal comparison result of the drift time is obtained, the average packet loss rate in the current collection period is obtained, and the average packet loss rate is compared with the preset packet loss rate threshold. If the average packet loss rate is less than the preset packet loss rate threshold, a link stability comparison result is obtained, corresponding to the first threshold verification result.

3. The method for real-time processing of multimodal data for an IoT gateway according to claim 1, characterized in that, The set of decision changes includes: The series variation term obtained based on numerical differences; The directional variation term obtained based on the directional difference; The decision variation set is composed of the decision parameters corresponding to the aforementioned number-level variation term and direction variation term.

4. The method for real-time processing of multimodal data for an IoT gateway according to claim 1, characterized in that, The process of comparing the corrected delay time with a preset normal baseline delay range includes: The corrected delay time is compared with the preset normal baseline delay range; If the corrected delay time is greater than or equal to the upper limit of the normal baseline delay range, it is determined that there is still an abnormal delay after compensation; If the corrected delay time is within the normal baseline delay range, there is no abnormal delay after compensation; If the corrected delay time is less than or equal to the lower limit of the normal baseline delay interval, it is determined that there is overcompensation, and step rollback is performed.

5. The method for real-time processing of multimodal data for an IoT gateway according to claim 1, characterized in that, The process for adjusting the current decision-making cycle includes: The delay deviation step size is calculated based on the difference between the corrected delay time and the normal reference delay interval. Based on the aforementioned delay deviation step size, the duration of the current decision cycle is adjusted to obtain the corrected decision cycle; The revised decision cycle is used as the duration of the next standard running dataset decision cycle.

6. The method for real-time processing of multimodal data for an IoT gateway according to claim 5, characterized in that, The process for calculating the delay deviation step size includes: Calculate the ratio of the difference between the corrected delay time and the normal baseline delay interval to the current decision cycle to obtain the intermediate ratio; The intermediate ratio is calculated using a preset compensation correction function to obtain the delay deviation step size.

Citation Information

Patent Citations

  • Data processing method based on API gateway and API gateway

    CN114374603A

  • Internet of Things data processing method and system based on edge computing

    CN118748676A