Real-time data processing method and system based on edge computing
Through edge computing methods in the Internet of Things system, device fingerprints and timestamps are used to mark data, network status and bandwidth are combined to dynamically cut data, routing decisions and resource allocation are optimized, structured result sets are generated and geometric distance measurements are performed, which solves the data integrity and load imbalance problems of the Internet of Things system in high-frequency sampling scenarios, and achieves efficient real-time data processing and control stability.
Patent Information
- Application Number
- CN202511097026.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing IoT systems have difficulty dynamically adapting to network bandwidth fluctuations in high-frequency sampling scenarios, resulting in damaged data packet integrity and a lack of dual arbitration between physical layer timeliness characteristics and business rules. This causes high-value data overload or critical data timeout and discarding, routing decisions and computing resource allocation to be uncoordinated, load imbalance, and control instructions to rely on noisy historical benchmarks and lack geometric distance measurement, frequently leading to miscontrol accidents.
By marking device fingerprints and timestamps on terminal devices, the edge gateway performs variable time window cutting, calculates data validity period based on bandwidth utilization and business rules, determines routing decision strategies based on network status snapshots and real-time bandwidth, allocates computing resources, generates structured result sets, and obtains historical benchmark data sets through device fingerprints to calculate geometric distance metrics and generate hierarchical control instructions.
It improves transmission efficiency and resource utilization accuracy in high-concurrency real-time data processing systems, avoids key data timeouts or resource crowding, optimizes edge node load distribution, and enhances control instruction stability and system robustness.
Smart Images

Figure CN120602549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a real-time data processing method and system based on edge computing. Background Art
[0002] Amid the rapid development of the Internet of Things (IoT) and edge computing, a large number of terminal devices continuously generate high-density real-time data streams. IoT systems face three core bottlenecks. First, traditional fixed-window data segmentation mechanisms struggle to dynamically adapt to network bandwidth fluctuations, resulting in severe loss of packet integrity in high-frequency sampling scenarios. Second, data validity period calculation relies solely on single business rules or device sampling constraints, lacking a dual arbitration model that considers physical layer timeliness characteristics and business rules. This results in overloaded processing of high-value data or discarded critical data due to timeouts. Third, routing decisions and computing resource allocation operate independently, lacking a quantitative mapping mechanism from network state snapshots to routing health, and failing to dynamically allocate computing resources based on data timeliness priorities. This leads to load imbalances at edge nodes. Fourth, control command generation relies on noisy historical benchmark data, and the difference between structured results and benchmarks lacks a quantitative grading mechanism based on geometric distance metrics, frequently leading to miscontrol incidents. While existing solutions partially optimize data transmission efficiency or resource utilization, they fail to address the technical bottlenecks of adaptive network segmentation, dual-factor validity period arbitration, coordinated routing resource scheduling, and closed-loop control of benchmark purification, severely hindering the accuracy and robustness of real-time systems.
[0003] Therefore, it is necessary to provide a real-time data processing method and system based on edge computing to solve the above technical problems. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides a real-time data processing method and system based on edge computing, which achieves the beneficial effects of optimizing the efficiency of real-time data transmission and processing and improving utilization.
[0005] The present invention provides a real-time data processing method based on edge computing, comprising:
[0006] S1: The terminal device continuously collects raw data streams based on the preset physical sampling period. After receiving the raw data streams, the edge gateway marks the raw data streams with device fingerprints and timestamps to obtain marked data packets.
[0007] S2: Perform variable time window segmentation on the marked data packets based on real-time network status data to generate segmented data blocks containing network status snapshots and timestamp sequences;
[0008] S3: Calculates the data validity period of the segmented data blocks based on the physical sampling frequency of the terminal device and preset business rules, and generates processing priorities based on bandwidth utilization;
[0009] S4: Determine a routing decision strategy based on the network status snapshot, the real-time bandwidth in the real-time network status data, and a preset decision rule mapping table, allocate computing resources based on processing priorities, and generate a structured result set based on the routing decision strategy and computing resource processing segmented data blocks;
[0010] S5: Obtain a historical benchmark data set based on the device fingerprint, calculate the geometric distance measurement between the structured result set and the historical benchmark data set, generate hierarchical control instructions based on the geometric distance measurement, and push them to the terminal device after encryption.
[0011] Preferably, in step S2, the window length of the variable time window cutting is determined by the following steps:
[0012] Calculate the network bandwidth standard deviation within the latest multiple consecutive preset bandwidth monitoring periods;
[0013] The bandwidth utilization term is obtained by dividing the real-time bandwidth in the real-time network status data by the preset maximum theoretical bandwidth of the network and then multiplying the result by the preset bandwidth weight factor;
[0014] Multiply the network bandwidth standard deviation by the preset bandwidth fluctuation adjustment coefficient to obtain the bandwidth fluctuation term;
[0015] The bandwidth utilization term and the bandwidth fluctuation term are added together to obtain the time window determination value;
[0016] The window length is determined based on the time window decision value and the preset window length constraint rules.
[0017] Preferably, in step S3, the step of calculating the data validity period includes:
[0018] Get the physical sampling frequency of the terminal device and divide it by 2 to get the physical validity period.
[0019] Identify the data type of the segmented data block and obtain the business validity period of the segmented data block based on the preset business rules;
[0020] Compare the physical validity period and the business validity period, and take the smaller value as the data validity period.
[0021] Preferably, in step S3, the step of calculating the processing priority includes:
[0022] Extract bandwidth utilization from real-time network status data;
[0023] Calculate the remaining valid time ratio of the segmented data block based on the survival time and data validity period of the segmented data block;
[0024] Generate processing priority based on bandwidth utilization, remaining effective time ratio and preset data value coefficient.
[0025] Preferably, in step S4, the step of determining the routing decision strategy includes:
[0026] Quantify the network status snapshot into a routing health score through a preset decision rule mapping table;
[0027] Extract real-time bandwidth from real-time network status data and calculate the bandwidth adaptability coefficient by combining the route health score and the preset route weight factor;
[0028] Select routing decision strategy based on broadband adaptability coefficient.
[0029] Preferably, in step S4, the routing decision strategy includes the target processing node type, the transmission path selection strategy and the disaster recovery backup mechanism.
[0030] Preferably, in step S4, the step of generating a structured result set includes:
[0031] Based on the timestamps in the segmented data blocks, the time dimension feature vector is output after correcting the data clock offset;
[0032] According to the device fingerprint in the segmented data block, the preset device database is queried to obtain the geographic coordinates and a spatial dimension distribution matrix is constructed;
[0033] Extracting multi-dimensional feature values from the data payload of the segmented data blocks;
[0034] The time dimension feature vector, space dimension distribution matrix and multi-dimensional feature value are fused into a structured feature tree, and the structured feature tree is encapsulated into a standardized structured result set.
[0035] Preferably, in step S5, the geometric distance metric is a weighted Euclidean distance between the structured result set and the historical benchmark dataset.
[0036] Preferably, in step S5, the steps of obtaining the historical benchmark data set are:
[0037] According to the device fingerprint in the segmented data block, a preset device database is queried to extract a historical operation data set with the same data type as the segmented data block;
[0038] Calculate the mean and standard deviation of all data points in the historical operation data set, and filter out data points in the historical operation data set that deviate from the mean by plus or minus 3 times the standard deviation. Use the filtered historical operation data set as the historical benchmark data set.
[0039] The present invention provides a real-time data processing system based on edge computing, which is applied to a real-time data processing method based on edge computing, including:
[0040] The data identification module is used for the terminal device to continuously collect the original data stream based on the preset physical sampling period. After the edge gateway receives the original data stream, it marks the device fingerprint and timestamp for the original data stream to obtain a marked data packet;
[0041] A dynamic segmentation module is used to perform variable time window segmentation on the marked data packets based on real-time network status data to generate segmented data blocks containing network status snapshots and timestamp sequences;
[0042] The resource decision module is used to calculate the data validity period of the segmented data blocks based on the physical sampling frequency of the terminal device and the preset business rules, and generate the processing priority based on the bandwidth utilization;
[0043] The routing scheduling module is used to determine the routing decision strategy based on the network status snapshot, the real-time bandwidth in the real-time network status data and the preset decision rule mapping table, allocate computing resources based on processing priority, and generate a structured result set based on the routing decision strategy and computing resource processing segmented data blocks;
[0044] The closed-loop control module is used to obtain a historical benchmark data set based on the device fingerprint, calculate the geometric distance measurement between the structured result set and the historical benchmark data set, generate hierarchical control instructions based on the geometric distance measurement, and push them to the terminal device after encryption.
[0045] Compared with related technologies, the real-time data processing method and system based on edge computing provided by the present invention have the following beneficial effects:
[0046] The present invention calculates the data validity period through a composite decision model based on the physical sampling frequency constraints of the terminal device and preset business rules, effectively balancing the signal restoration accuracy requirements with the urgency of the business scenario, avoiding the timeout of critical data or the crowding out of non-critical data resources; improving the data transmission integrity in a bandwidth fluctuation environment through a dynamic window cutting mechanism driven by real-time network status; optimizing the edge computing load distribution by combining routing decision mapping strategies with processing priority resource allocation; improving the reliability of terminal instructions based on a hierarchical control method based on historical data outlier purification and geometric distance measurement; forming a complete technology chain from data collection marking, adaptive transmission cutting, time-sensitive and precise arbitration, routing resource coordination to closed-loop control feedback, significantly enhancing the transmission efficiency, resource utilization accuracy and control instruction stability of high-concurrency real-time data processing systems in IoT scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of a real-time data processing method based on edge computing of the present invention;
[0048] Figure 2 This is a module structure diagram of a real-time data processing system based on edge computing of the present invention. DETAILED DESCRIPTION
[0049] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the structures. Furthermore, the embodiments of the present invention and the features of the embodiments may be combined with one another unless there is a conflict.
[0050] It should also be noted that, for ease of description, only portions relevant to the present invention are shown in the accompanying drawings, rather than all of the contents. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the various operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operations are completed, but may also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0051] Example 1
[0052] A real-time data processing method based on edge computing, in the specific implementation process, such as Figure 1 As shown, it shows a flow chart of a real-time data processing method based on edge computing, including:
[0053] Step S1: The terminal device continuously collects the original data stream based on the preset physical sampling period. After receiving the original data stream, the edge gateway marks the original data stream with a device fingerprint and a timestamp to obtain a marked data packet.
[0054] During the specific implementation process, the terminal device first continuously collects raw data streams according to a preset physical sampling period. The preset physical sampling period is set to millisecond-level accuracy based on sensor response time and industrial control requirements. For example, the temperature sensor collects environmental data every 100 milliseconds, and the vibration sensor collects device mechanical parameters every 50 milliseconds. After receiving the raw data stream via industrial Ethernet or 5G industrial private network, the edge gateway immediately performs a device fingerprint tagging operation. The device fingerprint consists of a hexadecimal-encoded string consisting of the device's unique identifier, hardware type code, and secure hash code to ensure that each device's identity is tamper-proof and traceable. At the same time, a satellite timing module is used to timestamp the data stream with nanosecond-level accuracy, accurately recording the time when the data arrived at the gateway. When the edge gateway detects multiple data streams being input concurrently, it generates a timestamp sequence in the order in which they are received to prevent clock skew between data from different devices. Finally, it outputs a tagged data packet that is doubly bound to the device fingerprint and timestamp. This tagged data packet is encapsulated in binary format, with a header containing a fingerprint identification area and a timestamp area, and a payload area storing the raw data stream, forming a unified and standardized transmission data unit that provides the input basis for subsequent variable time window cutting.
[0055] Step S2: Perform variable time window cutting on the marked data packet based on the real-time network status data to generate segmented data blocks containing network status snapshots and timestamp sequences.
[0056] Specifically, in step S2, the window length of the variable time window cutting is determined by the following steps:
[0057] Calculate the network bandwidth standard deviation within the latest multiple consecutive preset bandwidth monitoring periods;
[0058] The bandwidth utilization term is obtained by dividing the real-time bandwidth in the real-time network status data by the preset maximum theoretical bandwidth of the network and then multiplying the result by the preset bandwidth weight factor;
[0059] Multiply the network bandwidth standard deviation by the preset bandwidth fluctuation adjustment coefficient to obtain the bandwidth fluctuation term;
[0060] The bandwidth utilization term and the bandwidth fluctuation term are added together to obtain the time window determination value;
[0061] The window length is determined based on the time window decision value and the preset window length constraint rules.
[0062] During the specific implementation, for example, the network bandwidth value sequence within the latest three consecutive preset bandwidth monitoring periods is continuously monitored, each bandwidth monitoring period is a preset 200 millisecond time unit, and the bandwidth standard deviation of the network bandwidth value sequence is calculated to accurately quantify the network fluctuation intensity; at the same time, the current real-time bandwidth value is extracted from the real-time network status data, and it is divided by the preset maximum theoretical bandwidth of the network to obtain the original utilization ratio, and then the original utilization ratio is multiplied by the preset bandwidth weight factor. For example, the preset bandwidth weight factor is 0.8, and a bandwidth utilization item is generated. The preset bandwidth weight factor is used to strengthen the impact of the current load on the window decision; the calculated bandwidth standard deviation is simultaneously multiplied by the preset bandwidth fluctuation adjustment coefficient. For example, the preset bandwidth fluctuation coefficient is 0.5, forming a bandwidth fluctuation item. The preset bandwidth fluctuation coefficient is used to statistically fluctuate the value. Convert it into a time window adjustment amount; add the bandwidth utilization item and the bandwidth fluctuation item to obtain the time window decision value; dynamically adjust according to the preset window length constraint rule, for example, the preset minimum window length threshold is 1 second, and the maximum window length threshold is 5 seconds. If the time window decision value is less than 1 second, 1 second is used as the cutting window, and if it is greater than 5 seconds, 5 seconds is used. When it is in the range of 1 to 5 seconds, the decision value is directly taken as the actual window length. Based on this dynamically calculated window length, the marking data packet is cut. The generated segmented data block contains a network status snapshot and a timestamp sequence, wherein the network status snapshot records the real-time bandwidth utilization value, bandwidth fluctuation item value and standard deviation parameter at the cutting moment, and the timestamp sequence completely inherits the original data stream marking time stamp and maintains a strict increasing order, ensuring that each segmented data block carries both accurate network instantaneous state characteristics and device-level timing continuity.
[0063] Step S3: Calculate the data validity period of the segmented data block based on the physical sampling frequency of the terminal device and the preset business rules, and generate a processing priority in combination with the bandwidth utilization.
[0064] Specifically, in step S3, the step of calculating the data validity period includes:
[0065] Get the physical sampling frequency of the terminal device and divide it by 2 to get the physical validity period.
[0066] Identify the data type of the segmented data block and obtain the business validity period of the segmented data block based on the preset business rules;
[0067] Compare the physical validity period and the business validity period, and take the smaller value as the data validity period.
[0068] Specifically, in step S3, the step of calculating the processing priority includes:
[0069] Extract bandwidth utilization from real-time network status data;
[0070] Calculate the remaining valid time ratio of the segmented data block based on the survival time and data validity period of the segmented data block;
[0071] Generate processing priority based on bandwidth utilization, remaining effective time ratio and preset data value coefficient.
[0072] During the specific implementation process, the physical validity period is calculated based on the physical sampling frequency of the terminal device. The physical sampling frequency refers to the number of times the terminal device collects data per second. The physical validity period calculation formula is 2 divided by the sampling frequency value, where the value 2 is derived from the minimum two-cycle coverage principle required by the Nyquist sampling theorem, ensuring that the time window contains at least two sampling points to ensure the integrity of signal restoration; at the same time, the data type of the segmented data block is identified. For example, the data type includes three categories: device status data, environmental monitoring data, and security alarm data. The business validity period is obtained by querying the preset business rule mapping table to obtain the preset timeliness value corresponding to different data types; Compare the physical validity period and the business validity period and take the smaller value as the final data validity period; then extract the current bandwidth utilization parameter from the real-time network status data, which represents the ratio of the current network traffic to the maximum theoretical bandwidth; combine the survival time of the segmented data block with the data validity period to calculate the remaining valid time ratio, where the survival time is equal to the current time minus the data generation time, and the remaining valid time ratio is equal to the survival time divided by the data validity period; query the value weight value corresponding to the data type according to the preset data value coefficient mapping table, and perform a weighted combination of the bandwidth utilization, the remaining valid time ratio and the data value coefficient. The formula is: , get the processing priority, where To handle priority, is the bandwidth utilization, is the remaining effective time ratio, is the data value coefficient, The network load weight factor is fixed at 0.6. The timeliness urgency weight factor is fixed at 0.3. The business value weight factor is fixed at 0.1, and the calculation result is normalized to a numerical range of zero to one hundred. The processing priority value will directly determine the allocation gradient of subsequent computing resources and affect the processing order of data blocks.
[0073] Step S4: Determine the routing decision strategy based on the network status snapshot, the real-time bandwidth in the real-time network status data, and the preset decision rule mapping table, allocate computing resources based on processing priority, and the resource scheduler generates a structured result set based on the routing decision strategy and computing resource processing segmented data blocks.
[0074] Specifically, in step S4, the step of determining the routing decision strategy includes:
[0075] Quantify the network status snapshot into a routing health score through a preset decision rule mapping table;
[0076] Extract real-time bandwidth from real-time network status data and calculate the bandwidth adaptability coefficient by combining the route health score and the preset route weight factor;
[0077] Select routing decision strategy based on broadband adaptability coefficient.
[0078] Specifically, in step S4, the routing decision strategy includes the target processing node type, the transmission path selection strategy and the disaster recovery backup mechanism.
[0079] Specifically, in step S4, the steps of generating a structured result set include:
[0080] Based on the timestamps in the segmented data blocks, the time dimension feature vector is output after correcting the data clock offset;
[0081] According to the device fingerprint in the segmented data block, the preset device database is queried to obtain the geographic coordinates and a spatial dimension distribution matrix is constructed;
[0082] Extracting multi-dimensional feature values from the data payload of the segmented data blocks;
[0083] The time dimension feature vector, space dimension distribution matrix and multi-dimensional feature value are fused into a structured feature tree, and the structured feature tree is encapsulated into a standardized structured result set.
[0084] In the specific implementation process, the network status snapshot is first quantified into a routing health score through a preset decision rule mapping table. The preset decision rule mapping table is a two-dimensional matrix with the network status snapshot parameters as keys. The parameters include the real-time bandwidth utilization value, bandwidth fluctuation item value and standard deviation recorded at the cutting moment. The routing health score is mapped in a numerical range of 0 to 100 through linear weighting, where 0 represents that the network is completely unavailable and 100 represents an ideal state. At the same time, the current real-time bandwidth value is extracted from the real-time network status data, and the broadband adaptability coefficient is calculated by combining the routing health score and the preset routing weight factor. For example, the routing weight factor is preset according to the industrial network type, and the exemplary Ethernet setting is 0.8 or the wireless network setting is 0.5. The broadband adaptability coefficient calculation formula is: ,in, is the broadband adaptability coefficient, is the real-time bandwidth, is the route health score, The routing weight factor is the routing decision strategy selected based on the value of the broadband adaptability coefficient. For example, when the broadband adaptability coefficient is greater than 1.5, the target processing node type selects the edge computing node, when it is less than 1.0, the cloud computing center is selected, and in the range of 1.0 to 1.5, the hybrid processing mode is selected; the transmission path selection strategy is implemented for different node types. For example, the edge computing node adopts the shortest path algorithm within three hops based on topological hashing, and the cloud computing center adopts the optimal path of the border gateway protocol; the disaster recovery backup mechanism is set according to the network health score, when it is greater than 90 points, a single local copy is used, when it is between 60 points and 90 points, two copies are used for cross-node backup, and when it is less than 60 points, three copies are used for cloud storage; then, computing resources are allocated based on processing priority. For example, in the edge computing node mode, the number of computing cores is allocated according to the processing priority value divided by 20 and rounded up, and in the cloud computing center mode, two virtual central processing units and 4G bytes of memory resources are fixedly allocated; the resource scheduler calls the corresponding resource based on the routing decision strategy to process the segmented data blocks, and specifically generates a structured result set: first, according to the timestamp sequence in the segmented data blocks, the least squares time is used The clock offset correction algorithm outputs a time dimension feature vector, which is arranged in order by millisecond timestamps. The device fingerprint in the segmented data block is used to query the preset device database to obtain the geographic coordinate latitude and longitude values. The device fingerprint is a hash string that uniquely identifies the terminal device. The device database stores a mapping relationship table between the device fingerprint and the geographic coordinates, and constructs a spatial dimension distribution matrix. The spatial dimension distribution matrix is a two-dimensional Gaussian distribution weight matrix indexed by the geographic coordinates. Multi-dimensional eigenvalues are extracted from the data payload of the segmented data block. The data payload is a set of original sensor physical quantity values. The dimensional eigenvalues include maximum value, minimum value, average value and standard deviation statistics. The time dimension feature vector, spatial dimension distribution matrix and multi-dimensional eigenvalues are input into the feature fusion engine to generate a structured feature tree. The structured feature tree uses a binary prefix tree index structure to store spatiotemporal data relationships. Finally, the structured feature tree is encapsulated into a standardized structured result set encapsulation. The encapsulation format is a self-describing data object containing a time series vector field, a spatial distribution matrix field and an eigenvalue array field, completing the final output of edge computing data processing.
[0085] Step S5: Obtain a historical benchmark data set based on the device fingerprint, calculate the geometric distance measurement between the structured result set and the historical benchmark data set, generate a hierarchical control instruction based on the geometric distance measurement, and push it to the terminal device after encryption.
[0086] Specifically, in step S5, the geometric distance metric is the weighted Euclidean distance between the structured result set and the historical benchmark dataset.
[0087] Specifically, in step S5, the steps for obtaining the historical benchmark dataset are:
[0088] According to the device fingerprint in the segmented data block, a preset device database is queried to extract a historical operation data set with the same data type as the segmented data block;
[0089] Calculate the mean and standard deviation of all data points in the historical operation data set, and filter out data points in the historical operation data set that deviate from the mean by plus or minus 3 times the standard deviation. Use the filtered historical operation data set as the historical benchmark data set.
[0090] During the specific implementation, in the closed-loop control stage of the edge computing data processing flow, the historical operation data set is retrieved from the preset device database based on the device fingerprint as a benchmark reference. The device fingerprint is a hexadecimal hash string used to uniquely identify the terminal device. The device database associated with the device fingerprint stores the historical operation data set of the same type of device; the specific retrieval process is to filter the same type of historical operation data set in the device database according to the data type of the segmented data block, and the data type and business rule mapping classification matching include three standard categories: device status data, environmental monitoring data, and safety alarm data; perform statistical purification operations on the acquired historical operation data set, first calculate the arithmetic mean of all data points in the historical operation data set as the benchmark center value, and then calculate the average value of all data points in the historical operation data set as the benchmark center value. The standard deviation of the entire data set is calculated as a discrete measurement indicator. The arithmetic mean is calculated by adding up the values of all data points and dividing them by the total number. The standard deviation is calculated using the same root mean square algorithm as the network bandwidth monitoring statistics. Based on the mean and standard deviation, a data point filtering rule is constructed to accurately filter out data points that deviate from the mean value by more than three times the standard deviation in the positive direction or more than three times the standard deviation in the negative direction. The three times standard deviation boundary constitutes a statistically significant anomaly discrimination threshold. The data point set after purification is used as the final historical benchmark data set to participate in the subsequent geometric distance measurement calculation. The structured result set is then weighted with the historical benchmark data set to obtain a geometric distance measurement. The geometric distance measurement formula is the square root of the weighted sum of the squares of the differences in each feature dimension. The feature dimension includes the time dimension series. There are three core dimensions: vector, spatial dimension matrix coordinate value, and feature dimension statistic; the weight coefficient of each dimension is preset according to its feature importance. For example, the time dimension weight of the security alarm class is set to 0.7, the spatial dimension weight is set to 0.2, and the feature dimension weight is set to 0.1. For the environmental monitoring class, the time dimension weight is set to 0.3, the spatial dimension weight is set to 0.5, and the feature dimension weight is set to 0.2; the specific process of distance calculation is to subtract the eigenvalue of each dimension of the structured result set from the eigenvalue of the corresponding dimension of the historical benchmark data set to obtain the difference, square the difference and multiply it by the preset dimension weight coefficient, sum all the weighted square values and take the square root to get the final distance scalar value, that is, the geometric distance metric; based on this geometric distance metric Generate hierarchical control instructions, and the hierarchical rules are set according to the preset distance threshold interval. For example, when the geometric distance measurement is less than one standard deviation, a maintenance operation instruction is generated without processing. In the range of one to three standard deviations, an optimization adjustment instruction is generated to trigger fine-tuning of device parameters. When it exceeds three standard deviations, an emergency shutdown instruction is generated to immediately interrupt the operation of the device. The control instructions are encrypted using an asymmetric key derived from the device fingerprint. For example, the encryption algorithm adopts the block cipher mode of the national secret SM4 standard, and the encryption process is supplemented with an anti-replay signature enhanced by a quantum random number. Finally, the encrypted instructions are pushed to the terminal device through the secure transmission channel of the edge gateway to execute closed-loop control. After the terminal device executes the instruction, its operating status is backfilled into the device database as a new data point to complete the self-learning update cycle.
[0091] The working principle of the real-time data processing method based on edge computing provided by the present invention is as follows:
[0092] First, in terms of data collection, the present invention uses a terminal device to generate an original data stream based on a preset physical sampling period, and an edge gateway implements a unique identification of the data source and binding it to a time-space benchmark by embedding a device fingerprint and a timestamp; then, in terms of data transmission, the cutting window size is dynamically adjusted based on the real-time network status, the network fluctuation intensity is quantified by calculating the bandwidth standard deviation, and the window decision value is generated by combining the real-time bandwidth utilization weights to ensure the integrity and temporal continuity of high-frequency data streams in complex network environments; then, a dual arbitration mechanism is introduced, which calculates the minimum validity period under the Nyquist constraint based on the physical sampling frequency, and obtains the scenario-based timeliness requirements through business rule mapping, takes the minimum value of the two as the upper limit of the data validity period, and combines the bandwidth load, the remaining validity period ratio and the data value coefficient to generate a multi-dimensional processing. Priority is set to achieve dynamic matching of resource demand and supply; further, the network status snapshot is quantified into a routing health score, and the broadband adaptability coefficient is calculated in collaboration with the real-time bandwidth. The target node type and transmission path strategy are decided based on this, and the core computing resources are allocated based on the priority numerical gradient; clock offset correction, geographic coordinate mapping and multi-dimensional feature extraction technology are used to generate a structured result set of spatiotemporal feature fusion; finally, the weighted Euclidean geometric distance is calculated to measure the deviation between the real-time status and the benchmark through the purified historical data set after device fingerprint indexing, and hierarchical encryption instructions are generated based on the distance threshold and pushed to the terminal device, forming a full-process closed-loop system from data collection, dynamic transmission, intelligent arbitration, resource optimization to precise control, realizing efficient transmission of high-concurrency real-time data processing and improved resource utilization.
[0093] Example 2
[0094] A real-time data processing system based on edge computing is applied to a real-time data processing method based on edge computing. In the specific implementation process, Figure 2 As shown, it shows a module structure diagram of a real-time data processing system based on edge computing, including:
[0095] The data identification module 100 is used for the terminal device to continuously collect the original data stream based on the preset physical sampling period. After the edge gateway receives the original data stream, it marks the original data stream with the device fingerprint and timestamp to obtain a marked data packet;
[0096] A dynamic cutting module 200 is configured to perform variable time window cutting on the marked data packets based on the real-time network status data to generate segmented data blocks containing a network status snapshot and a timestamp sequence;
[0097] The resource decision module 300 is used to calculate the data validity period of the segmented data block based on the physical sampling frequency of the terminal device and the preset business rules, and generate a processing priority in combination with the bandwidth utilization;
[0098] The routing scheduling module 400 is used to determine the routing decision strategy based on the network status snapshot, the real-time bandwidth in the real-time network status data, and the preset decision rule mapping table, allocate computing resources based on processing priority, and generate a structured result set by the resource scheduler based on the routing decision strategy and computing resources to process the segmented data blocks;
[0099] The closed-loop control module 500 is used to obtain a historical benchmark data set based on the device fingerprint, calculate the geometric distance measurement between the structured result set and the historical benchmark data set, generate hierarchical control instructions based on the geometric distance measurement, and push them to the terminal device after encryption.
[0100] The working principle of the real-time data processing system based on edge computing provided by the present invention is as follows:
[0101] The data identification module 100 injects device fingerprints and precise timestamps into the original data stream collected by the terminal device at the edge gateway layer to establish a unique identity and time reference that can be traced back; the dynamic cutting module 200 calculates the dynamic window length by real-time monitoring of the network bandwidth standard deviation and current utilization, and performs variable time window cutting based on this and generates segmented data blocks that carry network status snapshots and time series features; the resource decision module 300 generates data validity period based on the dual constraints of the terminal device physical sampling frequency and preset business rules, and at the same time integrates bandwidth utilization, data survival time ratio and value weight coefficient to calculate the gradient processing priority; the routing scheduling module 400 quantifies the network status snapshot into a routing health indicator, and calculates the broadband adaptability coefficient in coordination with the real-time bandwidth, and determines the target processing type, transmission path strategy and disaster recovery backup mechanism of the edge node or cloud computing center based on this, and The level value allocates precise computing core resources; the resource scheduler calls the allocated computing resources to execute the spatiotemporal data processing pipeline: first correct the timestamp sequence to generate a time vector, then retrieve the geographic coordinates according to the device fingerprint to construct a spatial distribution matrix, and then extract the multi-dimensional eigenvalues of the data payload, and finally fuse them into a structured feature tree and encapsulate the output; the closed-loop control module 500 is based on the device fingerprint index device database, retrieves the same type of historical data set, and then uses the triple standard deviation rule to purify the benchmark data, calculates the weighted Euclidean geometric distance between the current structured result set and the purified benchmark, generates hierarchical control instructions according to the distance threshold and encrypts and pushes them to the terminal device, thereby forming a full-process self-optimization system from data identification, network dynamic cutting, resource intelligent decision-making, route optimization scheduling to closed-loop feedback control, realizing high-timeliness, high-integrity, and high-reliability real-time data processing full-chain collaboration in edge computing scenarios.
[0102] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0103] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0104] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
Claims
1. A real-time data processing method based on edge computing, characterized in that: The data processing method comprises the following steps: S1: The terminal device continuously collects raw data streams based on the preset physical sampling period. After receiving the raw data streams, the edge gateway marks the raw data streams with device fingerprints and timestamps to obtain marked data packets. S2: Perform variable time window segmentation on the marked data packets based on real-time network status data to generate segmented data blocks containing network status snapshots and timestamp sequences; S3: Calculates the data validity period of the segmented data blocks based on the physical sampling frequency of the terminal device and preset business rules, and generates processing priorities based on bandwidth utilization; S4: Determine the routing decision strategy based on the network status snapshot, the real-time bandwidth in the real-time network status data, and the preset decision rule mapping table, allocate computing resources based on processing priority, and the resource scheduler generates a structured result set based on the routing decision strategy and computing resource processing segmented data blocks; S5: Obtain a historical benchmark data set based on the device fingerprint, calculate the geometric distance measurement between the structured result set and the historical benchmark data set, generate hierarchical control instructions based on the geometric distance measurement, and push them to the terminal device after encryption.
2. A real-time data processing method based on edge computing according to claim 1, characterized in that: In step S2, the window length of the variable time window cutting is determined by the following steps: Calculate the network bandwidth standard deviation within the latest multiple consecutive preset bandwidth monitoring periods; The bandwidth utilization term is obtained by dividing the real-time bandwidth in the real-time network status data by the preset maximum theoretical bandwidth of the network and then multiplying the result by the preset bandwidth weight factor; Multiply the network bandwidth standard deviation by the preset bandwidth fluctuation adjustment coefficient to obtain the bandwidth fluctuation term; The bandwidth utilization term and the bandwidth fluctuation term are added together to obtain the time window determination value; The window length is determined based on the time window decision value and the preset window length constraint rules.
3. A real-time data processing method based on edge computing according to claim 2, characterized in that: In step S3, the step of calculating the data validity period includes: Get the physical sampling frequency of the terminal device and divide it by 2 to get the physical validity period. Identify the data type of the segmented data block and obtain the business validity period of the segmented data block based on the preset business rules; Compare the physical validity period and the business validity period, and take the smaller value as the data validity period.
4. A real-time data processing method based on edge computing according to claim 3, characterized in that: In step S3, the step of calculating the processing priority includes: Extract bandwidth utilization from real-time network status data; Calculate the remaining valid time ratio of the segmented data block based on the survival time and data validity period of the segmented data block; Generate processing priority based on bandwidth utilization, remaining effective time ratio and preset data value coefficient.
5. A real-time data processing method based on edge computing according to claim 4, characterized in that: In step S4, the step of determining the routing decision strategy includes: Quantify the network status snapshot into a routing health score through a preset decision rule mapping table; Extract real-time bandwidth from real-time network status data and calculate the bandwidth adaptability coefficient by combining the route health score and the preset route weight factor; Select routing decision strategy based on broadband adaptability coefficient.
6. A real-time data processing method based on edge computing according to claim 5, characterized in that: In step S4, the routing decision strategy includes the target processing node type, the transmission path selection strategy and the disaster recovery backup mechanism.
7. A real-time data processing method based on edge computing according to claim 6, characterized in that: In step S4, the steps of generating a structured result set include: Based on the timestamps in the segmented data blocks, the time dimension feature vector is output after correcting the data clock offset; According to the device fingerprint in the segmented data block, the preset device database is queried to obtain the geographic coordinates and a spatial dimension distribution matrix is constructed; Extracting multi-dimensional feature values from the data payload of the segmented data blocks; The time dimension feature vector, space dimension distribution matrix and multi-dimensional feature value are fused into a structured feature tree, and the structured feature tree is encapsulated into a standardized structured result set.
8. A real-time data processing method based on edge computing according to claim 7, characterized in that: In step S5 , the geometric distance metric is the weighted Euclidean distance between the structured result set and the historical benchmark dataset.
9. A real-time data processing method based on edge computing according to claim 8, characterized in that: In step S5, the steps for obtaining the historical benchmark dataset are: According to the device fingerprint in the segmented data block, a preset device database is queried to extract a historical operation data set with the same data type as the segmented data block; Calculate the mean and standard deviation of all data points in the historical operation data set, and filter out data points in the historical operation data set that deviate from the mean by plus or minus 3 times the standard deviation. Use the filtered historical operation data set as the historical benchmark data set.
10. A real-time data processing system based on edge computing, characterized in that: A real-time data processing method based on edge computing according to any one of claims 1 to 9, wherein the data processing system comprises: The data identification module is used for the terminal device to continuously collect the original data stream based on the preset physical sampling period. After the edge gateway receives the original data stream, it marks the device fingerprint and timestamp for the original data stream to obtain a marked data packet; A dynamic segmentation module is used to perform variable time window segmentation on the marked data packets based on real-time network status data to generate segmented data blocks containing network status snapshots and timestamp sequences; The resource decision module is used to calculate the data validity period of the segmented data blocks based on the physical sampling frequency of the terminal device and the preset business rules, and generate the processing priority based on the bandwidth utilization; The routing scheduling module is used to determine the routing decision strategy based on the network status snapshot, the real-time bandwidth in the real-time network status data and the preset decision rule mapping table, allocate computing resources based on processing priority, and generate a structured result set based on the routing decision strategy and computing resource processing segmented data blocks; The closed-loop control module is used to obtain a historical benchmark data set based on the device fingerprint, calculate the geometric distance measurement between the structured result set and the historical benchmark data set, generate hierarchical control instructions based on the geometric distance measurement, and push them to the terminal device after encryption.
Citation Information
Patent Citations
Industrial edge network system architecture and resource scheduling method
CN112235836A
Industrial network optimization system based on edge computing
CN118474100A
Intelligent decision management method and device for state monitoring and fault diagnosis of energy-saving equipment
CN120067772A
Cloud edge efficient synchronization mechanism and distributed management scheduling system based on crop water shortage data
CN120075264A
Industrial internet real-time cooperative control method based on edge computing
CN120276320A
Cited By
Building sale distributed house resource real-time synchronous management method and system based on cloud edge fusion computing
CN122019670A
A method and system for real-time synchronization management of distributed house resources based on cloud edge fusion computing
CN122019670B