Computer network data instant updating method based on cloud computing
Through the cloud-based computer network data instant update method, network data is collected and decomposed in real time into microbatch update operations, updated dependency graphs are built, and distributed updates are performed, which solves the immediacy and consistency of data processing in the existing technology, and improves the system's response speed and user satisfaction.
Patent Information
- Application Number
- CN202510631932.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-29
AI Technical Summary
The existing data processing methods have obvious limitations in immediacy, data processing efficiency, distributed consistency, and adapting to user needs and technological development. They cannot meet the needs of real-time updates such as social media, online advertising and logistics information, and consume too much computing resources.
The cloud-based computer network data instant update method is adopted, and network data is collected in real time, decomposed into microbatch update operations, and the update dependency graph is built, and the controller is used to perform distributed update operations. Combined with cache and incremental update strategies, monitoring tools are used to track the update process, and dynamically adjust the strategy to adapt to user needs and technological development.
It improves the real-time and scalability of data processing, reduces computing resource consumption, reduces update coordination time, ensures the consistency and efficiency of distributed updates, and improves user satisfaction.
Smart Images

Figure CN120560701A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data updating, and in particular to a method for instantly updating computer network data based on cloud computing. Background Art
[0002] In the current data processing landscape, the growing number of real-time application scenarios such as social media, online advertising, and logistics has placed higher demands on the immediacy and efficiency of data processing. Existing processing methods, especially in streaming data processing, typically employ a batch processing model that separates the input of the data source from the output of data processing. This means that at a fixed point in time, the data processing system stops accepting new data in order to process the accumulated batch data and store it on disk or other persistent storage media. While this processing method can effectively process large amounts of static data, it suffers from significant latency issues in scenarios with high real-time requirements, making it unable to meet users' demands for immediate data processing and feedback.
[0003] Specifically, the instantaneous updating of social media content, the precise delivery of online ads, and the real-time tracking of logistics information all require systems to rapidly respond to new data inputs and instantly update user interfaces or data models. However, traditional batch processing methods, due to their inherent processing cycles, cannot process data the instant it arrives, resulting in a poor user experience, especially in scenarios requiring real-time interaction or feedback.
[0004] Furthermore, as data volumes continue to grow, traditional batch processing methods are limited in their processing power and efficiency when handling large amounts of data. During data updates, if incremental data cannot be effectively filtered and processed, the system's computing resources will be consumed by a large amount of meaningless data processing, which not only wastes resources but also affects the overall system response speed.
[0005] Furthermore, controlling data flow and ensuring consistency during network data updates presents a major challenge in existing technologies. Traditional data update methods often rely on a centralized controller to coordinate data updates. This not only increases system complexity in large-scale distributed systems but also limits scalability and efficiency. The controller must process a large number of data update requests, which not only increases network latency but can also easily become a system bottleneck.
[0006] Finally, existing data update methods are also insufficient in adapting to user needs and technological developments. With increasing user demands for immediacy and personalization of data updates, coupled with the continuous advancement of technologies such as cloud computing, big data, and artificial intelligence, traditional data update methods struggle to respond quickly to these changes, thus failing to provide a high-quality service experience.
[0007] In summary, traditional data update processing methods have obvious limitations in terms of immediacy, data processing efficiency, distributed consistency, and adaptability to user needs and technological development. Summary of the Invention
[0008] The purpose of the present invention is to address the corresponding shortcomings of the existing technology and provide a cloud computing-based computer network data instant update method to improve the real-time and scalability of data processing, reduce the consumption of computing resources, reduce the time overhead of update coordination, and be able to respond to user needs and technological development trends in a timely manner, optimize update strategies, and solve the problems of delay time and data processing volume limitations in social media content updates, online advertising delivery, logistics process updates, etc.
[0009] The object of the present invention is to adopt the following scheme to achieve: a computer network data instant update method based on cloud computing, comprising the following steps:
[0010] 1) Collect network data information in real time from the data plane;
[0011] 2) Decompose network data information into micro-batch update operations according to time intervals;
[0012] 3) Convert the micro-batch update operation into an update dependency graph based on network data information;
[0013] 4) Constructing update messages based on the update dependency graph;
[0014] 5) The controller uses the update message to perform a distributed update operation. During the distributed update operation, the controller uses a monitoring tool to track the update operation process and data flow, and adjusts the update strategy based on the tracking results.
[0015] Preferably, in step 1), the method of collecting network data information from the data plane in real time includes using a Spark Streaming module to obtain the network data information.
[0016] Preferably, in step 2), the network data information is decomposed into micro-batch update operations according to time intervals, specifically including:
[0017] 2-1) Dividing the flow data in the network data information into several flow segments according to fixed time intervals;
[0018] 2-2) Extracting features from the first K data packets of each flow segment divided in step 2-1) to form an initial feature set;
[0019] 2-3) Forming the optimal feature set through feature selection;
[0020] 2-4) Divide the network data update into several micro-batch update operations (i.e., a series of update operations) based on the optimal feature set.
[0021] Preferably, in step 2-2), forming the initial feature set includes:
[0022] 2-2-1) Convert all data packets into a five-tuple consisting of "packet size sequence", "packet arrival time sequence", "packet difference sequence", "uplink speed sequence" and "downlink speed sequence";
[0023] 2-2-2) Use the quintuple to calculate the "statistical characteristics", "downlink conditional frequency characteristics" and "uplink conditional frequency characteristics" of each data packet respectively;
[0024] 2-2-3) Finally, the “statistical features”, “downlink conditional frequency features” and “uplink conditional frequency features” of each data packet are used to form an initial feature set.
[0025] Preferably, in step 2-3), forming an optimal feature set by feature selection includes:
[0026] 2-3-1) Calculate the Pearson correlation coefficient (PCC) between each feature vector in the initial feature set and each data type of the network data information as the importance score of each feature vector;
[0027] 2-3-2) Combine the feature vectors in the initial feature set in pairs to form several feature combinations;
[0028] 2-3-3) Calculate the PCC of each feature combination in step 2-3-2) and determine whether the PCC of each feature combination is greater than 0.9:
[0029] If PCC≤0.9, retain the feature combination in the initial feature set;
[0030] If PCC>0.9, the feature vectors with low correlation with the data type in the feature combination are deleted from the initial feature set;
[0031] 2-3-4) The remaining feature vectors in the initial feature set are sorted by importance using the extreme tree algorithm, and the stream data is classified several times according to the importance of each feature vector to obtain data feature sets corresponding to different classification situations;
[0032] 2-3-5) Finally, use RFC to judge the flow data classification accuracy of the data feature set corresponding to each classification situation, and select the data feature set with the highest flow data classification accuracy as the optimal feature set.
[0033] Preferably, in step 3), converting the micro-batch update operation into an update dependency graph according to the network data information includes:
[0034] 3-1) Based on the collected network data G0 and the updated network data G n , calculate the update from G0 to G n Several micro-batch update operations that need to be performed during the process;
[0035] 3-2) Based on the interdependence between each micro-batch update operation, network link bandwidth resources and the demand information of micro-batch update operation for link bandwidth resources, for each given collected network data G0 and updated network data G n The complete data update process builds an update dependency graph that meets consistency requirements.
[0036] Preferably, in step 4), an update message is constructed according to the update dependency graph, including: defining an update message structure, including "operation", "precondition" and "post-operation" fields, wherein the "precondition" is described in the form of a logical Boolean expression to ensure the dependency relationship between update operations.
[0037] Preferably, in step 4), the update message is constructed according to the update dependency graph, including defining a message delivery path in the update dependency graph, so that the micro-batch update operation can autonomously determine the execution order, reducing the scheduling burden of the controller.
[0038] Preferably, in step 5), the controller uses the update message to perform a distributed update operation, including the controller performing an incremental update operation according to the update dependency graph and the cache strategy, identifying the data segment to be updated from the Rdd sequence, and determining the specific data range of the incremental update according to the cache status.
[0039] Preferably, in step 5), the controller uses update messages to perform distributed update operations, including using two message modes, "query message" and "notification message", to coordinate the update sequence.
[0040] The beneficial effects of the present invention are as follows:
[0041] ① The present invention uses Spark-based near-real-time incremental data updates, combined with distributed interaction on the data plane, to effectively reduce the update coordination distance and the number of interactions required for update coordination, thereby reducing the update coordination time and significantly improving the speed of real-time data processing.
[0042] ② The present invention filters out simulated new classes from network data information, incorporates them into the known class set, and then gradually responds to the different new class data that appear in an incremental manner, thereby avoiding the situation where the update operation is difficult to adapt to the new class data and causes errors in the update.
[0043] ③ The present invention uses distributed updates performed by the controller, combined with cache and incremental update strategies, so that only necessary data is processed, thereby saving computing resources;
[0044] ④ The present invention utilizes a query and notification message mechanism in the data plane, which can autonomously coordinate the operation sequence and reduce dependence on the controller;
[0045] ⑤ By implementing real-time data updates and analyzing user feedback on a cloud service platform, the present invention can quickly adapt to changes in network data and user needs, thereby achieving higher user satisfaction. In other words, the present invention can utilize cloud computing resources to complete the calculation of network data information and RFC analysis, thereby providing update services to multiple users. Simultaneously, data processing in the cloud allows users greater flexibility in selecting updates, eliminating the need for excessive user interaction and minimizing the impact on normal user usage.
[0046] Glossary:
[0047] Spark: Spark is an efficient, general-purpose, open-source distributed computing system. Spark's characteristics are that it supports in-memory computing, iterative computing, and is faster than the Hadoop MapReduce system. In Spark, data processing can utilize distributed memory cache for efficient transformation and analysis. In stream processing, Spark supports the Spark Streaming module, which can be used independently or integrated with other Spark components. The core of Spark Streaming is DStream (discrete stream), where the batch processing time interval is determined by the time interval between RDD operations. In each time period, DStream collects real-time data within a specific time interval and divides it into batches for processing.
[0048] PCC: Pearson correlation coefficient.
[0049] RFC: Random Forest Model.
[0050] CM: sample confidence.
[0051] WebSocket: WebSocket is a protocol for full-duplex communication over a single TCP connection.
[0052] API: API (Application Programming Interface) is a set of rules and tools that define how software components interact with each other.
[0053] Spark RDD: Spark RDD (Resilient Distributed Dataset) is the core data structure in the Spark framework.
[0054] Dstream: discrete stream.
[0055] Hadoop MapReduce: Hadoop MapReduce is a programming model and software framework for large-scale data processing.
[0056] Controller: In this context, this refers to a program used by a computer to manage network data updates, or a functional module that stores such a program. The management and execution of network data updates within this functional module involves multiple programs and technologies, including the operating system (built-in update service), applications (built-in update mechanisms), version control tools, security update management, and cloud services. The specific method chosen depends on the updated network data content, computer environment requirements, and user preferences, and should be determined based on actual circumstances.
[0057] Micro-batch update operations: a series of tiny update operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0059] like Figure 1 As shown, a method for instantly updating computer network data based on cloud computing includes the following steps:
[0060] 1) Collect network data information in real time from the data plane;
[0061] Typically, real-time network data collection from the data plane involves using the Spark Streaming module. The selection of a collection method is based on the ability to establish a real-time interactive connection between the real-time streaming data API and the data source, as well as the immediacy, integrity, and efficiency of data collection.
[0062] 2) Decompose network data information into micro-batch update operations according to time intervals:
[0063] 2-1) Dividing the flow data in the network data information into several flow segments according to fixed time intervals;
[0064] 2-2) Extract features from the first K data packets of each flow segment divided in step 2-1) to form an initial feature set:
[0065] 2-2-1) Convert all data packets into a five-tuple consisting of "packet size sequence", "packet arrival time sequence", "packet difference sequence", "uplink speed sequence" and "downlink speed sequence";
[0066] 2-2-2) Use the quintuple to calculate the "statistical characteristics", "downlink conditional frequency characteristics" and "uplink conditional frequency characteristics" of each data packet respectively;
[0067] 2-2-3) Finally, the “statistical features”, “downlink conditional frequency features” and “uplink conditional frequency features” of each data packet are used to form an initial feature set.
[0068] 2-3) The optimal feature set is formed by feature selection in the following way:
[0069] 2-3-1) Calculate the Pearson correlation coefficient (PCC) between each feature vector in the initial feature set and each data type of the network data information as the importance score of each feature vector;
[0070] 2-3-2) Combine the feature vectors in the initial feature set in pairs to form several feature combinations;
[0071] 2-3-3) Calculate the PCC of each feature combination in step 2-3-2) and determine whether the PCC of each feature combination is greater than 0.9:
[0072] If PCC≤0.9, retain the feature combination in the initial feature set;
[0073] If PCC>0.9, the feature vectors with low correlation with the data type in the feature combination are deleted from the initial feature set;
[0074] 2-3-4) The remaining feature vectors in the initial feature set are sorted by importance using the extreme tree algorithm, and the stream data is classified several times according to the importance of each feature vector to obtain data feature sets corresponding to different classification situations;
[0075] 2-3-5) Finally, use RFC to judge the flow data classification accuracy of the data feature set corresponding to each classification situation, and select the data feature set with the highest flow data classification accuracy as the optimal feature set.
[0076] 2-4) Divide the network data update into several micro-batch update operations (i.e., a series of update operations) based on the optimal feature set.
[0077] 3) Convert the micro-batch update operation into an update dependency graph based on network data information:
[0078] 3-1) Based on the collected network data G0 and the updated network data G n , calculate the update from G0 to G n Several micro-batch update operations that need to be performed during the process;
[0079] 3-2) Based on the interdependence between each micro-batch update operation, network link bandwidth resources and the demand information of micro-batch update operation for link bandwidth resources, for each given collected network data G0 and updated network data G n The complete data update process builds an update dependency graph that meets consistency requirements.
[0080] 4) Construct update message based on update dependency graph:
[0081] 4-1) Define the update message structure, which includes the "operation", "precondition" and "postcondition" fields. The "precondition" is described in the form of a logical Boolean expression to ensure the dependency between update operations.
[0082] For example, by defining time windows between update operations, the system ensures that when executing an update, the current operation must wait 10 seconds after the previous operation to complete before starting. This prevents time conflicts during data processing and ensures the accuracy of data updates. Furthermore, the system also specifies specific network conditions through the "precondition" field, such as link utilization below 50%, to ensure that updates are executed under appropriate conditions, further improving system stability and reliability.
[0083] 4-2) Defining message delivery paths in the update dependency graph enables micro-batch update operations to autonomously determine their execution order, reducing the controller's scheduling burden, improving data processing speed and efficiency, and reducing data processing latency.
[0084] 5) The controller uses the update message to perform a distributed update operation. During the distributed update operation, it uses monitoring tools to track the update operation process and data flow, and adjusts the update strategy based on the tracking results:
[0085] 5-1) The controller performs incremental update operations based on the update dependency graph and cache strategy, identifies the data segments to be updated from the RDD sequence, and determines the specific data range of the incremental update based on the cache status, effectively processing large-scale network data and improving data processing speed and efficiency.
[0086] By identifying the data segments to be updated from the Rdd sequence, the present invention can effectively determine the specific data range of the incremental update, thereby improving data processing speed and efficiency, and effectively processing large-scale network data, improving data processing speed and efficiency, and reducing data processing delays.
[0087] 5-2) Use the "query message" and "notification message" message modes to coordinate the update sequence to ensure the consistency and efficiency of the update process.
[0088] This allows for autonomous coordination of micro-batch update operations through a query and notification message mechanism. In principle, query messages are used to confirm whether preconditions are met, while notification messages are used to inform the execution status of postconditions. This ensures the correct execution order of update operations, reduces the controller's scheduling burden, improves the system's concurrent processing capabilities, and reduces data processing latency.
[0089] According to the above method, the following embodiments are made:
[0090] 1) Collect network data information in real time from the data plane:
[0091] In this example, the Spark Streaming module is used to establish real-time connections with network data sources, such as those monitoring internet traffic, server communication data, and IoT device data, ensuring immediate data collection. The Spark Streaming module continuously monitors the data source's real-time streaming interface, converting the streaming data into a processable RDD sequence. It then performs preliminary data cleaning on this data, removing invalid or duplicate data to improve the efficiency and accuracy of subsequent processing.
[0092] The reason why this embodiment uses the Spark Streaming module's API for real-time streaming data to obtain network data information is because Spark is an efficient, general-purpose open source distributed computing system. The characteristics of Spark are that it supports in-memory computing, iterative computing, and is faster than the Hadoop MapReduce system. In Spark, data processing can use distributed memory cache for efficient conversion and analysis. In stream processing, Spark supports the SparkStreaming module, which can be used independently or integrated with other Spark components. The core of Spark Streaming is DStream (discrete stream), in which the batch processing time interval is determined by the time interval between RDD operations. In each time period, DStream collects real-time data within a specific time interval and divides it into batches for processing. The Spark Streaming module uses Spark Rdd at the bottom layer to generate Rdd sequences. Each Rdd sequence contains all data from the previous sequence time interval to the current time interval.
[0093] In summary, this embodiment utilizes the real-time streaming data processing capabilities of the Spark Streaming module and implements seamless integration between the data source and the Spark Streaming module through the real-time streaming data API, ensuring the immediacy and integrity of data in real-time transmission and processing, and improving data collection efficiency.
[0094] 2) Decompose the network data information update obtained in step 1) into a series of update operations in the following manner:
[0095] 2-1) Divide the collected flow data into flow segments according to fixed time intervals. For example, divide the flow data collected in real time into time intervals of 5 minutes to form multiple flow segments.
[0096] 2-2) Extract features from the first K data packets of each flow segment divided in step 2-1) in the following manner to form an initial feature set.
[0097] For example, from the first 100 data packets of each flow segment, we extract the packet size sequence, packet arrival time sequence, packet difference sequence, upstream speed sequence, and downstream speed sequence to construct a five-tuple. For example, suppose we have a simplified network flow dataset as follows:
[0098]
[0099] Then, the constructed quintuple includes the following parts:
[0100] A) A "packet size sequence" used to record the size of each packet, such as [400, 600, 300, 800, 1200];
[0101] B) a "packet arrival time series" used to record the time when each packet arrives, for example [1.0, 2.5, 3.5, 5.0, 6.0];
[0102] C) a "packet difference sequence" used to calculate the time difference between two adjacent packets (i.e., the difference in arrival time), for example, [1.5, 1.0, 1.5, 1.0] (i.e., 1.5 seconds, 1 second, 1.5 seconds, 1 second);
[0103] D) Uplink speed sequence:
[0104] In this embodiment, the uplink speed is the packet size divided by the packet arrival time difference, which is as follows:
[0105] Uplink speed from packet 1 to packet 2 = (600-400) / 1.5 = 133.33 bytes / sec;
[0106] Uplink speed from packet 2 to packet 3 = (300-600) / 1.0 = -300 bytes / sec;
[0107] Uplink speed from packet 3 to packet 4 = (800-300) / 1.5 = 333.33 bytes / sec;
[0108] Uplink speed from packet 4 to packet 5 = (1200-800) / 1.0 = 400 bytes / sec;
[0109] Uplink speed sequence: [133.33,-300,333.33,400];
[0110] E) Downstream speed sequence:
[0111] In this embodiment, the downlink speed is the same as the uplink speed.
[0112] Downlink speed sequence (same as uplink speed, assuming symmetric traffic): [133.33, -300, 333.33, 400];
[0113] Of course, in actual applications, the downlink speed can also be calculated based on other conditions.
[0114] This quintuple serialization processing can convert the original data into a form that is easy to analyze and process, facilitate subsequent feature extraction and data analysis, improve the accuracy and efficiency of feature extraction, and provide a solid foundation for subsequent data analysis and update strategy formulation.
[0115] Then, the five-tuple is used to calculate the statistical characteristics, downlink conditional frequency characteristics, and uplink conditional frequency characteristics of each data packet, including:
[0116] a) Calculate statistical features, including mean, standard deviation, maximum value, minimum value, etc. For example:
[0117] a-1) Mean packet size:
[0118]
[0119] a-2) Standard deviation of packet size:
[0120]
[0121] a-3) Average uplink speed:
[0122]
[0123] a-4) Standard deviation of uplink speed:
[0124]
[0125] The result is approximately 320.7.
[0126] b) Calculate conditional frequency features
[0127] In fact, it is to calculate the frequency of uplink and downlink speed under specific conditions, such as the frequency of speed greater than a certain threshold. For example, the frequency of uplink speed greater than 0:
[0128]
[0129] Frequency of downlink speed greater than 0:
[0130]
[0131] The final initial feature set includes:
[0132] Average packet size: 660 Bytes
[0133] Packet size standard deviation: 348.2 Bytes
[0134] Average upload speed: 141.67 bytes / sec
[0135] Uplink speed standard deviation: 320.7 bytes / sec
[0136] Frequency of uplink speed greater than 0: 75%
[0137] Frequency of downlink speed greater than 0: 75%
[0138] These features can be used as the basis for further analysis and modeling to provide a comprehensive feature description, providing the basis for subsequent feature optimization and updating operations.
[0139] It's particularly noteworthy that the K value can be dynamically adjusted during this step to balance data processing speed and accuracy, adapting to different scenarios. For example, the controller automatically adjusts the K value based on peak network traffic and data processing capacity. For example, during high-traffic periods, the K value might be reduced to 50 to ensure data processing speed; during low-traffic periods, the K value might be increased to 200 to improve feature extraction accuracy and enhance the efficiency and quality of data updates.
[0140] 2-3) Using the Pearson Correlation Coefficient (PCC) and Random Forest Classifier (RFC) to optimize the initial feature set generated in step 2-2). By calculating the correlation between features, PCC helps us identify which features are most critical for describing network data, thereby allowing for feature ranking. RFC is used for classification tasks, evaluating the classification accuracy of different feature combinations to determine the optimal feature set, reduce unnecessary data processing, and improve system performance. The details are as follows:
[0141] 2-3-1) Calculate the Pearson correlation coefficient (PCC) between each feature vector in the initial feature set and each data type of the network data information as the importance score of each feature vector;
[0142] 2-3-2) Combine the feature vectors in the initial feature set in pairs to form several feature combinations;
[0143] 2-3-3) Calculate the PCC of each feature combination in step 2-3-2) and determine whether the PCC of each feature combination is greater than 0.9:
[0144] If PCC≤0.9, retain the feature combination in the initial feature set;
[0145] If PCC>0.9, the feature vectors with low correlation with the data type in the feature combination are deleted from the initial feature set;
[0146] 2-3-4) The remaining feature vectors in the initial feature set are sorted by importance using the extreme tree algorithm, and the stream data is classified several times according to the importance of each feature vector to obtain data feature sets corresponding to different classification situations;
[0147] 2-3-5) Finally, use RFC to judge the flow data classification accuracy of the data feature set corresponding to each classification situation, and select the data feature set with the highest flow data classification accuracy as the optimal feature set. There is a logical correlation between the optimal feature set and the network data information;
[0148] This demonstrates that this embodiment employs two machine learning methods, PCC and RFC, to optimize the initial feature set, improving both the quality of the feature set and the accuracy of classification. PCC measures the linear correlation between feature vectors and the target variable, while RFC evaluates the classification performance of the feature set by constructing multiple decision trees. This approach selects the most relevant feature vectors, improving the accuracy and efficiency of data processing and reducing unnecessary computational resource consumption.
[0149] 2-4) Divide the network data information update into several micro-batch update operations (i.e., a series of update operations) based on the optimal feature set.
[0150] In this way, the real-time streaming data API enables Spark Streaming to receive data streams at fixed time intervals, and to achieve efficient processing of real-time data by dividing the streaming data into micro-batches and processing each micro-batch as an RDD. This technical solution improves the immediacy and integrity of data collection, reduces data latency, and enhances user experience. This embodiment facilitates batch processing of data by dividing streaming data into continuous stream segments, thereby improving the flexibility and efficiency of data processing. This method of "extracting features from the first K data packets of each stream segment to form an initial feature set containing statistical features of packet size, packet arrival time, packet difference, uplink speed, and downlink speed" provides a basis for subsequent data analysis and update strategy formulation, can process real-time data more effectively, avoids the computational pressure brought about by processing a large amount of data at one time, and improves the response speed of the system.
[0151] 3) Convert a series of update operations into an update dependency graph based on network data information in the following manner:
[0152] According to the collected network data G0 and the updated network data G n , calculate the update from G0 to G n A series of update operations need to be performed in the process, and then according to the mutual dependence between the update operations, the network link bandwidth resources and the demand information of the update operation for the link bandwidth resources, for each given collected network data G0 and the updated network data G n The complete data update process builds an update dependency graph that meets its consistency requirements.
[0153] For example, suppose that updating from A to B requires five update operations, namely OP1, OP2, OP3, OP4 and OP5, where OP2 and OP3 depend on OP1, OP4 depends on OP2 and OP3, and OP5 depends on OP4. Then, an update dependency graph containing these dependencies can be constructed.
[0154] In this embodiment, the construction of the update dependency graph also takes into account network link bandwidth resources and update operation requirements. This includes analyzing network link status, assessing the impact of update operations on bandwidth, and optimizing update operation scheduling. For example, the update dependency graph notates the timing relationships between update operations to ensure the consistency of distributed updates and improve data processing speed and efficiency.
[0155] In this way, not only is the scheduling of update operations optimized by considering network link bandwidth resources and update operation requirements (such as data packet size, network delay, etc.), but in the process of building the update dependency graph, the network link status and the impact of update operations on bandwidth are analyzed to reasonably arrange the execution order and time of update operations, avoiding update failures or delays caused by insufficient bandwidth, improving the success rate and efficiency of update operations, reducing data processing delays, and improving the stability and reliability of the system. Even in the case of tight bandwidth, data updates can be completed efficiently and orderly.
[0156] In summary, this embodiment can not only significantly reduce update coordination time, improve data processing speed, and ensure the consistency and efficiency of distributed updates, but also clearly depict the logical relationship between different micro-batch update operations by constructing an update dependency graph, providing a foundation for distributed updates.
[0157] 4) Construct an update message based on the update dependency graph in the following manner;
[0158] Improved update messages are used to update network data. These messages are updated through distributed coordination between micro-batches. The improved update messages contain information about "operations," "preconditions," and "post-operations," and all improved update messages are sent simultaneously to the corresponding micro-batches. In this embodiment, the improved update messages have the following meanings:
[0159] ① Operation: represents an arbitrary update operation (or, the operation of the update message describes an arbitrary update instruction);
[0160] ② Preconditions: These represent a combination of pre-update operations that must be executed before executing the update operation. These preconditions are presented as logical Boolean expressions, and the update operations are connected by logical relationships. The operation can only be executed after the preconditions are met, ensuring the orderly execution of the update operation. This embodiment takes into account the real-time nature and complexity of network data. The preconditions also define time windows between operations to ensure that update operations are sequential. For example, this ensures that the current micro-batch update operation can only be initiated within 30 seconds of the previous operation to avoid time overlap during data processing and ensure the accuracy of data updates.
[0161] In this way, the orderly execution of update operations is ensured through the update message's operation description and precondition logic expression. Furthermore, the update message's operation description provides specific update instructions, while the precondition logic expression specifies the precondition state. The update operation can only be executed when the precondition is met. This ensures the correct execution order of update operations, improves the accuracy and consistency of update operations, and avoids data inconsistencies caused by improper operation order.
[0162] ③ Post-operation: represents a set of update operations that depend on the execution of this update operation.
[0163] For example, the "Operation" field might describe updating traffic statistics for a specific network link; the "Precondition" field might be a logical expression, such as "The previous micro-batch update has completed and the packet sequence length is greater than or equal to 100"; and the "Post-Operation" field might specify the data analysis tasks to be performed after the update. This message structure ensures the orderly execution of update operations and avoids data conflicts.
[0164] As can be seen, this embodiment clarifies the execution order of update operations by defining the update message structure, ensuring the consistency and correctness of distributed updates. The operation field in the update message structure describes the specific update instructions, the precondition field specifies the precondition state through a logical Boolean expression, and the postcondition field indicates the subsequent actions that need to be triggered after the update operation completes. This allows update operations to be executed in an orderly manner when the preconditions are met, avoiding data inconsistencies caused by improper operation order.
[0165] In other words, this embodiment defines a new update message structure through the operation, precondition, and postcondition fields, which clearly describes the logical relationship between update operations, thereby optimizing the update process. Furthermore, by describing the preconditions as logical Boolean expressions, the dependencies between update operations are ensured, thereby improving the accuracy of update operations. This ensures the dependencies between update operations, improves the accuracy of update operations, reduces data processing delays, and further improves the speed and efficiency of message construction.
[0166] 5) The controller performs distributed updates based on the update message in the following manner:
[0167] Cache as needed. First, cache the data that needs to be calculated. Then define the data that needs to be incrementally updated. Filter the required data from the RDD sequence. When a new batch arrives, perform incremental updates on the filtered incremental data. Finally, cache the calculation results again for the next batch of processing.
[0168] In this embodiment, the execution of the incremental update operation also includes quickly comparing the data and only processing the differences to reduce data processing delay.
[0169] This not only reduces data processing latency during incremental updates through rapid comparison and difference processing, but also rapidly compares data during incremental updates, processing only the differences and avoiding repeated processing of unchanged data. This reduces data processing latency, increases data processing speed, reduces computing resource consumption, and improves system responsiveness and stability.
[0170] The implementation of the caching and incremental update strategy also includes regularly clearing expired cache to avoid excessive memory usage and ensure stable system operation. That is, during the implementation of the caching and incremental update strategy, this embodiment regularly clears expired cache to free up memory space, avoiding the problem of excessive memory usage caused by excessive cached data. This ensures stable system operation, greatly improves system stability and reliability, and avoids system crashes or performance degradation caused by excessive memory usage.
[0171] In this embodiment, when performing caching and incremental update operations, there is also a logical association relationship between the caching and incremental update operations and the update dependency graph.
[0172] That is to say, in this embodiment, the cache and incremental update strategy are used to process micro-batch update operations, and a logical association relationship is established between the cache and incremental update strategies. By using this cache mechanism to store the latest network data information and execute a change-based data update strategy, large-scale network data can be effectively processed, thereby improving data processing speed and efficiency.
[0173] In other words, this embodiment improves data processing performance and efficiency through a caching mechanism and incremental update strategy. The caching mechanism stores recent network data information, reducing data access latency, while the incremental update strategy processes data based on changes, avoiding repeated processing of unchanged data. This significantly increases data processing speed, reduces computing resource consumption, and improves system responsiveness and stability.
[0174] It is worth noting that when the controller performs distributed updates based on update messages, it uses query and notification messages to coordinate the update order in the following manner:
[0175] In distributed interactions on the data plane, micro-batches are efficiently coordinated through two models: query messages and notification messages. When a controller receives an update message, it sends query messages to the servers corresponding to all update operations in the preconditions to ensure that the preconditions are met. These query messages confirm whether the update operations have been executed. Only when all preconditions are met—that is, when a reply confirms that all precondition update operations have been executed—is the current update operation allowed to execute.
[0176] Once the operations in an update message are executed by its corresponding node, that node then sends notification messages to all servers corresponding to the update operations in its postconditions. These notification messages inform other relevant servers that the update operation has completed, allowing subsequent updates that depend on this operation to proceed. Through this mechanism, micro-batches in the data plane can autonomously coordinate the order of operations, ensuring the consistency and efficiency of the update process while reducing dependence on the controller.
[0177] It is worth noting that micro-batches in the data plane interact in a distributed manner through two message modes: "query message" and "notification message". The execution order between operations is completely completed by the data plane. The controller only needs to construct corresponding update messages for all update operations according to the update dependency graph and send them to the data plane, so that a logical association relationship can be established between the update messages and the micro-batch update operations.
[0178] In this embodiment, the specific meanings of the two message modes are as follows:
[0179] ① Query message: When a controller receives an update message, it will send a query message to the servers corresponding to all update operations in its preconditions, asking whether these update operations have been executed.
[0180] ② Notification message: After the operation in an update message is executed by its corresponding node, the node will send a notification message to the servers corresponding to all update operations in its post-conditions, thereby notifying the servers that the update operation has been executed.
[0181] The query message and notification message in the present embodiment are logically associated with the micro-batch update operation. In this way, the data plane can utilize the message coordination mechanism formed by the query message and the notification message to autonomously coordinate the update order and reduce the dependence on the controller. This message coordination mechanism can monitor the data update execution process, identify potential conflicts (such as data dependency conflicts or resource competition conflicts), and adjust the update order in time, thereby avoiding the update failure or data inconsistency problem caused by the conflict, improving the efficiency and accuracy of data update, reducing data processing delay, and improving the stability and reliability of the system.
[0182] In summary, this embodiment utilizes a controller and cache strategy to achieve efficient execution of incremental update operations. Based on the operation sequence and cache status in the update dependency graph, the controller intelligently identifies the data segments that need to be updated. This avoids repeated processing of the same data, conserves computing resources, improves data update efficiency, reduces computing resource waste, and shortens data processing time.
[0183] 7) Use monitoring tools to track the update process and data flow, and adjust the update strategy based on user feedback and technological development as follows:
[0184] Based on the design requirements, the update method is deployed to the selected cloud service platform. WebSocket is used to implement real-time data updates and display. Data analysis tools are used to identify trends in user feedback, assess the impact of technological developments on user needs, and dynamically adjust update strategies to adapt to technological trends. In other words, by implementing real-time data updates and analyzing user feedback on the cloud service platform, this embodiment can quickly adapt to changes in network data and user needs, thereby achieving higher user satisfaction.
[0185] It is worth noting that this embodiment chooses to use WebSocket and data analysis tools to achieve dynamic adjustment of the update strategy. This is because WebSocket can provide full-duplex real-time communication capabilities, allowing user feedback to be transmitted to the system immediately, and data analysis tools can identify trends in user feedback, providing a basis for adjusting the update strategy, quickly responding to user needs and technological development, thereby improving the system's adaptability and user satisfaction.
[0186] In this embodiment, the application of the cache strategy also includes estimating the future data status based on the historical network data change pattern, and loading data that may need to be updated into the cache in advance.
[0187] For example, on a cloud service platform that deploys the instant computer network data update method described in the present invention, WebSocket is used to implement real-time data updates and displays, allowing users to instantly see changes in network data. At the same time, data analysis tools are used to deeply explore and analyze user feedback to identify potential trends and patterns in user feedback. Specifically, by analyzing the frequency and duration of user use of specific technologies or services, users' preferences and needs for different technologies and services can be understood. Through questionnaires, user interviews, and other methods, user satisfaction and feedback on technologies and services can be collected, and user suggestions for improvements to existing technologies and services and future expectations can be analyzed. Furthermore, attention can be paid to user acceptance and reactions to new technologies and new products to understand user needs and expectations for new technologies. These trends may include changes in user preferences and needs regarding update frequency, update content, and update methods.
[0188] In assessing the impact of technological development on user needs, we closely monitor industry trends in technological development and the application of emerging technologies. Through questionnaires, user interviews, social media analysis, and other methods, we collect user experiences with existing technologies and services, suggestions for improvement, and future expectations, to understand changes and trends in user needs. We predict the direction of technological development and its potential impact by collecting and analyzing current and future technological development trends, including but not limited to the latest research results, application cases, and market reports in fields such as cloud computing, big data, and artificial intelligence. We compare and analyze technological development trends with the results of user needs surveys to identify the matching points and potential gaps between technology and user needs, providing a basis for adjusting update strategies (that is, the basis for update strategy adjustments includes user behavior analysis results and technology change assessment reports to continuously optimize the update process). Based on these analysis and prediction results, we promptly adjust and optimize update strategies to adapt to the continuous changes in technological development and user needs.
[0189] Experiments have shown that existing technologies often require extensive computing resources for data updates, resulting in low efficiency. The coordination process for data updates is complex, requiring frequent interactions and increasing time overhead. Furthermore, existing technologies lack a mechanism for responding promptly to user feedback and technological trends, making it difficult to quickly adjust update strategies to adapt to changes.
[0190] In other words, with the advent of the big data era, the rapid growth of massive amounts of data has placed higher demands on data processing capabilities. Traditional data processing methods are no longer able to meet the requirements of real-time performance, efficiency, and scalability. Existing technologies typically use batch processing to process real-time streaming data, processing data in batches. This approach suffers from latency, making it difficult to respond to users' real-time needs, impacting the user experience. Furthermore, existing technologies are poorly adaptable to new types of data and struggle to keep pace with changes in data types, leading to errors in update operations. In traditional streaming data processing, the input of the data source and the output of the data processing are generally separate. In other words, the input of the data source is typically an infinite stream of data, while the output of the data processing is batch results. Traditional methods may stop data processing at a certain point in time and write the processed batch data to a storage medium. This approach introduces latency and fails to meet the rapid processing requirements of users for real-time activities such as social media content updates, online advertising placements, and logistics progress updates, resulting in a poor user experience. When updating existing computer application software data, it often requires the establishment and operation of computer network data, terminals, and users, and updates based on data feedback. In addition, the real-time update process will cause low update efficiency due to network delays or instability. Such delay time and data processing volume limit traditional stream processing capabilities.
[0191] The present invention is aimed at these problems. On the basis of a near real-time incremental data update solution based on Spark, it utilizes the query and notification mechanism in distributed interaction and combines it with caching and incremental update strategies. Only necessary data is processed, thereby saving computing resources and performing data updates more efficiently and in real time, thus meeting the needs of modern data processing.
[0192] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications made to the present invention by those skilled in the art without departing from the spirit of the present invention shall fall within the scope of protection of the present invention.
Claims
1. A method for real-time updating of computer network data based on cloud computing, characterized in that: The following steps are involved: 1) Collect network data information in real time from the data plane; 2) Decompose network data information into micro-batch update operations according to time intervals; 3) Convert the micro-batch update operation into an update dependency graph based on network data information; 4) Constructing update messages based on the update dependency graph; 5) The controller uses the update message to perform a distributed update operation. During the distributed update operation, the controller uses a monitoring tool to track the update operation process and data flow, and adjusts the update strategy based on the tracking results.
2. The computer network data real-time updating method according to claim 1, characterized in that: In step 1), a method of collecting network data information from the data plane in real time includes using a Spark Streaming module to obtain the network data information.
3. The computer network data real-time updating method according to claim 1, characterized in that: In step 2), the network data information is decomposed into micro-batch update operations according to time intervals, specifically including: 2-1) Dividing the flow data in the network data information into several flow segments according to fixed time intervals; 2-2) Extracting features from the first K data packets of each flow segment divided in step 2-1) to form an initial feature set; 2-3) Forming the optimal feature set through feature selection; 2-4) Divide the network data update into several micro-batch update operations based on the optimal feature set.
4. The computer network data real-time updating method according to claim 3, characterized in that: In step 2-2), the formation of the initial feature set includes: 2-2-1) Convert all data packets into "packet size sequence", "packet arrival time sequence", A five-tuple consisting of "packet difference sequence", "uplink speed sequence" and "downlink speed sequence"; 2-2-2) Use the quintuple to calculate the "statistical characteristics", "downlink conditional frequency characteristics" and "uplink conditional frequency characteristics" of each data packet; 2-2-3) Finally, the initial feature set is formed using the "statistical features", "downlink conditional frequency features" and "uplink conditional frequency features" of each data packet.
5. The computer network data real-time updating method according to claim 3, characterized in that: In step 2-3), the optimal feature set is formed through feature selection, including: 2-3-1) Calculate the Pearson correlation coefficient (PCC) between each feature vector in the initial feature set and each data type of the network data information as the importance score of each feature vector; 2-3-2) Combine the feature vectors in the initial feature set in pairs to form several feature combinations; 2-3-3) Calculate the PCC of each feature combination in step 2-3-2) and determine whether the PCC of each feature combination is greater than 0.9: If PCC≤0.9, retain the feature combination in the initial feature set; If PCC>0.9, the feature vectors with low correlation with the data type in the feature combination are deleted from the initial feature set; 2-3-4) The remaining feature vectors in the initial feature set are sorted by importance using the extreme tree algorithm, and the stream data is classified several times according to the importance of each feature vector to obtain data feature sets corresponding to different classification situations; 2-3-5) Finally, use RFC to judge the flow data classification accuracy of the data feature set corresponding to each classification situation, and select the data feature set with the highest flow data classification accuracy as the optimal feature set.
6. The computer network data real-time updating method according to claim 1, characterized in that: In step 3), the micro-batch update operation is converted into an update dependency graph based on the network data information, including: 3-1) Based on the collected network data G0 and the updated network data G n , calculate the update from G0 to G n Several micro-batch update operations that need to be performed in the process; 3-2) Based on the interdependence between each micro-batch update operation, network link bandwidth resources and the demand information of micro-batch update operation for link bandwidth resources, for each given collected network data G0 and updated network data G n The complete data update process builds an update dependency graph that meets consistency requirements.
7. The computer network data real-time updating method according to claim 1, characterized in that: In step 4), an update message is constructed according to the update dependency graph, including: defining an update message structure, including "operation", "precondition" and "post-operation" fields, where "precondition" is described in the form of a logical Boolean expression to ensure the dependency relationship between update operations.
8. The computer network data real-time updating method according to claim 1, characterized in that: In step 4), the update message is constructed according to the update dependency graph, including defining the message delivery path in the update dependency graph, so that the micro-batch update operation can autonomously determine the execution order and reduce the scheduling burden of the controller.
9. The computer network data real-time updating method according to claim 1, characterized in that: In step 5), the controller uses the update message to perform distributed update operations, including the controller performing incremental update operations based on the update dependency graph and cache strategy, identifying the data segments to be updated from the Rdd sequence, and determining the specific data range of the incremental update based on the cache status.
10. The computer network data real-time updating method according to claim 1, characterized in that: In step 5), the controller uses the update message to perform a distributed update operation, including using two message modes, "query message" and "notification message", to coordinate the update sequence.