Edge calculation method for mass data analysis
By constructing a resource allocation list and behavioral time series features, and combining them with grey relational scoring, the edge computing method for massive data analysis is optimized. This solves the problems of coarse resource evaluation and uneven scheduling in traditional methods, and achieves more efficient resource utilization and task processing.
Patent Information
- Application Number
- CN202511058613.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Traditional edge computing methods lack the ability to dynamically match task resource requirements with the actual capabilities of nodes when processing massive data analysis. This leads to uneven node allocation, scheduling conflicts, and a coarse resource assessment. Consequently, they are unable to make efficient scheduling in resource-constrained scenarios, which affects the overall task processing efficiency.
By acquiring the log field structure and data volume, a resource configuration list is constructed. Combined with the time series characteristics of the edge computing server's behavior, gray relational scoring is used to filter scheduling candidate nodes, analyze resource usage, optimize scheduling paths, and achieve precise resource allocation and parallel processing.
It improves the adaptability of edge node scheduling and the accuracy of task assignment, reduces resource waste and processing delay caused by scheduling imbalance, and enhances response speed and resource utilization efficiency.
Smart Images

Figure CN120950245A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an edge computing method for massive data analysis. Background Technology
[0002] The field of data processing technology involves the collection, transformation, storage, and analysis of digital information, aiming to extract useful information from raw data to support intelligent decision-making and system optimization.
[0003] Traditional edge computing methods for massive data analysis involve performing preliminary data filtering, compression, and analysis on edge nodes closer to the data source. However, when the amount of data generated is too large, the processing capacity of the central node is insufficient or the bandwidth bottleneck limits the overall performance, thus reducing transmission load and response latency.
[0004] Traditional edge computing methods mainly rely on the geographical advantage of nodes being close to the data source for simple initial screening and compression. Although this reduces the load on the central node, it lacks the ability to dynamically match task resource requirements with the actual capabilities of nodes when data traffic surges or node operating status fluctuates. This can easily lead to uneven node allocation or scheduling conflicts. Without considering the complexity of log field structures and resource call characteristics, resource assessment is too rough and cannot accurately determine the load level required for task execution. For example, when log field types are highly diverse but node configurations are fixed, task execution time may significantly exceed expectations. In addition, node selection based solely on static performance or simplified metrics cannot make efficient scheduling strategies in resource-constrained scenarios, causing some nodes to be overloaded and interrupted or experience performance bottlenecks, affecting the overall task processing efficiency. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies and propose an edge computing method for massive data analysis.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an edge computing method for massive data analysis, comprising the following steps:
[0007] S1: Obtain the number of fields, log entries, and field types in the Web server access logs, and construct the resource configuration list for the corresponding log data analysis task;
[0008] S2: The node operation data reported by the edge computing server is weighted and processed to construct behavioral time series features.
[0009] S3: Filter the set of candidate scheduling nodes by calculating the grey relational score between the behavioral time series features of the edge computing server and the resource configuration list of the log data analysis task;
[0010] S4: Analyze the resource usage of each node in the scheduling candidate node set under the log data analysis task resource requirements, sort all nodes, and obtain the adjusted scheduling path set;
[0011] S5: Based on the adjusted set of scheduling paths, the Web server access log analysis task is assigned to each edge computing server to obtain the data edge computing results.
[0012] As a further embodiment of the present invention, the resource configuration list includes thread requirements, memory usage, and channel concurrency values; the behavioral time series features include response latency trends, completion rate sequences, and changes in the number of interruptions; the scheduling candidate node set includes node identifiers, correlation scores, and filtering order; the scheduling path set includes path numbers, node sorting positions, and resource usage comparison values; and the data edge computing results include log analysis output items, node processing results, and task execution markers.
[0013] As a further aspect of the present invention, the step of obtaining the resource configuration list specifically includes:
[0014] S111: Obtain the three data items in the Web server access log: the number of fields, the number of log entries, and the number of field types. Determine the field structure composition based on the number of fields and the number of field types. Combine this with the total number of records represented by the number of log entries to establish a field structure record.
[0015] S112: Identify the proportion of field types in all the field structure records, combine the number of fields and the distribution of types to form the field occupancy structure, and calculate the log structure complexity by field dimension;
[0016] S113: Combining the complexity of the log structure with the data volume reflected by the field structure records, summarize and normalize the requirements of the corresponding log task on three resources during the processing: CPU thread call volume, memory usage and IO scheduling frequency, and obtain the resource configuration list.
[0017] As a further aspect of the present invention, the step of obtaining the behavioral time series features specifically includes:
[0018] S211: Obtain node operation data reported by the edge computing server within multiple consecutive scheduling cycles, extract three behavioral indicators: task response latency, processing completion rate, and number of interruption rollbacks, and construct a periodic behavior fluctuation sequence based on the distribution of the indicators in each cycle.
[0019] S212: Based on the periodic behavior fluctuation sequence, the periodic increase or decrease of task response latency is compared with the continuous periodic average of processing completion rate. Combined with the variation of interrupt rollback times between periods, the trend differences of the three behavioral indicators in the time dimension are identified to obtain the indicator change trend characteristics.
[0020] S213: Based on the trend characteristics of the aforementioned indicators, a relative relationship is established between the processing completion rate and the number of interruptions and rollbacks. Combined with the task response latency, a multi-dimensional feature set is formed to generate behavioral time series features.
[0021] As a further aspect of the present invention, the step of obtaining the scheduling candidate node set specifically includes:
[0022] S311: Call the resource configuration list of the behavior time series features and log data analysis task of the edge computing server, and construct a behavior resource association group based on the correspondence of parameters in numerical distribution, change direction and relative position;
[0023] S312: Apply the grey relational analysis algorithm to calculate the grey relational score between each indicator item in the behavioral time series features and each resource item in the resource configuration list, extract the scores corresponding to all edge computing servers, and obtain the grey relational score results.
[0024] S313: Based on the gray relational degree scoring results, the median value is set as the screening criterion according to the distribution range of all scores, and server nodes with scores higher than the screening criterion are selected to obtain the scheduling candidate node set.
[0025] As a further aspect of the present invention, the step of obtaining the adjusted scheduling path set specifically includes:
[0026] S411: Obtain the running status of the scheduling candidate node set under the resource requirements of the log data analysis task, extract the CPU utilization, memory usage and IO channel usage of each edge computing server under the current task conditions, and construct a node resource usage status group.
[0027] S412: Compare and analyze the resource usage characteristics of each node in the node resource usage status group with the requirement parameters in the task resource configuration list, and evaluate and rank them using the analytic hierarchy process to obtain the node resource ranking result.
[0028] S413: Based on the node resource sorting results, extract the specified server node path and obtain the adjusted scheduling path set.
[0029] As a further aspect of the present invention, the step of obtaining the data edge computing result specifically includes:
[0030] S511: Based on the adjusted set of scheduling paths, the Web server access log analysis task is assigned to each edge computing server in the order of the paths, and a log fragment scheduling mapping result is constructed.
[0031] S512: Based on the log fragment scheduling and mapping results, each edge computing server performs the processing operation of the received data fragments, extracts log elements by combining the field information structure, classifies and processes access behavior, response status and request content, and generates local processing results of log fragments.
[0032] S513: Based on the local processing results of the log fragments, the scheduling center integrates and merges the data returned by all edge computing servers to uniformly construct the edge computing results of the log data analysis task after completion on the edge side.
[0033] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0034] In this invention, when processing massive amounts of web server access logs, resource configuration requirements are clearly defined by introducing a combination of log field structure, field complexity, and data volume. This enables accurate prediction of thread, memory, and IO calls. Furthermore, behavioral sequences are constructed based on server behavior characteristics such as response latency, completion rate, and interruption status, forming a dynamic evaluation standard for scheduling capabilities. In the candidate node selection process, a grey relational scoring method between resource requirements and behavioral capabilities is integrated to improve the matching of node selection. The sorting optimization is combined with the real-time resource occupancy status of nodes, resulting in a better balance in task allocation paths from a multi-dimensional resource utilization perspective. After task scheduling, parallel processing is achieved through log fragment mapping, and then the results are aggregated at the edge to form a structured analysis output. Overall, this improves the adaptability of edge node scheduling and the accuracy of task assignment, reduces resource waste and processing delays caused by scheduling imbalances, and significantly enhances response speed, resource utilization efficiency, and data processing integrity in log analysis execution. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the main steps of the present invention;
[0036] Figure 2 This is a flowchart of step S1 of the present invention;
[0037] Figure 3 This is a flowchart of step S2 of the present invention;
[0038] Figure 4 This is a flowchart of step S3 of the present invention;
[0039] Figure 5 This is a flowchart of step S4 of the present invention;
[0040] Figure 6 This is a flowchart of step S5 of the present invention. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0042] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0043] Please see Figure 1 This invention provides a technical solution: an edge computing method for massive data analysis, comprising the following steps:
[0044] S1: Obtain the number of fields, log entries, and field types in the Web server access logs, and construct the resource configuration list for the corresponding log data analysis task;
[0045] S2: The node operation data reported by the edge computing server is weighted and processed to construct behavioral time series features.
[0046] S3: Filter the set of candidate scheduling nodes by calculating the grey relational score between the time series characteristics of the edge computing server's behavior and the resource configuration list of the log data analysis task;
[0047] S4: Analyze the resource usage of each node in the candidate scheduling node set under the resource requirements of the log data analysis task, sort all nodes, and obtain the adjusted scheduling path set;
[0048] S5: Based on the adjusted set of scheduling paths, the Web server access log analysis task is assigned to each edge computing server to obtain the data edge computing results;
[0049] The resource configuration list includes thread requirements, memory usage, and channel concurrency values. The behavioral time series characteristics include response latency trends, completion rate sequences, and changes in the number of interruptions. The scheduling candidate node set includes node identifiers, correlation scores, and filtering order. The scheduling path set includes path numbers, node sorting positions, and resource usage comparison values. The data edge computing results include log analysis output items, node processing results, and task execution markers.
[0050] Please see Figure 2 The specific steps to obtain the resource configuration list are as follows:
[0051] S111: Obtain the three data items in the Web server access log: the number of fields, the number of log entries, and the number of field types. Determine the field structure composition based on the number of fields and the number of field types. Combine this with the total number of records represented by the number of log entries to establish a field structure record.
[0052] To obtain three data items from the web server access logs: the number of fields, the number of log entries, and the number of field types, we first extract each log record from the raw log text and perform standardized preprocessing. All log entries are then split into lines, with each line representing a log record instance. For each record, delimiters such as spaces, tabs, or commas are used to identify the number of fields. For example, a log entry like "192.168.1.1--[10 / Mar / 2025:13:55:36+0000]'GET / index.htmlHTTP / 1.1'2001024" has 10 fields when separated by spaces. The total number of fields can be calculated by accumulating all log entries. The total number of fields is determined by using pattern matching techniques (such as regular expressions) to identify the type of each field. For example, the IP address field is "xxx.xxx.xxx.xxx", the timestamp field is "[xx / Mon / yyyy:hh:mm:ss+zzzz]", and the HTTP method field is "GET", "POST", etc. The number of field types is the total number of different types of fields. If a total of 7 field types are identified, such as IP address, timestamp, method, URL, protocol, status code, and response size, then the number of field types is 7. The number of log entries is the number of lines after splitting. For example, if a file has 100,000 lines, then it is 100,000 records. The field structure record is constructed based on the above three 4E2A data.
[0053] S112: Identify the proportion of field types in all field structure records, combine the number of fields and the distribution of types to form the field occupancy structure, and calculate the log structure complexity by field dimension;
[0054] To identify the proportion of each field type in all field-structured records, we first need to extract the field type count for each field-structured record. Assuming a field set of 11 items, including 5 numeric fields, 3 string fields, 2 time fields, and 1 boolean field, then the proportion of each type is as follows: The proportion of numeric fields is... The proportion of string fields is The proportion of time-type fields is The proportion of Boolean fields is In this statistical process, the value format of each log record needs to be checked field by field. For example, if the field "status" is 200, it is an integer field; if "user_agent" is "Mozilla / 5.0", it is a string; if "request_time" is "0.215", it is a floating-point number; if "timestamp" is "2025-06-01T12:30:15", it is a time field; and if "is_cache_hit" is true or false, it is a boolean field. The number of similar fields in all field structure records is summed to obtain the field type ratio of the entire sample, which constitutes the field occupancy structure. Then, the field type ratio is used as the distribution probability to calculate the log structure complexity, which can be expressed using the information entropy formula:
[0055]
[0056] Where H: field structure complexity (unit: bits), representing the degree of uncertainty in the distribution of field types; n: total number of field types, for example, commonly 4 types (integer, string, time, boolean); p i : The proportion of the i-th field type among all fields (ratio value), for example, p1 = 0.45 means that the first type of field accounts for 45%; log2: a logarithmic function with base 2, used to measure information entropy in bits.
[0057] The percentages by field type are as follows: numeric type p1 = 0.45, string type p2 = 0.27, and time type p3.
[0058] Taking p4 = 0.18 and p4 = 0.09 as examples, substituting into the above formula yields:
[0059] H = -
[0060] (0.45·log20.45+0.27·log20.27+0.18·log20.18+0.09·log20.09)
[0061] ≈1.75;
[0062] Furthermore, if the proportions of another log source field are: numeric type p1 = 0.8, string type p2 = 0.1, and time type p3 = 0.1, its structural complexity is calculated as follows:
[0063] H=-(0.8·log20.8+0.1·log20.1+0.1·log20.1)≈0.92;
[0064] This indicates high field concentration and low structural complexity, with field usage concentrated in a few types. If the structural complexity range is set as follows: -H < 1.0: Low complexity (single field type, such as mostly integer or time types); -1.0 ≤ H < 1.5: Medium complexity (some field diversity); -H ≥ 1.5: High complexity (evenly distributed field types, diverse structures).
[0065] For example, if server A's log contains 12 fields, distributed as follows: 5 integer fields, 3 string fields, 2 floating-point fields, and 2 boolean fields, then... Substituting into the entropy formula, we get:
[0066] H = -
[0067] (0.417·log20.417+0.25·log20.25+0.167·log20.167+0.167·log20.167)
[0068] ≈1.89.
[0069] It belongs to a high-complexity structure. The execution process needs to be combined with the log field extraction module, field type determination module, field proportion calculation module, and structural complexity calculation module to gradually complete the analysis of the field structure and output the complexity.
[0070] S113: Combining the log structure complexity and the data volume reflected by the field structure records, summarize and normalize the requirements of the corresponding log task for three resources during the processing of CPU thread calls, memory usage and IO scheduling frequency, and obtain the resource configuration list.
[0071] Combining the log structure complexity and the data volume reflected by the field structure records, the first step is to obtain the structure complexity value and the number of log entries for each log. For example, server A has a log structure complexity of 1.92 and a total of 180,000 records, while server B has a complexity of 0.86 and 32,000 records. By combining the structure complexity and the number of records into data feature groups, for example, A corresponds to feature group (1.92, 180,000), B to (0.86, 32,000), and so on, a reference dataset is constructed by aggregating multiple server samples. Next, it is necessary to analyze the CPU thread call volume, memory usage, and IO scheduling frequency values corresponding to each set of features in the sample data. If server A has 8 threads, 3.8GB of memory, and 150 I / O operations per second, while server B has 3 threads, 1.2GB of memory, and 60 I / O operations per second, then we need to perform horizontal normalization on each dimension to standardize the subsequent list. For example, if we set the maximum range for the number of threads to 16 and the minimum to 2, then the normalized value for the number of threads for server A is (8-2) / (16-2) = 0.43, and for server B it is (3-2) / (16-2) = 0.07. If we set the maximum range for memory usage to 8GB and the minimum to 0.5GB, then for A it is (3.8-0.5) / (8-0.5) = 0.47, and for B it is (1.2-0.5) / ... Given (8-0.5) = 0.09, and an I / O scheduling frequency of a maximum of 200 times / second and a minimum of 20 times / second, then A is normalized to (150-20) / (200-20) = 0.72, and B to (60-20) / (200-20) = 0.22. The final resource configuration list is: A: Threads 0.43, Memory 0.47, I / O frequency 0.72; B: Threads 0.07, Memory 0.09, I / O frequency 0.22. The normalized resource indicators form a three-dimensional vector, representing the intensity of system resource demand for different log structures during task execution. If used for task scheduling, the normalized results can be converted into weighted criteria for resource allocation sorting, for example... When the resource pool has 2GB of remaining memory, 6 threads, and a disk I / O frequency of 100 times / second, it will be prioritized for tasks with lower overall resource requirements. For example, task B has a smaller total resource requirement and a lower proportion, so it can be executed first. The normalized values of multiple samples are combined to further construct a resource threshold reference. For example, a thread normalized value exceeding 0.75 can be considered a high concurrency requirement, a value below 0.25 can be considered a low thread dependency task, a memory normalized value exceeding 0.65 can be considered a high memory usage task, and an I / O frequency normalized value above 0.7 can be considered a intensive I / O operation task. In this way, the association mapping process between the log structure and resource requirements is completed, and a standardized resource configuration data table is finally output.
[0072] Please see Figure 3 The specific steps for obtaining behavioral time series features are as follows:
[0073] S211: Obtain node operation data reported by the edge computing server within multiple consecutive scheduling cycles, extract three behavioral indicators: task response latency, processing completion rate, and number of interruption rollbacks, and construct a periodic behavior fluctuation sequence based on the distribution of the indicators in each cycle.
[0074] To obtain node operation data reported by edge computing servers over multiple consecutive scheduling cycles, the first step is to collect operation logs or status reports for each scheduling cycle in chronological order. From the data in each cycle, three key performance indicators (KPIs) are extracted: task response latency, processing completion rate, and number of interruption rollbacks. Task response latency refers to the time elapsed from task initiation to node feedback, which is obtained by comparing the timestamp fields in the scheduling record. For example, if a task is initiated at 12:00:01 and completes at 12:00:04, the task response latency is 3 seconds. The processing completion rate is calculated by dividing the number of successfully processed tasks per cycle by the total number of tasks. For example, if 85 tasks are successfully processed within a cycle out of a total of 100 tasks, the completion rate is 85%. Interruption rollbacks... The number of times refers to the number of times a node reschedules after a task is terminated due to an anomaly. This can be accumulated and summed through the failure retry record field. For example, if there are 4 tasks that trigger the rollback process within a cycle, then the number of interruption rollbacks in that cycle is 4. All three indicators in the scheduling cycle are recorded and saved in time series form to form a cycle behavior data table. Then, the distribution of each indicator value within the cycle is sorted out to construct a cycle behavior fluctuation sequence. For example, the task response latency, processing completion rate, and number of interruption rollbacks in 10 consecutive scheduling cycles are arranged sequentially for subsequent trend judgment and behavior evaluation. The entire process relies on the consistency of cycle division and the integrity of node-reported data to ensure the comparability of behavior data, and finally, a structured cycle fluctuation sequence data set is obtained.
[0075] S212: Based on the periodic behavior fluctuation sequence, the periodic increase or decrease of task response latency is compared with the continuous periodic average of processing completion rate. Combined with the variation of interrupt rollback times between periods, the trend differences of the three behavioral indicators in the time dimension are identified, and the trend characteristics of indicator changes are obtained.
[0076] Based on the cyclical behavior fluctuation sequence, the periodic variation trend of task response latency needs to be analyzed. This requires calculating the difference between each cycle and the previous cycle. For example, if the response latency in a scheduling cycle is 4 seconds and the previous cycle was 3.6 seconds, the increase is 0.4 seconds. If the next cycle is 4.3 seconds, the increase continues. A sequence of increases and decreases in response latency is constructed. The processing completion rate needs to be averaged over consecutive cycles to smooth the trend direction. For example, if the completion rates for three consecutive cycles are 85%, 88%, and 86%, the average is (85+88+86) / 3 = 86.33%. By observing the trends in response latency in parallel, it can be identified whether the two indicators are rising synchronously. Alternatively, one might rise and the other fall. Then, the magnitude of the change in the number of interruption rollbacks is introduced to analyze its impact. The change in the number of interruption rollbacks can be recorded as the absolute value of the difference between adjacent weeks. For example, if there are 5 rollbacks in the current cycle and 3 rollbacks in the previous cycle, the change is 2. By observing whether the change is consistent with the fluctuation of response latency or opposite to the fluctuation of completion rate, the mutual trend relationship of the three on the time axis can be determined. If the response latency continues to rise while the completion rate fluctuation decreases and the number of rollbacks increases frequently, it can be regarded as the system's instability under the current resource or load scheduling strategy. Output the change curve characteristics and mutual differences of the three behavioral indicators on the time dimension.
[0077] S213: Based on the trend characteristics of indicator changes, the processing completion rate and the number of interruptions and rollbacks are constructed into a relative relationship, and combined with the task response latency to form a multi-dimensional feature set, generating behavioral time series features;
[0078] Based on the trend characteristics of indicator changes, a relative relationship is established between the completion rate and the number of interruptions and rollbacks. For example, the ratio of the completion rate value to the number of rollbacks in each cycle is compared. If the completion rate in a certain cycle is 90% and there are 3 rollbacks, the success rate per rollback is 90 / 3 = 30. If the completion rate in another cycle is 80% and there are 4 rollbacks, the ratio is 80 / 4 = 20, indicating that the cost of rollbacks has increased. This relative ratio can form a stability response coefficient. Combined with the task response latency data of the corresponding cycle, a three-dimensional feature set is constructed, where the completion rate represents system efficiency, and the number of rollbacks... The numbers reflect stability, and the response latency reflects processing speed. These three indicators are combined into a feature vector. For example, the data for period 1 is 85% completion rate, 3 rollbacks, and 4 seconds response latency; for period 2, it is 88%, 2 rollbacks, and 3.8 seconds. These three sets of values are concatenated in sequence to construct a continuous behavioral time series feature set, which is used for subsequent processing of system state modeling or anomaly detection. The entire construction process iterates step by step in each scheduling cycle, calculating the ratio of completion rate to rollbacks. By combining this with latency, a joint feature value sequence of the three periodic indicators is formed, realizing the time series expression of multi-period behavioral data.
[0079] Please see Figure 4The specific steps for obtaining the candidate node set for scheduling are as follows:
[0080] S311: The resource configuration list for the action time series features and log data analysis task of calling the edge computing server, and construct the action resource association group based on the correspondence of parameters in terms of numerical distribution, direction of change and relative position;
[0081] The resource configuration list provided by the call behavior time series feature and log data analysis task must first ensure consistency in the correspondence between the two in the scheduling cycle time dimension. For each server, its behavior time series feature dataset is extracted, including information such as task response latency, processing completion rate, and number of interrupt rollbacks within multiple scheduling cycles. Simultaneously, data items such as CPU thread call volume, memory usage, and IO scheduling frequency for that server within the same cycle are extracted from the resource configuration list. Next, the correspondence between each parameter needs to be analyzed in three dimensions: first, the numerical distribution dimension, such as whether response latency and CPU thread usage are both in a high range. If response latency repeatedly exceeds 5 seconds and the number of thread calls repeatedly exceeds 10, then they constitute a numerically similar distribution. Second, the direction of change dimension, i.e., observing whether the two data points increase or decrease simultaneously within the cycle, for example, within cycle 1 to cycle 3. The task response latency increased from 3.5 seconds to 4.2 seconds, then to 5.1 seconds, while the number of CPU thread calls increased from 6 to 8, then to 10, indicating that the two changes were in the same direction. The third dimension is relative ranking, which involves ranking all servers under this indicator within each period. For example, if the task response latency ranks 5th out of 20 servers and the number of thread calls ranks 4th, the close rankings indicate that the two are in similar relative positions in the current period. The correspondence of the above three dimensions is then aggregated. The three indicators in the behavioral characteristics of each server are paired with the three resource items in the resource configuration list, forming 9 sets of behavior-resource correspondences. The correlation of each pairing item in terms of value, direction of change, and ranking is determined through the above analysis. The results of all pairing items are aggregated in sequence according to the server number, and finally, a behavioral resource association group is constructed for each edge computing server.
[0082] S312: Apply the grey relational analysis algorithm to calculate the grey relational score between each indicator item in the behavioral time series features and each resource item in the resource configuration list, extract the scores corresponding to all edge computing servers, and obtain the grey relational score results.
[0083] Applying the grey relational analysis algorithm requires comparing the time-series characteristics of each edge computing server's behavior with each resource item in the resource configuration list pairwise to calculate the grey relational score between three behavioral indicators (task response latency, processing completion rate, and number of interrupt rollbacks) and three resource indicators (CPU thread call volume, memory usage, and IO scheduling frequency). In actual calculations, all behavioral and resource data must first be normalized. The original data is scaled to between 0 and 1 using the range method to ensure comparability between different units. For example, if the task response latency data for a server ranges from 3.5 seconds to 5.1 seconds over 10 scheduling cycles, the standard value for the response latency in the first cycle after normalization is calculated as follows:
[0084] (3.5-3.5) / (5.1-3.5) = 0. The number of thread calls in the resource configuration is 6 to 12. If the first cycle has 6 calls, the normalization is 0. After normalization, for each row and resource pair, calculate the absolute difference for each cycle to form a difference sequence. Then, calculate the grey relational coefficient for each cycle using the following formula:
[0085]
[0086] in, The gray relation coefficient represents the i1th behavioral indicator (e.g., task response latency) and the kth resource indicator (e.g., CPU threads) in the kth scheduling cycle. Δ refers to the absolute difference between the i1th behavioral indicator and the kth resource indicator after normalization during the period; min Δ: The minimum difference among all periods and paired combinations; max : The maximum difference among all periods and paired combinations; ρ: the resolution coefficient, usually taken as 0.5, used to adjust the discrimination sensitivity;
[0087] If the difference between the two normalized sequences of task response latency and CPU thread usage over 10 periods is 0.05, and both the maximum and minimum difference are 0.05, then:
[0088]
[0089] This indicates perfect correlation. If the difference between the other pairings is 0.1, 0.25, 0.3, etc., with a maximum difference of 0.4 and a minimum of 0.1, then the correlation coefficient for the first period is:
[0090]
[0091] The second cycle is:
[0092]
[0093] After completing the correlation coefficients for all periods, their average value is calculated as the final grey relational score for the behavior and resource indicator pair. For example, the average for 10 periods is 0.81. This calculation process pairs the three behavioral indicators with the three resource indicators into nine pairs, calculates the average score for each pair, and finally forms nine score values for each edge node as a quantitative representation of its behavior and resource relationship, thus completing the grey relational score extraction process.
[0094] S313: Based on the grey relational degree scoring results, the median value is set as the screening criterion according to the distribution range of all scores, and server nodes with scores higher than the screening criterion are selected to obtain the scheduling candidate node set;
[0095] First, all nine server ratings are collected and organized according to their pairing relationships with behavior and resource items, forming a rating set for each pairing. For example, the ratings for "response latency and CPU threads" are compiled into one list, the ratings for "processing completion rate and memory usage" into another list, and so on. Then, each rating set is sorted, and the median value of its distribution range is calculated to construct the screening criteria. For example, if the "response latency and CPU threads" rating set contains 20 values, after sorting them from smallest to largest, the average of the 10th and 11th values in the middle position is taken as the median value for screening. Select the criteria, and then compare the scores of all server nodes under the pairing item with the corresponding screening criteria. If the score is higher than the median value, it is considered to have a strong behavioral correlation in the resource relationship dimension. In this way, the nine pairing items are screened and marked respectively. Then, the pairing items marked as "higher than the median value" are merged and counted. If a server has six or more scores higher than the corresponding median standard in the nine items, it is considered that the overall behavioral resource correlation of the server is high and it is included in the scheduling candidate node set. Finally, several sets of server node numbers are formed as a resource scheduling candidate list with preferred characteristics.
[0096] Please see Figure 5 The specific steps for obtaining the adjusted set of scheduling paths are as follows:
[0097] S411: Obtain the running status of the scheduling candidate node set under the resource requirements of the log data analysis task, extract the CPU utilization, memory usage and IO channel utilization of each edge computing server under the current task conditions, and construct a node resource usage status group.
[0098] To obtain the operational status of the scheduling candidate node set under the resource requirements of the log data analysis task, the resource operation indicators of each edge computing server must first be collected in real time before task loading or at the initial stage of the task. These indicators include CPU utilization, memory usage, and I / O channel utilization. CPU utilization can be directly read as the total percentage of processor usage using system monitoring tools. Memory usage is calculated as the ratio of currently used memory to total memory in the operating system resource table. For example, if a server is currently using 3.2GB of memory and the total memory is 8GB, the utilization rate is 40%. I / O channel utilization can be estimated by the proportion of disk or network I / O interface activity during specific time periods. For example, if I / O activity time is 6 seconds out of 10 seconds, the I / O channel utilization rate is 60%. These three data points are collected from each server in the scheduling candidate set and recorded in a data structure based on server nodes, forming a node resource usage status group. The three indicator values of each node will serve as important criteria for judging its resource availability and task adaptability. This process must maintain consistency with the indicator types in the task resource requirement table and remain real-time to facilitate subsequent node adaptation and ranking evaluation. Finally, the standardized collection of node resource usage characteristics is completed.
[0099] S412: Compare and analyze the resource usage characteristics of each node in the node resource usage status group with the requirement parameters in the task resource configuration list, evaluate and rank them using the analytic hierarchy process, and obtain the node resource ranking results.
[0100] The resource usage characteristics of each node in the node resource usage group are compared and analyzed against the resource configuration parameters required by the log data analysis task. The Analytic Hierarchy Process (AHP) is used to evaluate and rank the server nodes to determine the priority order of resource suitability. Log data analysis tasks often involve decoding, extracting, filtering, structurally analyzing, and storing a large number of log lines. These tasks are typically CPU-intensive, with moderate memory requirements and lower requirements for I / O scheduling frequency. Based on these actual operational characteristics, the relative importance of the three resource indicators in this example is set as follows: CPU utilization is the most critical, followed by memory usage, and then I / O channel utilization.
[0101] The resource importance preferences are set as follows: CPU is 3 times more important than memory, CPU is 5 times more important than I / O, and memory is 2 times more important than I / O. A third-order judgment matrix is constructed as follows:
[0102]
[0103] Calculate the sum of each column separately: First column: 1 + 1 / 3 + 1 / 5 ≈ 1.867, Second column: 3 + 1 + 0.5 = 4.5, Third column: 5 + 2 + 1 = 8.
[0104] Divide each element in the matrix by the sum of its columns to form a normalized matrix:
[0105]
[0106] The weight vector is obtained by averaging the values of each row: CPU utilization weight:
[0107] (0.535+0.667+0.625) / 3≈0.609, Memory Usage Weight:
[0108] (0.179+0.222+0.25) / 3≈0.217, IO channel utilization weight:
[0109] (0.107+0.111+0.125) / 3≈0.114.
[0110] The current resource usage of each node is represented in a standardized form: Node A: CPU utilization 70%, memory usage 40%, IO usage 50%, represented as [0.7, 0.4, 0.5], Node B: CPU utilization 50%, memory usage 30%, IO usage 40%, represented as [0.5, 0.3, 0.4].
[0111] Using the weight of each resource as a coefficient, the feature values of the node resources are weighted and summed to form a comprehensive suitability score, calculated as follows:
[0112] Node A Score: S A =0.609×0.7+0.217×0.4+0.114×0.5=0.5701;
[0113] Node B Score: S B =0.609×0.5+0.217×0.3+0.114×0.4=0.4152.
[0114] Conclusion: Node A scores 0.5701, and Node B scores 0.4152, indicating that Node A's resource structure better meets the resource requirements of the log data analysis task, and therefore has a higher priority in the ranking. Using this method, resource suitability calculations can be performed on all candidate server nodes, and a ranking list can be generated.
[0115] S413: Based on the node resource sorting results, extract the specified server node path and obtain the adjusted scheduling path set;
[0116] First, all candidate server nodes are sorted from highest to lowest score based on their resource suitability rating. For example, if the system is configured to allow three selectable paths during scheduling, the top three nodes with the highest scores are selected as the target node path set. If nodes with the same score exist during the sorting process, secondary filtering can be performed according to backup rules. For example, nodes with lower CPU utilization or those physically closer to the current task's data source can be prioritized to improve scheduling efficiency and response speed. For each selected node, its physical path information, including node identifier, access address, service port, and network channel configuration, must be extracted and recorded to form complete path data. Finally, this path information is combined into a new set of scheduling paths.
[0117] Please see Figure 6 The specific steps for obtaining the data edge computing results are as follows:
[0118] S511: Based on the adjusted set of scheduling paths, the Web server access log analysis task is assigned to each edge computing server in the order of the paths, and the log fragment scheduling mapping result is constructed.
[0119] Based on the adjusted scheduling path set, the Web server access log analysis task is distributed to various edge computing servers according to the path order. First, the original log dataset needs to be sliced into equal portions or slices proportional to resource capacity. For example, if the total log data volume is 1.2 million records, and the target scheduling path set contains 4 edge nodes, with node A having the highest resource score and allocating 35% of the total, nodes B and C having similar scores and allocating 25% each, and node D having the lowest resource score and allocating 15%, then approximately 420,000, 300,000, 300,000, and 180,000 log records are extracted according to the corresponding proportions and distributed to each node. Simultaneously, the task number and node mapping relationship during the distribution process are recorded, constructing a log fragment scheduling mapping table. This mapping table needs to include information such as each node number, the start and end numbers of the received logs, the data distribution timestamp, and the structure type of the allocated fields. After the entire scheduling mapping is completed at the scheduling center, the task fragments are pushed to the edge nodes through a distributed transmission mechanism, thereby establishing the path allocation structure between task fragments and processing nodes.
[0120] S512: Based on the log fragment scheduling and mapping results, each edge computing server performs the processing operation of the received data fragments, extracts log elements by combining field information structure, classifies and processes access behavior, response status and request content, and generates local processing results of log fragments.
[0121] Based on the log fragment scheduling and mapping results, each edge computing server needs to independently complete the processing operation of its received data fragments. First, it loads the corresponding fragment log record and parses it according to the field structure format. For example, each log entry contains fields such as IP address, access time, request path, status code, and user agent. The field values are extracted and stored in a local data table. Next, the access behavior fields are categorized. Module types can be divided according to keywords contained in the access path; for example, " / api / " indicates an API call, and " / login" is categorized as user authentication. The status code field is categorized and statistically analyzed as success, redirection, client error, server error, etc. The request content field can be further parsed to extract parameter structures, for example, "GET..."
[0122] The query " / product? id=123" extracts requests with the GET method, path " / product", and parameter "id=123". All results are organized chronologically or by request source to form a node-level log fragment local processing result set. Each processing result should include access behavior category statistics, status code distribution, request type details, and a summary of key fields, and should be labeled with the node number and task number.
[0123] S513: Based on the local processing results of log fragments, the scheduling center integrates and merges the data returned by all edge computing servers to uniformly construct the edge computing results of the log data analysis task after it is completed on the edge side.
[0124] Based on the partial processing results of log fragments returned by each edge computing server, the scheduling center collects the output data of all processing nodes and concatenates and merges them according to the task number and log fragment number. First, the categorized statistics are summed. For example, if the number of accesses for "interface call type" by each node is 120,000, 100,000, 90,000, and 60,000 respectively, the merged result is 370,000. Then, the status code distribution is normalized and synthesized. For example, if the number of 2XX successful statuses processed by each node is 240,000, 220,000, 180,000, and 100,000 respectively, the sum is 740,000. Request content can be deduplicated and merged by concatenating fields and using unique identifier fields. If different nodes process duplicate request paths or IP access sequences, the latest record must be retained during merging, sorted by access timestamp and source. After integration, the data output by the scheduling center is the log analysis result data jointly generated by each edge server during task execution, representing the full completion status of the task in the distributed environment under the edge computing architecture.
[0125] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. An edge computing method for massive data analysis, characterized in that, Includes the following steps: S1: Obtain the number of fields, log entries, and field types in the Web server access logs, and construct the resource configuration list for the corresponding log data analysis task; S2: The node operation data reported by the edge computing server is weighted and processed to construct behavioral time series features. S3: Filter the set of candidate scheduling nodes by calculating the grey relational score between the behavioral time series features of the edge computing server and the resource configuration list of the log data analysis task; S4: Analyze the resource usage of each node in the scheduling candidate node set under the log data analysis task resource requirements, sort all nodes, and obtain the adjusted scheduling path set; S5: Based on the adjusted set of scheduling paths, the Web server access log analysis task is assigned to each edge computing server to obtain the data edge computing results.
2. The edge computing method for massive data analysis according to claim 1, characterized in that, The resource configuration list includes thread requirements, memory usage, and channel concurrency values. The behavioral time series features include response latency trends, completion rate sequences, and changes in the number of interruptions. The scheduling candidate node set includes node identifiers, correlation scores, and filtering order. The scheduling path set includes path numbers, node sorting positions, and resource usage comparison values. The data edge computing results include log analysis output items, node processing results, and task execution markers.
3. The edge computing method for massive data analysis according to claim 1, characterized in that, The specific steps for obtaining the resource configuration list are as follows: S111: Obtain the three data items in the Web server access log: the number of fields, the number of log entries, and the number of field types. Determine the field structure composition based on the number of fields and the number of field types. Combine this with the total number of records represented by the number of log entries to establish a field structure record. S112: Identify the proportion of field types in all the field structure records, combine the number of fields and the distribution of types to form the field occupancy structure, and calculate the log structure complexity by field dimension; S113: Combining the complexity of the log structure with the data volume reflected by the field structure records, summarize and normalize the requirements of the corresponding log task on three resources during the processing: CPU thread call volume, memory usage and IO scheduling frequency, and obtain the resource configuration list.
4. The edge computing method for massive data analysis according to claim 3, characterized in that, The specific steps for obtaining the behavioral time series features are as follows: S211: Obtain node operation data reported by the edge computing server within multiple consecutive scheduling cycles, extract three behavioral indicators: task response latency, processing completion rate, and number of interruption rollbacks, and construct a periodic behavior fluctuation sequence based on the distribution of the indicators in each cycle. S212: Based on the periodic behavior fluctuation sequence, the periodic increase or decrease of task response latency is compared with the continuous periodic average of processing completion rate. Combined with the variation of interrupt rollback times between periods, the trend differences of the three behavioral indicators in the time dimension are identified to obtain the indicator change trend characteristics. S213: Based on the trend characteristics of the aforementioned indicators, a relative relationship is established between the processing completion rate and the number of interruptions and rollbacks. Combined with the task response latency, a multi-dimensional feature set is formed to generate behavioral time series features.
5. The edge computing method for massive data analysis according to claim 4, characterized in that, The specific steps for obtaining the set of scheduling candidate nodes are as follows: S311: Call the resource configuration list of the behavior time series features and log data analysis task of the edge computing server, and construct a behavior resource association group based on the correspondence of parameters in numerical distribution, change direction and relative position; S312: Apply the grey relational analysis algorithm to calculate the grey relational score between each indicator item in the behavioral time series features and each resource item in the resource configuration list, extract the scores corresponding to all edge computing servers, and obtain the grey relational score results. S313: Based on the gray relational degree scoring results, the median value is set as the screening criterion according to the distribution range of all scores, and server nodes with scores higher than the screening criterion are selected to obtain the scheduling candidate node set.
6. The edge computing method for massive data analysis according to claim 5, characterized in that, The specific steps for obtaining the adjusted set of scheduling paths are as follows: S411: Obtain the running status of the scheduling candidate node set under the resource requirements of the log data analysis task, extract the CPU utilization, memory usage and IO channel usage of each edge computing server under the current task conditions, and construct a node resource usage status group. S412: Compare and analyze the resource usage characteristics of each node in the node resource usage status group with the requirement parameters in the task resource configuration list, and evaluate and rank them using the analytic hierarchy process to obtain the node resource ranking result. S413: Based on the node resource sorting results, extract the specified server node path and obtain the adjusted scheduling path set.
7. The edge computing method for massive data analysis according to claim 6, characterized in that, The specific steps for obtaining the data edge computing results are as follows: S511: Based on the adjusted set of scheduling paths, the Web server access log analysis task is assigned to each edge computing server in the order of the paths, and a log fragment scheduling mapping result is constructed. S512: Based on the log fragment scheduling and mapping results, each edge computing server performs the processing operation of the received data fragments, extracts log elements by combining the field information structure, classifies and processes access behavior, response status and request content, and generates local processing results of log fragments. S513: Based on the local processing results of the log fragments, the scheduling center integrates and merges the data returned by all edge computing servers to uniformly construct the edge computing results of the log data analysis task after completion on the edge side.
Citation Information
Patent Citations
Server resource scheduling system integrating AI and edge computing
CN119718682A
Edge computing scheduling method and system for heterogeneous multi-source sensor
CN119960950A
Cited By
Platform information low-delay concurrent communication transmission method and system
CN121125648A
Secure access method for privacy data of Internet of Things
CN121750380A