Data optimization method and system based on machine learning

Through the machine learning-based data optimization method, time stamp differences and amplitude fluctuations are used to identify data abnormalities, reconstruct data points, and optimize data storage and transmission, the problem of inefficient data processing in the existing technology is solved, and efficient data management and decision support is achieved.

CN120448900AInactive Publication Date: 2025-08-08ANHUI SCI & TECH UNIV
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510527408.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is unable to effectively respond to real-time changes in large-scale data sets, lack of flexibility, and fail to identify and adapt to changes in data patterns, resulting in delays or misjudgment in data analysis and decision-making processes, especially in inefficient processing in application scenarios that require continuous monitoring and immediate response.

Method used

Through machine learning-based methods, time stamp difference calculation, amplitude fluctuation screening exceptions, dispersion detection and path consistency evaluation are used to establish data feature maps, analyze data deviations, reconstruct abnormal data points, optimize data storage and transmission processes, classify access levels, and schedule transmission priorities.

Benefits of technology

Improve the efficiency of data during storage and transmission, ensure data quality, reduce storage space requirements, realize efficient data processing and analysis, and support rapid decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448900A_ABST
    Figure CN120448900A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data optimization, in particular to a data optimization method and system based on machine learning, and the method comprises the following steps: optimizing data based on machine learning, establishing feature mapping through timestamp difference calculation, amplitude fluctuation screening abnormity, dispersion detection and path consistency evaluation, analyzing data deviation, and marking trajectory deviation, and reconstructing data according to amplitudes, counting access path frequencies, classifying access levels, analyzing bandwidth occupation, and scheduling transmission priorities according to data frequency differences. According to the method, the data exception is accurately identified by analyzing the timestamp difference and the amplitude fluctuation, then the integrity of the data is enhanced through path consistency evaluation, the efficiency of the data in the storage and transmission process is optimized by reconstructing the abnormal data points, the data quality is ensured, meanwhile, the requirement for the storage space is reduced, and the data storage efficiency is improved. Through deep analysis of the data access mode, the access priority of the data is effectively classified and adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data optimization technology, and in particular to a data optimization method and system based on machine learning. Background Art

[0002] Data optimization mainly involves using various methods to improve data processing and storage efficiency, ensuring that data can flow efficiently and accurately and support the decision-making process. It includes data compression, data cleaning, data integration, data conversion, data compression and other methods. Through optimization technology, it can reduce redundant data, improve data quality and processing speed, optimize storage usage, and ensure that data is not lost or damaged during processing.

[0003] Among them, the data optimization method of machine learning is a technology that improves data processing efficiency and quality by applying machine learning algorithms. Its purpose is to automatically discover patterns, trends, anomalies, etc. in the data, thereby optimizing data storage, transmission, processing and analysis. Through machine learning, it can help identify redundant data, fill missing values, reduce noise, predict future data trends or behaviors, and even improve data processing speed and storage efficiency when the data set is large. It is widely used in financial analysis, recommendation systems, intelligent manufacturing, medical data analysis and other fields to help achieve efficient data management and decision support in a big data environment.

[0004] Existing technologies cannot effectively respond to real-time changes in large-scale data sets and lack sufficient flexibility to identify and adapt to changes in data patterns. During the data integration and conversion process, they cannot effectively identify all relevant data patterns, resulting in delays or misjudgments in the data analysis and decision-making process. Existing technologies are still insufficient in automated processing, especially in application scenarios that require continuous monitoring and immediate response. These limitations can easily lead to inefficient processing, thereby affecting overall data operation efficiency and decision-making quality. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a data optimization method and system based on machine learning.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: a data optimization method based on machine learning, comprising the following steps:

[0007] S1: Based on the time segments, amplitude distribution, identification sequence and path trajectory of the data, the time intervals of the data segments are calculated one by one using the timestamps. The amplitude distribution is screened for outliers using the amplitude fluctuation rate. The discreteness of the identification sequence is detected. The continuity and consistency of the data are then verified to establish a data feature map set.

[0008] S2: Based on the data feature map set, analyze the data boundary, measure the deviation of each data point from the historical trajectory standard, locate the position of data deviation, and mark the data points that deviate from the trajectory to obtain data trajectory deviation information;

[0009] S3: Analyze the amplitude information of the adjacent data of the deviated node according to the data trajectory offset information, sort the data according to amplitude similarity, and perform amplitude reconstruction on the sorted data. Replace the abnormal part of the original data with the reconstruction result to obtain the reconstructed data index;

[0010] S4: Using the reconstructed data index, according to the number of data accesses and the access interval time, the frequency of statistical access paths is counted, the concentration of the access frequencies is analyzed, and the access level categories of the data are divided according to the concentration, to obtain the access level classification result.

[0011] The present invention has the following improvements: the data feature mapping set includes a time difference index, an amplitude anomaly detection result, and a sequence discreteness score; the data trajectory offset information includes a trajectory deviation metric, a deviation position index, and an offset node identifier; the reconstructed data index includes an amplitude consistency sort, a data reconstruction effect, and an abnormal data replacement record; and the access level classification result includes an access frequency analysis result, an interval time statistical result, and a level division label.

[0012] The present invention is improved in that the steps of obtaining the data feature mapping set are specifically as follows:

[0013] S111: Based on the time segments, amplitude distribution, identification sequence, and path trajectory of the data, the start and end time intervals of each time segment in the timestamp sequence are compared, the interval difference between adjacent timestamps is calculated, each time interval is recorded, and it is determined whether the interval difference falls within the time stability interval to obtain the time stability verification interval;

[0014] S112: Based on the time stability verification interval, calculate the amplitude change rate of each point, compare it with its average amplitude change, filter out data points that exceed the benchmark, and obtain a sequence of abnormal amplitude fluctuation points;

[0015] S113: Call the abnormal amplitude fluctuation point sequence, check the consistency of the path segments associated with each identification point according to the distribution of the identification points in the path, analyze the deviation between the path offset and the identification density, determine the continuity and fluctuation matching degree of the path, and establish a data feature mapping set.

[0016] The present invention is improved in that the step of obtaining the data track offset information is specifically as follows:

[0017] S211: Based on the data feature map set, the data boundary is identified, the data point position is compared with the minimum envelope area of the current data set, and the relative position distance of the data point in the boundary area is calculated based on the feature position of each data point, the boundary map value and the distance between adjacent points to obtain the boundary positioning deviation;

[0018] S212: Calling the boundary positioning deviation, performing deviation analysis with the historical trajectory standard, obtaining the relative difference degree of each data point at the same time node, and identifying the change interval length and direction change amplitude of the trajectory point position to obtain the trajectory difference index;

[0019] S213: Identify key data points with varying amplitudes based on the trajectory difference index, calculate the degree of deviation of the data points, and mark the data points that deviate from the trajectory to obtain data trajectory deviation information.

[0020] The present invention is improved in that the steps of obtaining the reconstructed data index are specifically as follows:

[0021] S311: Based on the data trajectory offset information, call the original amplitude data of the adjacent moments before and after the deviation node, extract the amplitude difference between each pair of adjacent nodes, accumulate the differences, and generate a node amplitude difference sequence;

[0022] S312: Based on the node amplitude difference sequence, call the first n data items with the smallest difference, compare the fluctuation range to perform data screening, remove abnormal data that exceeds the normal fluctuation range, and retain the data group consistent with the original fluctuation, to obtain a filtered amplitude data group;

[0023] S313: Calculate the reconstructed amplitude of the node according to the filtered amplitude data group, replace the abnormal part in the original data, and obtain the reconstructed data index.

[0024] The present invention is improved in that the steps of obtaining the access level classification result are specifically as follows:

[0025] S411: Analyze each data entry using the reconstructed data index according to the number of data accesses and the access interval, calculate the frequency of access paths, and obtain a time access frequency data set;

[0026] S412: Analyze the access frequencies of data items according to the time access frequency data set, construct a frequency distribution interval division standard based on the number of data items and the frequency interval distribution ratio, and classify the time access frequencies into intervals to obtain frequency interval classification results;

[0027] S413: Call the frequency interval classification result, analyze the discrete degree of access frequency and the aggregation degree of the number of paths, calculate the access concentration of each type of access frequency interval, and divide the access level categories according to the concentration to obtain the access level classification result.

[0028] The present invention is improved in that the steps further include:

[0029] S5: Analyze the bandwidth usage of each type of data transmission based on the access level classification results, identify the data transmission frequency, sort the data transmission priorities based on the degree of data frequency difference, and schedule the data to be transmitted based on the priority to obtain a data priority scheduling result;

[0030] The data priority scheduling result includes bandwidth usage analysis results, transmission frequency comparison, and transmission sequence priority.

[0031] The present invention is improved in that the steps of obtaining the data priority scheduling result are specifically as follows:

[0032] S511: Based on the access level classification results, identify the data size and transmission time of each type of data information during network transmission, calculate the total transmission traffic and the number of occupied channels, and compare the total bandwidth usage with the current available bandwidth margin based on the network bandwidth upper limit in the data transmission path to determine the bandwidth usage of each type of data during transmission;

[0033] S512: Call the bandwidth usage information, combine the access trigger times and the average transmission period, calculate the average data transmission frequency within the period, filter the data categories with high transmission frequency, and obtain a high-frequency transmission data set;

[0034] S513: Based on the high-frequency transmission data set, the transmission frequencies are arranged from high to low, and the data categories whose bandwidth occupancy exceeds the bandwidth limit are excluded. Then, corresponding transmission priority levels are assigned according to the arrangement order to obtain a data priority scheduling result.

[0035] A data optimization system based on machine learning, the system comprising:

[0036] The time and amplitude analysis module is based on the time segments, amplitude distribution, identification sequence and path trajectory of the data. It uses timestamps to calculate the difference between the time intervals of the data segments one by one, screens the amplitude distribution for abnormal points through amplitude fluctuation rate, detects the degree of dispersion of the identification sequence, verifies the continuity and consistency of the data, and establishes a data feature mapping set.

[0037] The trajectory deviation identification module analyzes the data boundary based on the data feature mapping set, measures the deviation of each data point from the historical trajectory standard, locates the position of data deviation, and marks the data points that deviate from the trajectory to obtain data trajectory deviation information;

[0038] The data reconstruction module analyzes the amplitude information of the adjacent data of the deviation node according to the data trajectory offset information, sorts the data according to the amplitude similarity, and performs amplitude reconstruction on the sorted data to replace the abnormal part in the original data to obtain the reconstructed data index;

[0039] The access frequency analysis module uses the reconstructed data index to collect the frequency of data access paths based on the number of data accesses and the access interval time, analyzes the concentration of the access frequencies, and divides the data into access level categories according to the concentration, thereby obtaining the access level classification result;

[0040] The transmission priority sorting module analyzes the bandwidth occupancy of each type of data transmission according to the access level classification results, identifies the data transmission frequency, sorts the data transmission priority according to the degree of data frequency difference, and schedules the data to be transmitted according to the priority to obtain the data priority scheduling result.

[0041] Compared with the prior art, the advantages and positive effects of the present invention are:

[0042] In the present invention, by analyzing timestamp differences and amplitude fluctuations, data anomalies are accurately identified, and then the integrity of the data is enhanced through path consistency evaluation. By reconstructing abnormal data points, the efficiency of data in the storage and transmission process is optimized, ensuring data quality while reducing the demand for storage space. Through in-depth analysis of data access patterns, the access priority of data is effectively classified and adjusted, so that data streams can be more efficiently processed and analyzed in environments that require rapid decision support, thereby achieving data optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of the main steps of the present invention.

[0044] Figure 2 This is a flow chart for obtaining the data feature mapping set in the present invention.

[0045] Figure 3 This is a flow chart for obtaining data track offset information in the present invention.

[0046] Figure 4 This is a flowchart for obtaining the reconstructed data index in the present invention.

[0047] Figure 5 This is a flow chart for obtaining access level classification results in the present invention.

[0048] Figure 6 This is a flow chart for obtaining data priority scheduling results in the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0050] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the devices or elements referred to must have a specific direction, be constructed and operate in a specific direction, and therefore should not be understood as limiting the present invention. In addition, in the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0051] Example

[0052] See also Figure 1 , the present invention provides a technical solution: a data optimization method based on machine learning, comprising the following steps:

[0053] S1: Based on the time segments, amplitude distribution, identification sequence, and path trajectory of the data, the time intervals of the data segments are calculated one by one using timestamps. The amplitude distribution is screened for outliers using amplitude fluctuations, and the degree of dispersion of the identification sequence is detected. The continuity and consistency of the data are then verified through path consistency assessment to establish a data feature map set.

[0054] S2: Based on the data feature map set, analyze the data boundary, measure the deviation of each data point from the historical trajectory standard, locate the location of data deviation, and mark the data points that deviate from the trajectory to obtain data trajectory deviation information;

[0055] S3: Based on the data trajectory offset information, the amplitude information of the adjacent data of the deviated node is analyzed, the data is sorted according to the amplitude similarity, and the amplitude of the sorted data is reconstructed. The abnormal part of the original data is replaced with the reconstruction result to obtain the reconstructed data index;

[0056] S4: Using the reconstructed data index, the frequency of data access paths is calculated based on the number of data accesses and the access interval. The concentration of access frequencies is analyzed, and the data access levels are classified according to the concentration. These levels are marked one by one to obtain the access level classification results.

[0057] S5: Based on the access level classification results, analyze the bandwidth usage of each type of data transmission, identify the data transmission frequency, sort the data transmission priority according to the degree of data frequency difference, and schedule the data to be transmitted based on the priority to obtain the data priority scheduling result.

[0058] The data feature mapping set includes time difference indicators, amplitude anomaly detection results, and sequence discreteness scores. The data trajectory offset information includes trajectory deviation measurement, deviation position index, and offset node identification. The reconstructed data index includes amplitude consistency sorting, data reconstruction effect, and abnormal data replacement records. The access level classification results include access frequency analysis results, interval time statistics results, and level division labels. The data priority scheduling results include bandwidth usage analysis results, transmission frequency comparison, and transmission sequence priority.

[0059] Anomalies refer to points in a data set that appear abnormal or unusual compared to other data points. These points are caused by data errors, transmission interruptions, noise or other factors, and usually appear as values that deviate from the normal data pattern.

[0060] See also Figure 2 , the steps for obtaining the data feature mapping set are as follows:

[0061] S111: Based on the time segments, amplitude distribution, identification sequence, and path trajectory of the data, the start and end time intervals of each time segment in the timestamp sequence are compared, the interval difference between adjacent timestamps is calculated, each time interval is recorded, and it is determined whether the interval difference falls within the time stability interval to obtain the time stability verification interval. In the data optimization process, a path generally refers to the continuous sequence or trajectory of data points in time or space. For time series data, a path represents the changing trend of data points over time. Path analysis can help evaluate the continuity and consistency of data to ensure that there is no loss or abnormal fluctuation in the time series. Deviation from trajectory refers to the deviation of a data point from its expected trajectory or historical trajectory pattern. In time series analysis, the historical trajectory refers to the changing trend of data over a period of time, while deviation from trajectory refers to the deviation of certain data points from this trend. Amplitude refers to the range of change or numerical size of a data point. In data analysis, amplitude can be used to measure the degree of fluctuation or magnitude of change of a data point. For example, in time series data, amplitude represents the change between each data point and the previous point. By analyzing the amplitude fluctuation rate, it is possible to identify which data points have excessive amplitude changes, which are abnormal data or data that needs to be reconstructed.

[0062] Based on the time series, numerical amplitude, unique identification number and logical data group recorded in the data itself, the time field of the original data is extracted line by line, and a time segment is formed for every two adjacent data points. The start and end times are directly subtracted to obtain the time interval value. Subsequently, the time interval values of all adjacent time pairs are arranged to form a time difference set. Each time interval value in the set is then judged in turn. The judgment needs to refer to a pre-set standard time interval. For example, the allowed data generation time interval is defined as between 10 and 15 seconds. If the time interval value is 12 seconds, it falls within the allowed range; if it is 8 seconds or 18 seconds, it is marked as unstable. The time interval is determined; using this logic, all time differences are judged one by one, and all time segments that meet the stability standards are screened out. The time segment numbers are recorded as valid verification segments. For example, if there is a batch of data with time fields of 12:00:00, 12:00:10, 12:00:20, 12:00:32, and 12:00:42, the intervals are 10, 10, 12, and 10 seconds, respectively, which all meet the 10 to 15 second standard and are all included in the time stability verification interval. If a 30-second interval such as 12:01:00 to 12:01:30 appears, this segment will be eliminated, and only the stably generated data segments will be retained for subsequent analysis.

[0063] S112: Based on the time stability verification interval, the amplitude change rate of each point is calculated, compared with its average amplitude change, and the data points exceeding the benchmark are screened to obtain a sequence of abnormal amplitude fluctuation points;

[0064] Based on the data in the time stability verification interval, the numerical amplitude field is extracted from each record to construct a complete amplitude sequence. Then, the change ratio between each amplitude value in the sequence and its adjacent amplitude values is compared. The amplitude change at each position is calculated one by one, and all the change values are averaged to obtain the amplitude change benchmark level of the entire interval. After that, each group of specific change amplitudes is compared with this benchmark level. If the degree of change of a group exceeds the standard range of allowable change, the position is marked as an abnormal point; the setting of the allowable change range here needs to refer to the past The typical change behavior of normal data, for example, if the numerical fluctuation of a set of data in adjacent records is usually maintained at ±5%, the allowable range is set to ±7% to accept slight fluctuations. If a change exceeds 10%, it can be considered as a significant abnormality; for example, if the original amplitude sequence is 50, 52, 51, 78, 53, 54, then the amplitude of 78 jumps by more than 50% compared with the previous value, which is significantly different from the surrounding data jumps, constituting a mutation. The position of 78 will be recorded as an abnormal point. By comparing point by point, the record number, value and position of all fluctuation abnormal points in the original sequence are output to form an abnormal amplitude fluctuation point sequence.

[0065] S113: Calling the abnormal amplitude fluctuation point sequence, according to the distribution of the identification points in the path, checking the consistency of the path segments associated with each identification point, analyzing the deviation between the path offset and the identification density, determining the continuity and fluctuation matching degree of the path, and establishing a data feature mapping set;

[0066] The record number of each data point in the original data is located one by one, and its corresponding unique identification number is read. All identification numbers are sorted and classified, and data segments belonging to the same type of identification are extracted. Then, the number of identification numbers contained in each data segment is counted. Combined with the total number of data records in the segment, the concentration of the identification number distribution within the unit is calculated. The distribution degree of all data segments is then averaged to form a reference level of identification concentration. The identification distribution level of each data segment is then analyzed against the reference value. If the identification number density in a segment is significantly higher or lower, it is marked as a distribution anomaly segment. Sequence consistency analysis is further performed on data segments with high identification density to detect whether the identification numbers in adjacent data records are arranged regularly in an ascending or other preset order. If the order is chaotic or jumps widely, the data segment is judged to have a sequence anomaly. Through these steps, the stability and consistency of the data in the identification number dimension are simultaneously evaluated. Finally, the identification distribution characteristics, sequential coherence and fluctuation position of each data segment are integrated to construct a complete data feature mapping set, providing a structural foundation for data correction and optimization in the subsequent stages.

[0067] See also Figure 3 ,The specific steps for obtaining data trajectory offset information are:

[0068] S211: Based on the data feature map set, the data boundary is identified, the data point position is compared with the minimum envelope area of the current data set, and the relative position distance of the data point in the boundary area is calculated based on the feature position of each data point, the boundary map value and the distance between the adjacent points to obtain the boundary positioning deviation;

[0069] Identify the position parameters of all data points in the original data set, extract the characteristic position of each data point, and the position information can be represented by its sorting number or index position in the data record. Then, construct the boundary outline of the data point according to its distribution characteristics, define the boundary as the minimum outer area containing all data points, and generate the coordinate range or position range of the boundary. By directly comparing the position value of each data point with the maximum and minimum thresholds of the boundary, compare whether it is at the edge, close to the edge, or in the internal area. The judgment method is to subtract the coordinate value of each point from the upper and lower limits of the boundary interval respectively. If the value is less than the preset tolerance range, it is considered to be in the boundary area. For example, the boundary offset tolerance is set to 5 Position unit: if the maximum position of the boundary is 200 and the minimum position is 1, then the offset of a data point with a position value of 6 relative to the lower boundary is 5, and the point can be determined to be in the lower edge area. Furthermore, for the data points in the boundary area, the distance between them and their nearest neighboring data points is calculated. This distance is its relative density index, which is used to evaluate the density of the point at the boundary. For example, if the position of a point is 195 and the nearest neighbor position is 192, then its spacing is 3 units. Compare it with the average data spacing. If the average data spacing is 1 unit, then this distance is more than twice as high, indicating that the boundary point is relatively sparse. The offset distance of all boundary points is recorded as the boundary positioning deviation, which serves as a reference for the deviation of the point from the boundary contour.

[0070] S212: Calling the boundary positioning deviation, performing deviation analysis with the historical trajectory standard, obtaining the relative difference degree of each data point at the same time node, and identifying the change interval length and direction change amplitude of the trajectory point position to obtain the trajectory difference index;

[0071] Call the obtained boundary positioning deviation data, align the record values of the same time node in the historical data on the timestamp field of all data points one by one, extract the reference record with the same time as the current data point in the historical trajectory data set, and then calculate the numerical difference between the current data point and the data position of the corresponding historical record. The difference is the basic data of the relative difference degree, and the consistency performance of the data point on the time axis is judged accordingly. If the difference value exceeds the specified allowable error range, for example, the offset tolerance is set to 10 position units, the current data point position is 80, and the corresponding historical trajectory position is 95, then the difference is 15, which exceeds the allowable range, and the point is marked as severely offset. Then, in each data segment, the amplitude and direction of the position value changes of adjacent points are scanned in chronological order. The direction of the trajectory change is judged by the change trend of the previous and next points. If there is a continuous upward or downward trend in the position with the same amplitude, for example, the consecutive recorded positions are 80, 83, 86, and 89, indicating that the direction is positive and the amplitude of the change is 3 units. If it then turns to 86, 83, and 80, the direction change is negative and the amplitude remains unchanged, indicating that the direction is reversed. The difference before and after the direction change is calculated to obtain the direction change amplitude. Finally, the relative difference degree, change interval length, and direction change information of each data point are comprehensively counted to form a trajectory difference index reflecting the trajectory change behavior of each data point.

[0072] S213: Based on the trajectory difference index, identify the data points with key change amplitudes, calculate the degree of deviation of the data points, and mark the data points that deviate from the trajectory using the formula:

[0073]

[0074] The data trajectory offset information DV is obtained, which is used to quantitatively express the degree of deviation of the overall data point relative to the historical trajectory, where TQ i represents the actual trajectory of the i-th data point, which is the position or state of the current data point on the trajectory. TR represents the standard trajectory, which is used as a comparison benchmark to determine the degree of deviation of each data point from the past norm. SQ i Indicates the deviation sensitivity of the i-th data point, which is used to adjust the calculation of the deviation value to ensure that the influence of data points with different sensitivities in the calculation is appropriately reflected. i is the weight of the i-th data point, which affects the role of this point in the overall offset calculation, n dv represents the total number of data points;

[0075] The specific parameters of the three key data points are as follows:

[0076] Data point 1: TQ1 = 12.6, SQ1 = 0.4, WQ1 = 1.0;

[0077] Data point 2: TQ2 = 14.2, SQ2 = 0.9, WQ2 = 0.7;

[0078] Data point 3: TQ3 = 11.7, SQ3 = 0.5, WQ3 = 1.1;

[0079] Substitute the above parameters into the formula to calculate:

[0080] Point 1:

[0081]

[0082] Point 2:

[0083]

[0084] Point 3:

[0085]

[0086] Add the three terms and divide by 3:

[0087]

[0088] The trajectory deviation index DV≈0.561 is obtained. This value represents the average deviation degree of the current three key data points under the standard trajectory reference, providing data support for subsequent trajectory marking and anomaly detection, and can be marked as a trajectory deviation comprehensive mark.

[0089] See also Figure 4 , the specific steps for obtaining the reconstructed data index are:

[0090] S311: Based on the data trajectory offset information, the original amplitude data of the adjacent moments before and after the deviation node are called, and the amplitude difference between each pair of adjacent nodes is extracted and accumulated to generate a node amplitude difference sequence;

[0091] Identify which data points are determined to be offset nodes, and then for each offset node, extract the position index of the node in the data record, and find the two data records immediately before and after it, extract the amplitude field values of the previous node, current node and next node in turn, perform direct numerical subtraction on the amplitude between the previous node and the current node, and record the difference, and then do the same with the amplitude between the current node and the next node to obtain a set of two groups of amplitude differences centered on the offset node and spanning three points. Then, according to the arrangement order of the offset nodes, all such pairwise difference combinations are accumulated one by one to construct A node amplitude difference sequence based on the offset node is generated. The values of this sequence represent the cumulative degree of amplitude mutation before and after the offset node. Subsequently, a rolling window method can be introduced as needed to sum up adjacent differences in a longer interval to reflect the strength of the area where amplitude fluctuations are concentrated. For example, if the adjacent amplitude values of an offset node are 48, 55, and 60, the first difference is 7, the second difference is 5, and the total is 12, which means that the cumulative value of the fluctuation amplitude of this node in the original amplitude data is 12. Finally, the cumulative amplitude differences corresponding to all offset nodes are sorted and output to form a complete node amplitude difference sequence.

[0092] S312: Based on the node amplitude difference sequence, the first n items of data with the smallest difference are retrieved, and the data are screened by comparing the fluctuation range. Abnormal data that exceeds the normal fluctuation range is eliminated, and the data group consistent with the original fluctuation is retained to obtain the screened amplitude data group;

[0093] Sort all the cumulative differences from small to large, and select the first n data as the sample set with the smallest fluctuation. The value of n here should be set according to the overall data scale. For example, if the original sequence contains 100 data, n can be set to 20, and the 20 records with the smallest difference are taken as the preliminary data group. Then these 20 records are judged again to see whether their fluctuation range is within the normal amplitude variation range. The upper and lower limits of the normal fluctuation range should be set with reference to the amplitude variation statistics of similar historical data. For example, if the amplitude difference of the past 10 groups of data when stable mostly falls between 3 and 8, then This interval is defined as the normal fluctuation range, and values outside this range are regarded as abnormal fluctuations. During the continued processing, records with differences greater than 8 or less than 3 in the first n items are eliminated, and the remaining ones are data points consistent with the historical original amplitude change trend. The retained amplitude data are then numbered and recorded, and output as the filtered amplitude data group. For example, in the difference sorting result, the amplitude differences of the 3rd, 4th, 6th, and 9th digits are 4, 5, 6, and 7 respectively. These 4 records fall into the normal fluctuation range and are retained. If the difference of the 8th digit is 10, it is excluded from the final amplitude data group, completing the construction of the cleaned amplitude data group.

[0094] S313: Based on the filtered amplitude data set, use the formula:

[0095]

[0096] Calculate the reconstruction amplitude CR of the sth node s , is the new amplitude value obtained by the node through calculation, which is used to replace the corresponding abnormal amplitude data in the original data, replace the abnormal part in the original data, and obtain the reconstructed data index, where CA sj Indicates the original amplitude data corresponding to the sth node in the jth data group. The data was originally obtained by measuring the adjacent moments before and after the node, and is used in the calculation to determine the new amplitude of the node. CW j The weight assigned to the jth data group is used to adjust the influence of each group of data in the reconstruction calculation, N cr Indicates the total number of data groups;

[0097] According to the filtered amplitude data group, let the total number of data groups be N cr =3, recorded as data groups 1, 2, and 3, the target node is the s=5th node, and the original amplitude data of the node in the three groups of data are CA 51 =2.3, CA 52 =2.7, CA 53 =2.5, the weights of each data group are CW1=0.25, CW2=0.35, CW3=0.40, and the calculation is as follows:

[0098]

[0099] The products are calculated as:

[0100] 2.3×0.25=0.575,2.7×0.35=0.945,2.5×0.40=1.000;

[0101] Sum the products:

[0102] 0.575+0.945+1.000=2.520;

[0103] Normalization processing:

[0104]

[0105] Therefore, the reconstructed amplitude CR5 result of the s=5th node is 0.840, which will be used to replace the abnormal amplitude data of node 5 in the original data to correct the data anomaly caused by offset or noise interference. By introducing amplitude data from multiple data groups and combining different weight ratios, the results reflect the differences in credibility between different data sources, thereby improving the representativeness and balance of the reconstructed amplitude data and obtaining an updated amplitude index table.

[0106] See also Figure 5 ,The specific steps for obtaining the access level classification results are:

[0107] S411: Analyze each data entry based on the number of data accesses and access intervals using the reconstructed data index, calculate the frequency of access paths, and obtain a time access frequency dataset;

[0108] Each data entry recorded in the data index is identified one by one, and the access timestamp and access count fields carried by each record are read. On this basis, for each data entry, it is first sorted in ascending order by its access timestamp, and then the sorted time series is scanned, and the time difference between adjacent accesses is recorded to form an access time interval sequence. At the same time, the total number of times the data is accessed is statistically summarized to obtain the access frequency value of the entry in the overall access behavior. Subsequently, the access time interval and the number of accesses are recorded to form the access path information at the data entry level. Then, the access path information of each data entry is merged and analyzed to obtain the time interval frequency. Data that is frequently accessed within a fixed time period is marked as frequently accessed. For example, if a piece of data is accessed five times in a row within five minutes, while other items are accessed only once per hour, the former is classified as a frequently accessed record. The access frequencies of all items are then sorted according to the number of accesses and associated with their corresponding time intervals to form a complete time access frequency dataset. For example, if a piece of data is accessed three times at 12:00, 12:03, 12:07, and 12:12, the time intervals are 3 minutes, 4 minutes, and 5 minutes, corresponding to an access frequency of once every five minutes. This frequency is recorded as the access frequency of the current data item, and all items generate the dataset according to this rule.

[0109] S412: Analyze the access frequencies of data items based on the time access frequency data set, construct a frequency distribution interval division standard based on the number of data items and the frequency interval distribution ratio, and classify the time access frequencies into intervals to obtain frequency interval classification results;

[0110] Extract the access frequency values of all data items as the basis for analysis, count the ranges of all frequency values, obtain the minimum and maximum frequency values, and then divide the entire frequency interval into equal intervals or layers to construct multiple frequency interval segments. For example, if the maximum frequency is 1 time per minute and the minimum frequency is 1 time per hour, the entire frequency interval can be divided into three frequency segments: 1 time per minute to 1 time per 5 minutes, 1 time per 5 minutes to 1 time per 15 minutes, and 1 time per 15 minutes to 1 time per hour. Then, classify the frequency values of all data items according to their respective intervals. The judgment basis is to compare each frequency value with each frequency value. The upper and lower limits of the interval are compared. If a frequency is once every 10 minutes and falls into the second interval range, it is classified as a medium-frequency category. At the same time, the number of data items contained in each category is counted, and its proportion in the total data items is calculated. For example, the first frequency segment contains 120 data items, accounting for 20% of the total; the second category contains 300 data items, accounting for 50%; and the last category contains 180 data items, accounting for 30%. Based on this, the frequency interval division standard and proportion structure table are established, and each data item is marked as the corresponding frequency interval classification result according to the interval into which its frequency falls.

[0111] S413: Call the frequency interval classification result, analyze the discrete degree of access frequency and the aggregation degree of the number of paths, and use the formula:

[0112]

[0113] Calculate the access concentration HC of each access frequency interval r , represents the degree of aggregation of data access in this interval. High concentration indicates that access occurs intensively at a few time points. Access level categories are divided according to the concentration, and the access level classification results are obtained. Among them, HF rb is the time access frequency of the bth entry in the rth class, that is, the number of accesses to the data entry per unit time, is the average time access frequency of all data entries in the rth class, which is used to measure the average access level of the entire interval, n r is the total number of data entries in the rth access frequency interval;

[0114] If the frequency interval to be analyzed is a certain one (referred to as the rth category), which contains 4 items, its access frequency per unit time (i.e. HF rb )as follows:

[0115] The first entry has a visit frequency of 10, the second entry has a visit frequency of 12, the third entry has a visit frequency of 8, and the fourth entry has a visit frequency of 20.

[0116] Calculate average visit frequency

[0117]

[0118] Compute the squared deviation of each entry:

[0119] (10-12.5) 2 =6.25;

[0120] (12-12.5) 2 =0.25;

[0121] (8-12.5) 2 =20.25;

[0122] (20-12.5) 2 =56.25;

[0123] Find the total sum of squared deviations and divide by the number of entries:

[0124]

[0125] There are two entries with similar access frequencies (10 and 12), but due to the presence of an entry with a significantly higher access frequency (20) and an entry with a relatively low access frequency (8), the overall access distribution appears unbalanced, resulting in an access concentration HC r Reaching 20.75 indicates that data access within this frequency range is clearly concentrated. Through calculation, access levels can be further divided according to concentration. Categories with high concentration represent hot data, which are suitable for caching or priority processing, while categories with low concentration represent cold data and can be processed at a lower frequency.

[0126] See also Figure 6 ,The specific steps for obtaining data priority scheduling results are:

[0127] S511: Based on the access level classification results, identify the data size and transmission time of each type of data information during network transmission, calculate the total transmission traffic and the number of occupied channels, and compare the total bandwidth usage with the current available bandwidth margin based on the network bandwidth upper limit in the data transmission path to determine the bandwidth usage of each type of data during transmission;

[0128] Each type of data entry is classified into its own classification group, and the single transmission data size and the time required for the transmission of all data entries in the group are extracted. The transmission time and data size of each data are directly multiplied to obtain the traffic value consumed by the data entry in one transmission. The traffic values of all entries in this category are then summed up to obtain the total transmission traffic of this category of data. The number of concurrent connections required for this type of data transmission is continued to be counted. It is judged based on whether there is an overlapping time period for the transmission of each data entry on the network. If the transmission time of multiple data entries overlaps, it is judged that they occupy multiple channels, and the usage of such channels is recorded one by one to summarize the data usage of this category. The number of channels occupied by different types of data is then called, and the upper limit of the bandwidth that can be allocated in the transmission path is called to read the total bandwidth value. The transmission flow of the current type of data is divided by the corresponding transmission time period to obtain the instantaneous bandwidth demand value of this type of data, and then the value is compared with the remaining allocable value of the current network bandwidth. If the value is less than or equal to the current remaining bandwidth, the data of this type can be transmitted, otherwise it is judged as insufficient bandwidth. For example, the transmission size of a certain type of data is 100KB each time, the average transmission time is 2 seconds, and the bandwidth demand is 50KB / s. If the remaining bandwidth is 80KB / s, the data of this type can be transmitted normally, otherwise it is recorded as exceeding the limit. The bandwidth occupancy of each type of data is output and recorded.

[0129] S512: Invoke the bandwidth usage, combine the access trigger times and the average transmission period, calculate the average data transmission frequency within the period, filter the data categories with high transmission frequency, and obtain the high-frequency transmission data set;

[0130] The access trigger count for each data type is calculated per unit period. For example, taking one minute as a period, the total number of times a certain type of data is accessed or called within one minute is read. The average transmission time of this data type is also extracted. The length of one minute is then divided by the average transmission time to obtain the theoretical transmission count. The actual trigger count is then compared with the theoretical value. If the actual trigger count is greater than half of the theoretical value, the data type is considered to be in a high-frequency transmission state in that period and is added to the temporary high-frequency data set. This process is continued until all data types are calculated, forming a preliminary high-frequency transmission data set. The bandwidth usage value of each data type in the set is further determined, and data types whose bandwidth requirements exceed the maximum remaining network capacity are filtered out. For example, if the average bandwidth requirement of a certain data type is 120KB / s and the current remaining bandwidth is 100KB / s, this data will cause conflicts during transmission and is removed from the high-frequency set. If the bandwidth of a certain data type is 80KB / s, it is within the acceptable range and is retained, forming a purified high-frequency transmission data set.

[0131] S513: Based on the high-frequency transmission data set, the transmission frequencies are sorted from high to low, and the data categories whose bandwidth usage exceeds the bandwidth limit are excluded. Then, corresponding transmission priorities are assigned according to the sorting order to obtain the data priority scheduling result;

[0132] Based on the high-frequency transmission data set, all data categories are sorted in descending order according to the transmission frequency value within the statistical period. The data category with the largest frequency value is ranked first. If the frequency is the same, it is sorted according to the number of accesses. During the sorting process, the original index number of each data category must be marked for subsequent allocation tracking. After the sorting is completed, the bandwidth demand of the data is checked one by one to see if it exceeds the current remaining bandwidth limit. If it exceeds, the data category is removed from the queue and does not participate in the priority assessment. After the screening is completed, the corresponding priority number is assigned from top to bottom according to the arrangement order of the remaining data in the queue. For example, the first data category is assigned priority 1, the second data category is assigned priority 2, and so on. Each data category is paired with its priority to form a priority scheduling table. For example, the frequency of category A data is 50 times / minute and the bandwidth occupies 40KB / s, so it is listed as priority 1. Category B data is 45 times / minute and the bandwidth occupies 60KB / s, so it is listed as priority 2. The output scheduling result is the priority scheduling list.

[0133] A data optimization system based on machine learning, the system comprising:

[0134] The time and amplitude analysis module is based on the time segments, amplitude distribution, identification sequence and path trajectory of the data. It uses timestamps to calculate the difference between the time intervals of the data segments one by one, screens the amplitude distribution for abnormal points through amplitude fluctuation rate, detects the degree of dispersion of the identification sequence, verifies the continuity and consistency of the data, and establishes a data feature mapping set.

[0135] The trajectory deviation identification module analyzes the data boundary based on the data feature map set, measures the deviation of each data point from the historical trajectory standard, locates the location of data deviation, and marks the data points that deviate from the trajectory to obtain data trajectory deviation information;

[0136] The data reconstruction module analyzes the amplitude information of the adjacent data of the deviated node based on the data trajectory offset information, sorts the data according to the amplitude similarity, and performs amplitude reconstruction on the sorted data to replace the abnormal part in the original data to obtain the reconstructed data index;

[0137] The access frequency analysis module uses the reconstructed data index to collect the frequency of data access paths based on the number of data accesses and the access interval time, analyzes the concentration of access frequencies, and divides the data into access level categories based on the concentration to obtain the access level classification results;

[0138] The transmission priority sorting module analyzes the bandwidth usage of each type of data transmission according to the access level classification results, identifies the data transmission frequency, sorts the data transmission priority according to the degree of data frequency difference, and schedules the data to be transmitted according to the priority to obtain the data priority scheduling result.

[0139] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A data optimization method based on machine learning, characterized in that: The following steps are involved: S1: Based on the time segments, amplitude distribution, identification sequence and path trajectory of the data, the time intervals of the data segments are calculated one by one using the timestamps. The amplitude distribution is screened for outliers using the amplitude fluctuation rate. The discreteness of the identification sequence is detected. The continuity and consistency of the data are then verified to establish a data feature map set. S2: Based on the data feature map set, analyze the data boundary, measure the deviation of each data point from the historical trajectory standard, locate the position of data deviation, and mark the data points that deviate from the trajectory to obtain data trajectory deviation information; S3: Analyze the amplitude information of the adjacent data of the deviated node according to the data trajectory offset information, sort the data according to amplitude similarity, and perform amplitude reconstruction on the sorted data. Replace the abnormal part of the original data with the reconstruction result to obtain the reconstructed data index; S4: Using the reconstructed data index, according to the number of data accesses and the access interval time, the frequency of statistical access paths is counted, the concentration of the access frequencies is analyzed, and the access level categories of the data are divided according to the concentration, to obtain the access level classification result.

2. The data optimization method based on machine learning according to claim 1, characterized in that: The data feature mapping set includes time difference indicators, amplitude anomaly detection results, and sequence discreteness scores; the data trajectory offset information includes trajectory deviation metrics, deviation position indexes, and offset node identifiers; the reconstructed data index includes amplitude consistency sorting, data reconstruction effects, and abnormal data replacement records; and the access level classification results include access frequency analysis results, interval time statistics results, and level division labels.

3. The data optimization method based on machine learning according to claim 1, characterized in that: The steps for obtaining the data feature mapping set are specifically as follows: S111: Based on the time segments, amplitude distribution, identification sequence, and path trajectory of the data, compare the start and end time intervals of each time segment in the timestamp sequence, calculate the interval difference between adjacent timestamps, record each time interval, and determine whether the interval difference falls within the time stability interval to obtain the time stability verification interval; S112: Based on the time stability verification interval, calculate the amplitude change rate of each point, compare it with its average amplitude change, filter out data points that exceed the benchmark, and obtain a sequence of abnormal amplitude fluctuation points; S113: Call the abnormal amplitude fluctuation point sequence, check the consistency of the path segments associated with each identification point according to the distribution of the identification points in the path, analyze the deviation between the path offset and the identification density, determine the continuity and fluctuation matching degree of the path, and establish a data feature mapping set.

4. The data optimization method based on machine learning according to claim 1, characterized in that: The steps for obtaining the data track offset information are specifically as follows: S211: Based on the data feature map set, the data boundary is identified, the data point position is compared with the minimum envelope area of the current data set, and the relative position distance of the data point in the boundary area is calculated based on the feature position of each data point, the boundary map value and the distance between adjacent points to obtain the boundary positioning deviation; S212: Calling the boundary positioning deviation, performing deviation analysis with the historical trajectory standard, obtaining the relative difference degree of each data point at the same time node, and identifying the change interval length and direction change amplitude of the trajectory point position to obtain the trajectory difference index; S213: Identify key data points with varying amplitudes based on the trajectory difference index, calculate the degree of deviation of the data points, and mark the data points that deviate from the trajectory to obtain data trajectory deviation information.

5. The data optimization method based on machine learning according to claim 1, characterized in that: The steps for obtaining the reconstructed data index are specifically as follows: S311: Based on the data trajectory offset information, call the original amplitude data of the adjacent moments before and after the deviation node, extract the amplitude difference between each pair of adjacent nodes, accumulate the differences, and generate a node amplitude difference sequence; S312: Based on the node amplitude difference sequence, call the first n data items with the smallest difference, compare the fluctuation range to perform data screening, remove abnormal data that exceeds the normal fluctuation range, and retain the data group consistent with the original fluctuation, to obtain a filtered amplitude data group; S313: Calculate the reconstructed amplitude of the node according to the filtered amplitude data group, replace the abnormal part in the original data, and obtain the reconstructed data index.

6. The data optimization method based on machine learning according to claim 1, characterized in that: The steps for obtaining the access level classification result are specifically as follows: S411: Analyze each data entry using the reconstructed data index according to the number of data accesses and the access interval, calculate the frequency of access paths, and obtain a time access frequency data set; S412: Analyze the access frequencies of data items according to the time access frequency data set, construct a frequency distribution interval division standard based on the number of data items and the frequency interval distribution ratio, and classify the time access frequencies into intervals to obtain frequency interval classification results; S413: Call the frequency interval classification result, analyze the discrete degree of access frequency and the aggregation degree of the number of paths, calculate the access concentration of each type of access frequency interval, and divide the access level categories according to the concentration to obtain the access level classification result.

7. The data optimization method based on machine learning according to claim 1, characterized in that: The steps also include: S5: Analyze the bandwidth usage of each type of data transmission based on the access level classification results, identify the data transmission frequency, sort the data transmission priorities based on the degree of data frequency difference, and schedule the data to be transmitted based on the priority to obtain a data priority scheduling result; The data priority scheduling result includes bandwidth usage analysis results, transmission frequency comparison, and transmission sequence priority.

8. The data optimization method based on machine learning according to claim 7, characterized in that: The steps for obtaining the data priority scheduling result are specifically as follows: S511: Based on the access level classification results, identify the data size and transmission time of each type of data information during network transmission, calculate the total transmission traffic and the number of occupied channels, and compare the total bandwidth usage with the current available bandwidth margin based on the network bandwidth upper limit in the data transmission path to determine the bandwidth usage of each type of data during transmission; S512: Call the bandwidth usage information, combine the access trigger times and the average transmission period, calculate the average data transmission frequency within the period, filter the data categories with high transmission frequency, and obtain a high-frequency transmission data set; S513: Based on the high-frequency transmission data set, the transmission frequencies are arranged from high to low, and the data categories whose bandwidth occupancy exceeds the bandwidth limit are excluded. Then, corresponding transmission priority levels are assigned according to the arrangement order to obtain a data priority scheduling result.

9. A data optimization system based on machine learning, characterized in that: The system is used to implement the data optimization method based on machine learning according to any one of claims 1 to 8, and the system includes: The time and amplitude analysis module is based on the time segments, amplitude distribution, identification sequence and path trajectory of the data. It uses timestamps to calculate the difference between the time intervals of the data segments one by one, screens the amplitude distribution for abnormal points through amplitude fluctuation rate, detects the degree of dispersion of the identification sequence, verifies the continuity and consistency of the data, and establishes a data feature mapping set. The trajectory deviation identification module analyzes the data boundary based on the data feature mapping set, measures the deviation of each data point from the historical trajectory standard, locates the position of data deviation, and marks the data points that deviate from the trajectory to obtain data trajectory deviation information; The data reconstruction module analyzes the amplitude information of the adjacent data of the deviation node according to the data trajectory offset information, sorts the data according to the amplitude similarity, and performs amplitude reconstruction on the sorted data to replace the abnormal part in the original data to obtain the reconstructed data index; The access frequency analysis module uses the reconstructed data index to collect the frequency of data access paths based on the number of data accesses and the access interval time, analyzes the concentration of the access frequencies, and divides the data into access level categories according to the concentration, thereby obtaining the access level classification result; The transmission priority sorting module analyzes the bandwidth occupancy of each type of data transmission according to the access level classification results, identifies the data transmission frequency, sorts the data transmission priority according to the degree of data frequency difference, and schedules the data to be transmitted according to the priority to obtain the data priority scheduling result.

Citation Information

Cited By

  • Cloud-based teaching platform student learning behavior track analysis method

    CN121144727A

  • Cloud-based teaching platform student learning behavior trajectory analysis method

    CN121144727B

  • Control method and system for automatic receiving and sending of parting strips

    CN121500921A

  • Mobile application big data management method and system based on machine learning

    CN121579561A

  • Machine learning-based mobile application big data management method and system

    CN121579561B