Intelligent data processing method and system for data visualization platform

By calculating data activity and building a searchable data slice directory, the accuracy issues of data screening and access behavior are solved, efficient hierarchical storage and compression of data are achieved, and data reading efficiency and traceability are improved.

CN120744185AInactive Publication Date: 2025-10-03HANGZHOU HAIYIN RONGZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510790203.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies make it difficult to conduct a concrete assessment of data activity during the data screening stage, resulting in a lack of accuracy in data screening, a lack of dynamic association between access behavior records and the data entity, and a lack of differentiation in compression operations, which in turn leads to a decrease in the efficiency of reading high-value data.

Method used

By calculating the data activity of the data segment, generating a preliminary subset of data of interest, identifying the start and end location information and assigning a unique identifier, building a searchable data slice directory, monitoring the cumulative access frequency of retrieval instructions, and combining storage levels and compression parameters for hierarchical optimization processing.

Benefits of technology

It improves the addressability and traceability accuracy of data during the retrieval process, avoids access redundancy, improves compression efficiency and data reading and writing performance, and ensures efficient allocation of storage resources for high-frequency data slices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744185A_ABST
    Figure CN120744185A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information retrieval, in particular to an intelligent data processing method and system for a data visualization platform, and the method comprises the following steps: calculating the data activeness of each data segment based on an input data stream of the data visualization platform, screening the data segments with the data activeness higher than a preset threshold value, and generating a preliminary interest data subset. According to the method, activeness calculation and screening are carried out on each data segment in the input data stream, high-value data segments are actively identified to serve as a preliminary interest data subset, a unique identifier is generated in combination with starting and ending position information in the original data stream, and a description data unit is established, so that the original data has the capability of being accurately positioned and traced; the addressability and the tracing accuracy of the data in the retrieval process are improved; the unique identification and the position information are structurally integrated, a retrieval-oriented positioning structure is constructed, and the problem of access redundancy caused by unclear structure in traditional information storage is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval technology, and in particular to a data intelligent processing method and system for a data visualization platform. Background Art

[0002] The intelligent data processing method is an automatic processing mechanism that integrates data analysis, behavioral modeling, dynamic screening and hierarchical compression. It is mainly used to improve the searchability, manageability and access efficiency of massive streaming data in data visualization platforms.

[0003] While existing technologies have enabled dynamic analysis and management of streaming data, they struggle to concretely assess data activity during the data screening phase. This results in inaccurate data screening, which can easily miss high-value segments or introduce redundant content. Access behavior is recorded only at the statistical level, lacking dynamic association with the data itself, making it difficult to support hierarchical management. Compression operations often employ a unified strategy, ignoring differences in usage frequency and storage cost across different data slices, which can lead to decreased efficiency in reading high-value data. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a data intelligent processing method and system for a data visualization platform.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a data intelligent processing method for a data visualization platform, comprising the following steps: Based on the input data stream of the data visualization platform, the data activity of each data segment is calculated, and the data segments with data activity higher than the preset threshold are filtered to generate a preliminary subset of interesting data; Based on each data segment in the preliminary interest data subset, the start and end position information of the data segment in the original data stream is identified and assigned a unique identifier to obtain a descriptive data unit. Based on the descriptive data unit, all unique identifiers and corresponding start and end position information are integrated to construct a positioning structure for information retrieval and establish a searchable data slice directory. Based on the searchable data slice directory, the search instructions of each data slice in the directory are monitored, the number of searches for each data slice is accumulated to form an access frequency, and an access statistical data slice is obtained. Based on the access frequency of the access statistical data slice and the inherent size attribute of the data slice, a storage tier is assigned to each data slice, and a list of data to be tiered and optimized is obtained; Based on the storage level of each data slice in the list of data to be layered and optimized, the corresponding compression parameter set is matched to obtain a compressed parameterized data item. Based on the compressed parameterized data item, a compression operation is performed on the data item and stored in the target storage level. The storage location and compression status information of the data item in the retrievable data slice directory are synchronously updated to establish a layered compressed data view.

[0006] Preferably, the steps of obtaining the preliminary interest data subset are: Based on the data flow input from the data visualization platform, the duration of each data segment is set as a fixed time window, and the data is divided into several data segments in sequence. A sliding window is constructed for each data segment. Each group contains three consecutive data points and the pairwise difference slope, standard deviation within the window, and local maximum and minimum difference are calculated to form a set of original activity indicators for each data segment. Calculating an activity level value for each data segment based on the original activity indicator set; According to the activity level value of each data segment, it is determined whether it exceeds the activity threshold, and all data segments whose activity level values ​​exceed the activity threshold are retained as preliminary interest data sources to form a preliminary interest data subset.

[0007] Preferably, the steps of obtaining the descriptive data unit are: Based on each data segment that has been filtered in the preliminary interest data subset, traverse each corresponding data segment in the original data stream, retrieve the index sequence position in the original data stream and record the start point and end point indexes to obtain the start and end position information of each data segment in the original data stream; According to the start and end position information, the original data stream is segmented and numbered, a unique identifier is assigned to each data segment, and the start point index, end point index and unique identifier of each data segment are combined into a unified structure to generate an identification data mapping table; Based on the identified data mapping table, the complete data content corresponding to each data segment in the original data stream is extracted, and each data segment is combined and bound with the identification information and the position information to form a descriptive data unit.

[0008] Preferably, the steps for obtaining the searchable data slice directory are: Based on the unique identifier and start and end position information bound in the descriptive data unit, extract the identification number of each descriptive data unit and the corresponding index boundary in the original data stream, establish a one-to-one mapping relationship between the unique identification number and the start and end indexes, and generate an identification index matching list; According to the identification index matching list, the unique identifiers recorded in each descriptive data unit are classified and sorted, and encapsulated together with the corresponding starting position index and ending position index into a standard search structure format to generate a structured index set; Based on the structured index set, all information items of unique identifiers, starting position indexes and ending position indexes are uniformly organized to establish a searchable data slice directory.

[0009] Preferably, the steps of obtaining the access statistical data slice are: Based on the unique identifier, starting position index and ending position index of each data slice recorded in the searchable data slice directory, monitor the search request log issued, decode the search target contained in the search request log and identify the unique identifier referenced therein to obtain the set of unique identifiers of the accessed data slices; According to the set of unique identifiers of the accessed data slices, each occurrence of the unique identifier is recorded and counted, the number of occurrences within the target time range is accumulated one by one according to the unique identifier dimension, and the statistical results are bound to the corresponding unique identifier to generate a detailed table of data slice access frequency; Based on the data slice access frequency details table, the cumulative access frequency corresponding to each unique identifier is used as the basis for access behavior evaluation, all statistical data are integrated to form a complete access data overview structure, and access statistical data slices are generated.

[0010] Preferably, the steps for obtaining the list of data to be optimized by stratification are: Calculating an access density stratification factor value of each data slice based on the access statistical data slices; According to the access density stratification factor value, a comparison is made with the average value and median value of the access density stratification factor values ​​of all data slices. If the access density stratification factor value is greater than both the average value and the median value, it is classified as a high storage priority; if it is only higher than one of them, it is classified as a medium priority; otherwise, it is classified as a low priority, and a list of data to be stratified and optimized is generated.

[0011] Preferably, the steps of obtaining the compressed parameterized data item are: Based on the storage level information of each data slice recorded in the list of data to be layered and optimized, all data slices in the list are traversed, unique identifiers are extracted, and a corresponding relationship between the unique identifiers and the storage levels to which they belong is established to obtain a data slice mapping table with level identifiers; According to the data slice mapping table with the level identifier, the compression parameter templates preset for different storage levels are retrieved, and the storage level to which each data slice belongs is matched with the corresponding compression parameter template item in sequence to generate a compression adaptation list; Based on the compression adaptation list, the unique identifier of each data slice, the storage level to which it belongs, and the corresponding compression parameter template information are combined and encapsulated into structured compression control data to generate compression parameterized data items.

[0012] Preferably, the step of obtaining the layered compressed data view is: Based on the unique identifier, compression parameter template and storage level information contained in the compression parameterized data item, the original data content of each data item is read in sequence, and compression is performed according to the corresponding compression parameters to generate a set of compressed data items to be stored; According to the storage level information corresponding to each data item in the set of compressed data items to be stored, and according to the configured hierarchical storage path rules, each data item is written to the target storage path of the target, and the writing path and the compression status mark are recorded to generate a compressed storage record set; Based on the compressed storage record set, the corresponding unique identifier entry in the searchable data slice directory is located, the storage location field and the compression status field are updated to the latest values, and a hierarchical compressed data view is generated.

[0013] The present invention provides a data intelligent processing system, comprising: Data screening module: Based on the input data stream of the data visualization platform, it calculates the data activity of each data segment, filters the data segments with data activity higher than the preset threshold, and generates a preliminary subset of interesting data; A positioning structure construction module: Based on each data segment in the preliminary interest data subset, the starting and ending position information in the original data stream is identified and assigned a unique identifier to obtain a descriptive data unit; based on the descriptive data unit, all unique identifiers and corresponding starting and ending position information are integrated to construct a positioning structure for information retrieval and establish a searchable data slice directory; Access statistics module: Based on the searchable data slice directory, monitor the search instructions of each data slice in the directory, accumulate the number of searches for each data slice to form the access frequency, and obtain the access statistical data slice; based on the access frequency of the access statistical data slice and the inherent size attribute of the data slice, assign a storage tier to each data slice and obtain a list of data to be optimized by tiering; Compression processing module: Based on the storage level of each data slice in the list of data to be layered and optimized, the corresponding compression parameter set is matched to obtain a compressed parameterized data item; based on the compressed parameterized data item, a compression operation is performed on the data item and stored in the target storage level; the storage location and compression status information of the data item in the retrievable data slice directory are synchronously updated to establish a layered compressed data view.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: The present invention calculates and screens the activity of each data segment in the input data stream, actively identifies high-value data fragments as preliminary interest data subsets, generates unique identifiers based on the start and end position information in the original data stream, and establishes descriptive data units, so that the original data has the ability to be accurately located and traced, thereby improving the addressability and traceability accuracy of the data during the retrieval process; by structurally integrating the unique identifier with the location information, a retrieval-oriented positioning structure is constructed, which solves the access redundancy problem caused by unclear structure in traditional information storage; during the data access process, each retrieval instruction is dynamically monitored and the access frequency is accumulated, the access behavior is precipitated into a quantitative indicator, and a hierarchical factor is calculated based on the size of the data slice itself to divide the storage priority, so as to allocate more efficient storage resources to high-frequency and large-data-volume data slices, and avoid the performance bottleneck caused by frequent access to low-speed storage; by matching the compression parameters with the storage hierarchy, the compression strategy is made differentiable and adaptable, thereby improving compression efficiency and data reading and writing performance; after the compression is completed, the storage location and compression status are updated, the real-time accuracy of the searchable directory is continuously maintained, and a consistent path support is provided for subsequent retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Schematic diagram of the steps of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] See also Figure 1 The present invention provides a technical solution, a data intelligent processing method for a data visualization platform, comprising the following steps: Based on the input data stream of the data visualization platform, the data activity of each data segment is calculated, and the data segments with data activity higher than the preset threshold are filtered to generate a preliminary subset of interesting data; Based on each data segment in the preliminary interest data subset, the start and end position information in the original data stream is identified and assigned a unique identifier to obtain a descriptive data unit. Based on the descriptive data unit, all unique identifiers and corresponding start and end position information are integrated to construct a positioning structure for information retrieval and establish a searchable data slice directory. Based on the searchable data slice directory, monitor the search instructions of each data slice in the directory, accumulate the number of searches for each data slice to form the access frequency, and obtain the access statistical data slice. Based on the access frequency of the access statistical data slice and the inherent size attribute of the data slice, assign a storage tier to each data slice and obtain a list of data to be optimized by tiering; Based on the storage level of each data slice in the list of data to be tiered and optimized, the corresponding compression parameter set is matched to obtain the compressed parameterized data item. Based on the compressed parameterized data item, the data item is compressed and stored in the target storage level. The storage location and compression status information of the data item in the retrievable data slice directory are synchronously updated to establish a tiered compressed data view.

[0018] The steps to obtain the initial subset of interest data are: Based on the data flow input from the data visualization platform, the duration of each data segment is set as a fixed time window, and the data is divided into several data segments in sequence. A sliding window is constructed for each data segment. Each group contains three consecutive data points and the pairwise difference slope, standard deviation within the window, and local maximum and minimum difference are calculated to form a set of original activity indicators for each data segment. Based on the original activity indicator set, the activity level value of each data segment is calculated using the following formula: ; in, is the activity level value of the d-th data segment, is the slope of the difference before and after the i-th three-point window in the d-th segment, is the standard deviation within the corresponding window, is the difference between the maximum and minimum values ​​in the data segment. is the mean of all input data, is the global mean of the fluctuation span of all data segments, is the number of three-point sliding windows in each data segment; According to the activity level value of each data segment, it is determined whether it exceeds the activity threshold, and all data segments whose activity level values ​​exceed the activity threshold are retained as preliminary interest data sources to form a preliminary interest data subset.

[0019] Specifically, based on the data flow input of the data visualization platform, the specific operation is to first set the duration of each data segment to a fixed time window length. For example, for the data flow generated by the real-time monitoring system, if the data sampling frequency is 1 data point per second, then every 60 seconds of data can be set as a data segment, that is, the time window seconds, thereby sequentially dividing the continuous input data stream into multiple non-overlapping data segments, ensuring that each data segment contains data of the same time length. Then, for each divided data segment, a sliding window of 3 data points is constructed. The sliding window starts from the starting position of the data segment and slides backward with a step size of 1 data point until the window can no longer completely contain 3 data points. For each 3 consecutive data points in the sliding window, it is recorded as , calculate the difference between the three data points and then calculate the slope. For example, you can calculate the point with dot The slope between the points with dot The slope between, or calculate the point with dot The slope between these three data points is taken as the representative slope of the window. The standard deviation is used to measure the degree of dispersion of the data in the window, and the difference between the maximum and minimum values ​​of the three data points is calculated, that is, the local range. These calculated slope values, standard deviation values, and local maximum and minimum value differences together constitute the original feature indicators of each sliding window within each data segment. These feature indicators generated by all sliding windows in a data segment are summarized and organized to form a set of original activity indicators for each data segment.

[0020] formula: The benefit of the formula is that it can comprehensively evaluate the "interestingness" or "importance" of a data segment, not just relying on the overall fluctuation amplitude of the data segment, but also going deeper into the local change characteristics within the data segment, by combining the rate of local change (slope ) and local variations instability (standard deviation ), can more accurately capture potential abnormal fluctuations or rapid changes in the data, and the molecule The data is normalized by the term, eliminating the influence of the absolute value of the data on the activity calculation, making data of different magnitudes comparable. The logarithmic term in the denominator Introducing the fluctuation range of the data segment itself and the global average fluctuation amplitude When the fluctuation range of a data segment far exceeds the average level, the denominator will be increased accordingly, thereby smoothing the activity value to a certain extent. This prevents the activity of data segments with large overall fluctuations from being overly amplified simply due to changes in local details. This design makes it more likely that the selected high-activity data segments represent real patterns or events worthy of attention, rather than simple noise or large but stable changes. The parameter is In the data segment The difference slope before and after a three-point sliding window represents the local change rate and direction of the data in the window. The acquisition steps are: for the data segment The first A three-point sliding window, which contains the data points (For example, if the data points are equally spaced in time, with a time interval of 1), the slope of the difference between the before and after values ​​is calculated as , which reflects the changing trend and magnitude of the data points on both sides of the window center. For example, a data visualization platform monitors the number of online users. No. There are three-point windows, and the number of users at three consecutive sampling points is 100, 150, and 120 respectively. .

[0021] The parameter is the standard deviation within the corresponding window, indicating the In the data segment Data points in a three-point sliding window The larger the standard deviation, the more drastic or unstable the data fluctuation within the window. For example, for the above user number data (100, 150, 120), the average value is , then the standard deviation .

[0022] The parameter is The difference between the maximum and minimum values ​​of the data in a data segment, that is, the overall fluctuation range or range of the data segment, is obtained by traversing the first Find the maximum value among all data points in a data segment and minimum value ,but For example, if a data segment contains 1-minute CPU usage data (percentage): [10, 12, 15, 50, 45, 48, 20, 22, 25, 18], then the data segment , ,therefore .

[0023] The parameter is the mean of all input data, representing the average level of the entire data set, and is used to normalize the scale of local changes in activity calculations. It is obtained by averaging the values ​​of all historical data points input to the data visualization platform or all data points in a sufficiently large recent representative data set. For example, by averaging the CPU usage data for the past 24 hours (a total of Data points (for example, once per minute) are collected and the average value of all CPU usage is calculated to be 35.0. .

[0024] The parameter is the global mean of the fluctuation span of all data segments, which represents the average level of the overall fluctuation amplitude of each data segment in the data set. It is used to evaluate the relative size of the fluctuation amplitude of the current data segment in the activity calculation. The method of obtaining it is: first calculate the fluctuation span of each data segment , then for all data segments The average value is calculated as follows: ,in It is The fluctuation span of the data segment, is the total number of data segments. For example, if 100 data segments are analyzed, their respective values, the sum of these values ​​is 3000, then .

[0025] The parameter is the number of three-point sliding windows in each data segment, which depends on the length of the data segment and the definition of the sliding window. The method of obtaining is: if a data segment contains data points, and the step size of the three-point sliding window is 1 data point, then For example, if a data segment contains 60 data points, the number of three-point sliding windows in the data segment is .

[0026] Calculation process: Set a data segment , which contains 5 data points: .

[0027] Calculation data segment Parameters within: data segment length The number of three-point sliding windows . .

[0028] Calculate the value of each sliding window and : Window 1( ): data points [10, 20, 15]; ; mean ; ; Window 2( ): data points [20, 15, 30]; ; mean ; ; Window 3 ( ): data points [15, 30, 25]; ; mean ; ; Set global parameters (obtained through statistics of a large amount of historical data or current batch data): The mean of all input data .

[0029] The global mean of the fluctuation span of all data segments .

[0030] Calculate the molecular part : for : ; for : ; for : ; Sum: ; Prescription: ; Calculate the denominator : ; ; ; calculate : ; The results show that the data segment The activity level value is 5.3445. This value itself is a relative value and needs to be compared with the activity level values ​​of other data segments to determine its relative activity level. A higher value (for example, greater than a preset threshold or between all The higher percentile of the value distribution means that there are significant and unstable changes within the data segment, and its overall fluctuation range is also representative. Such a data segment may contain important information or abnormal events and deserves further attention and analysis. A lower value indicates that the data segment is relatively stable or the changes are not significant.

[0031] The activity level value of each data segment calculated in the previous step Next, we need to determine whether the activity of each data segment exceeds a pre-set activity threshold. The setting of the activity threshold is crucial and directly affects the quality and quantity of the preliminary interest data subsets screened out. There are various ways to set it. For example, one method is based on experience. By analyzing the distribution of activity level values ​​of data segments corresponding to known important events or high-value information in historical data, a fixed value that can effectively distinguish these data segments is selected as the threshold. For example, if the analysis finds that the CPU utilization has increased sharply in the past, resulting in performance problems, the threshold can be set. If the value is generally higher than 15.0, The initial setting is 15.0. Another more adaptable method is to use statistical methods to dynamically determine the threshold, for example, calculate the activity level value of all data segments in the current batch The mean and standard deviation , and then set the threshold to ,in is a tunable parameter, such as or , indicating that the selection exceeds the average activity The data segment of times the standard deviation, or the percentile method can be used, for example, all After sorting the values, take the top 20% The minimum value corresponding to the value is the threshold, which ensures that a fixed proportion of high-activity data segments are screened out. The specific threshold setting method and parameters should be determined according to the specific needs of the application scenario, data characteristics, and the sensitivity and specificity requirements of the screening results. After that, the system will traverse all data segments value, each and Compare, all that meets The data segments that meet the conditions are all considered to be data segments with high activity and may contain useful information. These filtered data segments are retained intact as the source of preliminary interest data. Finally, all these data segments that meet the conditions are brought together to form a preliminary interest data subset.

[0032] The steps to obtain the descriptive data unit are: Based on the filtered data segments in the preliminary interest data subset, traverse each corresponding data segment in the original data stream, retrieve the index sequence position in the original data stream and record the start and end point indexes to obtain the start and end position information of each data segment in the original data stream; According to the start and end position information, the original data stream is segmented and numbered, a unique identifier is assigned to each data segment, and the start point index, end point index and unique identifier of each data segment are combined into a unified structure to generate an identification data mapping table; Based on the identified data mapping table, the complete data content corresponding to each data segment in the original data stream is extracted, and each data segment is combined and bound with the identification information and location information to form a descriptive data unit.

[0033] Specifically, based on the filtered data segments in the preliminary interest data subset, the system will start an iterative process to process each member data segment in the subset one by one. For the current data segment selected in the preliminary interest data subset, it may only contain the data value itself or a reference to the approximate location in the original data stream. The system will use this information to accurately match and locate in the complete, undivided original data stream (such as an array or list of data points arranged in chronological order or serial number). This positioning process may involve comparing the features of the interest data segment (such as the starting timestamp, a brief fingerprint of the data pattern) with the corresponding part of the original data stream, or if the interest data segment has recorded its approximate index range in the original stream when it was generated, then the range is directly used. Once the exact corresponding area of ​​the data segment of interest in the original data stream is determined, the system will accurately retrieve and record the index number of the starting data point of the data segment in the original data stream sequence and the index number of the ending data point in the original data stream sequence. For example, if the original data stream is a sequence containing millions of continuous sampling points, an identified data segment of interest may correspond to the part from the 10532th sampling point to the 10891th sampling point in the original sequence. Then the starting point index recorded by the system is 10532 and the ending point index is 10891. By repeating this traversal, retrieval and recording operation for all data segments in the preliminary interest data subset, the start and end position information of each data segment in the original data stream is obtained.

[0034] Based on the start and end position information of each data segment in the original data stream obtained in the previous step, the system will then uniformly identify these data segments that have been screened out from the preliminary interest data subset and have been precisely located in the original data stream. The specific operation is that the system assigns a globally unique identifier to each such data segment. The generation strategy of this unique identifier can adopt a variety of mature mechanisms, such as generating a UUID (universally unique identifier) ​​of version 4 that complies with the RFC4122 standard, such as "f47ac10b-58cc-4372-a567-0e02b2c3d479", or adopting a method based on a combination of a high-precision timestamp and an atomically incremented serial number. , for example, "YYYYMMDDHHMMSSms_SEQUENCE", to ensure the uniqueness and traceability of identifiers in distributed environments or long-term operations. After completing the allocation of unique identifiers, the system will integrate the starting point index, ending point index and corresponding unique identifier of each data segment into a structured data record. The format of the record can be a tuple or object containing three fields, for example, (unique identifier, starting point index, ending point index), such as ("SEG_TS_20250517140000_001", 10532, 10891). Such structured records of all processed data segments are gathered together to build and generate an identified data mapping table.

[0035] Based on the identified data mapping table generated in the previous stage, which stores the unique identifier of each data segment of interest and its precise start and end index in the original data stream, the system will extract the actual data content of each data segment of interest based on this mapping table and construct a descriptive data unit containing metadata. During execution, the system will traverse each entry in the identified data mapping table. For any record in the table, such as a record containing the unique identifier "SEG_TS_20250517140000_001", a start point index of 10532, and an end point index of 10891, the system will use these two index values ​​to extract the complete data sequence of the corresponding range from the original data stream storage, that is, starting from the 10532th data point of the original data stream and extracting all data up to the 10891th data point (including this point). The extracted data content will be combined with the unique identifier of the data segment and its start and end position information (which may again contain the start point index 10532 and the end point index 10891, or the calculated data segment length, such as This binding forms a structured data object, which not only contains the actual data slice, but also comes with key metadata about the data slice. After each data segment of interest is processed in this way, a series of descriptive data units are finally formed.

[0036] The steps to obtain the searchable data slice directory are: Based on the unique identifier and start and end position information bound in the descriptive data unit, the identification number of each descriptive data unit and the corresponding index boundary in the original data stream are extracted, a one-to-one mapping relationship between the unique identification number and the start and end index is constructed, and an identification index matching list is generated; According to the identification index matching list, the unique identifiers recorded in each descriptive data unit are classified and sorted, and encapsulated together with the corresponding starting position index and ending position index into a standard retrieval structure format to generate a structured index set; Based on the structured index set, all information items of unique identifiers, starting position indexes and ending position indexes are uniformly organized to establish a searchable data slice directory.

[0037] Specifically, based on the various descriptive data units generated in the previous steps, each of which encapsulates the actual data content and the unique identifier bound to it and the start and end position information of the data segment in the original data stream, the system will first traverse this series of descriptive data units, and accurately extract the unique identifier assigned to each descriptive data unit, such as a string of the form "DU_20250517_143005_00001", and at the same time extract the corresponding starting point index value and ending point index value of the recorded data segment in the original data stream. These two index values ​​together define the data segment in the unprocessed original continuous data sequence. Precise boundaries, for example, the starting index is 50000 and the ending index is 50600. Then, for each group of extracted unique identifiers and corresponding start and end index pairs, the system organizes them to build a clear one-to-one correspondence. For example, "DU_20250517_143005_00001" is linked to its index interval [50000, 50600] to ensure that this identifier can be traced back to its specific position in the original data stream without ambiguity. The mapping relationships generated by all descriptive data units after the above extraction and association operations are summarized to form a detailed list or set, that is, an identification index matching list is generated.

[0038] According to the identification index matching list generated in the previous step, which contains the unique identifiers of each data segment and their starting and ending indexes in the original data stream, the system will further structure this information. First, the system will review each record in the identification index matching list. Although the correspondence between the unique identifier of the record and the index has been ensured in the previous step, the "classification and arrangement" here focuses on separating these core positioning information (i.e., unique identifier, starting position index, and ending position index) from the list entries that may also contain other temporary information, and encapsulating them according to a predefined "standard retrieval structure format". The design of this standard retrieval structure format will aim to support efficient retrieval. For example, a structure containing the following fields can be defined: data_slice_id (storing unique identifier, type is string), start_offset (storing starting position index, type is The system uses the following structures: data_slice_id: "DU_20250517_143005_00001", "start_offset": 50000, "end_offset": 50600, "data_length": 601}. After all entries are converted and encapsulated in this format, these standard structure instances are collected to generate a structured index set.

[0039] Based on the constructed structured index set, each element of which is a data slice location information encapsulated in a standard retrieval structure format, the system will finally integrate and solidify these information items. Specifically, the system will traverse all entries in the structured index set, and reconfirm the data integrity and format consistency of the core information such as the unique identifier, starting position index and ending position index contained in each entry. For example, check whether the unique identifier conforms to the preset naming convention, whether the index value is a valid non-negative integer and the ending index is not less than the starting index, and arrange these information items in a certain order (for example, according to the lexicographical order of the unique identifier or in ascending order of the starting position index). After the arrangement is completed, the system will Some key positioning information is persistently stored or loaded into an efficient search data structure in memory, thereby establishing a searchable data slice directory. The core of this directory is a mapping mechanism that allows applications to quickly query the precise starting position index and ending position index of a data slice in the original data stream by providing a unique identifier of the data slice. For example, this directory can be implemented as a hash table, where the key is a unique identifier and the value is a simple object or tuple containing the starting and ending indexes, such as {"DU_20250517_143005_00001": (50000, 50600), "DU_20250517_143010_00002": (65000, 65250)}.

[0040] The steps to access the statistics slice are: Based on the unique identifier, starting position index, and ending position index of each data slice recorded in the searchable data slice directory, monitor the search request log issued, decode the search target contained in the search request log, identify and match the unique identifier referenced therein, and obtain the set of unique identifiers of the accessed data slices; Based on the set of unique identifiers of the accessed data slices, each occurrence of the unique identifier is recorded and counted. The number of occurrences within the target time range is accumulated one by one by the unique identifier dimension. The statistical results are then bound to the corresponding unique identifier to generate a detailed table of data slice access frequency. Based on the data slice access frequency details table, the cumulative access frequency corresponding to each unique identifier is used as the basis for access behavior evaluation. All statistical data are integrated to form a complete access data overview structure, and access statistical data slices are generated.

[0041] Specifically, based on the searchable data slice directory established in the previous step, which records in detail the unique identifier of each data slice and its starting position index and ending position index in the original data stream, the system will start the monitoring process of data access behavior. Specifically, the system will access and continuously analyze the retrieval request logs generated by the data service interface or related data query components. These logs usually contain information such as request time, source IP, request method, and specific data target of the request. For example, a typical log entry may be recorded as "Time: 2025-05-17 14:05:30, User: app_user_X, Request: GET / data / slices / UID_0012 3". The system will parse each newly generated retrieval request log entry or the one collected within the scheduled batch processing cycle, focusing on extracting the part of the request target that indicates the data slice to be accessed, and decoding the unique identifier of the data slice from it, such as extracting "UID_00123" from the above example. Subsequently, the system will match and verify this extracted unique identifier with the record in the searchable data slice directory. Only when the unique identifier exists in the directory is it considered a legitimate access to a valid data slice, and this unique identifier is recorded. After processing all log entries within the specified time window, a set of unique identifiers of all data slices that have been successfully identified and confirmed to be effectively accessed is finally obtained.

[0042] Based on the unique identifiers of all successfully accessed data slices obtained in the previous step within the specified monitoring period (this stage should process a list of corresponding unique identifiers in all valid access events, which allows repetition), the system will perform frequency statistics on these unique identifiers to quantify the access popularity of each data slice. Specifically, the system will set a clear "target time range", such as "the last 24 hours" or "any time period from the last statistics to the present", and then traverse all valid data slice access events recorded within this target time range. For the unique identifier extracted from each access event, the system will use an internal A counting mechanism, such as a hash map structure with a unique identifier as the key and the number of accesses as the value, is used to accumulate the total number of times each different unique identifier appears. For example, if the data slice "UID_00123" was requested to be read 150 times within the time range, and the data slice "UID_00456" was requested to be read 35 times, then their respective cumulative number of times will be accurately recorded. After completing the statistics of all access records within the target time range, the system will organize and bind these statistical results, that is, each unique identifier and its corresponding cumulative number of accesses, to form a structured record list. This list is the data slice access frequency details table.

[0043] Based on the data slice access frequency details table generated in the previous link, the table clearly lists the unique identifier of each data slice and its corresponding cumulative number of accesses within a specific statistical period. The system will use this information to build a more comprehensive access behavior portrait. Specifically, the system uses each record in the data slice access frequency details table, that is, each unique identifier and its corresponding cumulative access frequency, as the core basis for evaluating the current access activity of the data slice. Subsequently, the system will integrate these core access frequency data and may combine them with other relevant metadata stored in the retrievable data slice directory, such as the actual storage size of each data slice, Information such as creation date, data type or business affiliation is integrated with access frequency data to form a richer and more complete access data overview structure. Each entry in this structure not only reflects the access popularity of the slice, but may also contain other contextual information that affects its access pattern. For example, a record may contain: unique identifier "UID_00123", cumulative access frequency 150 times, data size 5MB, and creation date "2025-01-10". Finally, this complete access data overview structure that integrates all relevant statistical data and metadata is output or saved as a whole to generate access statistical data slices.

[0044] The steps to obtain the list of data to be layered and optimized are: Based on the access statistical data slices, the access density stratification factor value of each data slice is calculated using the following formula: ; in, For the The access density stratification factor value of each data slice, For the The number of data points per data slice, is the global average number of data points, For the The access frequency of each data slice, is the global average of access frequency, For the The average access time interval of each data slice in the past month (unit: seconds), is the average access time interval of all data slices, For the The number of storage blocks for a data slice, is the average number of storage blocks for all data slices; According to the access density stratification factor value, it is compared with the average value and median value of the access density stratification factor values ​​of all data slices. If the access density stratification factor value is greater than both the average value and the median value, it is classified as high storage priority. If it is only higher than one of them, it is classified as medium priority. Otherwise, it is classified as low priority, and a list of data to be stratified and optimized is generated.

[0045] Specifically, the formula: The benefit of the formula is that it provides a comprehensive evaluation index for data tiered storage optimization by quantifying the characteristics of data slices in multiple dimensions, namely, the access density tiering factor value. , the numerator takes into account the size of the data slice itself (through It shows that larger slices have the potential to obtain higher factor values, and the square operation amplifies its influence), as well as the comprehensive performance of the activity and recency of access behavior (through Reflects the data slices that are frequently accessed and recently accessed. Larger and Small, so that the difference term contributes a positive factor, and the absolute value ensures that slices with access patterns deviating from the average state will be paid attention to), and the denominator introduces the consideration of storage cost (through The more data slices occupy the more storage blocks, the larger the denominator is, so the factor value is appropriately lowered to avoid occupying too much high-cost storage simply due to large or frequent access). This design makes the calculated The value can more evenly reflect the importance and storage efficiency of data slices, provide a more reasonable basis for subsequent storage tier allocation, and encourage high-value, high-utilization, and relatively good storage-efficiency data to be stored in the high-performance tier first.

[0046] The parameter is The number of data points in a data slice represents the number of entries in the original data records contained in the data slice, which directly reflects the size of the data slice. This value can be calculated from the start point index and end point index of each data slice recorded in the process of building the searchable data slice directory. The calculation formula is: , for example, a data slice The starting point index in the original data stream is 1001 and the ending point index is 2000, so the number of data points is data points.

[0047] The parameter is the global average of the number of data points, representing the average level of all data slice sizes in the system. It is used to make relative comparisons of the sizes of individual data slices. The method of obtaining it is: summing up all Data slices The value is then calculated as the average value. For example, if there are 1000 data slices in the system and the total number of data points is 50,000,000, then data points.

[0048] The parameter is The access frequency of a data slice represents the total number of times the data slice is retrieved or accessed within a specific statistical period (for example, the last month). This data is directly obtained from the "Data Slice Access Frequency Details Table" generated in the previous step. For example, according to the Data Slice Access Frequency Details Table, the data slice 120 visits in the past month, .

[0049] The parameter is the global average of access frequency, which represents the average number of times all data slices in the system are accessed in the same statistical period. The acquisition method is: summing up all The access frequency of each data slice , and then calculate its average value. For example, if the total number of accesses to 1000 data slices in the system in the past month is 15,000 times, then Second-rate.

[0050] The parameter is The average access time interval of each data slice in the past month, in seconds, reflects the recency and regularity of data slice access. Smaller values ​​indicate more frequent or more concentrated access in the near future. The acquisition steps are: First, collect data slices In the past month (e.g. The timestamp of all access events within seconds is recorded as (in ascending order), if the number of visits , then calculate the time interval between two adjacent visits , and then calculate the mean of these intervals ,like ,but It can be set to the total duration of the statistical period, that is, 2592000 seconds, which means that it is only accessed once in the period. , It is also set to the total duration of the statistical period or a large agreed value, for example, a data slice It has been visited 3 times in the past month, and the timestamps (converted to seconds) are , then the average visit interval , if calculated seconds (i.e., an average of once a day).

[0051] The parameter is the average access time interval of all data slices, which represents the average length of time required for a data slice in the system to be accessed again. The acquisition method is: summing up all The average access time interval of a data slice , and then calculate its average, for example, the average of 1000 data slices in the system The total value is 500,000,000 seconds, then Second.

[0052] The parameter is The number of storage blocks for a data slice indicates the number of standard storage units (blocks) actually occupied by the data slice on the physical storage medium. This depends on the size of the data slice and the block size configuration of the storage system. The method for obtaining it is: first determine the number of bytes occupied by each data point and the block size of the storage system (For example, 4KB, bytes), then the data slice The total number of bytes is , so the number of storage blocks , that is, round up, for example, a data slice Include data points, each data point occupies 4 bytes, the storage block size is 4096 bytes, then the total number of bytes is bytes, so storage blocks.

[0053] The parameter is the average number of storage blocks for all data slices, which represents the average physical storage space (in blocks) occupied by data slices in the system. The acquisition method is: Summarize all The number of storage blocks for a data slice , and then calculate its average value. For example, if 1000 data slices in the system occupy a total of 2500 storage blocks, then storage blocks.

[0054] Calculation process: Taking a specific data slice For example, calculate its access density stratification factor value . Given the following parameter values: data points, data points, Second-rate, Second-rate, Second, Second, storage blocks, storage blocks; Evaluate the terms in the numerator: ; ; ; ; ; The values ​​within the square root of the numerator are: ; molecular ; Evaluate the terms in the denominator: ; ; Denominator ; calculate : ; This result shows that data slicing The access density tiering factor value is 0.805, which will be used to compare with the factor values ​​of other data slices to determine their relative storage optimization priority. A value of means that the data slice has a high activity or importance due to its comprehensive evaluation of size, access frequency, access recency and storage overhead, and may be allocated to a faster storage tier first. A value of 0 indicates a lower priority.

[0055] The access density stratification factor value calculated for each data slice in the previous step , the system will then analyze these factor values ​​to determine the storage priority of each data slice. First, the system will collect the access density tiering factor values ​​of all data slices in the current batch. , forming a A set of values, based on which the system will calculate two key statistics: all The average value of the value is recorded as , and all The median of the values ​​is denoted as , when calculating the average value, all The sum of the values ​​is divided by the total number of data slices. When calculating the median, all The values ​​are sorted from small to large. If the total number of data slices is an odd number, the median is the middle one after sorting. If the value is an even number, the median is usually the middle two values ​​after sorting. The average of these two statistical benchmarks and After that, the system will traverse each data slice value and compare it with and Compare and judge: If a data slice Values ​​greater than And also greater than , then the data slice is judged as "high storage priority"; if its Value only greater than and If the value of the data slice is greater than the average value but less than or equal to the median value, or less than or equal to the average value but greater than the median value, the data slice is judged as "medium storage priority"; if its The value is not greater than No more than (i.e. less than or equal to the average and less than or equal to the median), the data slice is judged to have "low storage priority". By executing this classification logic on all data slices, a list of data to be optimized is finally generated. Each entry in the list clearly indicates the unique identifier of the corresponding data slice and its assigned storage priority (high, medium or low).

[0056] The steps to obtain compressed parameterized data items are: Based on the storage level information of each data slice recorded in the data list to be layered and optimized, all data slices in the list are traversed, unique identifiers are extracted, and a corresponding relationship between the unique identifiers and the storage levels to which they belong is established to obtain a data slice mapping table with level identifiers; According to the data slice mapping table with level identification, the compression parameter templates preset for different storage levels are retrieved, and the storage level to which each data slice belongs is matched with the corresponding compression parameter template item in turn to generate a compression adaptation list; Based on the compression adaptation list, the unique identifier of each data slice, the storage level to which it belongs, and the corresponding compression parameter template information are combined and encapsulated into structured compression control data to generate compression parameterized data items.

[0057] Specifically, based on the list of data to be tiered and optimized generated in the previous step, the list has assigned a preliminary storage priority to each data slice, such as "high storage priority", "medium storage priority" or "low storage priority". The system will first traverse each record in this list. For each data slice entry in the list, the system will extract its unique identifier, such as "SLICE_ID_0001", and at the same time extract its assigned storage level information, namely "high storage priority". Subsequently, the system will establish a direct mapping relationship from its unique identifier to its corresponding storage level for each data slice, for example, mapping "SLICE_ID_0001" to "high storage priority", mapping "SLICE_ID_0002" to "medium storage priority", and so on. By performing this extraction and mapping operation on all data slices in the list, all these corresponding relationships are collected and organized to form a clear data structure that uses the unique identifier as an index to quickly find the storage level to which it belongs. This is the data slice mapping table with level identification.

[0058] According to the data slice mapping table with hierarchical identification established in the previous step, the mapping table clearly defines the unique identifier of each data slice and its corresponding storage level (i.e., "high storage priority", "medium storage priority" or "low storage priority"). The system will then match the appropriate compression strategy for each data slice. This process relies on a set of pre-defined compression parameter templates. These templates are specially set for different storage level characteristics. For example, for the "high storage priority" storage level, considering its high access performance requirements, the preset compression parameter template may be to use the LZ4 compression algorithm, the compression level is set to 1 (indicating speed priority), or no compression (the template parameter is {"algorithm": "None"}); for the "medium storage priority", the preset template may be to use the ZSTD compression algorithm, the compression level is set to 5 (indicating speed priority). For "low storage priority", which is more sensitive to storage space, the preset template may be to use the XZ compression algorithm and set the compression level to 9 (pursuing a high compression ratio). These preset compression parameter templates are usually stored in a configuration library. The system will traverse each data slice entry in the data slice mapping table with a hierarchical identifier and retrieve the corresponding compression parameter template content from the configuration library based on the storage level recorded. For example, if the level of the data slice "SLICE_ID_0001" is "high storage priority", it will match the template {"algorithm": "LZ4", "level": 1}. By performing this matching process for each data slice, a compression adaptation list is generated, in which each item contains the unique identifier of the data slice, its storage level, and the matched complete compression parameter template.

[0059] Based on the compression adaptation list generated in the previous step, each record contains the unique identifier of the data slice, its assigned storage tier, and a matching preset compression parameter template. The system will further process this information, integrate and encapsulate it into a structured data format, namely structured compression control data. Specifically, the system will traverse each entry in the compression adaptation list and extract the unique identifier of the data slice from the entry, such as "SLICE_ID_0001", the storage tier to which it belongs, such as "high storage priority", and the specific compression instructions parsed from the matching compression parameter template, such as compression. The name of the compression algorithm (such as "LZ4"), the compression level (such as 1), and other possible compression options (such as window size, whether to enable multi-threading, etc., which should be included in the parameter template), these extracted information - unique identifier, target storage tier and detailed compression parameter set - are combined and encapsulated into a separate data object or record. This encapsulated data object is a compression parameterized data item, which accurately defines all the control parameters and target states required for compressing a specific data slice. After all items in the compression adaptation list are processed in this way, a series of compression parameterized data items are finally formed.

[0060] The steps to obtain the hierarchical compressed data view are: Based on the unique identifier, compression parameter template and storage level information contained in the compression parameterized data item, the original data content of each data item is read in turn, compression is performed according to the corresponding compression parameters, and a set of compressed data items to be stored is generated; According to the storage level information corresponding to each data item in the compressed data item set to be stored, and according to the configured hierarchical storage path rules, each data item is written to the target storage path of the target, and the write path and compression status mark are recorded to generate a compressed storage record set; Based on the compressed storage record set, the corresponding unique identifier entry in the retrievable data slice directory is located, the storage location field and the compression status field are updated to the latest values, and a hierarchical compressed data view is generated.

[0061] Specifically, based on a series of compression parameterized data items generated in the previous step, each of which clearly indicates the unique identifier of a data slice, the specific compression parameters to be adopted (for example, the compression algorithm such as LZ4, ZSTD or GZIP, and the corresponding compression level) and the target storage layer to which the data slice is assigned, the system will process these compression parameterized data items in sequence. For each data item, first, the unique identifier contained in it is used to read the data at the original data storage location (for example, by consulting the initial searchable data slice directory to obtain the start and end indexes of the original data corresponding to the unique identifier in the original data stream). The data slice is the complete uncompressed original data content. After obtaining the original data, the system strictly follows the compression parameters specified in the current compression parameterized data item and calls the corresponding compression function library to perform compression operations on the data content. For example, if the parameter specifies the use of the ZSTD algorithm and the level is 3, the original data will be converted into a compressed byte stream in the ZSTD format. After compression is completed, the system will encapsulate the compressed data together with its original unique identifier and the target storage level information to form a unit to be processed containing the compressed data. After traversing and processing all the compression parameterized data items, it will finally generate a set of compressed data items to be stored.

[0062] Based on the set of compressed data items to be stored generated in the previous step, each data item contains the unique identifier of the data slice, the compressed data content, and the storage tier information (such as "high storage priority", "medium storage priority" or "low storage priority") where the data item is scheduled to be stored. The system will determine the specific storage location of each compressed data item based on a set of pre-configured tiered storage path rules. These rules map abstract storage tiers to actual physical or logical storage paths. For example, the rules may define that "high storage priority" data is stored in a solid-state drive partition with a path prefix of / mnt / fast_tier / compressed_data / , while "low storage priority" data is stored in a specific bucket of an object storage service such as Amazon S3: / In / archive-storage-bucket / cold-slices / , the system will traverse each item in the set of compressed data items to be stored, and build a complete target storage path based on its corresponding storage level information and the above path rules. The file name is usually based on a unique identifier and appended with the corresponding compression format suffix (for example, SLICE_ID_0001.zst). Then, the system writes the compressed data stream in the data item to the calculated target storage path. After successful writing, the system will record the complete file path of the data slice actually stored this time and a clear compression status flag, which can be the name of the compression algorithm used (such as "ZSTD") or a Boolean value indicating that it is compressed. These record information is collected to form a compressed storage record set.

[0063] Based on the compressed storage record set generated in the previous step, which contains the unique identifier of each successfully compressed and stored data slice, its latest physical storage path, and the current compression status information, the system will use these records to update the core metadata index, namely the "searchable data slice directory". The specific operation is that the system will traverse each record in the compressed storage record set. For each record, it will first extract the unique identifier of the data slice, and then use this unique identifier to locate the corresponding entry in the searchable data slice directory. Once the entry is found, the system will use the information in the current record to update it. Specifically, the "storage location" field (or similarly named field) of the entry corresponding to the unique identifier in the directory will be updated. The "segment" will be updated to the latest actual storage path recorded in the compressed storage record set (for example, from the path pointing to the original uncompressed data to the path pointing to the new compressed file / mnt / fast_tier / compressed_data / SLICE_ID_0001.zst). At the same time, the "Compression Status" field (or a similarly named field) will also be updated to the latest compression status mark (for example, updated to "ZSTD" or marked as compressed). If these fields do not exist in the directory entry yet, they will be created and filled. After completing the update of all related entries, the retrievable data slice directory reflects the latest status of the data after tiered compression, forming a tiered compressed data view.

[0064] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A data intelligent processing method for a data visualization platform, characterized in that: The following steps are involved: Based on the input data stream of the data visualization platform, the data activity of each data segment is calculated, and the data segments with data activity higher than the preset threshold are filtered to generate a preliminary subset of interesting data; Based on each data segment in the preliminary interest data subset, the start and end position information of the data segment in the original data stream is identified and assigned a unique identifier to obtain a descriptive data unit. Based on the descriptive data unit, all unique identifiers and corresponding start and end position information are integrated to construct a positioning structure for information retrieval and establish a searchable data slice directory. Based on the searchable data slice directory, the search instructions of each data slice in the directory are monitored, the number of searches for each data slice is accumulated to form an access frequency, and an access statistical data slice is obtained. Based on the access frequency of the access statistical data slice and the inherent size attribute of the data slice, a storage tier is assigned to each data slice, and a list of data to be tiered and optimized is obtained; Based on the storage level of each data slice in the list of data to be layered and optimized, the corresponding compression parameter set is matched to obtain a compressed parameterized data item. Based on the compressed parameterized data item, a compression operation is performed on the data item and stored in the target storage level. The storage location and compression status information of the data item in the retrievable data slice directory are synchronously updated to establish a layered compressed data view.

2. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps for obtaining the preliminary interest data subset are: Based on the data flow input from the data visualization platform, the duration of each data segment is set as a fixed time window, and the data is divided into several data segments in sequence. A sliding window is constructed for each data segment. Each group contains three consecutive data points and the pairwise difference slope, standard deviation within the window, and local maximum and minimum difference are calculated to form a set of original activity indicators for each data segment. Calculating an activity level value for each data segment based on the original activity indicator set; According to the activity level value of each data segment, it is determined whether it exceeds the activity threshold, and all data segments whose activity level values ​​exceed the activity threshold are retained as preliminary interest data sources to form a preliminary interest data subset.

3. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps of obtaining the descriptive data unit are: Based on each data segment that has been filtered in the preliminary interest data subset, traverse each corresponding data segment in the original data stream, retrieve the index sequence position in the original data stream and record the start point and end point indexes to obtain the start and end position information of each data segment in the original data stream; According to the start and end position information, the original data stream is segmented and numbered, a unique identifier is assigned to each data segment, and the start point index, end point index and unique identifier of each data segment are combined into a unified structure to generate an identification data mapping table; Based on the identified data mapping table, the complete data content corresponding to each data segment in the original data stream is extracted, and each data segment is combined and bound with the identification information and the position information to form a descriptive data unit.

4. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps for obtaining the searchable data slice directory are: Based on the unique identifier and start and end position information bound in the descriptive data unit, extract the identification number of each descriptive data unit and the corresponding index boundary in the original data stream, establish a one-to-one mapping relationship between the unique identification number and the start and end indexes, and generate an identification index matching list; According to the identification index matching list, the unique identifiers recorded in each descriptive data unit are classified and sorted, and encapsulated together with the corresponding starting position index and ending position index into a standard search structure format to generate a structured index set; Based on the structured index set, all information items of unique identifiers, starting position indexes and ending position indexes are uniformly organized to establish a searchable data slice directory.

5. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps for obtaining the access statistics data slice are: Based on the unique identifier, starting position index and ending position index of each data slice recorded in the searchable data slice directory, monitor the search request log issued, decode the search target contained in the search request log and identify the unique identifier referenced therein to obtain the set of unique identifiers of the accessed data slices; According to the set of unique identifiers of the accessed data slices, each occurrence of the unique identifier is recorded and counted, the number of occurrences within the target time range is accumulated one by one according to the unique identifier dimension, and the statistical results are bound to the corresponding unique identifier to generate a detailed table of data slice access frequency; Based on the data slice access frequency details table, the cumulative access frequency corresponding to each unique identifier is used as the basis for access behavior evaluation, all statistical data are integrated to form a complete access data overview structure, and access statistical data slices are generated.

6. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps for obtaining the list of data to be optimized by stratification are as follows: Calculating an access density stratification factor value of each data slice based on the access statistical data slices; According to the access density stratification factor value, a comparison is made with the average value and median value of the access density stratification factor values ​​of all data slices. If the access density stratification factor value is greater than both the average value and the median value, it is classified as a high storage priority; if it is only higher than one of them, it is classified as a medium priority; otherwise, it is classified as a low priority, and a list of data to be stratified and optimized is generated.

7. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps for obtaining the compressed parameterized data item are: Based on the storage level information of each data slice recorded in the list of data to be layered and optimized, all data slices in the list are traversed, unique identifiers are extracted, and a corresponding relationship between the unique identifiers and the storage levels to which they belong is established to obtain a data slice mapping table with level identifiers; According to the data slice mapping table with the level identifier, the compression parameter templates preset for different storage levels are retrieved, and the storage level to which each data slice belongs is matched with the corresponding compression parameter template item in sequence to generate a compression adaptation list; Based on the compression adaptation list, the unique identifier of each data slice, the storage level to which it belongs, and the corresponding compression parameter template information are combined and encapsulated into structured compression control data to generate compression parameterized data items.

8. The data intelligent processing method of the data visualization platform according to claim 1 is characterized in that: The steps for obtaining the layered compressed data view are: Based on the unique identifier, compression parameter template and storage level information contained in the compression parameterized data item, the original data content of each data item is read in sequence, and compression is performed according to the corresponding compression parameters to generate a set of compressed data items to be stored; According to the storage level information corresponding to each data item in the set of compressed data items to be stored, and according to the configured hierarchical storage path rules, each data item is written to the target storage path of the target, and the writing path and the compression status mark are recorded to generate a compressed storage record set; Based on the compressed storage record set, the corresponding unique identifier entry in the searchable data slice directory is located, the storage location field and the compression status field are updated to the latest values, and a hierarchical compressed data view is generated.

9. The data intelligent processing system of the data intelligent processing method of the data visualization platform according to any one of claims 1 to 8, characterized in that: include: Data screening module: Based on the input data stream of the data visualization platform, it calculates the data activity of each data segment, filters the data segments with data activity higher than the preset threshold, and generates a preliminary subset of interesting data; Positioning structure building module: based on each data segment in the preliminary interest data subset, identifying the start and end position information in the original data stream, assigning a unique identifier, and obtaining a descriptive data unit; Based on the descriptive data units, all unique identifiers and corresponding start and end position information are integrated to construct a positioning structure for information retrieval and establish a searchable data slice directory; Access statistics module: Based on the searchable data slice directory, monitor the search instructions of each data slice in the directory, accumulate the number of searches for each data slice to form the access frequency, and obtain the access statistical data slice; Based on the access frequency of the access statistical data slice and the inherent size attribute of the data slice, a storage tier is assigned to each data slice, and a list of data to be optimized by tiering is obtained; Compression processing module: based on the storage level of each data slice in the list of data to be layered and optimized, matches the corresponding compression parameter set to obtain compression parameterized data items; Based on the compression parameterized data item, performing a compression operation on the data item and storing the data item in a target storage level; Synchronous updates can retrieve the storage location and compression status information of the data item in the data slice directory and establish a hierarchical compressed data view.