Processing method for data intelligent analysis based on a time series analysis model
Through the processing method based on the timing analysis model, time sequence data is collected and parsed in real time and time sequence change characteristics are identified, the traditional methods are solved in generalization ability, robustness and stability, and efficient and accurate timing data analysis is achieved.
Patent Information
- Application Number
- CN202411783082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Traditional timing data analysis methods rely on event sequence, lack generalization ability, robustness and stability, making it difficult to effectively process and analyze massive time series data.
Using a processing method based on the timing analysis model, the task data packets are collected and parsed in real time, the timing data under the timestamp is extracted, and the timing analysis model is used to identify the timing change characteristics to generate a data change feature table.
It improves the efficiency and accuracy of time-series data analysis, can adapt to online real-time analysis of massive time-series data, and provides higher data support capabilities.
Smart Images

Figure CN119272203B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a processing method for intelligent data analysis based on a time series analysis model, a processing device for intelligent data analysis based on a time series analysis model, an electronic device, and a computer-readable storage medium. Background Art
[0002] Time series data, also known as time series data, is data collected at different time points and is used to describe the situation of a phenomenon changing over time. Such data reflects the change state or degree of a certain thing or phenomenon over time, such as including stock prices, temperature changes, website traffic, server logs, etc. Its characteristics include timestamps, numerical values, and possible periodicity, trend, and seasonality.
[0003] Time series data can be further divided into quarterly data, monthly data, etc., and has the characteristic of periodic change. A time-series database is a database system specifically used to process data with time tags (changing in the order of time, i.e., time serialization). The typical characteristics of time series data include fast generation frequency, strong dependence on the collection time, many measurement points, and large amounts of information.
[0004] Time series data is widely used in various fields. For example, the Internet of Things (IoT) field needs to process a large amount of real-time data; in the industrial Internet, a large amount of sensor data with timestamps is generated during the production, testing, and operation stages; in the vehicle networking, real-time in-vehicle network quality monitoring can be achieved through the analysis of vehicle machine messages; in the power industry, the data generated in each link of power generation, transmission, transformation, distribution, and use is huge, and a professional time series database is required for processing.
[0005] Time series data analysis plays an important role in various fields, mainly including the following aspects:
[0006] 1. Predicting future trends: Through the analysis of historical data, time series data analysis can predict future market trends, user needs, etc., providing a basis for enterprise decision-making.
[0007] 2. Optimizing resource allocation: By analyzing historical data, time series data analysis can help enterprises optimize the allocation of resources such as production, inventory, and human resources, reduce costs, and improve efficiency.
[0008] 3. Quality control: Time series data analysis can monitor product quality, timely detect potential problems, avoid the production of defective products, and improve product quality.
[0009] 4. Risk warning: By analyzing risk factors such as financial markets and natural disasters, time series data analysis can provide risk warnings for enterprises and help avoid potential risks.
[0010] 5. Intelligent decision-making: Combining time series data analysis with artificial intelligence technology can provide intelligent decision-making support for enterprises and improve decision-making efficiency.
[0011] However, the current analysis and processing methods of traditional time series data mainly include the following methods:
[0012] One is to adopt the manual analysis method. Manual analysis requires analysts to have high data analysis capabilities and entry levels, and need to master a high level of data analysis professional capabilities. Therefore, the entry requirements are high and the labor cost is extremely high. Moreover, in the face of big data analysis, manual analysis is limited and cannot improve the data analysis efficiency of big data, and only adapts to small-scale data analysis scenarios.
[0013] The second is the traditional machine analysis methods including autoregressive (AR), moving average (MA), autoregressive moving average (ARMA), autoregressive integrated moving average (ARIMA), etc. Although it can improve the efficiency to a certain extent, these models rely on the order of event occurrence for prediction, so the analysis ability lacks breadth and accuracy. Because traditional models can only analyze and predict according to the generated events / logs of time series data, and cannot accurately and randomly find the characteristics of time series data from time series data, traditional models lack strong generalization ability, robustness and stability in the prediction analysis of time series data. Summary of the Invention
[0014] In order to solve the technical problems existing in the prior art, the present invention provides the following technical solutions:
[0015] On the one hand, a processing method for intelligent data analysis based on a time series analysis model is provided. This method is implemented by an electronic device, and the method includes:
[0016] S1. Real-time collect and parse the task data packets uploaded by the front end, and extract the time series data at several timestamps from the parsed data;
[0017] S2. Traverse the time series data at each timestamp, and import the time series data into a preset time series analysis model. Identify the time series change characteristics in the time series data through the time series analysis model and output them;
[0018] S3. Bind the time series change characteristics to the corresponding timestamps to generate a data change characteristic table corresponding to the task data packet;
[0019] S4. Feed back the data change characteristic table to the front end along the original path.
[0020] Preferably, S1: real-time collect and parse the task data packets uploaded by the front end, and extract the time-series data at several timestamps from the parsed data, including:
[0021] Receive the task data packets uploaded by the front end through the background server interface;
[0022] Identify the data format of the task data packets, and parse the task data packets according to the parsing method corresponding to the data format to obtain several time-series data marked with timestamps;
[0023] Traverse and identify the timestamps of each piece of the time-series data, and judge whether the timestamps of each piece of the time-series data are qualified based on the digital signature judgment rules pre-configured for the task data packets:
[0024] If qualified, store the time-series data under the qualified timestamps in sequence in the background MySQL database;
[0025] Otherwise, notify the front end to correct the data of the time-series data under the unqualified timestamps; after the correction is qualified, store it in the background MySQL database.
[0026] Preferably, the traversing and identifying the timestamps of each piece of the time-series data, and judging whether the timestamps of each piece of the time-series data are qualified based on the digital signature judgment rules pre-configured for the task data packets includes:
[0027] Extract the keywords and judgment logic in the digital signature judgment rules pre-configured for the task data packets, where the keywords include timestamp type, version, format, and digital signature identifier;
[0028] Use the keywords and the judgment logic as prompt words and input them into a pre-set large language model (LLM);
[0029] Traverse the timestamps of each piece of the time-series data, and identify and judge whether the timestamps of each piece of the time-series data are qualified through the LLM based on the keywords and the judgment logic.
[0030] Preferably, the method for constructing the time-series analysis model includes:
[0031] Collect and parse several historical task data packets, and construct a historical time-series data set composed of several historical time-series data;
[0032] Traverse the historical time-series data set, find out and remove the historical time-series data without timestamps, and sort the remaining historical time-series data according to timestamps to reorganize the historical time-series data set;
[0033] Perform feature engineering on the historical time series dataset to extract time series change features, including:
[0034] Extract the time features in the timestamps of each historical time series data;
[0035] Use a preset sliding window to extract data feature values in each historical time series data, including mean, variance, maximum value, minimum value, and / or median;
[0036] Identify and extract the trend / seasonal features between each time feature;
[0037] Statistically analyze each time feature, the data feature values, and the trend / seasonal features corresponding to each time feature to form a feature set;
[0038] Divide the feature set into a training set and a validation set according to a preset ratio;
[0039] Input the training set into a preset time series analysis model for feature learning, and optimize and train the time series analysis model according to the preset model optimization iteration conditions;
[0040] Use the validation set to verify the performance of the time series analysis model:
[0041] If the verification passes, deploy and apply the time series analysis model;
[0042] Otherwise, repeat the above steps.
[0043] Preferably, traverse the historical time series dataset, find and eliminate the historical time series data without timestamps, including:
[0044] Traverse the timestamps of each historical time series data in the historical time series dataset, and use the LLM large language model to identify the historical time series data without the keyword in the historical time series dataset based on the keyword;
[0045] Eliminate the historical time series data without the keyword identified by traversal from the historical time series dataset.
[0046] Preferably, in step S3, bind the time series change features to the corresponding timestamps to generate a data change feature table corresponding to the task data packet, including:
[0047] Read the time series change features at each timestamp: the time feature, the data feature value, and the trend / seasonal features corresponding to each time feature;
[0048] Sequentially write the timing change characteristics at each of the time stamps into a preset data fluctuation statistical table, and generate a data change characteristic table for the task data packet this time.
[0049] On the other hand, a processing device for intelligent data analysis based on a timing analysis model is provided. The processing device for intelligent data analysis based on the timing analysis model is used to implement the processing method for intelligent data analysis based on the timing analysis model. The device includes:
[0050] A front end for real-time collecting and uploading task data packets to a data analysis background;
[0051] The data analysis background includes:
[0052] A data receiving port for receiving task data packets uploaded by the front end;
[0053] A parsing module for parsing the task data packet and extracting timing data at several time stamps from the parsed data;
[0054] A timing feature recognition module for traversing the timing data at each time stamp, importing the timing data into a preset timing analysis model, identifying the timing change characteristics in the timing data through the timing analysis model, and outputting them;
[0055] A timing fluctuation recording module for binding the timing change characteristics with the corresponding time stamps to generate a data change characteristic table corresponding to the task data packet;
[0056] A message center for feedbacking the data change characteristic table back to the front end along the original path;
[0057] A MySQL database for storing data;
[0058] The front end is communicatively connected to the data analysis background.
[0059] On the other hand, an electronic device is provided. The electronic device includes: a processor; a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, any one of the methods in the above-mentioned processing method for intelligent data analysis based on the timing analysis model is implemented.
[0060] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned processing method for intelligent data analysis based on the timing analysis model.
[0061] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0062] The present invention parses and extracts time-series data from the task data packets uploaded by the front end based on timestamps, and obtains the time-series data at each timestamp in the current task. To facilitate the background's data analysis of the current task, the time-series data at each timestamp is subjected to AI recognition, and a time-series analysis model constructed by training with big data is used to recognize the time-series change characteristics (time characteristics, data characteristic values including mean, variance, maximum value, minimum value, and / or median, and trend / seasonal characteristics) in the time-series data. The time-series change characteristics of the time-series data at each timestamp in the current task are recorded, and a data change characteristic table for the current task is generated, facilitating the user to view the analysis and processing results of the time-series data in the current task. Therefore, based on time-series AI analysis, the change characteristics of task data can be quickly analyzed. Compared with manual analysis, the present invention can greatly improve the analysis efficiency of time-series data, and can adapt to the online real-time analysis of massive time-series data, providing higher data support capabilities.
[0063] Based on feature engineering, the present invention uses a time-series analysis model for feature training and learning, allowing the model to perform training and learning on the extracted feature set, and allowing the model to perform training and learning through the feature set of time-series change characteristics without sequential learning, enabling the model to obtain feature training and learning of a massive training set, facilitating the improvement of the robustness and randomness of the model's feature recognition, and improving the generalization ability of the model's prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 is a flowchart of a processing method for intelligent data analysis based on a time-series analysis model provided by an embodiment of the present invention;
[0066] Figure 2 is a schematic diagram of data parsing of a task data packet provided by an embodiment of the present invention;
[0067] Figure 3 is a schematic diagram of intelligent qualification judgment of timestamp data based on the LLM large language model provided by an embodiment of the present invention;
[0068] Figure 4 is a block diagram of a processing device for intelligent data analysis based on a time-series analysis model provided by an embodiment of the present invention;
[0069] Figure 5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0070] The technical solutions in the present invention will be described below with reference to the accompanying drawings.
[0071] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.
[0072] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, their intended meanings are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, their intended meanings are the same.
[0073] In the embodiments of the present invention, sometimes subscripts such as W1 may be miswritten as non-subscript forms such as W1. When their differences are not emphasized, their intended meanings are the same.
[0074] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0075] The embodiments of the present invention provide a processing method for intelligent data analysis based on a time series analysis model. This method can be implemented by an electronic device, and the electronic device can be a terminal or a server. As Figure 1 shown in the flowchart of the processing method for intelligent data analysis based on the time series analysis model, the processing flow of this method can include the following steps:
[0076] S1. Real-time collect and parse the task data packets uploaded by the front end, and extract the time series data at several time stamps from the parsed data;
[0077] S2. Traverse the time series data at each time stamp, import the time series data into a preset time series analysis model, identify the time series change characteristics in the time series data through the time series analysis model, and output them;
[0078] S3. Bind the time series change characteristics to the corresponding time stamps to generate a data change characteristic table corresponding to the task data packet;
[0079] S4. Feed back the data change characteristic table to the front end along the original path.
[0080] The task data packet solved by the present invention can be production task data (involving production data output and possible costs, profits, etc.) or industrial sensing data, etc., or other data that changes over time, such as stock data, power generation data (involving future power generation power prediction), etc.
[0081] The present invention uses artificial intelligence technology to perform intelligent analysis of continuous data, uses a deep learning model to learn the continuous change characteristics of time series data, and uses the trained time series analysis model to perform intelligent analysis of continuous data on the task data packet uploaded by the current front end, so as to analyze the change characteristics of each time series data by the model and feedback them to the front end. In this way, it replaces manual analysis and outputs the change characteristics of the corresponding time series data for the reference of front-end users. It can quickly display the change rules, change characteristics, etc. of the current continuous data for the front end and provide data analysis support for the front end.
[0082] As Figure 2 shown, preferably, in S1, the task data packet uploaded by the front end is collected and parsed in real time, and several time series data under several timestamps are extracted from the parsed data, including:
[0083] Receive the task data packet uploaded by the front end through the background server interface;
[0084] Identify the data format of the task data packet, and parse the task data packet according to the parsing method corresponding to the data format to obtain several time series data marked with timestamps;
[0085] Traverse and identify the timestamps of each time series data, and judge whether the timestamps of each time series data are qualified based on the digital signature judgment rule pre-configured for the task data packet:
[0086] If it is qualified, store the time series data under the qualified timestamp in sequence in the background MySQL database;
[0087] Otherwise, notify the front end to correct the time series data under the unqualified timestamp; after the correction is qualified, store it in the background MySQL database.
[0088] The task data packet uploaded by the front end, such as the business data packet that needs to be analyzed for a certain task currently (for example, for the analysis of the change in the power generation output power at the current stage, the data here can include the power generation operation data at different time periods / points (power generation voltage, current, temperature, wind speed, pitch angle, high-speed characteristic number, rotor diameter, rated power, conversion efficiency, and power generation coefficient of each unit, etc.)) or other data form task data packets. The front end needs to let the back end analyze the data change rule. Therefore, the current task data packet can be uploaded, received by the back-end service server interface, and the data packet is parsed to obtain the task data (time series data, such as the power generation unit operation data at a certain point and minute or a certain hour) at each time point / segment. And there is time series data under several time periods in the task data, and the time series data at different time points needs to be parsed. And in this aspect, the time stamp method is used to mark the time series data such as different time points and time periods, so as to perform the encrypted digital transmission and analysis authentication of the time series data.
[0089] The method of adding a time stamp: Add a new column named "time stamp" to the data record; Use the computer system time or the actual time point of the data record to convert each time point / segment into the time stamp format; If the data already contains the date and time, the timestamp() function in the datetime library of Python can be used to convert the date and time object into a time stamp; If the data only contains the date, it needs to be converted into a datetime object first and then into a time stamp; If the data contains different time formats, the time formats need to be unified first and then the time stamp conversion is performed; Add the converted time stamp to the corresponding data record. It can be specifically added by the front end.
[0090] Adding a time stamp to the time series data can accurately record the time point when the data is generated, facilitate the subsequent analysis of the time series characteristics and trend changes of the data, and at the same time contribute to data synchronization and event sorting. And it can use the time stamp to perform digital authentication on the time series data, identify the generation time of specific time series data, facilitate the system to prove the integrity of the data, and the back end can perform electronic verification on the time series data based on the digital signature judgment rule of the time stamp to verify whether the data is complete. And the digital signature judgment rule consists of keywords and judgment logic. Among them, the keywords include time stamp type, version, format, and digital signature identifier. The back end can pre-configure the digital signature judgment rules required for different tasks to perform digital signature judgment on the time series data in each task, so as to screen out the time series data with unqualified time stamps or not belonging to this task, and avoid the mixing of time series error data in this task.
[0091] The background traverses each piece of parsed timing data based on the digital signature judgment rule, determines whether the timestamps of each piece of the timing data are qualified, and eliminates the unqualified timing data. The timestamp recording information in the timing data can refer to the above timestamp adding method.
[0092] In order to improve efficiency, the present invention uses a large language model (LLM) to assist in completing the above digital signature judgment link, replacing the background administrator, improving the screening efficiency of timing data, and realizing intelligent and automated judgment work.
[0093] As Figure 3 shown, preferably, traversing and identifying the timestamps of each piece of the timing data, and determining whether the timestamps of each piece of the timing data are qualified based on the digital signature judgment rule pre-configured for the task data packet, includes:
[0094] Extracting the keywords and judgment logic in the digital signature judgment rule pre-configured for the task data packet, wherein the keywords include timestamp type, version, format, and digital signature identifier;
[0095] Taking the keywords and the judgment logic as prompt words and inputting them into a pre-set large language model (LLM);
[0096] Traversing the timestamps of each piece of the timing data, and identifying and determining whether the timestamps of each piece of the timing data are qualified by the large language model (LLM) based on the keywords and the judgment logic.
[0097] By receiving prompt words, the large language model (LLM) can perform data screening and analysis, thereby generating relevant content or providing decision support.
[0098] The steps for the large language model (LLM) to traverse and screen data that conforms to the prompt words based on the input prompt words may include the following steps:
[0099] 1. Tokenization and Embedding
[0100] Tokenization: The LLM first uses a tokenizer to split the input prompt words into several small text blocks (tokens), and these tokens can be words, partial words, or character sequences.
[0101] Embedding: The split tokens are mapped to specific token indices and converted into numerical representations of high-dimensional vectors, i.e., embeddings, which capture the semantic relationships and context information between the words.
[0102] 2. Prediction and Screening
[0103] Prediction: After word segmentation and embedding are completed, the LLM enters the prediction stage, generating grammatically and semantically correct text through multi-layer neural networks and attention mechanisms.
[0104] Filtering: Based on the generated text, the LLM traverses its internal data and filters out the most relevant and suitable data related to the prompt as the output.
[0105] This process ensures that the LLM can accurately understand and respond to the input prompt, thus providing the information required by the user.
[0106] Combined with the attached Figure 4 As shown, the timing feature recognition module on the background can log in to the third-party AI platform through the large language model API interface, request to call the LLM large language model, and thus complete the judgment work of the present invention. And the administrator can construct a prompt and input it into the large language model, so that the large language model makes judgments on each timing data according to the keywords and judgment logic configured in the present invention regarding the digital judgment of timing data.
[0107] The third-party LLM large language model, such as Wenxin Yiyan of Baidu, etc. It is up to the user to decide.
[0108] Using the model to traverse the timestamps of each timing data, based on the keywords and the judgment logic, identify and judge whether the timestamps of each timing data are qualified. If qualified, store the timing data under the qualified timestamp in sequence into the background MySQL database;
[0109] Conversely, notify the front end to correct the timing data under the unqualified timestamp. After the correction is qualified, store it into the background MySQL database. During the correction, the background can generate corresponding correction instructions according to the model recognition results and send them to the front end, allowing the front end to re-add the timestamp to the unqualified data. The instruction needs to include the timing of the timing data to be corrected to facilitate reminding the front end to perform corresponding data processing.
[0110] The training and construction scheme of the timing analysis model will be described below.
[0111] Preferably, the construction method of the timing analysis model includes:
[0112] Collect and parse a number of historical task data packets to construct a historical timing data set composed of a number of historical timing data;
[0113] Traverse the historical timing data set, find and eliminate the historical timing data without timestamps, and sort the remaining historical timing data according to timestamps to reorganize the historical timing data set;
[0114] Perform feature engineering on the historical time series dataset to extract time series change features, including:
[0115] Extract the time features in the timestamps of each historical time series data;
[0116] Use a preset sliding window to extract data feature values in each historical time series data, including mean, variance, maximum value, minimum value, and / or median;
[0117] Identify and extract the trend / seasonal features between each time feature;
[0118] Statistically analyze each time feature, the data feature values, and the trend / seasonal features corresponding to each time feature to form a feature set;
[0119] Divide the feature set into a training set and a validation set according to a preset ratio;
[0120] Input the training set into a preset time series analysis model for feature learning, and optimize and train the time series analysis model according to the preset model optimization iteration conditions;
[0121] Use the validation set to verify the performance of the time series analysis model:
[0122] If the verification passes, deploy and apply the time series analysis model;
[0123] Otherwise, repeat the above steps.
[0124] 1. Data collection
[0125] Historical task data packets can be collected from the background database (MySQL database). Specifically, several historical task data packets can be collected, where each data packet contains several historical time series data (such as historical generator set operation data, etc.).
[0126] After collecting the historical task data packets, it is necessary to parse them to obtain the historical time series data under several tasks, divide them according to the task time series, and form a historical time series dataset (historical time series data under each time series task). Here, taking one task as an example, for instance, historical power generation data at different time series can be collected, and other analysis tasks such as power generation cost analysis can be based on the analysis model of the present invention, which will not be elaborated here.
[0127] After collecting each historical time series data, preprocessing can be performed, such as:
[0128] Data understanding: Analyze the nature of the data, including timestamp format, data frequency (such as per second, per minute, etc.), missing value situation, outliers, etc.
[0129] 2. Data Preprocessing
[0130] Timestamp processing: unify the timestamp format, ensure that all timestamps are in the same time zone, and convert to the appropriate granularity (such as day, hour, minute, etc.) as needed;
[0131] Missing value processing: Based on the characteristics of the data and business needs, choose to fill in missing values (such as mean, median, previous value, etc.) or delete records containing missing values;
[0132] Outlier detection and processing: Use statistical methods or machine learning algorithms to detect outliers and correct or delete them as appropriate;
[0133] Data normalization / standardization: According to the distribution characteristics of the data, select an appropriate normalization or standardization method to eliminate dimensional differences.
[0134] Data timestamp verification, elimination of data with unqualified timestamp marks: traverse the historical time series data set, find and eliminate the historical time series data without timestamps, and sort the remaining historical time series data according to timestamps to reconstruct the historical time series data set. In order to avoid data loss affecting the accuracy of subsequent model recognition, it is necessary to verify the historical time series data set. Here, based on the previous method of traversing and identifying data using the LLM large language model, the model training big data (historical time series data set) is verified and identified, and the historical time series data without timestamps is found and eliminated, so as to quickly clean the training big data and improve the efficiency of model training.
[0135] Preferably, traversing the historical time series data set to find and remove the historical time series data without a timestamp includes:
[0136] Traversing the timestamps of each of the historical time series data in the historical time series data set, and identifying the historical time series data in the historical time series data set that does not have the keyword based on the keyword through the LLM large language model;
[0137] The historical time series data that is traversed and identified and does not have the keyword is removed from the historical time series data set.
[0138] The method of using a large language model to traverse and identify historical time series data sets can refer to the application of the previous large model. The model prompt words prepared for traversing and identifying the data in the historical time series data sets required for model training only need to use keywords to identify and clean the data, thereby reducing the difficulty of cleaning the training data here, and the prepared keywords can be determined by the user.
[0139] 3. Feature engineering, extracting time series change features
[0140] Next, it is necessary to extract the temporal features in the historical time-series data. The following temporal change features are mainly extracted:
[0141] 3.1 Extract the time features in the timestamps of each of the historical time-series data;
[0142] Time features, such as extracting time-related information such as date, time, day of the week, month, quarter, etc.
[0143] 3.2 Data feature values. Statistical methods can be used. Using a preset sliding window, calculate statistics such as the mean, variance, maximum value, minimum value, median, etc. within the sliding window, so as to extract the data feature values in each of the historical time-series data, including the mean, variance, maximum value, minimum value, and / or median; that is, the data change feature values of the time-series data at each timestamp, such as the data feature values of the unit power generation operation data in a certain time period.
[0144] The sliding window is used to process linear data such as arrays / strings, and examines or calculates the changes in the data within the window by moving the boundaries. The following are the basic steps and key points for constructing a sliding window:
[0145] Define the window:
[0146] According to the requirements of the problem, determine the initial size, starting position, and end condition of the window.
[0147] The window size can be fixed or dynamically changing.
[0148] Initialize variables:
[0149] Set the left and right boundary pointers of the window (usually initialized to 0 or a specific starting position).
[0150] Initialize variables for storing or calculating results (such as maximum value, minimum value, sum, frequency, etc.).
[0151] Window movement and update:
[0152] According to the requirements of the problem, traverse the data by moving the left and right boundaries of the window.
[0153] During the window movement, update the calculation results or status of the data within the window.
[0154] Note to maintain the validity and correctness of the data within the window.
[0155] Handle boundary conditions:
[0156] When the window moves to the boundary of the array / string, ensure the correctness of the processing logic.
[0157] Consider the influence of window size, starting position, and end conditions on boundary processing.
[0158] Output result:
[0159] According to the requirements of the problem, output the results obtained during the window movement (such as maximum / minimum values, sum, frequency statistics, etc.).
[0160] If necessary, record the starting and ending positions of the window for subsequent analysis or verification.
[0161] 3.3 Trend features: mainly identify and extract the trend change features between each of the time features, and use methods such as linear regression and polynomial fitting to extract the long-term trend and short-term fluctuations of the time series.
[0162] 3.4 Seasonal features: mainly identify and extract the seasonal changes between each of the time features. Methods such as Fourier transform and autocorrelation function can be used to capture the seasonal changes and periodic patterns in the time series.
[0163] In addition, the following can also be provided: Lag features: use historical data as features of the current time to capture the autocorrelation and lag effects of the time series; Rate of change features: calculate the first-order difference, second-order difference, etc. of the time series to capture the change speed and acceleration of the data.
[0164] The above-mentioned various feature processing methods can be processed by the administrator with reference to the corresponding technical means. This embodiment will not be elaborated.
[0165] The trend change features between the above-mentioned time features mainly analyze the change rules between the data feature values under each time feature, so as to display the change trend and seasonal change features of the data. The time series change features at each timestamp will be bound to the corresponding timestamp after feature extraction, which is convenient for the next step of feature engineering or feature recording.
[0166] Therefore, finally, count each of the time features, the data feature values, and the trend / seasonal features corresponding to each of the time features to form a feature set.
[0167] After feature extraction, feature selection and dimensionality reduction can also be performed:
[0168] Feature importance assessment: Use algorithms such as random forest and gradient boosting tree to evaluate the importance of each feature for trend analysis;
[0169] Correlation analysis: Calculate the correlation coefficients between features and remove redundant and highly correlated features;
[0170] Dimensionality reduction processing: Use methods such as PCA (Principal Component Analysis) and LDA (Linear Discriminant Analysis) to reduce the high-dimensional feature space to a low-dimensional space, improving the computational efficiency and generalization ability of the model.
[0171] 4. Model training
[0172] Divide the feature set into a training set and a validation set according to a preset ratio;
[0173] Input the training set into a preset time series analysis model for feature learning, and optimize and train the time series analysis model according to the preset model optimization iteration conditions;
[0174] Use the validation set to verify the performance of the time series analysis model:
[0175] If the verification passes, deploy and apply the time series analysis model;
[0176] Otherwise, repeat the above steps.
[0177] Model selection: Select a suitable time series analysis model according to the characteristics of the data and business requirements, such as ARIMA, LSTM, GRU, Prophet, XGBoost, etc.
[0178] Model training: Train the model using the training data set and adjust the model parameters as needed.
[0179] Cross-validation: Use cross-validation (such as K-fold cross-validation, the rolling window validation unique to time series) to evaluate the performance of the model. Evaluation metrics: Select suitable evaluation metrics, such as mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), etc. For traditional model performance verification methods, such as precision, recall, and F1 score: used for classification problems to evaluate the prediction accuracy of the model for positive classes, and can also be used for model verification. Specifically, the administrator can perform model verification.
[0180] Model optimization: According to the evaluation results, adjust steps such as model parameters, feature selection, and data preprocessing to improve the model performance.
[0181] Model deployment: Deploy the trained model to the production environment for real-time or batch prediction.
[0182] Performance monitoring: Continuously monitor the performance of the model, and promptly detect and handle problems such as model drift and data changes.
[0183] Preferably, step S3 of binding the timing change feature with the corresponding timestamp to generate a data change feature table corresponding to the task data packet includes:
[0184] Read the timing change features at each of the timestamps: the time feature, the data feature value, and the trend / seasonal feature corresponding to each of the time features;
[0185] Write the timing change features at each of the timestamps into a preset data fluctuation statistical table in sequence according to the time sequence, and generate the current data change feature table for the task data packet.
[0186] The data fluctuation statistical table, such as using a line chart or other change charts, records the data change features of the time points recorded at different timestamps, and imports the feature data to generate the corresponding data fluctuation change situation. The production method of this table can be:
[0187] 1. Determine statistical indicators
[0188] Select key data: clarify the data types and indicators to be statistically analyzed. The data change feature values, change trends, etc. at each time point / segment.
[0189] Set the time range: determine the statistical time period, such as daily, weekly, monthly, etc. Determine according to the time sequence of the timestamps.
[0190] 2. Collect data
[0191] Data source: collect data from a database, API, or manually. The timing change features at each of the timestamps statistically analyzed above.
[0192] 3. Calculate fluctuations
[0193] Fluctuation formula: select a suitable fluctuation calculation formula, such as standard deviation, variance, etc. Here, only the timing change features at each of the timestamps need to be imported.
[0194] Software tools: use tools such as Excel, Python, etc. for calculation.
[0195] 4. Make a statistical table
[0196] Table header design: include columns such as time, data indicators, fluctuation values, etc. The timing change features at each of the timestamps under each task respectively occupy a column in the table.
[0197] Data entry: enter the calculated fluctuation values into the table.
[0198] Chart display: line charts, bar charts, etc. can be added to visually display the fluctuations.
[0199] 5. Analysis and Interpretation
[0200] Reasons for fluctuations: Analyze the reasons and trends of data fluctuations.
[0201] Propose suggestions: Propose corresponding suggestions or measures based on the analysis results.
[0202] Therefore, the present invention can quickly analyze the change characteristics of task data based on time-series AI analysis. Compared with manual analysis, the present invention can greatly improve the analysis efficiency of time-series data, and can adapt to the online real-time analysis of massive time-series data, providing higher data support capabilities. The present invention is based on feature engineering, uses a time-series analysis model for feature training and learning, allows the model to perform training and learning on the extracted feature set, and allows the model to perform training and learning through the feature set of time-series change characteristics without sequential learning, so that the model obtains feature training and learning of a massive training set, which is convenient for improving the robustness and randomness of the model's feature recognition and the generalization ability of the model prediction.
[0203] Figure 3 It is a block diagram of a processing device for intelligent data analysis based on a time-series analysis model shown according to an exemplary embodiment. This device is used for a processing method of intelligent data analysis based on a time-series analysis model. Refer to Figure 3 This device includes:
[0204] A front end, used to collect and upload task data packets to the data analysis background in real time;
[0205] The data analysis background includes:
[0206] A data receiving port, used to receive the task data packets uploaded by the front end;
[0207] A parsing module, used to parse the task data packets and extract time-series data at several timestamps from the parsed data;
[0208] A time-series feature recognition module, used to traverse the time-series data at each timestamp, import the time-series data into a preset time-series analysis model, identify the time-series change characteristics in the time-series data through the time-series analysis model, and output them;
[0209] A time-series fluctuation recording module, used to bind the time-series change characteristics with the corresponding timestamps to generate a data change characteristic table corresponding to the task data packets;
[0210] A message center, used to feedback the data change characteristic table back to the front end along the original path;
[0211] A MySQL database, used to store data;
[0212] The front end is communicatively connected to the data analysis background.
[0213] For the interaction between the front end and the back end above, please understand it in combination with the above method steps. For each functional service module on the back end, it should also be understood with reference to the corresponding description. Details are not elaborated here.
[0214] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 5 shown, the electronic device 410 may include a first processor 2001.
[0215] Optionally, the electronic device 410 may further include a memory 2002 and a transceiver 2003.
[0216] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.
[0217] Next, in combination with Figure 5 specific components of the electronic device 410 will be introduced in detail:
[0218] Among them, the first processor 2001 is the control center of the electronic device 410, which may be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0219] Optionally, the first processor 2001 may execute various functions of the electronic device 410 by running or executing software programs stored in the memory 2002 and by calling data stored in the memory 2002.
[0220] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 5 the CPU0 and CPU1 shown in
[0221] In a specific implementation, as an embodiment, the electronic device 410 may also include multiple processors, such as Figure 5The first processor 2001 and the second processor 2004 shown in []. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0222] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0223] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit of the electronic device 410 ( Figure 5 not shown in []). The embodiments of the present invention do not make specific limitations on this.
[0224] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0225] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 5 not shown separately in []). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0226] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit of the electronic device 410 ( Figure 5 not shown in []). The embodiments of the present invention do not make specific limitations on this.
[0227] It should be noted thatFigure 5 The structure of the electronic device 410 shown does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have a different component arrangement.
[0228] In addition, for the technical effects of the electronic device 410, reference may be made to the technical effects of the data intelligent analysis processing method based on the timing analysis model described in the foregoing method embodiments, which will not be elaborated herein.
[0229] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0230] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0231] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0232] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0233] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0234] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0235] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0236] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0237] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0238] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0240] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0241] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A processing method for intelligent data analysis based on a time series analysis model, characterized in that: The method comprises: S1. Collect and parse the task data packets uploaded by the front end in real time, and extract time series data under several timestamps from the parsed data, including: Receiving the task data packet uploaded by the front end through the backend server interface; Identify the data format of the task data packet, and parse the task data packet according to the parsing method of the corresponding data format to obtain a plurality of time series data marked with a timestamp; Traversing and identifying the timestamps of each of the time series data, and judging whether the timestamps of each of the time series data are qualified based on the digital signature judgment rule pre-configured for the task data packet, comprises: Extracting keywords and judgment logic from the digital signature judgment rule pre-configured for the task data packet, wherein the keywords include timestamp type, version, format and digital signature identifier; Input the keywords and the judgment logic as prompt words into a preset LLM large language model; Traverse the timestamps of each of the time series data, and identify and judge whether the timestamps of each of the time series data are qualified based on the keywords and the judgment logic by the LLM large language model: If qualified, the time series data under the qualified timestamp is stored in the backend MySQL database in order; Otherwise, the front end is notified to correct the time series data under the unqualified timestamp; after the correction is qualified, it is stored in the backend MySQL database; S2, traversing the time series data at each time stamp, and importing the time series data into a preset time series analysis model, identifying the time series change characteristics in the time series data through the time series analysis model, and outputting the time series change characteristics; the method for constructing the time series analysis model includes: Collect and parse several historical task data packets to construct a historical time series data set consisting of several historical time series data; Traversing the historical time series data set, finding and removing the historical time series data without timestamps, and sorting the remaining historical time series data according to timestamps to reconstruct the historical time series data set; Perform feature engineering on the historical time series data set to extract time series change features, including: Extracting time features from the timestamps of each of the historical time series data; Utilizing a preset sliding window, extracting data characteristic values including mean, variance, maximum value, minimum value and / or median from each of the historical time series data; Identify and extract trend / seasonal features between each of the time features; Counting each of the time features, the data feature values, and the trend / seasonal features corresponding to each of the time features to form a feature set; Dividing the feature set into a training set and a validation set according to a preset ratio; Inputting the training set into a preset time series analysis model to perform feature learning, and optimizing the time series analysis model according to preset model optimization iteration conditions; Use the verification set to perform performance verification on the timing analysis model: If the verification passes, the timing analysis model is deployed and applied; Otherwise, repeat the above steps; S3, binding the time series change feature with the corresponding timestamp to generate a data change feature table corresponding to the task data packet, including: Read the time series variation characteristics under each of the timestamps: the time characteristics, the data characteristic value, and the trend / seasonality characteristics corresponding to each of the time characteristics; According to the time sequence, the time sequence change characteristics under each time stamp are sequentially written into a preset data fluctuation statistical table, and a data change characteristic table of this time is generated for the task data packet; S4, feeding back the data change characteristic table to the front end in the original path.
2. The method for performing intelligent data analysis based on a time series analysis model according to claim 1, characterized in that: The traversing the historical time series data set to find and remove the historical time series data without a timestamp includes: Traversing the timestamps of each of the historical time series data in the historical time series data set, and identifying the historical time series data in the historical time series data set that does not have the keyword based on the keyword through the LLM large language model; The historical time series data that is traversed and identified and does not have the keyword is removed from the historical time series data set.
3. A processing device for performing intelligent data analysis based on a time series analysis model, wherein the processing device for performing intelligent data analysis based on a time series analysis model is used to implement the processing method for performing intelligent data analysis based on a time series analysis model as described in any one of claims 1-2, characterized in that: The device comprises: The front end is used to collect and upload task data packets to the data analysis backend in real time; Data analysis backend, including: Data receiving port, used to receive task data packets uploaded by the front end; A parsing module, used for parsing the task data packet and extracting time series data under a plurality of timestamps from the parsed data; A time series feature recognition module, used to traverse the time series data at each time stamp, import the time series data into a preset time series analysis model, identify the time series change features in the time series data through the time series analysis model, and output; A timing fluctuation recording module, used for binding the timing variation characteristics with the corresponding timestamps to generate a data variation characteristic table corresponding to the task data packet; The message center is used to feed back the data change characteristic table to the front end in the original path; MySQL database, used to store data; The front end is communicatively connected with the data analysis back end.
4. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 2 is implemented.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 2.
Citation Information
Patent Citations
Load reduction method and system for real-time streaming data predictive analysis
CN113535527A
Data acquisition and analysis method for AI recognition of facial expressions
CN116597497A