Data preprocessing method and processing terminal based on time sequence in data center
By adopting a time series-based data preprocessing method in the data in Taiwan and using the OSRCNN network model to extract and fuse feature data, the problems of data missing and outliers in the power system are solved, and the data quality and accuracy of power prediction are improved.
Patent Information
- Application Number
- CN202510091323.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, there are data missing and outliers in the original data in the power system, resulting in poor data quality and affecting the accuracy of power prediction.
In the data, the data preprocessing method based on time series is adopted in the Taiwan data. The initial data matrix is formed by obtaining the to-processed data, different sample matrices are extracted and inputted to the OSRCNN network model, extract and fuse feature data, and generate the preprocessed target data matrix.
By extracting and fusion of rich feature data, the data quality is improved, making the power prediction more accurate and the prediction results more accurate.
Smart Images

Figure CN119988835A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data preprocessing method and a processing terminal based on time series in a data center. Background Art
[0002] To achieve safe and stable operation of the power system, it is necessary to maintain a real-time dynamic balance in all aspects of power generation, transmission and consumption. In order to provide better power supply services to users in low-voltage areas, understand user needs more systematically and improve the power supply quality of the system, the data center needs to effectively predict the power generation and power consumption of the area. Therefore, it is necessary to collect power data for prediction.
[0003] In the prior art, due to the presence of a certain amount of missing data and outliers in the original data, the data quality is poor and the accuracy of the prediction is affected. Therefore, it is necessary to provide a data preprocessing method to preprocess the original power data in the data center to improve the accuracy of the prediction. Summary of the invention
[0004] The embodiment of the present invention provides a data preprocessing method and a processing terminal based on time series in a data center to solve the problems of lack of data preprocessing method, poor data quality and low accuracy of power prediction.
[0005] In a first aspect, an embodiment of the present invention provides a data preprocessing method based on time series in a data center, comprising:
[0006] Obtain the data to be processed and form an initial data matrix;
[0007] Extracting a first sample matrix and a second sample matrix from an initial data matrix; wherein the first sample matrix and the second sample matrix are different;
[0008] The first sample matrix is input into the OSRCNN network model to obtain a first feature data matrix; and the second sample matrix is input into the OSRCNN network model to obtain a second feature data matrix;
[0009] A preprocessed target data matrix is obtained according to the first characteristic data matrix and the second characteristic data matrix.
[0010] In a second aspect, an embodiment of the present invention provides a processing terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor implements the steps of a time series-based data preprocessing method in a data center provided in the first aspect or any possible implementation of the first aspect.
[0011] The embodiment of the present invention provides a data preprocessing method and processing terminal based on time series in a data center. The data preprocessing method based on time series in the above-mentioned data center includes: obtaining the data to be processed and forming an initial data matrix; extracting the first sample matrix and the second sample matrix of the initial data matrix; wherein the first sample matrix and the second sample matrix are different; inputting the first sample matrix into the OSRCNN network model to obtain the first characteristic data matrix; and inputting the second sample matrix into the OSRCNN network model to obtain the second characteristic data matrix; according to the first characteristic data matrix and the second characteristic data matrix, a preprocessed target data matrix is obtained. The embodiment of the present invention extracts and fuses data features from two dimensions based on the OSRCNN network model, so that the features are richer, the target data matrix is closer to the real data, and the data quality is effectively improved. When it is applied to power forecasting, the prediction accuracy is higher and the prediction results are more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0013] Figure 1 It is a flow chart of an implementation of a data preprocessing method based on time series in a data center provided by an embodiment of the present invention;
[0014] Figure 2 is an input image provided by an embodiment of the present invention;
[0015] Figure 3 yes Figure 2 Image after 3*3 convolution;
[0016] Figure 4 yes Figure 2 Images after 3*3, 5*5 and 7*7 convolutions;
[0017] Figure 5 is an image before mapping provided by an embodiment of the present invention;
[0018] Figure 6 yes Figure 5 An image obtained by mapping using the prior art;
[0019] Figure 7 yes Figure 5 An image obtained by mapping using an embodiment of the present invention;
[0020] Figure 8It is a structural schematic diagram of a data preprocessing device based on time series in a data center provided by an embodiment of the present invention;
[0021] Fig. 9 It is a schematic diagram of a processing terminal provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.
[0023] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below in conjunction with the accompanying drawings.
[0024] See also Figure 1 , which shows a flow chart of a data preprocessing method based on time series in a data center provided by an embodiment of the present invention, and is described in detail as follows:
[0025] The data preprocessing method based on time series in the above data includes:
[0026] S101: Acquire data to be processed and form an initial data matrix;
[0027] The data to be processed is data based on time series. For example, it can be historical power data of M days. The number of complete power data points per day can be N. Specifically, if the power data acquisition cycle is 1 minute, the complete power data per day should have 1440 points, and N is 1440.
[0028] Since the embodiment of the present invention is based on the concept of image features and constructs an OSRCNN model, in order to facilitate subsequent data processing, it needs to be simply processed to form an initial data matrix.
[0029] In a possible implementation, S101 may include:
[0030] S1011: Forming an original data matrix from the data to be processed in chronological order;
[0031] S1012: Normalize the original data matrix to obtain an initial data matrix.
[0032] The data to be processed are formed into an initial data matrix in chronological order, as follows:
[0033]
[0034] Among them, d1t1 is the first power data of the first day, d M t N It is the Nth power data of the Mth day.
[0035] Different data to be processed may have different dimensions and orders of magnitude, which may affect subsequent data analysis and model building. The purpose of normalization is to eliminate these differences, make various types of data to be processed have the same importance and comparability, and avoid deviations caused by different dimensions.
[0036] Specifically, the Min-Max normalization method can be used to linearly map each data in the original data matrix to the interval [0,1]. Specifically, the calculation formula can be:
[0037]
[0038] Among them, Z ij is the normalized data, X ij is the data in the i-th row and j-th column of the original data matrix, X avgi is the average value of the data in the i-th row, X imax is the maximum value of the data in the i-th row, X imin is the minimum value of the i-th row of data.
[0039] Based on the above steps, each data in the original data matrix can be converted into a form that is more suitable for data analysis and modeling, providing a reliable and accurate data basis for subsequent feature extraction.
[0040] S102: extracting a first sample matrix and a second sample matrix from an initial data matrix; wherein the first sample matrix and the second sample matrix are different;
[0041] The embodiment of the present invention extracts features from two dimensions. The first sample matrix and the second sample matrix are different, reflecting features of different dimensions.
[0042] In a possible implementation, S102 may include:
[0043] S1021: Perform 3*3 and 5*5 convolutions on the initial data matrix in sequence to obtain a first sample matrix;
[0044] S1022: Perform 3*3, 5*5 and 7*7 convolutions on the initial data matrix in sequence to obtain a second sample matrix.
[0045] Convolution operations are often used for feature extraction, and convolution kernels of different sizes can capture features of different scales. Small convolution kernels such as 3*3 can capture details and local features in the image, while large convolution kernels such as 5*5 and 7*7 can capture more macroscopic and global features. By using convolution kernels of different sizes for convolution operations in sequence, rich feature information can be extracted, thereby better understanding and processing data.
[0046] Therefore, in the embodiment of the present invention, 3*3 and 5*5 convolutions are performed on the initial data matrix to extract features of one scale to better retain local features. At the same time, 3*3, 5*5 and 7*7 convolutions are performed on the initial matrix to extract features of another scale, retaining more global information, enhancing the capture of global patterns, including more adjacent data information, and helping to reduce sensitivity to local noise.
[0047] For example, Figure 2 is the input image, Figure 3 is the image after 3*3 convolution. Figure 4 is the image after 3*3, 5*5 and 7*7 convolution. Figure 2 , Figure 3 and Figure 4 It can be seen that the images after 3*3, 5*5 and 7*7 convolution can obtain more detailed features and edge information.
[0048] At the same time, convolution can also reduce the data dimension of the initial data matrix, reduce the amount of calculation in the sample operation process, and reduce accidental and random errors caused by acquisition failures.
[0049] The convolution operation requires a convolution kernel and a data matrix. The convolution kernel is a small matrix, usually with odd rows and columns, such as 3*3, 5*5, 7*7, etc., and its element values are set through training or manually. The data matrix is the initial data that needs to be processed, usually image data or other two-dimensional data, etc., here refers to the initial data matrix.
[0050] Based on the application scenario of the present application, in the embodiment of the present invention, 3*3 and 5*5 convolutions are used to obtain the first sample matrix, and 3*3, 5*5 and 7*7 convolutions are used to obtain the second sample matrix, which can better extract sample features.
[0051] More specifically, when the data to be processed is power generation data, the 3*3 convolution kernel can be:
[0052]
[0053] The 5*5 convolution kernel can be:
[0054]
[0055] The 7*7 convolution kernel can be:
[0056]
[0057] When the data to be processed is electricity consumption data, the 3*3 convolution kernel can be:
[0058]
[0059] The 5*5 convolution kernel can be:
[0060]
[0061] The 7*7 convolution kernel can be:
[0062]
[0063] The convolution kernel can be determined according to the actual application scenario, including but not limited to the above two.
[0064] Convolution process: Slide the convolution kernel on the data matrix. Each time it slides to a position, multiply the convolution kernel with the element at the corresponding position of the data matrix, and then add these products to get a result value, which is an element in the new matrix obtained after convolution. By continuously sliding the convolution kernel, the entire convolution matrix can be obtained.
[0065] Exemplarily, the step of S1021 may be: use a 3*3 convolution kernel to perform a convolution operation on the initial data matrix. The 3*3 convolution kernel is slid from the upper left corner on the initial data matrix to the right and downward in sequence, one pixel position at a time, and the sum of the products of the convolution kernel and the corresponding data matrix area is calculated to obtain a new matrix, which is recorded as matrix A.
[0066] Then, use the 5*5 convolution kernel to perform convolution operation on matrix A. Similarly, slide the 5*5 convolution kernel on matrix A to calculate another new matrix, which is the first sample matrix.
[0067] The steps of S1022 are the same as above and will not be described in detail here.
[0068] It should be noted that when the convolution size is not appropriate, padding operation is usually required. The size of the matrix after convolution is:
[0069]
[0070] Where I is the number of rows of the initial data matrix. For example, if the initial data matrix is a 30*1440 matrix, then I is 30. F is the size of the convolution kernel, which is 3 and 5 respectively. Pad is the number of padded pixels. If N is 1440, the value of Pad is 0. If Pad is not 0, padding is required.
[0071] The specific filling method is as follows:
[0072] Define the boundary filling similarity coefficient Sim Clif , coefficient upper and lower thresholds Sim upthr 、Sim downthr Cohort Sim with Coefficients vect If a column of data is missing, a column in the initial data matrix corresponding to the time of the column to be filled is used as a benchmark to calculate the vector cosine similarity between this column and other columns, that is, the boundary filling similarity coefficient. The boundary filling similarity coefficients in the range of 0.6 to 0.9 are sorted, and the column with the highest boundary filling similarity coefficient is selected as the filling sequence for filling.
[0073] For example, if a column to be filled is the data of 24:15min, and 24:15min should actually be 0:15min of the next day, then the boundary filling similarity coefficients of the 0:15min column and other columns in the initial data matrix are calculated to obtain the filling sequence.
[0074] By using the above method to fill the matrix, the filled data is more in line with reality. At the same time, the filled data has certain differences from the actual data, which increases the data diversity.
[0075] S103: inputting the first sample matrix into the OSRCNN network model to obtain a first feature data matrix; and inputting the second sample matrix into the OSRCNN network model to obtain a second feature data matrix;
[0076] OSRCNN utilizes representations transferred from pre-trained models on large-scale object and scene recognition datasets for event recognition, mainly for event recognition in images, by viewing OS-CNNs as “end-to-end event predictors” or “universal feature extractors”.
[0077] In order to further increase the feature information in the low-resolution image (i.e., the first sample matrix and the second sample matrix) generated by convolution, in an embodiment of the present invention, an OSRCNN network model is constructed based on the concept of image features to enhance the features and obtain richer detail features.
[0078] The specific principle of the OSRCNN network model is:
[0079] 1) Feature extraction: Extract patches from the low-resolution image Y. Each patch is a high-dimensional vector. These vectors form a feature map whose size is equal to the dimension of these vectors. The formula is defined as follows:
[0080] F1(Y)=max(0,W1*Y+B1)
[0081] Among them, W1 and B1 are the weight matrix and bias respectively, the size of W1 is c*f1*f1*n1, c is the number of channels of the input image, f1 is the spatial size of the weight matrix, and n1 is the number of matrices. Intuitively, W1 uses n1 convolutions, and the size of each convolution kernel is c*f1*f1. The output is n1 feature maps. B1 is an n1-dimensional vector, each element is related to a matrix, and the Rectified Linear Unit (ReLU, max(0,x)) is used in the embodiment of the present invention.
[0082] 2) Nonlinear mapping: Map a high-dimensional vector to another high-dimensional vector. Each mapped vector represents a high-resolution patch, and these vectors form another feature map. The formula is defined as follows:
[0083] F2(Y)=max(0,W2*F1(Y)+B2)
[0084] Among them, the size of W2 is n1*1*1n2, B2 is an n2-dimensional vector, and each output n2-dimensional vector represents a high-resolution patch for subsequent reconstruction.
[0085] 3) Reconstruction: All high-resolution patches are combined to form the final high-resolution image. Ideally, the output image is the same as the real high-resolution image. The formula is as follows:
[0086] F(Y)=W3*F2(Y)+B3*W3
[0087] Among them, the size of W3 is n2*f3*f3*c, n2 is the number of filters, f3 is the filter size, and B3 is a c-dimensional vector.
[0088] In the embodiment of the present invention, the spatial size of the weight matrix f1=3 is defined, and the detailed texture features of the input image can be obtained. In order to add more detailed features, we add 5*5 and 7*7 convolution kernels (step S102) on the basis of f1=3 for convolution, which can not only obtain more detailed features, but also extract richer edge information.
[0089] In the prior art, simple average concatenation is usually used in mapping, and the formula is as follows:
[0090]
[0091] Among them, M a and M b These are two high-dimensional vectors that need to be concatenated in the linear mapping.
[0092] In the embodiment of the present invention, the coverage ratio C0 defined by two adjacent mappings represents the repetition rate of the high-dimensional vector obtained after two adjacent mappings, which replaces the simple splicing in the original mapping, so that the features at the beginning and end of the mapped high-dimensional vector are strengthened, and the features will not be weakened due to averaging in the reconstruction in the next step. The specific mapping formula is:
[0093]
[0094] Among them, R' e is the mapped data, M a and M b are two high-dimensional vectors that need to be concatenated, and C0 is the coverage.
[0095] Figure 5 is the image before mapping, Figure 6 For an image obtained by mapping using the prior art, Figure 7 is an image obtained by mapping using an embodiment of the present invention. Figure 5 , Figure 6 and Figure 7 It can be seen that the image mapped by the embodiment of the present invention has more prominent details and richer detail features.
[0096] In a possible implementation, the coverage ratio may be 0.45.
[0097] The coverage rate C0 ranges from 0 to 1, and the coverage rate Co is defined with a step size of 0.1. After nonlinear mapping experiments, it is found that the SSIM is optimal when the C0 value is between 0.3 and 0.6.
[0098] Table 1 Comparison table of coverage 0-1 and SSIM
[0099] Co 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 SSIM(dB) 34.85 36.23 38.01 39.05 39.03 37.73 35.95 35.82 34.24 33.78
[0100] Furthermore, the coverage rate C0 is narrowed to 0.3-0.6, and the coverage rate is defined with a step size of 0.05. After experiments, the final output image effect is normally distributed. According to the 3σ principle of normal distribution, the SSIM effect obtained when the C0 value is between 0.35 and 0.55 is relatively good. Therefore, we define C0 as 0.45, which is the best effect.
[0101] Table 2 Comparison table of coverage 0.3-0.6 and SSIM
[0102] Co 0.3 0.35 0.40 0.45 0.50 0.55 0.60 SSIM(dB) 38.01 38.41 39.05 40 39.03 38.65 37.73
[0103] S104: Obtain a preprocessed target data matrix according to the first characteristic data matrix and the second characteristic data matrix.
[0104] In a possible implementation, S104 may include:
[0105] S1041: Perform 7*7 inverse convolution on the second feature data matrix to obtain a third reconstructed data matrix.
[0106] Since the convolution kernels of the two convolutions in the above feature extraction process are different, the sizes of the first feature data matrix and the second feature data matrix are different. Therefore, in an embodiment of the present invention, a 7*7 inverse convolution is performed on the second feature data matrix to unify the sizes of the two matrices.
[0107] S1042: performing weighted summation of the first characteristic data matrix and the third reconstruction data matrix to obtain a comprehensive data matrix;
[0108] The first feature data matrix and the third reconstruction data matrix are weightedly summed to perform feature fusion, and the global features, details and texture features are integrated to obtain richer features.
[0109] S1043: Obtain a target data matrix based on the comprehensive data matrix.
[0110] In a possible implementation, the calculation formula of the comprehensive data matrix P" may be:
[0111] P”=W1*P1”+W2*P2”
[0112] Wherein, W1 is the first weight, W2 is the second weight, P1" is the first characteristic data matrix, and P2" is the third reconstruction data matrix;
[0113] The value range of the first weight is 0.6-0.9, and the value range of the second weight is 0.1-0.4.
[0114] Considering that the embodiment of the present invention considers details and texture features more and pays less attention to capturing global features, the value range of the first weight is defined as 0.6-0.9, and the value range of the second weight is defined as 0.1-0.4.
[0115] More specifically, after model training and testing, the first weight may be 0.83 and the second weight may be 0.17.
[0116] In a possible implementation, S1043 may include:
[0117] 1. Perform 3*3 and 5*5 deconvolution on the comprehensive data matrix to obtain the target data matrix.
[0118] Corresponding to S1021~S1022, 3*3 and 5*5 deconvolutions are performed on the comprehensive data matrix to map the comprehensive data matrix back to a relatively high-dimensional space to obtain the target data matrix.
[0119] In a possible implementation manner, S1043 may further include:
[0120] 2. Perform inverse normalization on the matrix after deconvolution to obtain the target data matrix.
[0121] Corresponding to the normalization in S1012, the embodiment of the present invention needs to be inversely normalized to restore the original data format to obtain a target data matrix. The format of the target data matrix is the same as the initial data matrix, which is a matrix based on a time series. The data in the target data matrix are arranged in chronological order to obtain preprocessed data.
[0122] In the embodiment of the present invention, different convolution kernels are used to extract the global features and detail features of the original data matrix respectively, and the OSRCNN model is used to enrich the feature information, and then the features are fused, so that the fused features are richer, while retaining the local features, the capture of the global features will not be weakened, and the target data matrix is closer to the original data, which effectively improves the data quality. When the data is applied to power forecasting, the prediction accuracy is higher and the prediction results are more accurate.
[0123] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0124] The following is an embodiment of the device of the present invention. For details not described in detail, reference may be made to the corresponding method embodiments described above.
[0125] Figure 8 The schematic diagram of the structure of the data preprocessing device based on time series in the data center provided by the embodiment of the present invention is shown. For the convenience of explanation, only the part related to the embodiment of the present invention is shown, which is described in detail as follows:
[0126] like Figure 8 As shown, the data preprocessing device based on time series in the data center includes:
[0127] The matrix building module 21 is used to obtain the data to be processed and form an initial data matrix;
[0128] A feature extraction module 22, used to extract a first sample matrix and a second sample matrix from an initial data matrix; wherein the first sample matrix and the second sample matrix are different;
[0129] The feature enhancement module 23 is used to input the first sample matrix into the OSRCNN network model to obtain a first feature data matrix; and input the second sample matrix into the OSRCNN network model to obtain a second feature data matrix;
[0130] The feature fusion module 24 is used to obtain a preprocessed target data matrix according to the first feature data matrix and the second feature data matrix.
[0131] In a possible implementation, the feature extraction module 22 may include:
[0132] A first convolution unit is used to perform 3*3 and 5*5 convolutions on the initial data matrix in sequence to obtain a first sample matrix;
[0133] The second convolution unit is used to perform 3*3, 5*5 and 7*7 convolutions on the initial data matrix in sequence to obtain a second sample matrix.
[0134] In a possible implementation, the feature fusion module 24 may include:
[0135] A deconvolution unit, used for performing a 7*7 deconvolution on the second feature data matrix to obtain a third reconstructed data matrix;
[0136] A weighting unit, used for weighted summing the first characteristic data matrix and the third reconstruction data matrix to obtain a comprehensive data matrix;
[0137] The target data output unit is used to obtain the target data matrix according to the comprehensive data matrix.
[0138] In a possible implementation, the calculation formula of the comprehensive data matrix P" may be:
[0139] P”=W1*P1”+W2*P2”
[0140] Wherein, W1 is the first weight, W2 is the second weight, P1" is the first characteristic data matrix, and P2" is the third reconstruction data matrix;
[0141] The value range of the first weight may be 0.6-0.9, and the value range of the second weight may be 0.1-0.4.
[0142] In a possible implementation, the target data output unit may be specifically used to: perform 3*3 and 5*5 deconvolution on the comprehensive data matrix to obtain the target data matrix.
[0143] In a possible implementation, the matrix establishing module 21 may include:
[0144] A matrix forming unit, used for forming an original data matrix from the data to be processed in time sequence;
[0145] The normalization unit is used to normalize the original data matrix to obtain an initial data matrix.
[0146] In a possible implementation, the target data output unit may also be used to: perform inverse normalization on the matrix after deconvolution to obtain a target data matrix.
[0147] In a possible implementation, the mapping of the OSRCNN network model may be:
[0148]
[0149] Among them, R' e is the mapped data, M a and M b are two high-dimensional vectors that need to be concatenated, and C0 is the coverage.
[0150] In a possible implementation, the coverage ratio may be 0.45.
[0151] Fig. 9 Schematic diagram of the processing terminal 3 provided in the embodiment of the present invention. Fig. 9 As shown, the processing terminal 3 of this embodiment includes: a processor 30 and a memory 31. The memory 31 is used to store a computer program 32, and the processor 30 is used to call and run the computer program 32 stored in the memory 31 to execute the steps in the above-mentioned data preprocessing method based on time series in each data center, such as Figure 1 Alternatively, the processor 30 is used to call and run the computer program 32 stored in the memory 31 to implement the functions of each module / unit in the above-mentioned device embodiments, such as Figure 8 The functions of modules 21 to 24 are shown.
[0152] Exemplarily, the computer program 32 may be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the processing terminal 3. For example, the computer program 32 may be divided into Figure 8 Modules / units 21 to 24 are shown.
[0153] The processing terminal 3 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The processing terminal 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will appreciate that Fig. 9It is only an example of the processing terminal 3 and does not constitute a limitation on the processing terminal 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal may also include input and output devices, network access devices, buses, etc.
[0154] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0155] The memory 31 may be an internal storage unit of the processing terminal 3, such as a hard disk or memory of the processing terminal 3. The memory 31 may also be an external storage device of the processing terminal 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the processing terminal 3. Further, the memory 31 may also include both an internal storage unit of the processing terminal 3 and an external storage device. The memory 31 is used to store computer programs and other programs and data required by the terminal. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0156] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0157] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0158] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0159] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0160] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0161] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0162] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0163] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.
Claims
1. A data preprocessing method based on time series in data center, characterized in that: include: Obtain the data to be processed and form an initial data matrix; Extracting a first sample matrix and a second sample matrix of the initial data matrix; wherein the first sample matrix and the second sample matrix are different; Inputting the first sample matrix into the OSRCNN network model to obtain a first feature data matrix; and inputting the second sample matrix into the OSRCNN network model to obtain a second feature data matrix; A preprocessed target data matrix is obtained according to the first characteristic data matrix and the second characteristic data matrix.
2. The data preprocessing method based on time series in the data center according to claim 1 is characterized in that: The step of extracting a first sample matrix and a second sample matrix from the initial data matrix comprises: Performing 3*3 and 5*5 convolutions on the initial data matrix in sequence to obtain the first sample matrix; The initial data matrix is sequentially convolved by 3*3, 5*5 and 7*7 to obtain the second sample matrix.
3. The data preprocessing method based on time series in the data center according to claim 2 is characterized in that: The step of obtaining preprocessed target data according to the first feature data matrix and the second feature data matrix comprises: Performing a 7*7 inverse convolution on the second feature data matrix to obtain a third reconstructed data matrix; Performing weighted summation of the first characteristic data matrix and the third reconstruction data matrix to obtain a comprehensive data matrix; The target data matrix is obtained according to the comprehensive data matrix.
4. The data preprocessing method based on time series in the data center according to claim 3 is characterized in that: The calculation formula of the comprehensive data matrix P" is: P”=W1*P1”+W2*P2” Wherein, W1 is the first weight, W2 is the second weight, P1" is the first feature data matrix, and P2" is the third reconstruction data matrix; The value range of the first weight is 0.6-0.9, and the value range of the second weight is 0.1-0.
4.
5. The data preprocessing method based on time series in the data center according to claim 3 is characterized in that: The step of obtaining the target data matrix according to the comprehensive data matrix comprises: Perform 3*3 and 5*5 deconvolution on the comprehensive data matrix to obtain the target data matrix.
6. The data preprocessing method based on time series in the data center according to claim 5 is characterized in that: The step of obtaining the data to be processed and forming an initial data matrix includes: The data to be processed are formed into an original data matrix in chronological order; The original data matrix is normalized to obtain the initial data matrix.
7. The data preprocessing method based on time series in the data center according to claim 6 is characterized in that: After performing 3*3 and 5*5 deconvolution on the comprehensive data matrix, obtaining the target data matrix according to the comprehensive data matrix further includes: The matrix after deconvolution is inversely normalized to obtain the target data matrix.
8. The data preprocessing method based on time series in the data center according to any one of claims 1 to 7, characterized in that: The mapping of the OSRCNN network model is: Among them, R' e is the mapped data, M a and M b are two high-dimensional vectors that need to be concatenated, and C0 is the coverage.
9. The data preprocessing method based on time series in the data center according to claim 8 is characterized in that: The coverage value is 0.
45.
10. A processing terminal, characterized in that: It includes a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the time series-based data preprocessing method in the data center as described in any one of claims 1 to 9.