Noise data ShardingJDBC dynamic library table extension system based on time sequence
Through dynamic database table planning, noise data separation, and intelligent query and feature mining modules, the expansion and query efficiency problems of traditional databases when processing massive noise data are solved, efficient management of sensor data and accurate distinction of noise data are achieved, and the accuracy of data utilization and query efficiency are improved.
Patent Information
- Application Number
- CN202510932326.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional relational databases face difficulties in horizontal expansion, degraded write performance, and high storage costs when processing massive, high-frequency time series data; NoSQL databases are insufficient in complex queries and consistency assurance; dedicated time series databases have limited support for classified storage of noisy data and high integration costs; ShardingJDBC needs further optimization to cope with the dynamic growth of noise data and differentiated storage.
The dynamic library table planning module is used to establish the mapping relationship between sensors and databases. The wavelet analysis method is used to filter noise data and classify and store them. The intelligent query and feature mining module analyzes user intentions and distinguishes between hot and cold valid data. The principal component analysis algorithm is used to extract the main characteristic components of pseudo-noise data.
It realizes independent management and query of sensor data, accurately distinguishes noise and valid data, optimizes storage resource utilization, improves query efficiency and data utilization accuracy, and adapts to high-frequency writing and complex query needs.
Smart Images

Figure CN120763159A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of data processing, in particular to a time series-based noise data ShardingJDBC dynamic library table extension system. BACKGROUND
[0002] In the fields of industrial Internet of Things, financial transactions, environmental monitoring, etc., time series data presents the characteristics of mass, high frequency and continuous growth, and is often mixed with noise data generated by factors such as sensor hardware errors, environmental electromagnetic interference and transmission abnormalities;
[0003] Traditional relational databases have problems such as horizontal expansion difficulty, significant decline in write performance with data volume growth, high storage cost, etc., and are difficult to meet the high-frequency write and complex query requirements; although some NoSQL databases have high-throughput write capability, they have deficiencies in complex query, data consistency guarantee and ecological perfection; although special time series databases optimize time dimension query, they have limited support for noise data and effective data classification storage, and have high integration cost with existing enterprise systems;
[0004] In addition, the feature boundary of noise data and effective data is fuzzy, and simple filtering methods cannot accurately distinguish them, and noise data is not completely worthless, so the storage strategy needs to balance resource utilization and potential analysis needs. Although ShardingJDBC, as a distributed database middleware, has advantages in library table management and read-write separation, it still needs further optimization and expansion in the face of dynamic growth of time series noise data, differentiated storage and complex query scenarios. Therefore, we propose a time series-based noise data ShardingJDBC dynamic library table extension system. SUMMARY
[0005] The purpose of the present application is to solve the problems of traditional relational databases, such as horizontal expansion difficulty, significant decline in write performance with data volume growth, and difficulty in meeting high-frequency write and complex query requirements; the deficiencies of some NoSQL databases in complex query, consistency guarantee and ecological perfection; the limited support of special time series databases for noise data and effective data classification storage, and the high integration cost; the fuzzy feature boundary of noise data, which makes it difficult to accurately distinguish them by simple filtering methods, and the need to balance the storage strategy; and the need to further optimize ShardingJDBC in the face of noise data dynamic growth, differentiated storage and other scenarios.
[0006] To achieve the above purpose, the present application provides a time series-based noise data ShardingJDBC dynamic library table extension system, which comprises a dynamic library table planning module, a noise data and effective data separation module and an intelligent query and feature mining module, wherein:
[0007] The dynamic library table planning module establishes the mapping relationship between each database and the sensor, and dynamically creates a library table; the noise data effective data separation module filters the noise data in the time series data by using a wavelet analysis method, the time series data after filtering the noise data is effective data, and the distinguished wavelet coefficients are reconstructed, when the dynamic library table planning module classifies and stores the time series data, the effective data and the noise data are respectively stored in independent databases, and a double-database mapping association is constructed;
[0008] When a user queries the database, the intelligent query and feature mining module parses the user intent by using a natural language processing method, and counts the access frequency of each user to the effective data to distinguish the effective data as hot and cold effective data; if the user queries the effective data, the noise data is synchronously queried, the synchronous query of the noise data is defined as pseudo noise data, the user queries the effective data again, the pseudo noise data is synchronously called out for the user to view, the main characteristic components of the pseudo noise data are extracted by using a principal component analysis algorithm, and the main characteristic components are used as the common characteristics of the pseudo noise data; for the noise data that is not actively accessed by the user, the pseudo noise data is pre-judged through the common characteristics.
[0009] The dynamic library table planning module perceives the time series data detected by different sensors, creates a database corresponding to each sensor, and allocates a unique identifier to each sensor and the corresponding database, so as to establish the mapping relationship between each database and the sensor.
[0010] As a further improvement of the technical solution, the dynamic library table planning module dynamically creates a library table again according to the time series of the time series data: defining a time granularity, dividing the time series data into non-overlapping intervals according to the time stamp in units of the time granularity.
[0011] The beneficial effects of the above further solution are that the dynamic library table planning module independently builds a library for each sensor and maps a unique identifier, so that "one type of sensor data is dedicated to one library", the data logic of different sensors is isolated from the source, and management and query confusion caused by mixed data is avoided; based on the time granularity, a library table is dynamically created, and data is stored in non-overlapping intervals according to the time stamp, which on the one hand adapts to the characteristics of time series data "continuously generated over time and strong time sequence", so that the data is naturally sorted by time, and the problem of slow query after the data volume of a single table in the traditional database is expanded is solved.
[0012] On the basis of the above technical solution, the present application can also be improved as follows.
[0013] As a further improvement of the technical solution, the wavelet analysis method in the noise data effective data separation module includes a group of wavelet basis function decomposition time series data: the shape and position of the wavelet basis function are controlled through a scale factor and a translation factor;
[0014] The time series data is compared with the wavelet basis function point by point, and the sum is calculated to calculate the fit between the wavelet basis function and the time series data at different scales and positions to obtain the wavelet coefficients. The scale factor and translation factor are adjusted to split the time series data into a series of wavelet coefficients.
[0015] The beneficial effect of the above further solution is that the wavelet analysis method of the noise data effective data separation module, with the help of a set of wavelet basis functions, flexibly controls the "stretching" of the basis functions with the scale factor and accurately adjusts their "sliding" on the time axis with the translation factor, thus achieving fine control over the shape and position of the basis functions. By comparing the time series data with the wavelet basis functions point by point and performing summation operations, the degree of fit between the two at different scales and positions can be accurately calculated to obtain the wavelet coefficients, and then a series of coefficients can be split out by adjusting the factors.
[0016] On the one hand, this process adapts to the non-stationary and noisy characteristics of time series data, overcomes the limitation of traditional Fourier transform that is only suitable for stationary signals, retains local characteristics in the time domain, and can perform multi-scale analysis in the frequency domain; on the other hand, it lays a solid foundation for the subsequent accurate distinction between noise data and valid data. The wavelet coefficients can clearly define data with different frequency components, making noise data filtering more accurate, helping the system to efficiently mine effective information and optimize data quality in massive and complex time series data scenarios, making up for the shortcomings of traditional databases and tools in the fine processing of noise data, and promoting the upgrade of the whole process efficiency from data storage to in-depth analysis.
[0017] On the basis of the above technical solution, the present invention can also be improved as follows.
[0018] As a further improvement of the present technical solution, the wavelet analysis method in the noise data and valid data separation module sets a noise data threshold to distinguish noise data from valid data:
[0019] If the absolute value of the wavelet coefficient is within the noise data threshold, the corresponding wavelet coefficient is judged to be a valid data wavelet coefficient; otherwise, it is a noise data wavelet coefficient.
[0020] As a further improvement of the present technical solution, the noise data and effective data separation module reconstructs the distinguished wavelet coefficients: by traversing all scales and translating, the retained effective wavelet coefficients and the eliminated noise data wavelet coefficients are respectively weighted integrated with the corresponding wavelet basis functions, and finally the effective signal and the noise data signal are restored.
[0021] The beneficial effect of the above further solution is that it achieves accurate separation of noise data and valid data through the noise data threshold, allowing subsequent storage and analysis to focus on valid information, and relies on reconstruction to ensure the reversibility of data, leaving traces for the traceability of noise data, and allowing the system to clearly distinguish the storage and processing logic of the two types of data.
[0022] Based on the technical solution, the application can be further improved as follows.
[0023] As a further improvement of the technical solution, the natural language processing method in the intelligent query and feature mining module analyzes the user's intention, and the specific working principle is as follows: perceive the user query statement, preprocess the user query statement, the preprocessing includes cleaning and standardization; after standardization, the query statement is segmented, and the part of speech of each word is labeled, the key entity is recognized, and each key data is matched with the sensor name, so as to call out the corresponding effective data and respond to the user's data demand.
[0024] As a further improvement of the technical solution, the intelligent query and feature mining module counts the access frequency of each user to the effective data, sets a frequency threshold, if the access frequency of the user to the effective data A is greater than the frequency threshold, the corresponding effective data is defined as hot effective data, otherwise the effective data is defined as cold effective data, and then a control signal is output to the database corresponding to the effective data to adjust the query weight of hot effective data and cold effective data, and the query weight of hot data is greater than that of cold data.
[0025] The beneficial effects of the above further scheme are that by counting the user access frequency to distinguish hot and cold effective data, the frequency threshold is accurately defined, and the query weight is adjusted to let the hot data respond first, which adapts to the actual scene and meets the user's demand for fast acquisition of high-frequency attention data (such as real-time data of key indicators of equipment), the high weight of hot data ensures the query efficiency, and the reasonable weight of cold data takes into account the historical data backtracking, which optimizes the balance between storage resources and query performance, avoids the low efficiency caused by the traditional "one-size-fits-all" query mode, and makes the intelligent query and feature mining module in massive time series data not only efficiently respond to real-time demand, but also properly retain historical value, helping the industrial Internet of Things, financial transactions and other fields to upgrade from "data storage" to "intelligent service", and improving the accuracy and experience of data utilization.
[0026] Based on the technical solution, the application can be further improved as follows.
[0027] As a further improvement of the technical solution, the principal component analysis algorithm in the intelligent query and feature mining module is used to extract the main characteristic components of the pseudo-noise data as the common features of the pseudo-noise data, and the specific working principle is as follows: perceive the pseudo-noise data and construct a pseudo-noise data set, each row in the pseudo-noise data set represents a noise data record, and each column represents a feature of the noise data.
[0028] Each feature is standardized, including the mean and standard deviation. Mean: calculate the average value of all data on the feature: add all the values of each feature and divide it by the total number of data; Standard deviation: calculate the square of the difference between each value and the mean, find the average and take the square root, then subtract the mean from the feature value of each data and divide it by the standard deviation to convert each feature into a dimensionless standard value;
[0029] For the standardized pseudo-noise data, the degree of correlation between different feature columns is calculated: for any two feature columns, the standardized value of each sample in the feature column is multiplied, the products of all samples are added, and finally divided by the number of samples - 1. The result is the covariance of the two features; if the two features are the same feature column, the calculation result is the variance of the feature itself, and then the covariances between all features, including their own variances, are arranged in row and column order to form a square matrix; and the covariance matrix is subjected to eigendecomposition and singular value decomposition to obtain a set of eigenvalues and corresponding eigenvectors. The eigenvalue is a set of numbers arranged from large to small. The larger the value, the more fluctuation information of the pseudo-noise data can be retained in the corresponding direction; the eigenvector is a set of direction vectors, each vector corresponds to an eigenvalue, representing the direction of the main component. According to the size of the eigenvalue, the corresponding eigenvectors are sorted from large to small, and the proportion of each eigenvalue to the total sum of all eigenvalues is calculated. A selection threshold is set, and the eigenvalue greater than the selection threshold is selected, which is the main characteristic component of the noise data.
[0030] As a further improvement of the present technical solution, when a user accesses noise data, the intelligent query and feature mining module quickly compares the main characteristic components: calculates the main characteristic components of the new noise data, marks the similarity of the main characteristic components of the pseudo-noise data, and prioritizes the pseudo-noise data with high similarity, allowing the user to view the pseudo-noise data first.
[0031] The beneficial effect of the above further scheme is that, through the principal component analysis algorithm, a pseudo-noise data set is first constructed, and the feature dimension differences are standardized to eliminate them, and then the main feature components are accurately extracted through the covariance matrix and matrix decomposition. This process breaks through the analysis dilemma of noise data due to high dimensionality and heterogeneity, extracts core rules from complex features, and "refines" the noise data; when subsequent users visit, the similarity is quickly compared based on the main feature components, and pseudo-noise data with high correlation is called out first. On the one hand, users do not need to blindly search in massive noise data, and can accurately reach key noise data information, solving the problem of low efficiency of traditional passive queries; on the other hand, with the help of feature extraction and similarity matching, the potential correlation of noise data is excavated, allowing the seemingly irregular noise data to "actively emerge" with value, making up for the lack of in-depth utilization of noise data by traditional databases, and promoting the transition from data management to value mining.
[0032] In addition to the purposes, features and advantages described above, the present application has other purposes, features and advantages. The present application will be further described below with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 The overall module schematic diagram of the present application;
[0034] Figure 2 The flow chart of the intelligent query and feature mining module of the present application;
[0035] Figure 3 The overall flow of data flow of the present application.
[0036] The meanings of various labels in the drawings are as follows:
[0037] 100, dynamic library table planning module; 200, noise data and effective data separation module; 300, intelligent query and feature mining module. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0039] The noise data ShardingJDBC dynamic library table extension system based on time series includes a dynamic library table planning module 100, a noise data and effective data separation module 200, and an intelligent query and feature mining module 300, wherein:
[0040] The dynamic library table planning module 100 perceives time series data detected by different sensors, creates a database corresponding to each sensor, and assigns a unique identifier to each sensor and the corresponding database, establishes a mapping relationship between each database and sensor, and is used to store time series data generated by the sensor in the mapping relationship, realizing "one type of data in one database"; an example mapping table structure is as follows:
[0041] Sensor ID Database Name Data Type
[0042] S001 DB_S001_Temp Temperature sensor data
[0043] S002 DB_S002_Pressure Pressure sensor data
[0044] And again according to the time series of time series data dynamically create library table: time series data has time sequence (such as continuous data generated by day, hour), need to dynamically table according to time dimension; Define time granularity T (such as day, hour), in time granularity unit, time series data is divided into non overlapping interval according to timestamp, its corresponding expression is as follows: for timestamp t, its interval is determined by time granularity G, the formula is as follows: if the time granularity G is day, the interval start time is t-(t mod 86400), that is, the current 0 point.
[0045] The application further considers that noise data and effective data coexist in time series data, noise data is worthless interference information generated by sensor hardware error, environmental interference (such as electromagnetic noise data, data jitter) or transmission anomaly; Effective data is information that truly reflects the state of the monitored object and can be directly used for analysis and decision; When the dynamic library table planning module 100 classifies and stores the time series data into different databases, if the noise data and effective information in the time series data cannot be effectively distinguished, it will further lead to the situation that the user analyzes the existence of noise data interference of each sensor corresponding to the time series data after searching the time series time through different types of databases, and at the same time, since the effective data in the time series data is usually concentrated in a specific frequency interval, and the noise data is usually expressed as high frequency component (such as white noise data) or low frequency drift, therefore, the noise data effective data separation module 200 uses wavelet analysis method to filter the noise data in the time series data, the time series data after filtering the noise data is effective data, and the dynamic library table planning module 100 classifies and stores the effective data and noise data in independent databases respectively, and constructs double library mapping association, which realizes data classification storage, avoids noise data interference to effective data analysis, and retains data traceability path through mapping association, facilitates subsequent verification and backtracking, improves the scientific nature of time series data storage management, builds a solid data foundation for accurate analysis and decision making, and makes the value mining of effective data more efficient and the management of noise data more clear.
[0046] The wavelet analysis method in the noise data effective data separation module 200 uses multi-resolution analysis idea, relies on a group of "local oscillation, energy limited" wavelet basis functions ψ a,b (t) to decompose time series data: the wavelet basis function ψ a,b (t) changes the shape and position through two key factors, one is the scale factor a which controls "stretching", the scale factor a determines the width of the wavelet basis function on the time axis, the larger the value, the wider the basis function, which corresponds to the slow changing trend in the analysis data; The other is the translation factor b which controls "sliding", the translation factor b is used to determine the specific position of the basis function in the time domain.
[0047] In the decomposition process, a cluster of basis functions is generated by scale factor a and translation factor b: Then the time series data f(t) is compared with the wavelet basis function ψ a,b (t) point by point, summed up, and the wavelet basis function ψ a,b (t) and the fitting degree of the time series data f(t) at different scales and positions are calculated to obtain the wavelet coefficient c (a,b) , as follows: The physical meaning of the wavelet coefficient c (a,b)
[0048] is: the greater the absolute value of the correlation between the time series data f(t) and the basis function ψ a,b (t) at the scale factor a and the translation factor b, the more similar the waveforms of the two are;
[0049] The scale factor and the translation factor are continuously adjusted to traverse various possible scales and positions, so as to split the time series data into a series of wavelet coefficients reflecting slow change trends (large scale) to details reflecting rapid fluctuations (small scale), to prepare for subsequent differentiation of effective data and noise data, and to realize detailed and multi-dimensional feature mining and analysis of the time series data. The corresponding working principle is: starting from the maximum scale a0 (corresponding to the lowest frequency), the global trend component is calculated; the scale is gradually reduced to extract higher frequency detail components in turn; the entire time axis is traversed by the translation factor b at each scale to ensure that the time series data f(t) is not missed in the decomposition;
[0050] Effective data (such as stable running trends of equipment, slowly changing environmental parameters) correspond to large scale (low frequency) wavelet coefficients (large amplitude, stable fluctuation); noise data (such as electromagnetic interference, sensor glitches) correspond to small scale (high frequency) wavelet coefficients (small amplitude, random fluctuation); therefore, a noise data threshold is set to distinguish noise data and effective data: the noise data threshold is [c min ,c max ], if c min ≤|c (a,b) |≤c max , the corresponding wavelet coefficient is judged to be an effective data wavelet coefficient; otherwise, it is a noise data wavelet coefficient. Then the differentiated wavelet coefficients are reconstructed: where C 有效 (a,b) and C 噪声 (a,b) are the retained and removed wavelet coefficients, respectively, and ∫ 有效 (t) and ∫ 噪声 (t) represent the effective data signal and the noise data signal, respectively, which are the results in the time domain after reconstruction, For double integral operation, traverse all possible scales a (from 0 to +∞) and all possible translations b (from -∞ to +∞), make sure to cover the features of the full time-frequency range of the signal; For scale integral measure (introduce this measure to ensure energy conservation, reversibility of the transform, and make the analysis at different scales "fairly comparable" due to the similarity of wavelet transform); db is the translation integral measure, which traverses all positions on the time axis;
[0051] Specifically: by traversing all scales a, translations b, with the preserved effective wavelet coefficients C 有效 (a,b) or the discarded noise data wavelet coefficients C 噪声 (a,b), respectively, and the corresponding wavelet basis function ψ a,b (t) do weighted integration (weights include coefficient values and measures db), finally restore the effective signal ∫ 有效 (t) and the noise data signal ∫ 噪声 (t), realize the process of "backtracking time-domain signal from frequency-domain coefficients", and ensure the reversibility of the wavelet transform.
[0052] The intelligent query and feature mining module 300 perceives the user query statement and uses natural language processing to analyze the user's intention: the user query statement is preprocessed, including cleaning and standardization; after standardization, the query statement is segmented, and the part of speech of each word is labeled, key entities are identified, and each relevant data is matched with the sensor name, so that the corresponding effective data is called out to respond to the user's data demand.
[0053] The access frequency of each user to the effective data is counted, and a frequency threshold is set. If the user's access frequency to the effective data A is greater than the frequency threshold, the corresponding effective data is defined as hot effective data, otherwise the effective data is defined as cold effective data, and then a control signal is output to the database corresponding to the effective data to adjust the query weights of hot effective data and cold effective data, thereby optimizing the access strategy of storing effective data to the database. When the user outputs the query statement again, the hot effective data is preferentially called out for the user to view, and its corresponding expression is as follows: set the hot data query weight as w hot , and the cold data weight as w cold , where w hot > w cold , and then output a control signal to adjust the database query priority: query sorting = w hot · hot data + w cold·Cold data, through intelligent query and feature mining module 300, when the user queries again, the hot effective data is preferentially called out, the high-frequency data access efficiency is improved, which not only conforms to the user's use habit, but also adjusts the storage access mechanism dynamically, so that the database resource allocation is more reasonable, helps the user to quickly obtain the required data, enhances the practicability and response speed of the system, and realizes the individualization and high efficiency optimization of data service.
[0054] Because the feature boundary of effective data and noise data is fuzzy, part of the effective data may present the characteristics of noise data (such as weak vibration caused by early failure of equipment, which has similar frequency and amplitude to environmental noise data), and noise data may also contain similar characteristics of effective data (such as periodic electromagnetic interference to sensors, forming a similar signal waveform), which will cause errors when the noise data effective data separation module 200 distinguishes effective data from noise data. In order to avoid the need for the user to query noise data twice due to errors when the noise data effective data separation module 200 distinguishes effective data from noise data, if the user queries effective data in the intelligent query and feature mining module 300, the noise data queried is defined as pseudo-noise data, and when the user queries effective data again, the pseudo-noise data is called out for the user to view: if the user queries effective data D sig , the noise data D noide is queried synchronously, then D noide is marked as pseudo-noise data D fake ; when the user queries effective data D sig again, the pseudo-noise data D fake is forced to be associated and displayed, so that when the user queries effective data again, the corresponding pseudo-noise data is automatically associated and called out, solving the trouble of the user querying noise data twice after misjudgment, without the need for the user to repeatedly search in the massive noise database, directly presenting the suspected misjudgment data to the user for verification, which simplifies the query process, improves the efficiency, retains misjudgment data samples, and also provides a real basis for subsequent algorithm iteration optimization and correction of noise data distinguishing rules.
[0055] Then, the principal component analysis algorithm is used to extract the main characteristic components of the pseudo-noise data as the common characteristics of the pseudo-noise data; if the noise data effective data separation module 200 distinguishes noise data according to the common characteristics of the noise data, the corresponding pseudo-noise data is called out, even if the user does not actively access the noise data, the corresponding pseudo-noise data can be automatically called out for the user to preferentially view, which upgrades "passive waiting for query" to "active prediction and supply", so that the user can directly touch the key noise data without groping in the ocean of noise data, and the working principle of the principal component analysis algorithm is as follows: the pseudo-noise data is perceived and constructed as a pseudo-noise data set There are n samples in the pseudo-noise data set (i.e., n noise data records), and each sample contains p features (such as noise data amplitude, duration, relative value of occurrence time, etc.);
[0056] In order to eliminate the influence of the dimension difference of different features on the results, each feature j is normalized: for each feature j, first calculate the average value of all data on feature j (the average value of all pseudo-noise data on the jth feature), add up all the values of the jth feature, and divide it by the total number of data n: And the standard deviation (the "volatility" of the j-th feature), first calculate the square of the difference between each value and the mean, find the average and then take the square root to get the sample standard deviation of the j-th feature: Then subtract the mean from the eigenvalue of each data point and divide it by the sample standard deviation: Thus each feature j is converted into a dimensionless standard value;
[0057] After standardization, construct the j×j dimensional covariance matrix ∑, element ∑ jk (jth row, kth column) calculation formula: When j=k:∑ jj is the variance of the jth feature, reflecting the fluctuation intensity of the noise data feature itself; when j≠k:∑ jj is the covariance of the j-th and k-th features. A positive value indicates that the two features "rise and fall together" (positive correlation), and a negative value indicates that "one rises while the other falls" (negative correlation). The larger the absolute value, the closer the correlation.
[0058] Eigendecomposition of the covariance matrix ∑: ∑=U∧U T , decompose the principal component direction of the noise data (eigenvector U) and the variance contribution of each direction (eigenvalue ^), and obtain m groups of eigenvalues and the corresponding eigenvector And the eigenvector satisfies the covariance matrix multiplied by the eigenvector, which is equal to the eigenvalue multiplied by the eigenvector; the eigenvalue λ i is the variance (importance) of the corresponding principal component, the eigenvector u i The direction (weight) of the principal component;
[0059] Then according to the eigenvalue λ i Sort from large to small, and calculate the proportion of each eigenvalue to the total of all eigenvalues (variance contribution), set the selection threshold k, and select the first k principal components (corresponding to the first k eigenvalues) as the retained main characteristic components for subsequent noise data classification and matching (when users access noise data, similar pseudo-noise data can be quickly retrieved through the principal component features).
[0060] Through the main characteristic components, multiple original features of the pseudo-noise data (such as frequency, amplitude, duration, etc.) are compressed into a few main characteristic components. The main characteristic components are the common characteristics of the pseudo-noise data (for example, the core characteristics of a certain type of pseudo-noise data may be hidden in the first and second principal components).
[0061] When users access noise data, they can do a "quick comparison" through the main characteristic components: calculate the "similarity" between the main characteristic components of the new noise data and the main characteristic components of the previously marked pseudo-noise data. Pseudo-noise data with high similarity will be called out first, allowing users to view the pseudo-noise data first.
[0062] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. ShardingJDBC dynamic library table expansion system based on time series noise data, characterized by: It includes a dynamic library table planning module (100), a noise data and valid data separation module (200) and an intelligent query and feature mining module (300), wherein: The dynamic library table planning module (100) establishes a mapping relationship between each database and a sensor, and dynamically creates a library table; the noise data and effective data separation module (200) uses a wavelet analysis method to filter noise data in time series data, and the time series data after filtering the noise data is effective data, and the wavelet coefficients after the separation are reconstructed; when the dynamic library table planning module (100) classifies and stores the time series data, the effective data and the noise data are respectively stored in independent databases, and a dual-library mapping association is constructed; When a user queries a database, the intelligent query and feature mining module (300) uses a natural language processing method to analyze the user's intention, and counts the frequency of each user's access to valid data, and distinguishes the valid data into hot and cold valid data; if the user queries valid data, the noise data is simultaneously queried, and the synchronously queried noise data is defined as pseudo-noise data. When the user queries for valid data again, the pseudo-noise data is synchronously called out for the user to view, and the principal component analysis algorithm is used to extract the main characteristic components of the pseudo-noise data as the common characteristics of the pseudo-noise data. For noise data that the user has not actively accessed, the pseudo-noise data is predicted based on the common characteristics.
2. The noise data ShardingJDBC dynamic library and table expansion system based on time series according to claim 1 is characterized by: The dynamic library table planning module (100) senses the time series data detected by different sensors, creates a database corresponding to each sensor, and assigns a unique identifier to each sensor and the corresponding database, thereby establishing a mapping relationship between each database and the sensor.
3. The noise data ShardingJDBC dynamic library and table expansion system based on time series according to claim 2 is characterized by: The dynamic library table planning module (100) again dynamically creates a library table according to the time series of the time series data: defines a time granularity, and divides the time series data into non-overlapping intervals according to the time stamp using the time granularity as a unit.
4. The noise data Sharding JDBC dynamic library table expansion system based on time series according to claim 1 is characterized in that: The wavelet analysis method in the noise data effective data separation module (200) includes a set of wavelet basis functions to decompose time series data: the shape and position of the wavelet basis functions are controlled by scale factors and translation factors; The time series data is compared with the wavelet basis function point by point, and the sum is calculated. The degree of fit between the wavelet basis function and the time series data at different scales and positions is calculated to obtain the wavelet coefficients. The scale factor and translation factor are adjusted to split the time series data into a series of wavelet coefficients.
5. The noise data ShardingJDBC dynamic library and table expansion system based on time series according to claim 4 is characterized in that: The noise data and valid data separation module (200) uses a wavelet analysis method to set a noise data threshold to distinguish noise data from valid data: If the absolute value of the wavelet coefficient is within the noise data threshold, the corresponding wavelet coefficient is judged to be a valid data wavelet coefficient; otherwise, it is a noise data wavelet coefficient.
6. The noise data ShardingJDBC dynamic library and table expansion system based on time series according to claim 5 is characterized by: The noise data effective data separation module (200) reconstructs the differentiated wavelet coefficients: By traversing all scales and translations, the retained effective wavelet coefficients and the eliminated noise data wavelet coefficients are used to perform weighted integration with the corresponding wavelet basis functions, and finally the effective signal and noise data signal are restored.
7. The noise data Sharding JDBC dynamic library table expansion system based on time series according to claim 1 is characterized by: The natural language processing method in the intelligent query and feature mining module (300) analyzes user intentions, and the specific working principle is as follows: perceive the user query sentence, pre-process the user query sentence, and the pre-processing includes cleaning and standardization; after standardization, the query sentence is segmented and the part of speech of each word is marked, the key entity is identified, and the sensor name corresponding to each key valid data is matched, so as to call out the corresponding valid data and respond to the user's data needs.
8. The noise data Sharding JDBC dynamic library table expansion system based on time series according to claim 7 is characterized in that: The intelligent query and feature mining module (300) counts the access frequency of each user to valid data and sets a frequency threshold. If the user's access frequency to valid data A is greater than the frequency threshold, the corresponding valid data is defined as hot valid data, otherwise the valid data is defined as cold valid data. Then, a control signal is output to the database corresponding to the valid data to adjust the query weights of hot valid data and cold valid data, and the hot data query weight is greater than the cold data query weight.
9. The time series-based noise data ShardingJDBC dynamic library and table expansion system according to claim 8, characterized in that: The principal component analysis algorithm in the intelligent query and feature mining module (300) is used to extract the main characteristic components of the pseudo noise data as the common characteristics of the pseudo noise data. The specific working principle is as follows: the pseudo noise data is sensed and constructed into a pseudo noise data set, where each row in the pseudo noise data set represents a noise data record and each column represents a feature of the noise data; And standardize each feature. Standardization includes mean and standard deviation. Mean: calculate the average value of all data on the feature: add up all the values of each feature and divide by the total number of data; Standard deviation: Calculate the square of the difference between each value and the mean, find the average and take the square root, then subtract the mean from the characteristic value of each data, divide by the standard deviation, and convert each feature into a dimensionless standard value; For the standardized pseudo-noise data, the degree of correlation between different feature columns is calculated: for any two feature columns, the standardized value of each sample in the feature column is multiplied, the products of all samples are added, and finally divided by the number of samples - 1. The result is the covariance of the two features; if the two features are the same feature column, the calculation result is the variance of the feature itself, and then the covariances between all features, including their own variances, are arranged in row and column order to form a square matrix; and the covariance matrix is subjected to eigendecomposition and singular value decomposition to obtain a set of eigenvalues and corresponding eigenvectors. The eigenvalue is a set of numbers arranged from large to small. The larger the value, the more fluctuation information of the pseudo-noise data can be retained in the corresponding direction; the eigenvector is a set of direction vectors, each vector corresponds to an eigenvalue, representing the direction of the main component. According to the size of the eigenvalue, the corresponding eigenvectors are sorted from large to small, and the proportion of each eigenvalue to the total sum of all eigenvalues is calculated. A selection threshold is set, and the eigenvalue greater than the selection threshold is selected, which is the main characteristic component of the noise data.
10. The noise data ShardingJDBC dynamic library table expansion system based on time series according to claim 9, characterized in that: When a user accesses noise data, the intelligent query and feature mining module (300) performs a quick comparison based on the main feature components: calculates the main feature components of the new noise data, marks the similarity of the main feature components of the pseudo noise data, and preferentially calls out the pseudo noise data with high similarity, so that the user can view the pseudo noise data preferentially.