Log data quick query method and device based on big data analysis and medium
By constructing text data similarity sequence and time weights, and combining time characteristics for clustering and dimensionality reduction, the problem that traditional PCA methods cannot accurately reflect log data semantics is solved, and efficient and accurate log data query is achieved.
Patent Information
- Application Number
- CN202510576238.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-15
AI Technical Summary
When processing large-scale log data in the prior art, the traditional PCA dimensionality reduction method cannot accurately reflect the semantic information of text data, resulting in a decrease in the accuracy of query results and the keyword-based query method is inefficient.
By constructing the similarity sequence and time weight of text data, fitting function correction, clustering and dimensionality reduction combined with time characteristics, improving the accuracy and stability of the similarity measurement of text data, enhancing the accuracy of clustering, and querying the data after dimensionality reduction.
While narrowing the query space, it improves the accuracy and efficiency of log data query, reduces information loss, and enhances the accuracy of clustering.
Smart Images

Figure CN120492609A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database query technology, and specifically to a method, device, and medium for quickly querying log data based on big data analysis. Background Art
[0002] With the development of information technology, the amount of log data generated by businesses and social organizations is rapidly increasing, and its scale and complexity are also continuously climbing. However, faced with such large and rapidly growing log data sets, traditional data management and analysis methods are struggling, especially in terms of fast querying. Therefore, how to achieve fast data query has become a key issue that needs to be addressed in the current big data field.
[0003] Existing technologies perform PCA (Principal Component Analysis) dimensionality reduction on data in log databases, thereby reducing the subsequent query space and improving query efficiency. However, the PCA dimensionality reduction method is mainly applicable to numerical data. There is a large amount of non-numerical text data in existing log data. Although existing technologies can convert text data into numerical data, the converted numerical values are difficult to accurately reflect the actual semantic information of the text data, resulting in the loss of a large amount of valuable information after dimensionality reduction, which in turn affects the query results of log data. Summary of the Invention
[0004] In view of the above, it is necessary to provide a method, device, and medium for rapid log data query based on big data analysis. Compared with traditional log data query methods, this method can improve the accuracy of query results while narrowing the query space:
[0005] In a first aspect, an embodiment of the present application provides a method for quickly querying log data based on big data analysis, the method comprising the following steps:
[0006] Export each tuple from the log database;
[0007] Using a word vector model, a text data vector of each text data in each tuple is obtained respectively; a similarity sequence of any column of text data in any tuple is constructed by similarity between the text data in the same column of text data in the remaining tuples, and a curve fitting is performed to obtain a first fitting function; a time weight between any two tuples is obtained by a timestamp difference between the tuples and a total time span of a log database; the time weights between all any two tuples are arranged by an arrangement order of elements in the similarity sequence and a curve fitting is performed to obtain a second fitting function, and the first fitting function is corrected;
[0008] Obtaining a text similarity coefficient of the any column of text data through the correction result and the similarity of the text data vectors between the any column of text data and the rest of the text data in the tuple to which it belongs;
[0009] Constructing each total data vector using the text similarity coefficient of the numerical data in each tuple and all text data, and clustering all tuples using the total data vector, wherein the metric distance between each tuple and the cluster center of each cluster is obtained based on the timestamp difference between each tuple and the cluster center and the similarity of the total data vector;
[0010] Reduce the dimension of each cluster and perform data query based on the dimension reduction results.
[0011] In one embodiment, the elements in the similarity sequence are arranged according to numerical values.
[0012] In one embodiment, the time weight is expressed as:
[0013] Where, T ia represents the time weight between the i-th tuple and the a-th tuple; e represents a natural constant; Δt ia represents the timestamp difference between the i-th tuple and the a-th tuple; ΔT represents the total time span of the log database; λ represents a parameter constant that is preset to be greater than 0.
[0014] In one embodiment, arranging the time weights between all arbitrary two tuples and performing curve fitting includes:
[0015] According to the order of the remaining tuples corresponding to the elements in the similarity sequence, the time weights corresponding to the remaining tuples are arranged to form a time weight sequence, and curve fitting is performed on the time weight sequence.
[0016] In one embodiment, the method for correcting the first fitting function is: calculating the product of the first fitting function and the second fitting function.
[0017] In one embodiment, the method for obtaining the text similarity coefficient is:
[0018] Obtaining an area enclosed by the correction result, a straight line with a ordinate of 0, a straight line with a ordinate equal to the length of the similarity sequence, and a straight line with a horizontal coordinate of 0;
[0019] Calculate the mean similarity between the text data vectors of any column of text data and all other columns of text data in the tuple to which it belongs;
[0020] The text similarity coefficient is the product of the area and the mean.
[0021] In one embodiment, the method for obtaining the metric distance is:
[0022] The distance between each tuple and the total data vector of each cluster center is recorded as cluster distance;
[0023] Calculate the cosine similarity of the total data vector between each tuple and the cluster center of each cluster;
[0024] The metric distance is positively correlated with the timestamp difference and the cluster distance, respectively, and negatively correlated with the cosine similarity.
[0025] In one embodiment, the distance measurement expression is:
[0026] Where D iE represents the metric distance between the i-th tuple and the cluster center of cluster E; e represents a natural constant; w t_iE represents the timestamp difference between the i-th tuple and the cluster center of cluster E; c iE represents the cosine similarity of the total data vector between the i-th tuple and the cluster center of cluster E; d iE Represents the distance of the total data vector between the i-th tuple and the cluster center of cluster E.
[0027] In a second aspect, an embodiment of the present application provides a log data quick query device based on big data analysis, wherein the computer device includes a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the log data quick query method based on big data analysis as described in the first aspect are implemented.
[0028] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method for quickly querying log data based on big data analysis as described in the first aspect.
[0029] This application has at least the following beneficial effects:
[0030] This application constructs a similarity sequence based on the similarity of the same column of text data in different tuples, and obtains a first fitting function through curve fitting, which can smooth the fluctuations in the similarity sequence, better capture the overall trend of similarity between text data, and enhance the accuracy and stability of the text data similarity measurement;
[0031] Furthermore, by calculating the time weight and fitting a second fitting function, the first fitting function is corrected by the second fitting function, so that the time characteristics of the log data can be integrated into the similarity analysis of the text data. This makes the calculation of the text similarity coefficient not only consider the semantic information of the text data itself, but also fully considers the influence of the time factor on the text similarity, thereby more comprehensively and accurately calculating the actual similarity of the text data, enhancing the accuracy of the quantification of the text data, and providing a higher quality data foundation for subsequent clustering and query operations.
[0032] Furthermore, during clustering, the accuracy of log data clustering is improved by introducing time characteristics; dimensionality reduction is performed on each cluster separately, which can better retain the characteristics and information of the data in each cluster, avoiding the loss of a large amount of valuable information due to unified dimensionality reduction; data query is performed on the data after dimensionality reduction, which can improve the accuracy of the query results on the basis of reducing the query space. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0034] Figure 1 A flowchart of the steps of a method for quickly querying log data based on big data analysis provided by one embodiment of the present application;
[0035] Figure 2 Schematic diagram of the process of obtaining clusters. DETAILED DESCRIPTION
[0036] In the description of the embodiments of this application, words such as "exemplary," "or," and "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "or," and "for example" is intended to present the relevant concepts in a concrete manner.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application relates. The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise indicated, " / " represents or.
[0038] It should also be noted that the terms "first" and "second" in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0039] The following describes in detail the specific solutions of the log data rapid query method, device and medium based on big data analysis provided by this application with reference to the accompanying drawings.
[0040] See also Figure 1 , which shows a flowchart of a method for quickly querying log data based on big data analysis provided by an embodiment of the present application, the method comprising the following steps:
[0041] Step 1: Export each tuple from the log database.
[0042] Export each stored tuple from the log database and preprocess the original log data in each tuple, including deleting blank tuples and tuples with duplicate records; converting the timestamp data in the tuple to seconds; encoding the text data in the tuple into UTF-8; and normalizing all numerical data in all tuples.
[0043] In this embodiment, the Min-Max normalization method is used to normalize the numerical data. The Min-Max normalization method is a well-known technology and will not be described in detail in this application.
[0044] Step 2: Calculate the text similarity coefficient of each column of text data in each tuple based on the semantic features and time characteristics of the text data in the log data; when clustering all tuples, calculate the metric distance between tuples by introducing the time characteristics to obtain each cluster.
[0045] Traditional methods for querying log data require writing complex query statements, such as SQL statements. To reduce query complexity, keyword-based data query methods have been proposed, allowing users to query a single data item or a category of data by entering a keyword. However, keyword-based query methods suffer from similar drawbacks to traditional methods: when the tables and tuples in a log database are complex, the entire log database must be traversed, severely impacting query efficiency. Furthermore, because the keywords entered by users can be text data, they may not perfectly match the text stored in the log database, further impacting query efficiency when performing a full traversal query. Therefore, to improve query efficiency, PCA (Principal Component Analysis) dimensionality reduction is employed to reduce the query space. However, PCA dimensionality reduction has poor processing capabilities for text data, and existing text data quantization techniques struggle to accurately represent the actual semantics of text data. Consequently, information loss is exacerbated after dimensionality reduction, significantly impacting the accuracy of keyword-based data queries.
[0046] When using keywords to query log data, there are certain differences between the text data of each tuple, and the differences in text data cannot be measured by simple quantification to measure the actual meaning. Therefore, it is necessary to measure the similarity of text data of the same attribute between different tuples based on the semantic information of each text data; at the same time, since there may also be semantic similarities between different text data within each tuple, it is also necessary to measure the similarity between different text data within each tuple; in addition, since log data has a strong timeliness, the semantics of text data in different periods may also have certain differences. Therefore, when measuring the similarity of text data in different tuples, the above-mentioned time characteristics need to be taken into account. The smaller the time interval between two tuples, the greater the significance of the text similarity between the two tuples.
[0047] Step 2.1, use the word vector model to obtain the text data vector of each text data in each tuple respectively, and through the similarity of the text data vectors between any column of text data in any tuple and the same column of text data in the remaining tuples, construct the similarity sequence of any column of text data and perform curve fitting to obtain the first fitting function.
[0048] Each text data in each tuple is taken as input, and the word vector model is used to output the text data vector.
[0049] In this embodiment, the Word2Vec word vector model is used to obtain text data vectors. The Word2Vec word vector model is a well-known technology and will not be described in detail in this application. The implementer can choose other feasible word vector models on his own.
[0050] Taking the j-th column of text data in the i-th tuple in a log database as an example, the similarity between the text data vectors of the j-th column of text data in the remaining tuples and the j-th column of text data in the i-th tuple is first calculated. All the calculated similarities are arranged in descending order to form a similarity sequence of the j-th column of text data in the i-th tuple. A curve fitting is performed using the elements in the similarity sequence as the vertical coordinates and the index of the elements as the horizontal coordinates, and the first fitting function is output. Since the similarity measurement results may occasionally have values that are too large or too small, these values are not conducive to measuring the overall change trend of the similarity between the remaining tuples in the log data and the i-th tuple with respect to the j-th column of text data, thereby affecting the subsequent quantification of text data, the use of a fitting curve function can enhance the accuracy of measuring the overall change trend.
[0051] In this embodiment, the similarity between text data vectors is cosine similarity. The calculation of cosine similarity is a well-known technology and will not be described in detail in this application. On the basis of being able to measure the similarity between text data vectors, the implementer may adopt other existing technologies, such as the reciprocal of the Euclidean distance, etc., and this application does not impose any special restrictions.
[0052] In this embodiment, the least squares method is used to obtain the first fitting function. The specific process of using the least squares method to obtain the first fitting function is a well-known technology and will not be described in detail in this application. On the basis of being able to obtain the first fitting function, the implementer can select other existing feasible technologies on his own.
[0053] Step 2.2: Obtain the time weight between any two tuples based on the timestamp difference between the two tuples and the total time span of the log database; arrange the time weights between all any two tuples based on the order of the elements in the similarity sequence and perform curve fitting to obtain a second fitting function;
[0054] The time weight between any two tuples is obtained by using the timestamp difference between them and the total time span of the log database. The expression is:
[0055] T ia represents the time weight between the i-th tuple and the a-th tuple; e represents a natural constant; Δt ia represents the timestamp difference between the i-th tuple and the a-th tuple; ΔT represents the total time span of the log database; λ represents a parameter constant preset to be greater than 0. By controlling the value of λ, the influence of the timestamp difference on the semantic similarity analysis of the text data is adjusted. When λ is smaller, the influence on the semantic similarity analysis is greater. In this embodiment, the value of λ is 5. When the value of λ is too large, the influence of the timestamp difference on the semantic similarity becomes too weak. Therefore, the value of λ should not be too large. On the basis of satisfying the value range of λ being (0,10], the implementer can set the value of λ at will. Among them, the purpose of introducing the exponential function is to enhance the similarity between different text data and help improve the quality of text data quantification.
[0056] In this embodiment, the difference value between the timestamps is the absolute value of the difference.
[0057] It should be noted that due to the temporal characteristics of text data, the longer the time span between two tuples, the weaker the semantic association between the text data in the two tuples, and the smaller the time weight between the two tuples; the longer the total time span, the smaller the impact of the timestamp difference between the two tuples on semantic similarity, and the larger the time weight.
[0058] The reason for introducing the total time span is that the length of the total time span affects the degree of influence of timestamp differences on semantic similarity. When the total time span is shorter, the temporal distribution of the data is more concentrated, and the impact of timestamp differences on semantic similarity is more significant. Therefore, it is necessary to enhance the ability to capture the impact of temporal changes on semantic similarity, thereby improving the accuracy of distinguishing different tuples and ensuring the accuracy of subsequent queries. Conversely, when the total time span is longer, the impact of timestamp differences on semantic similarity is smaller. In this case, it is necessary to focus on the changes in semantic differences caused by long time differences and avoid over-emphasizing semantic differences within a short time. Otherwise, semantically related data with slight temporal differences may be missed during queries.
[0059] The time weights corresponding to the remaining tuples are arranged according to the order of the remaining tuples corresponding to the elements in the similarity sequence to form a time weight sequence, and a curve fitting is performed on the time weight sequence to obtain a second fitting function.
[0060] In this embodiment, the least squares method is used to obtain the second fitting function. The specific process of using the least squares method to obtain the second fitting function is a well-known technology and will not be described in detail in this application. On the basis of being able to obtain the second fitting function, the implementer can select other existing feasible technologies on his own.
[0061] Step 2.3, using the second fitting function to correct the first fitting function, and obtaining the text similarity coefficient of the any column of text data through the correction result and the similarity of the text data vectors between the any column of text data and the rest of the text data in its tuple.
[0062] Still taking the j-th column of text data in the i-th tuple in the log database as an example, the expression of the text similarity coefficient of the j-th column of text data in the i-th tuple is:
[0063] A ij =F[f(x)·t(x)]×σ ij ; A ij represents the text similarity coefficient of the j-th column text data in the i-th tuple; f(x) and t(x) represent the first fitting function and the second fitting function of the j-th column text data in the i-th tuple respectively; σ ij It represents the mean similarity between the text data vectors of the j-th column text data in the i-th tuple and all other columns of text data; F[*] represents the integral operation, and the upper and lower limits of the integral are the similarity sequence length and 0, respectively.
[0064] In this embodiment, the similarity between text data vectors is cosine similarity. On the basis of being able to measure the similarity between text data vectors, the implementer may adopt other existing technologies, such as the inverse of the Euclidean distance, etc., and this application does not impose any special restrictions.
[0065] It should be noted that: when the semantic similarity between the j-th column text data in the i-th tuple and the j-th column text data of the remaining tuples is greater, and the semantic similarity between the j-th column text data in the i-th tuple and the remaining text data is greater, the randomness of the j-th column text data in the i-th tuple is weaker, the text distribution characteristics of the j-th column text data in the i-th tuple are stronger, and the text similarity coefficient is greater.
[0066] By calculating the text similarity coefficient of the j-th column of the i-th tuple, the j-th column of the i-th tuple is quantified. Traditional methods use semantic vectors to quantify text data. While this takes into account the semantic characteristics of the text data, it does not consider the temporal characteristics of the text data within all log data. This can lead to errors in subsequent clustering, affecting the accuracy of log data queries. Quantifying text data by calculating the text similarity coefficient not only considers semantic characteristics but also incorporates the temporal characteristics of the text data, helping to enhance the accuracy of subsequent clustering. This, in turn, helps reduce information loss during subsequent dimensionality reduction, ensuring the accuracy of data queries while narrowing the query space.
[0067] According to the calculation method of the text similarity coefficient of the j-th column of text data in the i-th tuple, the text similarity coefficient of each column of text data in each tuple is calculated to achieve quantification of each column of text data in each tuple.
[0068] Step 2.4, construct each total data vector through the numerical data in each tuple and the text similarity coefficient of all text data, cluster all tuples through the total data vector, wherein the metric distance between each tuple and the cluster center of each cluster is obtained through the timestamp difference between each tuple and the cluster center and the similarity of the total data vector.
[0069] Normalize the quantized text data, that is, normalize the text similarity coefficients of all column text data in all tuples. Arrange the normalized values of all numerical data in each tuple and the normalized values of the text similarity coefficients of all column text data according to the original positions of the numerical data and text data in each tuple to form the total data vector of each tuple. Based on the total data vectors of all tuples, cluster all tuples to obtain clusters. The schematic diagram of the cluster acquisition process is shown in the figure. Figure 2 shown.
[0070] In this embodiment, the Min-Max normalization method is used to normalize the text similarity coefficient.
[0071] In this embodiment, the K-means clustering algorithm is used to cluster all tuples, and the number of clusters is determined by the elbow rule. The value range of K in the elbow rule is an integer in [5, 20]. Among them, the K-means clustering algorithm and the elbow rule are both well-known technologies and will not be described in detail in this application. On the basis of being able to cluster all tuples, the implementer can select other existing feasible clustering algorithms at his own discretion.
[0072] Since log data has time characteristics, time characteristics need to be taken into account when clustering. Therefore, the expression for the metric distance between each tuple and the cluster center of each cluster is:
[0073] Where D iE represents the metric distance between the i-th tuple and the cluster center of cluster E; e represents a natural constant; wt _iE represents the timestamp difference between the i-th tuple and the cluster center of cluster E; c iE represents the cosine similarity of the total data vector between the i-th tuple and the cluster center of cluster E; d iE Represents the distance between the total data vector of the i-th tuple and the cluster center of cluster E. iE It is recorded as cluster distance.
[0074] In this embodiment, Δs i , Δs E are the time intervals between the timestamp corresponding to the i-th tuple and the cluster center of cluster E and the current moment.
[0075] In this embodiment, the distance between the total data vector of the i-th tuple and the cluster center of cluster E is the Euclidean distance. As other implementation methods, on the basis of being able to measure the numerical differences between the total data vectors, the implementer may adopt other existing technologies, such as Manhattan distance, etc., and this application does not impose any special restrictions.
[0076] It should be noted that: Calculate w t_iE The purpose is to introduce time characteristics to improve the accuracy of log data clustering. If the time weight calculated in step 2.2 is still used, the exponential part will cause the influence of time characteristics to be too large, thereby reducing the clustering quality. Therefore, in the clustering process, as time weight;
[0077] When the difference in the total data vector between the i-th tuple and the cluster center of cluster E is smaller, the distance is smaller. Traditional distance calculation methods, such as Euclidean distance, can only measure the numerical difference between the total data vectors. However, since the total data vector of each tuple is a vector, it is necessary to combine cosine similarity to measure the directional difference between the total data vectors. At the same time, since log data has a strong time characteristic, when the time interval between two tuples is larger, the correlation between the tuples is weaker, and in the clustering process, the metric distance between the tuples is larger.
[0078] Step 3: Reduce the dimension of each cluster and perform data query based on the dimension reduction results.
[0079] Generate a log data matrix for each cluster cluster, wherein the row of the log data matrix represents a tuple, i.e., the total data vector of the tuple, and the column of the log data matrix represents each attribute in the tuple, i.e., each dimension of the total data vector of the tuple. The log data matrix of each cluster cluster is used as input respectively, PCA dimensionality reduction is adopted, and the data after dimensionality reduction is output. Then the data after dimensionality reduction, the keywords to be queried, and the number of queries k to be returned are used as input, and a Top-k query algorithm is adopted to output a query result, wherein the number of queries k to be returned can be set according to actual needs, representing the first k elements with the highest correlation with the keyword. PCA dimensionality reduction and Top-k query algorithm are both well-known technologies, and this application will not repeat them in detail.
[0080] Based on the same inventive concept as the above method, an embodiment of the present application also provides a log data quick query device based on big data analysis, wherein the computer device includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned log data quick query methods based on big data analysis are implemented.
[0081] Based on the same inventive concept as the above method, a computer-readable storage medium is proposed, which stores a computer program. When the computer program is executed by the processor, it implements the log data rapid query method based on big data analysis as described in the first aspect. Its specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0082] In summary, this application constructs a similarity sequence based on the similarity of the same column of text data in different tuples, and obtains a first fitting function through curve fitting. This can smooth out the fluctuations in the similarity sequence, better capture the overall trend of similarity between text data, and enhance the accuracy and stability of the text data similarity measurement.
[0083] Furthermore, by calculating the time weight and fitting a second fitting function, the first fitting function is corrected by the second fitting function, so that the time characteristics of the log data can be integrated into the similarity analysis of the text data. This makes the calculation of the text similarity coefficient not only consider the semantic information of the text data itself, but also fully considers the influence of the time factor on the text similarity, thereby more comprehensively and accurately calculating the actual similarity of the text data, enhancing the accuracy of the quantification of the text data, and providing a higher quality data foundation for subsequent clustering and query operations.
[0084] Furthermore, during clustering, the accuracy of log data clustering is improved by introducing time characteristics; dimensionality reduction is performed on each cluster separately, which can better retain the characteristics and information of the data in each cluster, avoiding the loss of a large amount of valuable information due to unified dimensionality reduction; data query is performed on the data after dimensionality reduction, which can improve the accuracy of the query results on the basis of reducing the query space.
[0085] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.
[0086] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the basic characteristics of the present application. Therefore, from all perspectives, the above embodiments of the present application should be regarded as exemplary and non-restrictive.
Claims
1. A log data fast query method based on big data analysis, characterized in that: The method comprises the following steps: Export each tuple from the log database; Using a word vector model, a text data vector of each text data in each tuple is obtained respectively; a similarity sequence of any column of text data in any tuple is constructed by similarity between the text data in the same column of text data in the remaining tuples, and a curve fitting is performed to obtain a first fitting function; a time weight between any two tuples is obtained by a timestamp difference between the tuples and a total time span of a log database; the time weights between all any two tuples are arranged by an arrangement order of elements in the similarity sequence and a curve fitting is performed to obtain a second fitting function, and the first fitting function is corrected; Obtaining a text similarity coefficient of the any column of text data through the correction result and the similarity of the text data vectors between the any column of text data and the rest of the text data in the tuple to which it belongs; Constructing each total data vector using the text similarity coefficient of the numerical data in each tuple and all text data, and clustering all tuples using the total data vector, wherein the metric distance between each tuple and the cluster center of each cluster is obtained based on the timestamp difference between each tuple and the cluster center and the similarity of the total data vector; Reduce the dimension of each cluster and perform data query based on the dimension reduction results.
2. The method for quickly querying log data based on big data analysis according to claim 1, characterized in that: The elements in the similarity sequence are arranged according to numerical values.
3. The method for rapid query of log data based on big data analysis according to claim 1, characterized in that: The expression of the time weight is: Where, T ia represents the time weight between the i-th tuple and the a-th tuple; e represents a natural constant; Δt ia represents the timestamp difference between the i-th tuple and the a-th tuple; ΔT represents the total time span of the log database; λ represents a parameter constant that is preset to be greater than 0.
4. The method for quickly querying log data based on big data analysis according to claim 1, characterized in that: The method of arranging the time weights between any two tuples and performing curve fitting includes: According to the order of the remaining tuples corresponding to the elements in the similarity sequence, the time weights corresponding to the remaining tuples are arranged to form a time weight sequence, and curve fitting is performed on the time weight sequence.
5. The method for rapid query of log data based on big data analysis according to claim 1, characterized in that: The method for correcting the first fitting function is: calculating the product of the first fitting function and the second fitting function.
6. The method for rapid query of log data based on big data analysis according to claim 1, characterized in that: The method for obtaining the text similarity coefficient is: Obtaining an area enclosed by the correction result, a straight line with a ordinate of 0, a straight line with a ordinate equal to the length of the similarity sequence, and a straight line with a horizontal coordinate of 0; Calculate the mean similarity between the text data vectors of any column of text data and all other columns of text data in the tuple to which it belongs; The text similarity coefficient is the product of the area and the mean.
7. The method for rapid query of log data based on big data analysis according to claim 1, characterized in that: The method for obtaining the metric distance is: The distance between each tuple and the total data vector of each cluster center is recorded as cluster distance; Calculate the cosine similarity of the total data vector between each tuple and the cluster center of each cluster; The metric distance is positively correlated with the timestamp difference and the cluster distance, respectively, and negatively correlated with the cosine similarity.
8. The method for rapid query of log data based on big data analysis according to claim 7, characterized in that: The expression of the metric distance is: Where D iE represents the metric distance between the i-th tuple and the cluster center of cluster E; e represents a natural constant; w t_iE represents the timestamp difference between the i-th tuple and the cluster center of cluster E; c iE represents the cosine similarity of the total data vector between the i-th tuple and the cluster center of cluster E; d iE Represents the distance of the total data vector between the i-th tuple and the cluster center of cluster E.
9. A log data quick query device based on big data analysis, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for rapid query of log data based on big data analysis as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for quickly querying log data based on big data analysis as described in any one of claims 1 to 8 is implemented.