User query behavior prediction method and device based on SQL query
By predicting the user's SQL query behavior and generating a target behavior prediction model, the problems of high resource consumption and low query efficiency in the existing technology from the data lake to the edge database are solved, and more efficient query prediction is achieved.
Patent Information
- Application Number
- CN202510057232.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the prior art updates from the data lake to the edge database synchronously, it consumes huge network communication resources and is inefficient in querying, making it difficult to solve the problem of instant query.
By determining the sample data query information list of each sample user, segmenting these lists to obtain the training set and label information, training is used to generate the target behavior prediction model, and then predicting the SQL query results of the target user.
It improves the matching degree between the predicted query results and the target user, improves the accuracy and stability of query prediction, and reduces the consumption of network communication resources.
Smart Images

Figure CN119938724A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database technology, and in particular to a method and device for predicting user query behavior based on SQL query. Background Art
[0002] Currently, various mobile devices, network sensors, social media, etc. are constantly generating large amounts of data. The data storage method has changed from traditional centralized storage to lake + end form. The issue of updating from data lake to database has become an issue that must be considered. The traditional method of synchronously updating from data lake to edge database in full update consumes huge network communication resources, but cannot solve the problem of instant query at critical moments. The time cost of returning query results is extremely high, and the query efficiency is low.
[0003] In the related technology, machine learning methods can be used to obtain estimated results and actual query results of historical queries in an offline manner based on historical query data. The estimated results and actual query results are used as features, and the actual query results are used as labels for offline training to obtain the underlying data distribution model, which can then be used to predict the user's query behavior and obtain query results.
[0004] However, in the above method, the accuracy of query results obtained by different users varies greatly, and the query prediction accuracy and stability of the above method are poor. Summary of the invention
[0005] In view of the above problems, embodiments of the present application provide a method, device, electronic device and readable storage medium for predicting user query behavior based on SQL query, so as to overcome the above problems or at least partially solve the above problems.
[0006] In the first aspect, the embodiment of the present application provides a method for predicting user query behavior based on SQL query, the method comprising; Determine a sample data query information list corresponding to each sample user; Segmenting the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set; Inputting the sample query information training set into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model; Determining a model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information; Based on the model loss value, adjusting the model parameters of the first behavior prediction network model to obtain a target behavior prediction model; Get the SQL query instructions of the target user; The SQL query instruction is input into the target behavior prediction model to obtain the prediction query result output by the target behavior prediction model.
[0007] Optionally, the segmenting of the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set includes: In the sample data query information list, dividing the sample data query information list into a plurality of sample query information sub-lists with a first preset number as a division scale; Determining the sample query information sublist as a sample query information training set; The next piece of sample data query information of the sample query information training set in the sample data query information list is determined as the sample annotation query information of the sample query information training set; wherein, the next piece of sample data query information of the sample annotation query information in the sample data query information list is the first piece of sample data query information of another sample query information sublist.
[0008] Optionally, the model loss value is a cross entropy loss value, and determining the model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information includes: Based on the predicted query result and the sample annotation query information, determining a predicted probability that the predicted query result is a category corresponding to the sample annotation query information; Based on the predicted probability, a cross entropy loss value of the first behavior prediction network model is determined.
[0009] Optionally, the step of inputting the sample query information training set into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model includes: Performing word embedding processing on each piece of sample data query information in the sample query information training set to obtain a first embedding representation; Performing knowledge graph embedding processing on each piece of sample data query information in the sample query information training set to obtain a second embedding representation; Combining the first embedding representation and the second embedding representation respectively corresponding to the pieces of sample data query information to obtain the first fused embedding representation respectively corresponding to the pieces of sample data query information; The first fusion embedding representation is input into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model.
[0010] Optionally, the first behavior prediction network model is composed of a long short-term memory network model and a Transformer-based encoder model, and the first fused embedding representation is input into the first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model, including: Inputting the first fused embedding representation into the long short-term memory network model and the Transformer-based encoder model respectively, to obtain a first prediction feature representation output by the long short-term memory network model and a first feature vector representation output by the Transformer-based encoder model respectively; The first prediction feature representation and the first feature vector representation are concatenated to obtain a prediction query result output by the first behavior prediction network model.
[0011] Optionally, the determining of the sample data query information list corresponding to each sample user includes: Based on a preset regular expression, the data query information in the sample database is screened to obtain a first data query information set; From the first data query information set, a sample data query information list corresponding to each sample user is determined.
[0012] Optionally, the target behavior prediction model is composed of a long short-term memory network model and a Transformer-based encoder model, and the SQL query instruction is input into the target behavior prediction model to obtain a prediction query result set output by the target behavior prediction model, including: Inputting the SQL query instruction into the long short-term memory network model and the Transformer-based encoder model respectively, and obtaining a first prediction feature representation output by the long short-term memory network model and a first feature vector representation output by the Transformer-based encoder model respectively; The first prediction feature representation and the first feature vector representation are concatenated to obtain a prediction query result output by the target behavior prediction network model.
[0013] In a second aspect, an embodiment of the present application provides a user query behavior prediction device based on SQL query, the device comprising: A first determination module is used to determine a sample data query information list corresponding to each sample user; A segmentation module, used to segment the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set; A first input-output module, configured to input the sample query information training set into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model; A second determination module, configured to determine a model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information; An adjustment module, used to adjust the model parameters of the first behavior prediction network model based on the model loss value to obtain a target behavior prediction model; The acquisition module is used to obtain the SQL query instructions of the target user; The second input-output module is used to input the SQL query instruction into the target behavior prediction model to obtain the prediction query result output by the target behavior prediction model.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement a user query behavior prediction method based on SQL query as described in any one of the above.
[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, a user query behavior prediction method based on SQL query as described in any one of the above items is implemented.
[0016] The specific beneficial effects are: The embodiment of the present application determines the sample data query information list corresponding to each sample user, divides the sample data query information list, obtains the sample query information training set and the sample annotation query information corresponding to the sample query information training set, inputs the sample query information training set into the first behavior prediction network model, obtains the prediction query result output by the first behavior prediction network model, determines the model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information, adjusts the model parameters of the first behavior prediction network model based on the model loss value, obtains the target behavior prediction model, obtains the SQL query instruction of the target user, inputs the SQL query instruction into the target behavior prediction model, obtains the prediction query result output by the target behavior prediction model, can train the first behavior prediction network model with the training data corresponding to different sample users, thereby obtaining the target behavior prediction model, and when using the target behavior prediction model to predict the data query of the target user, can obtain the prediction query result matching the target user, can improve the matching degree between the prediction query result and the target user, thereby improving the accuracy and stability of the query prediction to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 It is a flowchart of a method for predicting user query behavior based on SQL query provided in an embodiment of the present application; Figure 2 It is a flowchart of a specific implementation method of a user query behavior prediction method based on SQL query provided in an embodiment of the present application; Figure 3 It is a logic block diagram of a user query behavior prediction device based on SQL query provided by an embodiment of the present application; Figure 4 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to enable the scope of the present application to be fully communicated to those skilled in the art.
[0020] Reference Figure 1 , Figure 1 A flowchart of a method for predicting user query behavior based on SQL query provided in an embodiment of the present application, the method may include: Step 101: determine a sample data query information list corresponding to each sample user.
[0021] In an embodiment of the present application, each user may have his or her own unique data query behavior habits, and these data query behaviors may be stored in a database on a user-by-user basis, and the data in the database may be used as sample data, so that sample data query information corresponding to each sample user may be obtained in the database. Each piece of sample data query information corresponds to a data query behavior of a sample user. At the same time, since each sample user may correspond to multiple pieces of sample data query information, the sample data query information may be sorted according to the generation time of the sample data query information, so that a list of sample data query information corresponding to each sample user may be obtained.
[0022] Optionally, step 101 may include the following sub-steps: Sub-step 1011, based on a preset regular expression, the data query information in the sample database is screened to obtain a first data query information set.
[0023] In an embodiment of the present application, the data query method may be a SQL (Structured Query Language) query. SQL is a special-purpose programming language, a database query and programming language, used to access data and query, update and manage relational database systems. When obtaining data query information in a sample database, a preset regular expression may be used for screening. Data query information in a format that conforms to a preset regular expression may be added to the first data query information set.
[0024] For example, a preset regular expression could be "SELECT <columns>FROM <column> <value> <operation> <column> <value>", where "SELECT <columns> " indicates the name of the column object to be queried, which cannot be omitted; "FROM< / columns> < / value> < / column> < / operation> < / value> < / column> Where
[0025]
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072] Figure 2 Figure 2
[0073] Figure 3 Figure 3
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082] Figure 1 Figure 2
[0083] Figure 4
[0084]
[0085]
[0086]
[0087] Figure 1 Figure 1
[0088] Figure 1 Figure 1
[0089] Figure 1 Figure 1
[0090]
[0091]
[0092] " indicates the name of the data source list, which cannot be omitted; "Where" and the fields after it are query conditions, which can be omitted. According to the above preset regular expression, data query information that meets the expression can be screened out. For example, if a piece of data query information does not contain a column object or a data source list in the regular expression, the piece of data query information will be eliminated. Sub-step 1012, from the first data query information set, determine the sample data query information list corresponding to each sample user. In an embodiment of the present application, the data query information in the first data query information set can be classified according to the sample user, so that the sample data query information corresponding to each sample user can be obtained, and the sample data query information corresponding to each sample user can be generated. Sample data query information list. In an embodiment of the present application, by screening the data query information in the sample database based on a preset regular expression, a first data query information set is obtained, and from the first data query information set, the sample data query information lists corresponding to each sample user are determined, and a sample data query information list matching the preset regular expression can be obtained, which can improve the accuracy and reliability of the data contained in the sample data query information list to a certain extent. Step 102, the sample data query information list is segmented to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set. In an embodiment of the present application, the sample data query information can be segmented. The list is segmented to obtain multiple sub-lists, and a part of the sub-lists can be used as the sample query information training set, and the other part of the sub-lists can be used as the sample annotation query information corresponding to the sample query information training set. Optionally, step 102 may include the following sub-steps; sub-step 1021, in the sample data query information list, the sample data query information list is segmented into multiple sample query information sub-lists with a first preset number as the segmentation scale. In an embodiment of the present application, when the sample data query information list is segmented, the sample data query information list can be segmented into multiple sample query information sub-lists with a first preset number as the segmentation scale. In this way, multiple rows with the first preset number of rows can be obtained. number of sample query information sublists. Sub-step 1022, determining the sample query information sublist as a sample query information training set. In an embodiment of the present application, the sample query information sublist can be determined as a sample query information training set to train the first behavior prediction model through the sample query information training set. Sub-step 1023, determining the next sample data query information of the sample query information training set in the sample data query information list as the sample annotation query information of the sample query information training set; wherein, the next sample data query information of the sample annotation query information in the sample data query information list is the first sample data query information of another sample query information sublist.In an embodiment of the present application, the position of each sample query information training set in the sample data query information list can be determined, and the next sample data query information at the position can be determined as the sample annotation query information of the sample query information training set. At the same time, the sample annotation query information can also be made the next sample data query information in the sample data query information list as the first sample data query information of another sample query information sublist, that is, when segmenting, the first preset number of continuous sample data query information can be used as a sublist, and the first sample data query information after the sublist can be used as the sample annotation query information of the adjacent previously generated sublist. After the sample annotation query information, the first preset number of continuous sample data query information is used as another sublist, and so on, so as to segment the sample data query information list. In an embodiment of the present application, by dividing the sample data query information list into a plurality of sample query information sublists with a first preset number as a segmentation scale in the sample data query information list, determining the sample query information sublist as a sample query information training set, and determining the next sample data query information of the sample query information training set in the sample data query information list as the sample annotation query information of the sample query information training set; wherein the next sample data query information of the sample annotation query information in the sample data query information list is the first sample data query information of another sample query information sublist, the sample data query information list can be accurately segmented according to a certain method, thereby obtaining the sample query information training set and the sample annotation query information, which can improve the matching degree between the sample query information training set and the sample annotation query information to a certain extent, and improve the accuracy and reliability of the sample query information training set and the sample annotation query information. Step 103, input the sample query information training set into the first behavior prediction network model, and obtain the prediction query result output by the first behavior prediction network model. In an embodiment of the present application, the first behavior prediction network model may be some models with semantic learning and reasoning functions, such as a long short-term memory network model (Long Short-Term Memory, LSTM), BERT (Bidirectional Encoder Representations from Transformers), etc. Since the data query information in the sample query information training set may contain the query behavior habits of the sample user, the sample query information training set may be input into the first behavior prediction network model, so that the predicted query result output by the first behavior prediction network model may be obtained, so that the first behavior prediction network model may learn the query behavior habits of the sample user. Optionally, step 103 may include the following sub-steps: Sub-step 1031, word embedding processing is performed on each piece of sample data query information in the sample query information training set to obtain a first embedding representation.In an embodiment of the present application, word embedding is a technique in natural language processing, which is used to map words in a text into a low-dimensional vector space and vectorize them into a real-valued vector. This method embeds words into a continuous vector space so that words with similar semantics are closer in the vector space, thereby better expressing the semantic relationship between words. Each piece of sample data query information in the sample query information training set can be word-embedded to obtain a low-dimensional first embedding representation. Sub-step 1032, each piece of sample data query information in the sample query information training set is subjected to knowledge graph embedding processing to obtain a second embedding representation. In an embodiment of the present application, knowledge graph embedding is a method of mapping entities and relationships into a continuous vector space, which aims to simplify operations and retain structural information in the knowledge graph. Knowledge graph embedding converts entities and relationships in the knowledge graph into low-dimensional dense vectors, so that the originally discrete symbolic knowledge can be represented and calculated in a continuous vector space. This transformation not only retains the structural information in the knowledge graph, but also enables the knowledge graph to be combined with modern machine learning technologies such as deep learning, thereby improving the efficiency and effectiveness of tasks such as knowledge reasoning, completion, and question-answering systems. Each piece of sample data query information in the sample query information training set can be processed by knowledge graph embedding, so that a second embedding representation can be obtained. Sub-step 1033, the first embedding representation and the second embedding representation corresponding to each piece of sample data query information are combined to obtain the first fused embedding representation corresponding to each piece of sample data query information. In an embodiment of the present application, the first embedding representation and the second embedding representation can be expressed in the form of vector encoding, that is, each element in the embedding representation vector is digitally encoded. The first embedding representation and the second embedding representation corresponding to each piece of sample data query information can be combined, so that the first fused embedding representation corresponding to each piece of sample data query information can be obtained. When combined, the second embedding representation can be directly spliced after the first embedding representation to obtain a new first fused embedding representation. Taking into account that there may be default information in the sample data query information, it is also possible to fill the first fused embedded representation that is less than the maximum length of the first fused embedded representation with 0 elements, so that the fused embedded representation can reflect the characteristics corresponding to this default query method. Sub-step 1034, input the first fused embedded representation into the first behavior prediction network model to obtain the predicted query result output by the first behavior prediction network model. In an embodiment of the present application, the first fused embedded representation can be input into the first behavior prediction network model, so as to obtain the predicted query result output by the first behavior prediction network model.Optionally, sub-step 1034 may include the following sub-steps: Sub-step A1, inputting the first fused embedding representation into the long short-term memory network model and the Transformer-based encoder model respectively, and obtaining the first prediction feature representation output by the long short-term memory network model and the first feature vector representation output by the Transformer-based encoder model respectively. In an embodiment of the present application, the first behavior prediction network model may be composed of a combination of a long short-term memory network model and a Transformer-based encoder model. When combined, only the input ports and output ports of the long short-term memory network model and the Transformer-based encoder model may be connected, while maintaining the model structure and model parameters of the long short-term memory network model and the Transformer-based encoder model themselves. The long short-term memory network model can learn the semantic information of the fused embedding representation, and the Transformer-based encoder model can learn more attention-worthy information in the first fused embedding representation through the self-attention mechanism. On this basis, the first fused embedded representation can be input into the long short-term memory network model and the Transformer-based encoder model of the first behavior prediction network model respectively, so as to obtain the first prediction feature representation output by the long short-term memory network model and the first feature vector representation output by the Transformer-based encoder model respectively. Sub-step A2, splicing the first prediction feature representation and the first feature vector representation to obtain the prediction query result output by the first behavior prediction network model. In an embodiment of the present application, the first prediction feature representation and the first feature vector representation can be spliced to obtain the prediction query result output by the first behavior prediction network model. Among them, the first feature vector representation can be directly spliced after the first prediction feature representation, and after splicing, the spliced result can be decoded to obtain the prediction query result. In this way, the data format of the prediction query result obtained after decoding can be consistent with the data format of the sample data query information.In an embodiment of the present application, by respectively inputting the first fused embedding representation into a long short-term memory network model and a Transformer-based encoder model, a first prediction feature representation output by the long short-term memory network model and a first feature vector representation output by the Transformer-based encoder model are respectively obtained, the first prediction feature representation and the first feature vector representation are spliced to obtain a prediction query result output by the first behavior prediction network model, the first fused embedding representation can be respectively processed by the long short-term memory network model and the Transformer-based encoder model, and the splicing result obtained by splicing the respectively output first prediction feature representation and the first feature vector representation can be used as the prediction query result. Since the long short-term memory network model and the Transformer-based encoder model can respectively extract the semantic features and attention features of the input data, the accuracy of the prediction query results can be improved to a certain extent. In an embodiment of the present application, by performing word embedding processing on each piece of sample data query information in the sample query information training set, a first embedding representation is obtained, and a knowledge graph embedding processing is performed on each piece of sample data query information in the sample query information training set to obtain a second embedding representation. The first embedding representation and the second embedding representation corresponding to each piece of sample data query information are combined to obtain the first fused embedding representation corresponding to each piece of sample data query information. The first fused embedding representation is input into the first behavior prediction network model to obtain the predicted query result output by the first behavior prediction network model. The semantic features and associated features of the input data can be extracted using the word embedding and knowledge graph embedding methods, so that the above features can be reflected in the first fused embedding representation in a more significant way, which can improve the training efficiency of the first behavior prediction network model to a certain extent. Step 104, based on the predicted query result and the sample annotation query information, the model loss value of the first behavior prediction network model is determined. In an embodiment of the present application, the model loss value of the first behavior prediction network model can be calculated based on the predicted query result and the sample annotation query information. Among them, the model loss value can be a cross entropy loss value or a contrast loss value (i.e., the cosine similarity between the predicted query result and the sample annotation query information). Optionally, step 104 may include the following sub-steps: Sub-step 1041, based on the predicted query result and the sample annotation query information, determining the predicted probability that the predicted query result is the category corresponding to the sample annotation query information. In an embodiment of the present application, the model loss value may be a cross entropy loss value, and the information similarity between the predicted query result and the sample annotation query information corresponding to the input sample query information training set data may be calculated based on the predicted query result and the sample annotation query information, and the information similarity may be used as the predicted probability that the predicted query result is the category corresponding to each sample annotation query information.Sub-step 1042, based on the predicted probability, determine the cross entropy loss value of the first behavior prediction network model. In an embodiment of the present application, the cross entropy loss value of the first behavior prediction network model can be determined based on the predicted probability, and the calculation method is shown in Formula 1 below: (Formula 1) In Formula 1, i represents the number of the training data contained in the sample query information training set, indicating that the predicted query result corresponding to the training data i is inconsistent with the sample annotation query information of the training data, and 1 represents that the predicted query result corresponding to the training data i is consistent with the sample annotation query information of the training data, which is the predicted probability corresponding to the training data i. In an embodiment of the present application, by determining the predicted probability that the predicted query result is the corresponding category of the sample annotation query information based on the predicted query result and the sample annotation query information, and determining the cross entropy loss value of the first behavior prediction network model based on the predicted probability, the cross entropy loss value of the first behavior prediction network model can be calculated as the model loss value, which can improve the accuracy and reliability of the model loss value to a certain extent. Step 105, based on the model loss value, adjust the model parameters of the first behavior prediction network model to obtain the target behavior prediction model. In an embodiment of the present application, the model parameters of the first behavior prediction network model can be adjusted according to the model loss value of the first behavior prediction network model, so that the target behavior prediction model can be obtained. The direction of adjusting the model parameters can be the direction in which the model loss value is reduced. The process of adjusting the model parameters can be repeated for multiple times, and after each adjustment, the first behavior prediction network model can be repeatedly trained to adjust the model parameters again. When the model loss value meets the preset convergence condition, or the number of training times for the first behavior prediction network model reaches the preset number of times, the training of the first behavior prediction network model can be stopped, and the first behavior prediction network model obtained after the last adjustment of the model parameters is determined as the target behavior prediction model. The target behavior prediction model can be used to output the corresponding prediction query result based on the SQL query instruction of the target user. Step 106, obtain the SQL query instruction of the target user. In an embodiment of the present application, the SQL query instruction sent by the target user can be obtained. The SQL query instruction can contain all information types in the sample data query information. Step 107, input the SQL query instruction into the target behavior prediction model to obtain the prediction query result output by the target behavior prediction model. In an embodiment of the present application, the SQL query instruction sent by the target user can be input into the target behavior prediction model, so as to obtain the prediction query result output by the target behavior prediction model. The target behavior prediction model can be obtained by any of the above-mentioned user query behavior prediction methods based on SQL query.Optionally, step 107 may include the following sub-steps: Sub-step 1071, inputting the SQL query instruction into the long short-term memory network model and the Transformer-based encoder model respectively, and obtaining the first prediction feature representation output by the long short-term memory network model and the first feature vector representation output by the Transformer-based encoder model respectively. In an embodiment of the present application, the SQL query instruction may be input into the long short-term memory network model and the Transformer-based encoder model contained in the target behavior prediction model respectively, so as to obtain the first prediction feature representation output by the long short-term memory network model and the first feature vector representation output by the Transformer-based encoder model respectively. Sub-step 1072, splicing the first prediction feature representation and the first feature vector representation to obtain the prediction query result output by the target behavior prediction network model. In an embodiment of the present application, the first prediction feature representation and the first feature vector representation may be spliced to obtain the prediction query result output by the target behavior prediction network model. When splicing, the first feature vector representation may be spliced behind the first prediction feature representation. In an embodiment of the present application, by respectively inputting SQL query instructions into a long short-term memory network model and a Transformer-based encoder model, a first prediction feature representation output by the long short-term memory network model and a first feature vector representation output by the Transformer-based encoder model are obtained, and the first prediction feature representation and the first feature vector representation are spliced to obtain a prediction query result output by the target behavior prediction network model. The prediction query result can be obtained by splicing the first prediction feature representation output by the long short-term memory network model part of the target behavior prediction model and the first feature vector representation output by the Transformer-based encoder model part, which can improve the data visualization of the prediction query result to a certain extent.In an embodiment of the present application, by determining the sample data query information list corresponding to each sample user, the sample data query information list is segmented to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set, the sample query information training set is input into the first behavior prediction network model, and the prediction query result output by the first behavior prediction network model is obtained. Based on the prediction query result and the sample annotation query information, the model loss value of the first behavior prediction network model is determined, and based on the model loss value, the model parameters of the first behavior prediction network model are adjusted to obtain the target behavior prediction model, the SQL query instruction of the target user is obtained, and the SQL query instruction is input into the target behavior prediction model to obtain the prediction query result output by the target behavior prediction model. The first behavior prediction network model can be trained by the training data corresponding to different sample users, so as to obtain the target behavior prediction model. When the target behavior prediction model is used to predict the data query of the target user, the prediction query result matching the target user can be obtained, and the matching degree between the prediction query result and the target user can be improved, so as to improve the accuracy and stability of the query prediction to a certain extent. Referring to the flowchart of a specific implementation method of a user query behavior prediction method based on SQL query provided in an embodiment of the present application. In the figure, the first behavior prediction model can be trained using training data. During training, the training data can be first processed by word embedding and knowledge graph embedding, and then the embedded representations obtained after processing can be fused to obtain a fused embedded representation. Next, the fused embedded representation can be input into the long short-term memory network model part and the Transformer encoder model part of the first behavior prediction model, and the predicted feature representation output by the long short-term memory network model and the feature vector representation output by the Transformer encoder model can be obtained respectively. After that, the predicted feature representation and the feature vector representation can be fused, and the model loss value of the first behavior prediction model can be calculated by the prediction query result obtained after fusion. Finally, the model parameters of the long short-term memory network model and the model parameters of the Transformer encoder model in the first behavior prediction model can be adjusted by the model loss value to complete a training process of the first behavior prediction model. After the training process of the first behavior prediction model is completed, the target behavior prediction model can be obtained. The SQL query instructions sent by the target user can be input into the target behavior prediction model. The SQL query instructions are processed by the long short-term memory network model part and the Transformer encoder model part of the target behavior prediction model respectively, and the outputs of the above two model parts are fused to obtain the predicted query result corresponding to the SQL query instruction.Referring to the logical block diagram of a user query behavior prediction device based on SQL query provided in an embodiment of the present application, the device 300 may include: a first determination module 301, used to determine a list of sample data query information corresponding to each sample user; a segmentation module 302, used to segment the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set; a first input-output module 303, used to input the sample query information training set into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model; a second determination module 304, used to determine a model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information; an adjustment module 305, used to adjust the model parameters of the first behavior prediction network model based on the model loss value to obtain a target behavior prediction model; an acquisition module 306, used to obtain an SQL query instruction of a target user; a second input-output module 307, used to input the SQL query instruction into the target behavior prediction model to obtain a prediction query result output by the target behavior prediction model. Optionally, the segmentation module 302 includes: a segmentation submodule, used to segment the sample data query information list into multiple sample query information sublists with a first preset number as the segmentation scale in the sample data query information list; a first determination submodule, used to determine the sample query information sublist as a sample query information training set; a second determination submodule, used to determine the next sample data query information of the sample query information training set in the sample data query information list as the sample annotation query information of the sample query information training set; wherein the next sample data query information of the sample annotation query information in the sample data query information list is the first sample data query information of another sample query information sublist. Optionally, the model loss value is a cross entropy loss value, and the second determination module 304 includes: a third determination submodule, used to determine the predicted probability that the predicted query result is the corresponding category of the sample annotation query information based on the predicted query result and the sample annotation query information; and a fourth determination submodule, used to determine the cross entropy loss value of the first behavior prediction network model based on the predicted probability.Optionally, the first input-output module 303 includes: a word embedding submodule, used to perform word embedding processing on each piece of sample data query information in the sample query information training set to obtain a first embedding representation; a knowledge graph embedding submodule, used to perform knowledge graph embedding processing on each piece of sample data query information in the sample query information training set to obtain a second embedding representation; a combination submodule, used to combine the first embedding representation and the second embedding representation corresponding to each piece of sample data query information, respectively, to obtain a first fused embedding representation corresponding to each piece of sample data query information; an input-output submodule, used to input the first fused embedding representation into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model. Optionally, the first behavior prediction network model is composed of a long short-term memory network model and a transformer-based encoder model, and the input-output submodule includes: an input-output unit, which is used to input the first fusion embedding representation into the long short-term memory network model and the transformer-based encoder model respectively, and obtain the first prediction feature representation output by the long short-term memory network model and the first feature vector representation output by the transformer-based encoder model respectively; a splicing unit, which is used to splice the first prediction feature representation and the first feature vector representation to obtain the prediction query result output by the first behavior prediction network model. Optionally, the first determination module 301 includes: a screening submodule, which is used to screen the data query information in the sample database based on a preset regular expression to obtain a first data query information set; a fifth determination submodule, which is used to determine the sample data query information list corresponding to each sample user from the first data query information set. Optionally, the second input-output module 307 includes: an input-output submodule, which is used to input the SQL query instruction into the long short-term memory network model and the Transformer-based encoder model, respectively, to obtain the first prediction feature representation output by the long short-term memory network model and the first feature vector representation output by the Transformer-based encoder model, respectively; a splicing submodule, which is used to splice the first prediction feature representation and the first feature vector representation to obtain the prediction query result output by the target behavior prediction network model. The user query behavior prediction device based on SQL query in the embodiment of the present application can be an electronic device, or it can be a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be a device other than a terminal.Exemplarily, the electronic device may be a GPUBOX, a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It may also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiments of the present application. The user query behavior prediction device based on SQL query in the embodiments of the present application may be a device having an operating system. The operating system may be an Android operating system, a Linux, Windows operating system, etc., or other possible operating systems, which are not specifically limited in the embodiments of the present application. The user query behavior prediction device based on SQL query provided in the embodiments of the present application can implement the various processes implemented in the embodiments of the method, and to avoid repetition, they are not described here. The embodiment of the present application provides an electronic device, see, the electronic device 40 includes: a processor 401, a memory 402, and a computer program 4021 stored on the memory 402 and executable on the processor 401, and the processor 401 implements the user query behavior prediction method based on SQL query of the aforementioned embodiment when executing the program. The embodiment of the present application also provides a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by the processor, the steps in the user query behavior prediction method based on SQL query disclosed in the embodiment of the present application are implemented. The embodiment of the present application also provides a computer program product, when the computer program product is run on the electronic device, the processor implements the steps in the user query behavior prediction method based on SQL query disclosed in the embodiment of the present application when executing. Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the embodiments can be referred to each other. The embodiment of the present application is described with reference to the flowchart and / or block diagram of the method, device, electronic device and computer program product according to the embodiment of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and a combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions.These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for implementing the functions specified in the process, process or multiple processes and / or box, box or multiple boxes. These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process, process or multiple processes and / or box, box or multiple boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in the process, process or multiple processes and / or box, box or multiple boxes. Although the preferred embodiments of the present application have been described, once the basic creative concept is known to those skilled in the art, additional changes and modifications can be made to these embodiments. Therefore, the attached claims are intended to be interpreted as including preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application. Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the statement "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. The above is a detailed introduction to a method and device for predicting user query behavior based on SQL query provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application. < / columns>
Claims
1. A method for predicting user query behavior based on SQL query, characterized in that: The method comprises: Determine a sample data query information list corresponding to each sample user; Segmenting the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set; Inputting the sample query information training set into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model; Determining a model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information; Based on the model loss value, adjusting the model parameters of the first behavior prediction network model to obtain a target behavior prediction model; Get the SQL query instructions of the target user; The SQL query instruction is input into the target behavior prediction model to obtain the prediction query result output by the target behavior prediction model.
2. The method according to claim 1, characterized in that The step of segmenting the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set includes: In the sample data query information list, dividing the sample data query information list into a plurality of sample query information sub-lists with a first preset number as a division scale; Determining the sample query information sublist as a sample query information training set; The next piece of sample data query information of the sample query information training set in the sample data query information list is determined as the sample annotation query information of the sample query information training set; wherein, the next piece of sample data query information of the sample annotation query information in the sample data query information list is the first piece of sample data query information of another sample query information sublist.
3. The method according to claim 1, characterized in that The model loss value is a cross entropy loss value, and determining the model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information includes: Based on the predicted query result and the sample annotation query information, determining a predicted probability that the predicted query result is a category corresponding to the sample annotation query information; Based on the predicted probability, a cross entropy loss value of the first behavior prediction network model is determined.
4. The method according to claim 1, characterized in that The step of inputting the sample query information training set into the first behavior prediction network model to obtain the prediction query result output by the first behavior prediction network model includes: Performing word embedding processing on each piece of sample data query information in the sample query information training set to obtain a first embedding representation; Performing knowledge graph embedding processing on each piece of sample data query information in the sample query information training set to obtain a second embedding representation; Combining the first embedding representation and the second embedding representation respectively corresponding to the pieces of sample data query information to obtain the first fused embedding representation respectively corresponding to the pieces of sample data query information; The first fusion embedding representation is input into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model.
5. The method according to claim 4, characterized in that The first behavior prediction network model is composed of a long short-term memory network model and a Transformer-based encoder model, and the first fusion embedding representation is input into the first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model, including: Inputting the first fused embedding representation into the long short-term memory network model and the Transformer-based encoder model respectively, to obtain a first prediction feature representation output by the long short-term memory network model and a first feature vector representation output by the Transformer-based encoder model respectively; The first prediction feature representation and the first feature vector representation are concatenated to obtain a prediction query result output by the first behavior prediction network model.
6. The method according to claim 1, characterized in that The step of determining the sample data query information list corresponding to each sample user includes: Based on a preset regular expression, the data query information in the sample database is screened to obtain a first data query information set; From the first data query information set, a sample data query information list corresponding to each sample user is determined.
7. The method according to claim 1, characterized in that The target behavior prediction model is composed of a long short-term memory network model and a Transformer-based encoder model. The SQL query instruction is input into the target behavior prediction model to obtain a prediction query result set output by the target behavior prediction model, including: Inputting the SQL query instruction into the long short-term memory network model and the Transformer-based encoder model respectively, and obtaining a first prediction feature representation output by the long short-term memory network model and a first feature vector representation output by the Transformer-based encoder model respectively; The first prediction feature representation and the first feature vector representation are concatenated to obtain a prediction query result output by the target behavior prediction network model.
8. A user query behavior prediction device based on SQL query, characterized in that: The device comprises: A first determination module is used to determine a sample data query information list corresponding to each sample user; A segmentation module, used to segment the sample data query information list to obtain a sample query information training set and sample annotation query information corresponding to the sample query information training set; A first input-output module, configured to input the sample query information training set into a first behavior prediction network model to obtain a prediction query result output by the first behavior prediction network model; A second determination module, configured to determine a model loss value of the first behavior prediction network model based on the prediction query result and the sample annotation query information; An adjustment module, used to adjust the model parameters of the first behavior prediction network model based on the model loss value to obtain a target behavior prediction model; The acquisition module is used to obtain the SQL query instructions of the target user; The second input-output module is used to input the SQL query instruction into the target behavior prediction model to obtain the prediction query result output by the target behavior prediction model.
Citation Information
Patent Citations
Knowledge map optimal path query system and method based on depth reinforcement learning
CN109241291A
Model training method, information prompting method and related equipment
CN118820312A
Small-sample knowledge graph reasoning method based on graph neural network
CN118916494A
Object detection method and apparatus, device, and storage medium
WO2024183181A1