Method and related device for predicting HQL execution statement based on user characteristics

Through the HQL execution statement prediction method based on user characteristics, the problems of cumbersome data processing and large code volume in the financial technology field are solved, and light code and automated data processing are realized, reducing the time consumption of repeated construction.

CN115495471BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211163367.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-06-24
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

In the field of financial technology, especially in insurance and banking, due to the large number of user groups and complex data volume, the data processing process is cumbersome, long links and huge code volume. It requires processing of data from the data warehouse layer to the data refinement processing layer to the data mart, involving a large amount of code writing and consuming a lot of manpower.

Method used

A method for predicting HQL execution statements based on user characteristics is proposed. By obtaining source data and storing it in the data fact table and dimension table, the correlation topology diagram between tables is constructed, the optimal HQL execution statement is assembled based on user behavior characteristics, and the prediction model is pre-trained to achieve the optimal HQL execution statement prediction for new users.

Benefits of technology

It realizes light code-based and automated data processing, reduces the time and consumption of repeated construction, and improves the efficiency and automation of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115495471B_ABST
    Figure CN115495471B_ABST
Patent Text Reader

Abstract

The embodiments of the present application belong to the fields of artificial intelligence and big data, and are applied to the field of data warehouse construction. It relates to a method for predicting HQL execution statements based on user characteristics and related devices, including storing source data into each data fact table and each data dimension table; obtaining foreign key fields and primary key fields of each data fact table, and obtaining the inter-table distance between each data fact table through preset algorithm rules, foreign key fields and primary key fields; constructing an inter-table association topology graph; assembling the optimal HQL execution statement; obtaining the optimal HQL execution statements of several first users for pre-training of the prediction model; receiving an optimal HQL execution statement assembly request sent by a second user, and obtaining the optimal HQL execution statement of the second user based on the pre-trained prediction model. This solution realizes lightweight coding and automated data processing, binds this automation with user characteristics, generates based on user-driven data processing models, and reduces the time consumption of repeated construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of artificial intelligence, big data, and data warehouse construction, and particularly relates to a method for predicting HQL execution statements based on user characteristics and related devices. Background Art

[0002] As time goes by, the amount of business data is getting larger and larger. We can no longer directly query and count our metrics. So data layering emerged. From the detail layer to the summary layer and then to the deep summary layer, it is processed layer by layer like this, involving a lot of table associations and data processing steps to obtain the result set of the small amount of data we need. Finally, this part of the data will be synchronized to the relational database to improve query capabilities.

[0003] Especially in the field of fintech, such as insurance business and banking business, the user groups involved are often large in quantity and the data is also quite complex. This leads to a very cumbersome and long-link process when processing such data using the manual processing mode, with a huge amount of code. Currently, the implementation of personalized data applications often requires processing data layer by layer from the data warehouse layer (DWD) to the data fine processing layer (DWS) and then to the data mart (DMI), involving a large amount of code writing and consuming a lot of manpower. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose a method for predicting HQL execution statements based on user characteristics and related devices, so as to facilitate lightweight coding and automated data processing, bundle this automation with user characteristics, generate based on a user-driven data processing model, and reduce the time consumption of repeated construction.

[0005] To solve the above technical problems, the embodiments of the present application provide a method for predicting HQL execution statements based on user characteristics, and adopt the following technical solutions:

[0006] A method for predicting HQL execution statements based on user characteristics includes the following steps:

[0007] Obtain source data, store the source data according to a preset modeling method and storage rules, and store the source data into corresponding data fact tables and data dimension tables, where the preset modeling methods include star schema, snowflake schema, and galaxy schema;

[0008] Obtain the primary key fields of each of the data fact tables, and obtain the foreign key fields from the corresponding data dimension tables, and obtain the inter-table distances between the data fact tables through preset algorithm rules, the foreign key fields, and the primary key fields;

[0009] Construct an inter-table association topology graph for each data fact table based on the inter-table distance and the primary key field, where the inter-table association topology graph uses the primary key field as the topology node;

[0010] Assemble the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph, where the behavior characteristics include the HQL execution statement assembly request issued by the user;

[0011] Obtain the optimal HQL execution statements corresponding to a number of first users, use the optimal HQL execution statements corresponding to the number of first users as a result set, obtain the portrait features corresponding to the number of first users based on a preset feature form, use the portrait features as an input set, and perform pre-training of the prediction model;

[0012] Receive the HQL execution statement assembly request issued by the second user, and obtain the portrait features corresponding to the second user based on the feature form;

[0013] Input the portrait features into the pre-trained prediction model to obtain the optimal HQL execution statement corresponding to the second user.

[0014] Further, the step of storing the source data into the corresponding data fact tables and data dimension tables according to the preset modeling method and storage rules specifically includes:

[0015] Classify the source data according to the preset data type classification rules, and divide the unit data in the source data into two categories: fact table data and dimension table data;

[0016] Generate corresponding data fact tables and data dimension tables in the preset data storage area according to the corresponding table creation instructions;

[0017] Store the unit data with the classification category of the fact table data into the corresponding data fact table according to the storage rules;

[0018] Store the unit data with the classification category of the dimension table data into the corresponding data dimension table according to the storage rules;

[0019] Among them, the storage rules specifically include: the primary key field of the current data fact table, the data index field corresponding to the primary key field, and the dimension information of the data index field are stored in the data fact table, the primary key fields of each data fact table and the dimension information of the primary key field are stored in the data dimension table, each primary key field corresponds to a data fact table, and the column used to store the primary key field of the data fact table in the data dimension table is used as the foreign key field.

[0020] Further, the step of obtaining the primary key fields of each of the data fact tables, obtaining the foreign key fields from the corresponding data dimension tables, and obtaining the inter-table distances between the data fact tables through a preset algorithm rule, the foreign key fields, and the primary key fields specifically includes:

[0021] Step A: Arbitrarily select one of the data fact tables among each of the data fact tables, obtain the primary key field of the data fact table, and set the primary key field as the distance point i n ;

[0022] Step B: Query all the data dimension tables that store the current primary key field based on the current primary key field;

[0023] Step C: Obtain the foreign key fields used as the primary key fields of the data fact tables and the dimension information of the foreign key fields in all the data dimension tables;

[0024] Step D: According to the foreign key fields and the dimension information of the foreign key fields, obtain the upper-level or lower-level primary key fields associated with the current primary key field, and set the primary key field as the distance point i n+1 , and update the primary key field to the current primary key field;

[0025] Step E: Execute Step B to Step D in a loop, update the distance point for the obtained primary key fields until there is no associated lower-level primary key field for the current primary key field, where the initial value of i n is 0, and i n+1 = i n + 1.

[0026] Further, the step of constructing an inter-table association topology graph for each of the data fact tables based on the inter-table distances and the primary key fields specifically includes:

[0027] Obtain and traverse the distance points corresponding to the primary key fields of each of the data fact tables;

[0028] Using the primary key field with a distance point of 0 as the starting topology node, sequentially search for the corresponding primary key fields as the next-level topology nodes in ascending order of the distance points until the topology node corresponding to the primary key field when the distance point is the maximum value is found, and sequentially connect all the found topology nodes to complete the construction of the inter-table association topology graph.

[0029] Further, it is characterized in that the step of assembling the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph specifically includes:

[0030] Receive an HQL execution statement assembly request sent by the first user. Among them, the HQL execution statement assembly request includes the target data index information required by the first user. Among them, the index information is the primary key field of the data fact table corresponding to the target data;

[0031] Parse the HQL execution statement assembly request, parse out the target data index information, and assemble and obtain the optimal HQL execution statement for the target data according to the inter-table association topology graph and the target data index information;

[0032] Use the optimal HQL execution statement as the return value and return it to the first user.

[0033] Furthermore, the step of assembling and obtaining the optimal HQL execution statement for the target data according to the inter-table association topology graph and the target data index information specifically includes:

[0034] Step a: Obtain the primary key field of the data fact table for caching user information;

[0035] Step b: Find the topology node corresponding to the primary key field from the inter-table association topology graph, mark it as the first topology node, and obtain the distance point corresponding to the first topology node, denoted as the first distance point;

[0036] Step c: According to the primary key field of the data fact table of the target data, find the topology node corresponding to the primary key field from the inter-table association topology graph, mark it as the target topology node, and obtain the distance point corresponding to the target topology node, denoted as the second distance point;

[0037] Step d: According to the inter-table association topology graph, construct a topology node path from the first distance point to the second distance point, and obtain the number of nodes included in the topology node path;

[0038] Step e: Determine whether the difference between the first distance point and the second distance point satisfies the preset equation formula: |i a -i b | = c - 1, where i a is the first distance point, i b is the second distance point, |i a -i b | is the absolute value corresponding to the difference between the first distance point and the second distance point, and c is the number of nodes included in the topology node path;

[0039] Step f: If the difference between the first distance point and the second distance point satisfies the preset equation formula, the determination of the shortest path is completed; otherwise, repeat steps d to e;

[0040] Step g: After the determination of the shortest path is completed, traverse to obtain all topological nodes corresponding to the shortest path, and identify the primary key fields corresponding to all the topological nodes according to the inter-table association topological graph. Assemble the HQL execution statements between multiple tables through the primary key fields, and use the assembled HQL execution statements between multiple tables as the optimal HQL execution statement.

[0041] Further, the step of using the optimal HQL execution statements corresponding to the several first users as a result set, obtaining portrait features corresponding to the several first users based on a preset feature form, and using the portrait features as an input set for pre-training a prediction model specifically includes:

[0042] Using the portrait features corresponding to the several first users as prediction conditions, where the portrait features include the data operation departments to which the users belong, and the feature form includes information on the data operation departments corresponding to each user;

[0043] Using the optimal HQL execution statements corresponding to the several first users as prediction results;

[0044] Constructing a one-to-one correspondence between the prediction conditions and the prediction results according to a preset user identifier, and generating an associated binary tuple, where the data format of the associated binary tuple is: [prediction condition: prediction result];

[0045] Inputting the prediction conditions and the prediction results into an initial prediction model, and performing grouped annotation on the corresponding prediction conditions and prediction results according to the associated binary tuple;

[0046] Obtaining a grouped annotation result, and obtaining all prediction results corresponding to the prediction conditions under the same prediction conditions according to the grouped annotation result;

[0047] Statistically analyzing all the prediction results, and screening out the same prediction result with the largest probability value among all the prediction results as the model output result under the prediction conditions, thereby obtaining a pre-trained prediction model.

[0048] To solve the above technical problems, an embodiment of the present application further provides a device for predicting HQL execution statements based on user characteristics, which adopts the following technical solutions:

[0049] A device for predicting HQL execution statements based on user characteristics includes:

[0050] A source data storage module, configured to obtain source data, store the source data according to a preset modeling method and storage rules, and store the source data into corresponding data fact tables and data dimension tables, where the preset modeling methods include star schema, snowflake schema, and constellation schema;

[0051] An inter-table distance acquisition module, configured to acquire the primary key fields of each of the data fact tables, and acquire the foreign key fields from the corresponding data dimension tables, and obtain the inter-table distances between the data fact tables through a preset algorithm rule, the foreign key fields, and the primary key fields;

[0052] An inter-table association topology graph construction module, configured to construct an inter-table association topology graph for each of the data fact tables based on the inter-table distances and the primary key fields, wherein the inter-table association topology graph uses the primary key fields as topology nodes;

[0053] An optimal HQL execution statement construction module, configured to assemble the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph, wherein the behavior characteristics include an HQL execution statement assembly request sent by the user;

[0054] A prediction model pre-training module, configured to acquire the optimal HQL execution statements corresponding to a plurality of first users, use the optimal HQL execution statements corresponding to the plurality of first users as a result set, acquire the portrait features corresponding to the plurality of first users based on a preset feature form, and use the portrait features as an input set to perform pre-training of a prediction model;

[0055] An input field acquisition module, configured to receive an HQL execution statement assembly request sent by a second user, and acquire the portrait features corresponding to the second user based on the feature form;

[0056] A prediction model prediction module, configured to input the portrait features into a pre-trained prediction model to obtain the optimal HQL execution statement corresponding to the second user.

[0057] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0058] A computer device, including a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the method for predicting an HQL execution statement based on user characteristics as described above are implemented.

[0059] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0060] A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the steps of the method for predicting an HQL execution statement based on user characteristics as described above are implemented.

[0061] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0062] In the method for predicting HQL execution statements based on user characteristics according to the embodiments of the present application, source data is stored in corresponding data fact tables and data dimension tables; foreign key fields and primary key fields of each data fact table are obtained, and the inter-table distances between each data fact table are obtained through preset algorithm rules, foreign key fields, and primary key fields; an inter-table association topology graph is constructed; an optimal HQL execution statement is assembled; the optimal HQL execution statements of several first users are obtained for pre-training of a prediction model; and a request for assembling the optimal HQL execution statement sent by a second user is received, and the optimal HQL execution statement of the second user is obtained based on the pre-trained prediction model. The present invention combines the idea of a dimension model, realizes direct access to data from the source ODS layer, constructs an optimal HQL execution statement assembly model, and then predicts the optimal HQL execution statement corresponding to a new user through a prediction model, realizes data processing in a light-code and automated manner, binds this automation with user characteristics, generates a data processing model driven by users, and reduces the time consumption during repeated construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the following-described drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 is an exemplary system architecture diagram in which the present application can be applied;

[0065] Figure 2 A flowchart of an embodiment of the method for predicting HQL execution statements based on user characteristics according to the present application;

[0066] Figure 3 is Figure 2 a flowchart of a specific implementation manner of step 201 shown;

[0067] Figure 4 is Figure 2 a flowchart of a specific implementation manner of step 202 shown;

[0068] Figure 5 is Figure 2 a flowchart of a specific implementation manner of step 204 shown;

[0069] Figure 6 is Figure 2 a flowchart of a specific implementation manner of step 205 shown;

[0070] Figure 7 Structural schematic diagram of an embodiment of an apparatus for predicting HQL execution statements based on user characteristics according to the present application;

[0071] Figure 8 Structural schematic diagram of an embodiment of a computer device according to the present application. Detailed implementation manners

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0073] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0074] To enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0075] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0076] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0077] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.

[0078] The server 105 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal devices 101, 102, and 103.

[0079] It should be noted that the method for predicting HQL execution statements based on user characteristics provided in the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the device for predicting HQL execution statements based on user characteristics is generally set in the server / terminal device.

[0080] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0081] Continuing to refer to Figure 2 , a flowchart of an embodiment of the method for predicting HQL execution statements based on user characteristics according to the present application is shown. The method for predicting HQL execution statements based on user characteristics includes the following steps:

[0082] Step 201, obtain source data, store the source data according to a preset modeling method and storage rules, and store the source data into corresponding data fact tables and data dimension tables, where the preset modeling method includes star schema, snowflake schema, and constellation schema.

[0083] In this embodiment, the step of storing the source data according to the preset modeling method and storage rules and storing the source data into the corresponding data fact tables and data dimension tables specifically includes: classifying the source data according to the preset data type classification rules, and dividing the unit data in the source data into two categories: fact table data and dimension table data; generating corresponding data fact tables and data dimension tables respectively in the preset data storage area according to the corresponding table creation instructions; storing the unit data with the classification type of the fact table data into the corresponding data fact tables according to the storage rules; storing the unit data with the classification type of the dimension table data into the corresponding data dimension tables according to the storage rules, where the storage rules specifically include: storing the primary key field of the current data fact table, the data index field corresponding to the primary key field, and the dimension information of the data index field in the data fact table, storing the primary key fields of each data fact table and the dimension information of the primary key fields in the data dimension table, each primary key field corresponds to a data fact table, and the column for storing the primary key field of the data fact table in the data dimension table is the foreign key field.

[0084] Taking the hospital data platform as an example, the hospital involves personnel information of various departments, as well as data such as charges, expenditures, and medical consumable costs of various departments, which is relatively complicated. At this time, if you want to manage the above data in a unified manner, you need to create a data warehouse, store the expenses and personnel information into several data fact tables, and store the names of each department into several data dimension tables as foreign keys for querying the corresponding personnel information and expenses.

[0085] By obtaining the source data, storing the source data according to the preset modeling method and storage rules, and storing the source data into the corresponding data fact tables and data dimension tables, it is convenient to build a data warehouse and manage data in a unified manner on the big data platform, and assist personnel in each department to quickly process data.

[0086] Continue to refer to Figure 3 , Figure 3 is Figure 2 a flowchart of a specific implementation manner of step 201 shown in

[0087] Step 301, classifying the source data according to the preset data type classification rules, and dividing the unit data in the source data into two categories: fact table data and dimension table data;

[0088] Step 302, generating corresponding data fact tables and data dimension tables respectively in the preset data storage area according to the corresponding table creation instructions;

[0089] Step 303, storing the unit data with the classification type of the fact table data into the corresponding data fact tables according to the storage rules;

[0090] Step 304: Store the unit data with the classification type of the dimension table data into the corresponding data dimension table according to the storage rule.

[0091] In this embodiment, the storage of the source data according to the preset modeling method and storage rule is in the stage of directly accessing the data from the ODS layer at the data source end during the construction of the data warehouse.

[0092] Step 202: Obtain the primary key fields of each of the data fact tables, and obtain the foreign key fields from the corresponding data dimension tables, and obtain the inter-table distances between the data fact tables through a preset algorithm rule, the foreign key fields, and the primary key fields.

[0093] In this embodiment, the step of obtaining the primary key fields of each of the data fact tables, obtaining the foreign key fields from the corresponding data dimension tables, and obtaining the inter-table distances between the data fact tables through a preset algorithm rule, the foreign key fields, and the primary key fields specifically includes: Step A: Arbitrarily select one of the data fact tables among each of the data fact tables, obtain the primary key field of the data fact table, and set the primary key field as the distance point i n ; Step B: Query all the data dimension tables that store the current primary key field based on the current primary key field; Step C: Obtain the foreign key fields used as the primary key fields of the data fact tables and the dimension information of the foreign key fields in all the data dimension tables; Step D: According to the foreign key fields and the dimension information of the foreign key fields, obtain the upper-level or lower-level primary key fields associated with the current primary key field, and set the primary key field as the distance point i n+1 , update the primary key field to the current primary key field; Step E: Execute Steps B to D in a loop manner, update the distance point for the obtained primary key fields until there is no associated lower-level primary key field for the current primary key field, where the initial value of i n is 0, and i n+1 = i n + 1.

[0094] Continuing with the above hospital data platform as an example, assume that the data fact table corresponding to the charging department is the initial distance point 0, respectively obtain the data fact tables corresponding to other departments that have data circulation with the charging department, and obtain the distance points between the other data fact tables and the charging department.

[0095] By selecting any data fact table as the initial distance point 0, based on all the data dimension tables corresponding to the data fact table, all the superior and inferior data fact tables of the data fact table are obtained, and distance points are set for all the superior and inferior data fact tables to obtain the inter-table distance between each data fact table. This facilitates preparing for generating an inter-table association topology graph and normalizing and standardizing the data in the data warehouse.

[0096] Continue to refer to Figure 4 , Figure 4 is Figure 2 a flowchart of a specific implementation manner of step 202 shown, including steps:

[0097] Step 401: Pre-select any one of the data fact tables among each of the data fact tables, obtain the primary key field of the data fact table, and set the primary key field as the distance point i n ;

[0098] Step 402: Query all the data dimension tables that store the current primary key field based on the current primary key field;

[0099] Step 403: Obtain the foreign key fields used as the primary key fields of the data fact table and the dimension information of the foreign key fields in all the data dimension tables;

[0100] Step 404: According to the foreign key fields and the dimension information of the foreign key fields, obtain the upper-level or lower-level primary key fields associated with the current primary key field, and set the primary key field as the distance point i n+1 , and update the primary key field to the current primary key field;

[0101] Step 405: Execute steps 402 to 404 in a loop to update the distance point for the obtained primary key fields until there are no associated lower-level primary key fields for the current primary key field, where the initial value of i n is 0, and i n+1 = i n + 1.

[0102] Step 203, based on the inter-table distance and the primary key field, construct an inter-table association topology graph for each of the data fact tables, where the inter-table association topology graph uses the primary key field as the topology node.

[0103] In this embodiment, the step of constructing an inter-table association topology graph for each data fact table based on the inter-table distance and the primary key field specifically includes: obtaining and traversing the distance points corresponding to the primary key fields of each data fact table; using the primary key field with a distance point of 0 as the starting topology node, and sequentially searching for the corresponding primary key field as the next-level topology node in ascending order of the distance point until the topology node corresponding to the primary key field when the distance point reaches the maximum value is found, and then connecting all the found topology nodes in sequence to complete the construction of the inter-table association topology graph.

[0104] Continuing with the above hospital data platform as an example, based on the distance points between each department and the charging department obtained in step 202, construct an inter-table association topology graph, which is convenient for the hospital data platform to perform refined processing on data at the DWS layer and quickly obtain the optimal HQL execution statement corresponding to each department.

[0105] By using the primary key field as the topology node and taking the distance point corresponding to the primary key field as the distance length between the current topology node and the starting topology node, construct an inter-table association topology graph. Complete the construction of the inter-table association topology graph at the DWM layer during the construction of the data warehouse, which is convenient for assisting the DWS layer to perform refined processing on data and quickly obtain the optimal HQL execution statement.

[0106] Step 204, assemble the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph, where the behavior characteristics include the HQL execution statement assembly request sent by the user.

[0107] In this embodiment, the optimal HQL execution statement corresponding to the first user can be regarded as a multi-dimensional data model when the first user uses the optimal HQL execution statement to perform specific data operations.

[0108] In this embodiment, the step of assembling the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph specifically includes: receiving the HQL execution statement assembly request sent by the first user, where the HQL execution statement assembly request includes the target data index information required by the first user, and the index information is the primary key field of the data fact table corresponding to the target data; parsing the HQL execution statement assembly request to parse out the target data index information, and assembling and obtaining the optimal HQL execution statement for the target data according to the inter-table association topology graph and the target data index information; using the optimal HQL execution statement as the return value and returning it to the first user.

[0109] Continuing with the above hospital data platform as an example, assume that a certain employee in the charging department is the first user. At this time, obtain the optimal HQL execution statement assembly request sent by the employee, and through the inter-table association topology diagram, assemble and obtain the optimal HQL statement for the target data as a multi-dimensional data model.

[0110] By parsing the optimal HQL execution statement assembly request sent by the first user, clarify the target data corresponding to the first user, and through the inter-table association topology diagram, assemble and obtain the optimal HQL statement for the target data. The optimal HQL statement is the multi-dimensional data model corresponding to the first user. Using the inter-table association topology diagram facilitates the data warehouse to quickly assemble the optimal HQL statement for obtaining the target data, improves the standardization and scientificity of data warehouse construction, and reduces the query and assembly time consumption in the data warehouse to a certain extent.

[0111] Continue to refer to Figure 5 , Figure 5 Yes Figure 2 is a flowchart of a specific implementation manner of step 204 shown in

[0112] Step 501, receive the HQL execution statement assembly request sent by the first user. Among them, the HQL execution statement assembly request includes the target data index information required by the first user. Among them, the index information is the primary key field of the data fact table corresponding to the target data;

[0113] Step 502, parse the HQL execution statement assembly request, parse out the target data index information, and according to the inter-table association topology diagram and the target data index information, assemble and obtain the optimal HQL execution statement for the target data;

[0114] Step 503, use the optimal HQL execution statement as the return value and return it to the first user.

[0115] In this embodiment, based on the behavioral characteristics of the first user and the inter-table association topology diagram, multi-dimensional data model modeling occurs between the DWM layer and the DWS layer during data warehouse construction.

[0116] In this embodiment, the step of assembling and obtaining the optimal HQL execution statement of the target data according to the inter-table association topology diagram and the target data index information specifically includes: Step a: Obtain the primary key field of the data fact table for caching user information; Step b: Find the topology node corresponding to the primary key field from the inter-table association topology diagram, mark it as the first topology node, and obtain the distance point corresponding to the first topology node, denoted as the first distance point; Step c: According to the primary key field of the data fact table of the target data, find the topology node corresponding to the primary key field from the inter-table association topology diagram, mark it as the target topology node, and obtain the distance point corresponding to the target topology node, denoted as the second distance point; Step d: According to the inter-table association topology diagram, construct a topology node path from the first distance point to the second distance point, and obtain the number of nodes included in the topology node path; Step e: Determine whether the difference between the first distance point and the second distance point satisfies the preset equation formula: |i a -i b | = c - 1, where i a is the first distance point, i b is the second distance point, |i a -i b | is the absolute value corresponding to the difference between the first distance point and the second distance point, and c is the number of nodes included in the topology node path; Step f: If the difference between the first distance point and the second distance point satisfies the preset equation formula, the determination of the shortest path is completed, otherwise, repeat Steps d to e; Step g: After the determination of the shortest path is completed, traverse and obtain all the topology nodes corresponding to the shortest path, and according to the inter-table association topology diagram, identify the primary key fields corresponding to all the topology nodes respectively, and assemble the HQL execution statements between multiple tables through the primary key fields, and use the assembled HQL execution statements between multiple tables as the optimal HQL execution statement.

[0117] By constructing the topology node path and the preset equation formula in a loop, it is judged whether the shortest path is screened out, and the HQL execution statement is constructed according to the shortest path, so as to ensure that when performing data operations, as few forms as possible are queried to ensure the query speed.

[0118] Step 205, obtain the optimal HQL execution statements corresponding to several first users, use the optimal HQL execution statements corresponding to the several first users as the result set, obtain the portrait features corresponding to the several first users based on the preset feature form, and use the portrait features as the input set for pre-training of the prediction model.

[0119] In this embodiment, the several first users may represent the scenario where the same login account used by employees in the same department performs data operations multiple times, or the scenario where employees in the same department use different login accounts of the department to perform data operations. The optimal HQL execution statements corresponding to the several first users are the optimal HQL execution statements corresponding to the department's multiple uses of the department-owned accounts to operate data.

[0120] In this embodiment, the step of using the optimal HQL execution statements corresponding to the several first users as a result set, obtaining the portrait features corresponding to the several first users based on a preset feature form, and using the portrait features as an input set for pre-training a prediction model specifically includes: using the portrait features corresponding to the several first users as prediction conditions, where the portrait features include the data operation department to which the user belongs, and the feature form includes the data operation department information corresponding to each user; using the optimal HQL execution statements corresponding to the several first users as prediction results; constructing a one-to-one correspondence between the prediction conditions and the prediction results according to a preset user identifier, and generating an associated binary tuple, where the data format of the associated binary tuple is: [prediction condition: prediction result]; inputting the prediction conditions and the prediction results into an initial prediction model, and performing grouped annotation on the corresponding prediction conditions and prediction results according to the associated binary tuple; obtaining the grouped annotation result, and obtaining all the prediction results corresponding to the prediction conditions under the same prediction conditions according to the grouped annotation result; performing statistics on all the prediction results, and screening out the same prediction result with the largest probability value among all the prediction results as the model output result under the prediction conditions, thereby obtaining a pre-trained prediction model.

[0121] Continue to refer to Figure 6 , Figure 6 Yes Figure 2 is a flowchart of a specific implementation manner of step 205 shown in

[0122] Step 601: Use the portrait features corresponding to the several first users as prediction conditions, where the portrait features include the data operation department to which the user belongs, and the feature form includes the data operation department information corresponding to each user;

[0123] Step 602: Use the optimal HQL execution statements corresponding to the several first users as prediction results;

[0124] Step 603: Construct a one-to-one correspondence between the prediction conditions and the prediction results according to a preset user identifier, and generate an associated binary tuple, where the data format of the associated binary tuple is: [prediction condition: prediction result];

[0125] Step 604: Input the prediction conditions and the prediction results into the initial prediction model, and perform grouped annotation on the corresponding prediction conditions and prediction results according to the associated binary tuples.

[0126] Step 605: Obtain the grouped annotation results, and obtain all the prediction results corresponding to the prediction conditions under the same prediction conditions according to the grouped annotation results.

[0127] Step 606: Statistically analyze all the prediction results, and screen out the same prediction result with the largest probability value among all the prediction results as the model output result under the prediction conditions, so as to obtain a pre-trained prediction model.

[0128] Continuing with the above hospital data platform as an example, by obtaining the portrait features of multiple first users, that is, the department information they belong to, a multi-dimensional data model commonly used in each department is trained. When new employees in the department need to obtain the multi-dimensional data model at this time, they can directly reuse the corresponding multi-dimensional data model according to the department information, without having to reconstruct the multi-dimensional data model again, reducing the steps of repeatedly constructing the multi-dimensional data model.

[0129] By obtaining the multi-dimensional data models constructed between the DWM layer and the DWS layer of the data warehouse for several first users, and inputting the multi-dimensional data models and the portrait features of the several first users into the prediction model for pre-training, it is convenient to directly use the pre-trained prediction model after the DWS layer of the data warehouse to obtain the multi-dimensional data model corresponding to the new user, simplifying the steps of repeatedly constructing the multi-dimensional data model in the data warehouse, directly reusing the multi-dimensional data model corresponding to the historical first user, which is more efficient.

[0130] Step 206: Receive the HQL execution statement assembly request sent by the second user, and obtain the portrait features corresponding to the second user based on the feature form.

[0131] Step 207: Input the portrait features into the pre-trained prediction model to obtain the optimal HQL execution statement corresponding to the second user.

[0132] By directly using the pre-trained prediction model for prediction after the second user sends the HQL execution statement assembly request to obtain the optimal HQL execution statement corresponding to the second user, it reduces the steps of reassembling the optimal HQL execution statement in the data warehouse again, and saves data assembly time to a certain extent.

[0133] In this application, the source data is stored in the corresponding data fact tables and data dimension tables; the foreign key fields and the primary key fields of each data fact table are obtained, and the inter-table distances between the data fact tables are obtained through a preset algorithm rule, the foreign key fields, and the primary key fields; an inter-table association topology graph is constructed; the optimal HQL execution statement is assembled; the optimal HQL execution statements of several first users are obtained for pre-training of the prediction model; a request for assembling the optimal HQL execution statement sent by the second user is received, and based on the pre-trained prediction model, the optimal HQL execution statement of the second user is obtained. The present invention combines the idea of the dimension model, realizes direct access to data from the source ODS layer, constructs an optimal HQL execution statement assembly model, and then predicts the optimal HQL execution statement corresponding to a new user through the prediction model, realizing lightweight coding and automated data processing. This automation is bundled with user characteristics, and based on user-driven generation of the data processing model, the time consumption during repeated construction is reduced.

[0134] Embodiments of this application can obtain and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0135] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0136] In the embodiments of this application, by training the prediction model, it is convenient to automatically reuse the corresponding optimal HQL execution statement from the constructed optimal HQL execution statement model when using the optimal HQL execution statement later, which is convenient for helping data analysis and operation personnel to quickly obtain the corresponding optimal HQL execution statement and reduce the steps of re-construction.

[0137] For further reference Figure 7 As an implementation of the method shown above Figure 2 In one embodiment of a device for predicting an HQL execution statement based on user characteristics provided by this application, this device embodiment corresponds to the Figure 2 shown method embodiment, and this device can be specifically applied to various electronic devices.

[0138] As Figure 7As shown in the figure, the device 700 for predicting HQL execution statements based on user characteristics in this embodiment includes: a source data storage module 701, an inter-table distance acquisition module 702, an inter-table association topology graph construction module 703, an optimal HQL execution statement construction module 704, a prediction model pre-training module 705, an input field acquisition module 706, and a prediction model prediction module 707. Among them:

[0139] The source data storage module 701 is used to obtain source data, store the source data according to a preset modeling method and storage rules, and store the source data into corresponding data fact tables and data dimension tables. Among them, the preset modeling methods include star schema, snowflake schema, and constellation schema;

[0140] The inter-table distance acquisition module 702 is used to obtain the primary key fields of each data fact table, obtain the foreign key fields from the corresponding data dimension tables, and obtain the inter-table distances between each data fact table through preset algorithm rules, the foreign key fields, and the primary key fields;

[0141] The inter-table association topology graph construction module 703 is used to construct an inter-table association topology graph for each data fact table based on the inter-table distance and the primary key fields. Among them, the inter-table association topology graph uses the primary key fields as topology nodes;

[0142] The optimal HQL execution statement construction module 704 is used to assemble the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph. Among them, the behavior characteristics include the HQL execution statement assembly request issued by the user;

[0143] The prediction model pre-training module 705 is used to obtain the optimal HQL execution statements corresponding to a number of first users, use the optimal HQL execution statements corresponding to the number of first users as a result set, obtain the portrait characteristics corresponding to the number of first users based on a preset feature form, use the portrait characteristics as an input set, and perform pre-training of the prediction model;

[0144] The input field acquisition module 706 is used to receive the HQL execution statement assembly request issued by the second user and obtain the portrait characteristics corresponding to the second user based on the feature form;

[0145] The prediction model prediction module 707 is used to input the portrait characteristics into the pre-trained prediction model to obtain the optimal HQL execution statement corresponding to the second user.

[0146] This application stores the source data into the corresponding data fact tables and data dimension tables respectively; obtains the foreign key fields and the primary key fields of each data fact table, and obtains the inter-table distances between each data fact table through a preset algorithm rule, foreign key fields and primary key fields; constructs an inter-table association topology graph; assembles the optimal HQL execution statement; obtains the optimal HQL execution statements of several first users, and performs pre-training of the prediction model; receives the optimal HQL execution statement assembly request sent by the second user, and obtains the optimal HQL execution statement of the second user based on the pre-trained prediction model. The present invention combines the idea of the dimension model, realizes direct access to data from the source ODS layer, constructs an optimal HQL execution statement assembly model, and then predicts the optimal HQL execution statement corresponding to the new user through the prediction model, realizing lightweight coding and automated data processing. This automation is bundled with user characteristics, and based on user-driven data processing model generation, it reduces the time consumption during repeated construction.

[0147] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0148] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times, and their execution order does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0149] To solve the above technical problems, the embodiments of this application also provide a computer device. For details, please refer to Figure 8 , Figure 8 which is the basic structural block diagram of the computer device in this embodiment.

[0150] The computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 8 with components 81 - 83 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0151] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0152] The memory 81 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 81 can be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 can also be an external storage device of the computer device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 8. Of course, the memory 81 can also include both the internal storage unit and the external storage device of the computer device 8. In this embodiment, the memory 81 is generally used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions for the method of predicting HQL execution statements based on user characteristics. In addition, the memory 81 can also be used to temporarily store various types of data that have been output or will be output.

[0153] In some embodiments, the processor 82 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 82 is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to run the computer-readable instructions stored in the memory 81 or process data, such as running the computer-readable instructions of the method for predicting HQL execution statements based on user characteristics.

[0154] The network interface 83 may include a wireless network interface or a wired network interface, which is generally used to establish a communication connection between the computer device 8 and other electronic devices.

[0155] The computer device proposed in this embodiment belongs to the field of big data technology. In this application, source data is stored in corresponding data fact tables and data dimension tables; foreign key fields and primary key fields of each data fact table are obtained, and the inter-table distances between each data fact table are obtained through preset algorithm rules, foreign key fields, and primary key fields; an inter-table association topology graph is constructed; an optimal HQL execution statement is assembled; the optimal HQL execution statements of several first users are obtained for pre-training of the prediction model; a request for assembling the optimal HQL execution statement sent by the second user is received, and based on the pre-trained prediction model, the optimal HQL execution statement of the second user is obtained. The present invention combines the idea of the dimension model, realizes direct access to data from the source ODS layer, constructs an optimal HQL execution statement assembly model, and then predicts the optimal HQL execution statement corresponding to a new user through the prediction model, realizes data processing with less code and automation, binds this automation with user characteristics, and generates a data processing model driven by users, reducing the time consumption during repeated construction.

[0156] This application also provides another implementation manner, that is, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by a processor to enable the processor to execute the steps of the method for predicting HQL execution statements based on user characteristics as described above.

[0157] The computer-readable storage medium proposed in this embodiment belongs to the field of big data technology. In this application, source data is stored in corresponding data fact tables and data dimension tables; foreign key fields and primary key fields of each data fact table are obtained, and the inter-table distances between each data fact table are obtained through a preset algorithm rule, foreign key fields, and primary key fields; an inter-table association topology graph is constructed; an optimal HQL execution statement is assembled; the optimal HQL execution statements of several first users are obtained for pre-training of a prediction model; a request for assembling the optimal HQL execution statement sent by a second user is received, and based on the pre-trained prediction model, the optimal HQL execution statement of the second user is obtained. The present invention combines the idea of a dimensional model, realizes direct access to data from the source ODS layer, constructs an optimal HQL execution statement assembly model, and then predicts the optimal HQL execution statement corresponding to a new user through a prediction model, realizing lightweight coding and automated data processing. This automation is bundled with user characteristics, and based on user-driven generation of a data processing model, the time consumption during repeated construction is reduced.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.

[0159] Obviously, the above-described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of this application's specification and drawings in other related technical fields is equally within the scope of this application's patent protection.

Claims

1. A method for predicting HQL execution statements based on user characteristics, characterized in that, Including the following steps: Obtain source data, store the source data according to a preset modeling method and storage rules, and store the source data into corresponding data fact tables and data dimension tables respectively, where the preset modeling method includes star schema, snowflake schema, and constellation schema; Obtain the primary key fields of each of the data fact tables, and obtain the foreign key fields from the corresponding data dimension tables. Obtain the inter-table distances between the data fact tables through a preset algorithm rule, the foreign key fields, and the primary key fields. The step of obtaining the primary key fields of each of the data fact tables, obtaining the foreign key fields from the corresponding data dimension tables, and obtaining the inter-table distances between the data fact tables through a preset algorithm rule, the foreign key fields, and the primary key fields specifically includes: Step A: Pre-select any one of the data fact tables, obtain the primary key field of the data fact table, and set the primary key field as the distance point ; Step B: Query all data dimension tables that store the current primary key field based on the current primary key field; Step C: Obtain the foreign key fields used as the primary key fields of the data fact tables and the dimension information of the foreign key fields in all the data dimension tables; Step D: Obtain the upper-level or lower-level primary key field associated with the current primary key field according to the foreign key field and the dimension information of the foreign key field, and set the primary key field as the distance point , and update the primary key field to the current primary key field; Step E: Execute steps B to D in a loop to update the distance point for the obtained primary key field until there is no associated next-level primary key field for the current primary key field, where the initial value of is 0; Construct an inter-table association topology graph for each of the data fact tables based on the inter-table distance and the primary key field, where the inter-table association topology graph uses the primary key field as a topology node; Assemble the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph, where the behavior characteristics include the HQL execution statement assembly request sent by the user; Obtain the optimal HQL execution statements corresponding to several first users, use the optimal HQL execution statements corresponding to the several first users as a result set, obtain the portrait features corresponding to the several first users based on a preset feature form, use the portrait features as an input set, and perform pre-training of a prediction model; Receive the HQL execution statement assembly request sent by the second user, and obtain the portrait features corresponding to the second user based on the feature form; Input the portrait features into the pre-trained prediction model to obtain the optimal HQL execution statement corresponding to the second user.

2. The method for predicting an HQL execution statement based on user characteristics according to claim 1, wherein The step of storing the source data into corresponding data fact tables and data dimension tables respectively according to a preset modeling method and storage rules specifically includes: Classify the source data according to a preset data type classification rule, and divide the unit data in the source data into two categories: fact table data and dimension table data; Generate corresponding data fact tables and data dimension tables respectively in a preset data storage area according to corresponding table creation instructions; Store the unit data with the classification category of the fact table data into the corresponding data fact table according to the storage rules; Store the unit data with the classification category of the dimension table data into the corresponding data dimension table according to the storage rules; Among them, the storage rules specifically include: in the data fact table, the primary key field of the current data fact table, the data index field corresponding to the primary key field, and the dimension information of the data index field are stored; in the data dimension table, the primary key fields of each data fact table and the dimension information of the primary key fields are stored. Each primary key field corresponds to a data fact table. The column used to store the primary key field of the data fact table in the data dimension table is used as the foreign key field.

3. The method for predicting HQL execution statements based on user characteristics according to claim 1, wherein The step of constructing an inter-table association topology graph for each data fact table based on the inter-table distance and the primary key field specifically includes: Obtaining and traversing the distance points corresponding to the primary key fields of each data fact table; Taking the primary key field with a distance point of 0 as the starting topology node, sequentially searching for the corresponding primary key field as the next-level topology node in ascending order of the distance point until the topology node corresponding to the primary key field when the distance point reaches the maximum value is found, and sequentially connecting all the found topology nodes to complete the construction of the inter-table association topology graph.

4. The method for predicting HQL execution statements based on user characteristics according to claim 3, wherein The step of assembling the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph specifically includes: Receiving an HQL execution statement assembly request sent by the first user, where the HQL execution statement assembly request includes the target data index information required by the first user, and the index information is the primary key field of the data fact table corresponding to the target data; Parsing the HQL execution statement assembly request to parse out the target data index information, and assembling and obtaining the optimal HQL execution statement for the target data according to the inter-table association topology graph and the target data index information; Taking the optimal HQL execution statement as the return value and returning it to the first user.

5. The method for predicting an HQL execution statement based on user characteristics according to claim 4, wherein The step of assembling and obtaining the optimal HQL execution statement for the target data according to the inter-table association topology graph and the target data index information specifically includes: Step a: Obtaining the primary key field of the data fact table for caching user information; Step b: Searching for the topology node corresponding to the primary key field in the inter-table association topology graph, marking it as the first topology node, and obtaining the distance point corresponding to the first topology node, denoted as the first distance point; Step c: According to the primary key field of the data fact table of the target data, searching for the topology node corresponding to the primary key field in the inter-table association topology graph, marking it as the target topology node, and obtaining the distance point corresponding to the target topology node, denoted as the second distance point; Step d: Constructing a topology node path from the first distance point to the second distance point according to the inter-table association topology graph, and obtaining the number of nodes included in the topology node path; Step e: Determine whether the difference between the first distance point and the second distance point satisfies a preset equation formula: , where is the first distance point, is the second distance point, is the absolute value corresponding to the difference between the first distance point and the second distance point, is the number of nodes included in the topological node path; Step f: If the difference between the first distance point and the second distance point satisfies the preset equation formula, the determination of the shortest path is completed; otherwise, steps d to e are repeatedly executed. Step g: After the determination of the shortest path is completed, traverse to obtain all topological nodes corresponding to the shortest path, and identify the primary key fields corresponding to all the topological nodes according to the inter-table association topology graph. Assemble the HQL execution statements between multiple tables through the primary key fields, and use the assembled HQL execution statements between multiple tables as the optimal HQL execution statement.

6. The method for predicting an HQL execution statement based on user characteristics according to claim 1, wherein The step of using the optimal HQL execution statements corresponding to the several first users as the result set, obtaining the portrait features corresponding to the several first users based on a preset feature form, and using the portrait features as the input set for pre-training the prediction model specifically includes: Using the portrait features corresponding to the several first users as prediction conditions, where the portrait features include the data operation departments to which the users belong, and the feature form includes the data operation department information corresponding to each user; Using the optimal HQL execution statements corresponding to the several first users as prediction results; Construct a one-to-one correspondence between the prediction conditions and the prediction results according to a preset user identifier, and generate an associated binary tuple, where the data format of the associated binary tuple is: [prediction condition: prediction result]; Input the prediction conditions and the prediction results into an initial prediction model, and perform grouped annotation on the corresponding prediction conditions and prediction results according to the associated binary tuple; Obtain the grouped annotation result, and obtain all the prediction results corresponding to the prediction conditions under the same prediction conditions according to the grouped annotation result; Statistically analyze all the prediction results, and screen out the same prediction result with the largest probability value among all the prediction results as the model output result under the prediction conditions, and obtain the pre-trained prediction model.

7. An apparatus for predicting HQL execution statements based on user characteristics, characterized in that including: A source data storage module, configured to obtain source data, store the source data according to a preset modeling method and storage rules, and store the source data into the corresponding data fact tables and data dimension tables, where the preset modeling methods include star schema, snowflake schema, and galaxy schema; An inter-table distance acquisition module, configured to obtain the primary key fields of each data fact table, obtain the foreign key fields from the corresponding data dimension tables, and obtain the inter-table distances between each data fact table through a preset algorithm rule, the foreign key fields, and the primary key fields. The step of obtaining the primary key fields of each data fact table, obtaining the foreign key fields from the corresponding data dimension tables, and obtaining the inter-table distances between each data fact table through a preset algorithm rule, the foreign key fields, and the primary key fields specifically includes: Step A: Pre-select any one of the data fact tables, obtain the primary key field of the data fact table, and set the primary key field as the distance point ; Step B: Query all the data dimension tables that store the current primary key field based on the current primary key field; Step C: Obtain the foreign key fields used as the primary key fields of the data fact table and the dimension information of the foreign key fields in all the data dimension tables; Step D: Obtain the upper-level or lower-level primary key field associated with the current primary key field according to the foreign key field and the dimension information of the foreign key field, and set the primary key field as the distance point , and update the primary key field to the current primary key field; Step E: Execute steps B to D in a loop to update the distance points of the obtained primary key fields until there are no associated next-level primary key fields for the current primary key field, where, The initial value of is 0; An inter-table association topology graph construction module, configured to construct an inter-table association topology graph for each data fact table based on the inter-table distance and the primary key field, where the inter-table association topology graph uses the primary key field as the topological node; Optimal HQL execution statement construction module, which is used to assemble the optimal HQL execution statement corresponding to the first user based on the behavior characteristics of the first user and the inter-table association topology graph, wherein the behavior characteristics include the HQL execution statement assembly request issued by the user; Prediction model pre-training module, which is used to obtain the optimal HQL execution statements corresponding to a number of first users, use the optimal HQL execution statements corresponding to the number of first users as the result set, obtain the portrait features corresponding to the number of first users based on a preset feature form, and use the portrait features as the input set to perform pre-training of the prediction model; Input field acquisition module, which is used to receive the HQL execution statement assembly request issued by the second user and obtain the portrait features corresponding to the second user based on the feature form; Prediction model prediction module, which is used to input the portrait features into the pre-trained prediction model to obtain the optimal HQL execution statement corresponding to the second user.

8. A computer device, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the method for predicting an HQL execution statement based on user characteristics according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor, the steps of the method for predicting an HQL execution statement based on user characteristics according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Data table generation method and device, equipment and storage medium

    CN113760891A

  • Systems and methods for detecting and mitigating threats to a structured data storage system

    US20140201838A1