Data acquisition method and device, computer equipment, storage medium and program product
By structuring the data query parameters, data query templates and prompt description information are generated, and data query characteristics are used to determine the degree of data availability, the problem of data acquisition in the existing technology is solved, and an efficient and personalized data acquisition process is realized, which significantly improves the data acquisition effect.
Patent Information
- Application Number
- CN202510205913.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has a high reliance on manual quality and coverage in the data acquisition process, and lacks user personalized information, resulting in poor data acquisition results.
By obtaining the query information of the data, the structured processing generates a data query template, combining the template information to generate prompt description information, analyze the data query characteristics of the query information, determine the degree of data availability, and obtain data based on the availability and prompt description information to obtain the target data matching the query information.
It reduces the user's understanding cost, improves the quality and coverage of data acquisition, improves the correlation between data query and query information, effectively improves the efficiency and quality of data acquisition, and significantly improves the data acquisition effect.
Smart Images

Figure CN120144604A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a data acquisition method, apparatus, computer device, storage medium and program product. Background Art
[0002] With the development of computer technology, data has become the core asset for strategic decision-making in various industries and scientific and technological fields. Currently, the acquisition of data mainly relies on data extraction caliber auxiliary tools such as high-heat chart generation systems. However, the data obtained in this way highly depends on manual work in terms of quality and coverage, and lacks user personalized information, resulting in poor data acquisition effects. Summary of the Invention
[0003] In view of this, the present disclosure provides a data acquisition method, apparatus, computer device, storage medium and program product to solve the problem of poor data acquisition effects.
[0004] In a first aspect, the present disclosure provides a data acquisition method, including: obtaining query information of data; performing structured processing on data query parameters according to the query information to generate a data query template; generating prompt description information based on the template information of the data query template; analyzing the data query characteristics of the query information to determine the availability of the data queried by the data query template; and performing data acquisition based on the availability and the prompt description information to obtain target data matching the query information.
[0005] In a second aspect, the present disclosure provides a data acquisition apparatus, including: an obtaining module for obtaining query information of data; a processing module for performing structured processing on data query parameters according to the query information to generate a data query template; a generating module for generating prompt description information based on the template information of the data query template; an analyzing module for analyzing the data query characteristics of the query information to determine the availability of the data queried by the data query template; and a data acquisition module for performing data acquisition based on the availability and the prompt description information to obtain target data matching the query information.
[0006] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other, wherein the memory stores computer instructions, and the processor executes the computer instructions to execute the data acquisition method according to the first aspect or any corresponding implementation manner thereof.
[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the data acquisition method according to the first aspect or any corresponding implementation manner thereof.
[0008] Fifth aspect, the present disclosure provides a computer program product, including computer instructions for causing a computer to execute the data acquisition method according to the first aspect or any corresponding embodiment thereof as described above.
[0009] The data acquisition method, device, computer device, storage medium and program product provided by the present disclosure can abstract the data query information from the data structure level by structuring the data query parameters according to the query information to generate corresponding data query templates, so as to support various storage objects such as data sets and data tables. Combining the template information of the data query template, generating prompt description information for user data acquisition, determining the availability of the queried data by analyzing the data query characteristics of the query information, and guiding the data acquisition process by combining the availability with the prompt description information to obtain target data matching the query information. Therefore, by guiding the data acquisition using the prompt description information, the understanding cost of the user is reduced, which is beneficial to improving the quality and coverage of the acquired data. At the same time, incorporating the data query characteristics into the data acquisition process is beneficial to improving the relevance between the queried data and the query information, effectively improving the data acquisition efficiency and data acquisition quality, and greatly enhancing the data acquisition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of data extraction using a high-heat chart generation system;
[0012] Figure 2 It is a schematic flowchart of the data acquisition method according to an embodiment of the present disclosure;
[0013] Figure 3 It is a schematic diagram of the query DSL according to an embodiment of the present disclosure;
[0014] Figure 4 It is a schematic diagram of the template structuring result according to an embodiment of the present disclosure;
[0015] Figure 5 It is a schematic flowchart of another data acquisition method according to an embodiment of the present disclosure;
[0016] Figure 6 It is a specific generation schematic diagram of the data query template according to an embodiment of the present disclosure;
[0017] Figure 7 It is a schematic flowchart of another data acquisition method according to an embodiment of the present disclosure;
[0018] Figure 8 It is a structural block diagram of a data acquisition device according to an embodiment of the present disclosure;
[0019] Figure 9 It is a schematic hardware structure diagram of a computer device according to an embodiment of the present disclosure. Specific embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0022] For example, when receiving a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0023] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the computer device.
[0024] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0025] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations, and related provisions.
[0026] The application of random computer technology, whether in traditional industries or emerging technology fields, relies on data to support daily operations, optimize management processes, and enhance the customer experience. Through in-depth analysis of big data, it is convenient to accurately grasp the industry development trend, understand user needs, and then adjust corresponding development strategies, improve operational efficiency, and reduce costs. Data has become the core asset for strategic decision-making in all industries to promote business growth and innovation, that is, data is the key factor for all enterprises to gain a competitive advantage, improve efficiency, and drive innovation.
[0027] Currently, three stages are required in the process of using data: searching for data tables / sets, confirming field calibers, and confirming data extraction calibers (the data extraction caliber is the way to obtain data). Among them, searching for data sets / tables is the first stage of using data, which is used to determine the specific location (library / table) where the data exists; confirming field calibers occurs in the second stage, which is used to determine the specific indicators, dimensions, and expression fields in the data set / table that can meet business requirements; confirming data extraction calibers is the last link in using data, which is used to combine the required fields such as indicators and dimensions into actual query logic, including dimension drilling down, indicator aggregation, and dimension filtering, etc.
[0028] As Figure 1 shown, the commonly used data acquisition method is a high-heat chart generation system. Specifically, first, confirm the information of the data set (fields) to be used, which can be completed by methods such as searching for library / table meta-information; secondly, filter out all the charts in the chart library that depend on the above-mentioned required data sets from the chart library; finally, sort the filtered charts in descending order according to the user access heat of the charts, and output the top N charts with the highest heat, such as 10 charts, as a guide to help users complete the final data confirmation.
[0029] However, the high-heat chart generation system has the following disadvantages: First, the data extraction caliber of the chart is not intuitive. The data extraction caliber is mainly represented by the chart title and description, which highly depends on manual work in terms of quality and coverage, and the title and description often have a highly abstract phenomenon, requiring users to query the specific chart content to determine the actual data extraction caliber of the chart, with a high usage cost; Second, the chart relevance is insufficient. It only considers the global heat information of the chart, lacks user personalization information, and the query results are fixed; Third, the coverage is not comprehensive. The chart only serves storage objects such as data sets and cannot cover storage objects such as data tables.
[0030] Based on this, the technical solution of the present disclosure relies on a large language model to generate prompt description information for data queries, reducing the user's understanding cost and improving the comprehensibility of the data acquisition process; incorporating personalized features into the data query process to enhance the relevance between the data query result and the query request; and structurally abstracting data query parameters through a domain-specific language, which can support common storage objects such as data sets and data tables. Thus, after the user confirms the data table / set and fields, data strongly related to the user's business scenario can be returned, effectively improving the data acquisition efficiency.
[0031] According to an embodiment of the present disclosure, an embodiment of a data acquisition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0032] In this embodiment, a data acquisition method is provided, which can be used in computer devices such as computers and servers. Figure 2 is a flowchart of the data acquisition method according to an embodiment of the present disclosure, as Figure 2 shown, the process includes the following steps:
[0033] Step S101, obtain the query information of the data.
[0034] The query information is the core information required for data queries, and specifically may include a query identification ID, a query object identification ID, a storage object identification ID, a query structured language DSL, and a query time, etc. Among them, the query identification ID is a mark used to uniquely determine a query; the query object identification ID is the unique identifier of the query object, and other attributes of the query object can be determined through the query object identification ID, such as the department where the query object is located, the position it belongs to, etc.; the storage object identification ID is the unique identifier of the data set or data table, and the attributes corresponding to the storage object can be determined through this storage object identification ID, such as name, field details, storage capacity, etc.; the query structured language DSL represents the specific data query logic, which is the underlying unified representation of high-level languages such as SQL by the data query engine and has better versatility. The query structured language DSL usually includes key information such as metrics, dimensions, aggregations, filters, and sorting; the query time is the time of the current data query.
[0035] Specifically, the data query is implemented through a data query engine deployed in the computer device. When a data query is performed, the data query engine will generate corresponding query logs for the data query operation. The query information of the data can be obtained by accessing the query logs.
[0036] Step S102: Structurally process the data query parameters according to the query information to generate a data query template.
[0037] The data query parameters are the parameters required in the data query acquisition process, which may specifically include index parameters, dimension parameters, aggregation parameters, filtering parameters, sorting information parameters, etc. As Figure 3 shown in the query DSL example, where selectidlist is the index parameter, dimensions is the dimension parameter, groupbyidlist is the aggregation parameter, wherelist is the filtering parameter, and sort is the sorting information parameter. Extract the data query parameters required in the data query process from the query information, and structurally process the data query parameters using a domain-specific language to obtain a structured query language for data query. Use this structured query language to process the data query parameters into a data query template.
[0038] Step S103: Generate prompt description information based on the template information of the data query template.
[0039] The prompt description information is used to convert the data query template into natural description language for data query and data acquisition in natural description language. Specifically, after obtaining the DSL data query template, structurally organize the template information in the data query template, including field name filling, function operator name conversion logic, etc., to generate structured template information, as Figure 4 shown in the template structuring result example. Then, use the prompt engineering and few-shot prompting of the large language model to generate prompt description information for the data query template, and use this prompt description information to convert the data query template into a natural language description. Among them, the prompt description information includes rule description information and example description information, and the example description information is implemented in a few-shot manner, consisting of an information pair of structured information and a natural language data acquisition problem.
[0040] Step S104: Analyze the data query characteristics of the query information to determine the availability of the data queried by the data query template.
[0041] The data query characteristics are the personalized characteristics possessed by the query object during data query, such as the characteristics of data query for the same department, the characteristics of data query for the data analysis object group, the characteristics of data query for the general public, etc. The availability indicates the correlation between the data queried using the data query template and the query information. Among them, the greater the correlation, the higher the availability of the data query template.
[0042] Specifically, by parsing the query information, it is possible to determine the frequency of data query and data usage of each query object for each storage object, and determine the data query characteristics of different query objects for data query or data usage according to the frequency of data query and data usage. The availability of the data queried by the data query template is evaluated through the data query characteristics, and the available degree of the data queried by the data query template is determined.
[0043] Step S105, based on the available degree and the prompt description information, perform data acquisition to obtain target data that matches the query information.
[0044] When the data query template combines the query information to query data from the storage object, multiple query data can be obtained, and each query data has a corresponding available degree. Sort the multiple query data according to the available degree, extract the query data with the highest available degree, and determine the prompt description information corresponding to the query data with the highest available degree. Subsequently, data query can be performed according to the query information according to the prompt description information corresponding to the high available degree, so as to obtain target data that highly matches the query information from the storage object.
[0045] The data acquisition method provided in this embodiment can abstract the data query information from the data structure layer by structuring the data query parameters according to the query information, so as to generate a corresponding data query template, which can support various storage objects such as data sets and data tables. Combine the template information of the data query template to generate the prompt description information for user data acquisition. By analyzing the data query characteristics of the query information, determine the available degree of the queried data, and combine the available degree with the prompt description information to guide the data acquisition process, so as to obtain target data that matches the query information. Thus, by using the prompt description information to guide the data acquisition, the understanding cost of the user is reduced, which is beneficial to improving the quality and coverage of the acquired data. At the same time, incorporating the data query characteristics into the data acquisition process is beneficial to improving the relevance between the query data and the query information, effectively improving the data acquisition efficiency and data acquisition quality, and greatly enhancing the data acquisition effect.
[0046] In this embodiment, a data acquisition method is provided, which can be used in computer devices such as computers and servers. Figure 5 is a flowchart of the data acquisition method according to an embodiment of the present disclosure, as Figure 5 shown, the process includes the following steps:
[0047] Step S201, obtain the query information of the data.
[0048] Specifically, the above step S201 includes:
[0049] Step S2011, obtain the log data generated for data query.
[0050] The log data is the data generated during the data query operation. The relevant information generated for the data query event, such as query time, query object, storage object, query result, etc., is recorded in the log data.
[0051] Specifically, the data query event is implemented by the data query engine, and the data query engine can record the query log data generated by the data query events initiated by each query object. Correspondingly, the computer device can access the log data recorded in the data query engine for the data query event.
[0052] Step S2012, perform data cleaning on the log data and extract the query information of the data from the log data.
[0053] Since all the data generated by the data query event is recorded in the log data, at this time, it is necessary to perform data cleaning on the log data to filter out the non-core data in the log data, so as to extract the core query information such as the query identification ID, query object identification ID, storage object identification ID, query structured language DSL (including index / dimension / aggregation / filtering / sorting information), and query time corresponding to the data query event from the log data.
[0054] Step S202, perform structured processing on the data query parameters according to the query information to generate a data query template.
[0055] Specifically, the above step S202 includes:
[0056] Step S2021, identify the constant parameters in the data query parameters and determine the parameter placeholders corresponding to the constant parameters.
[0057] The constant parameter represents the attribute parameter that remains unchanged during the data query process; the parameter placeholder is a placeholder used in the database query to replace the actual value, that is, the symbol used when writing the structured query statement, which is used to safely insert the specific value during the execution of the data query, and these specific values will be replaced at the position of the placeholder during the execution of the data query. Specifically, all the data query parameters are identified in turn to determine the constant parameters existing in the data query parameters, and the parameter placeholders are used for replacement at the positions where each constant parameter is located.
[0058] Step S2022, perform data aggregation according to the constant parameters and determine the set of constant parameters corresponding to the parameter placeholders.
[0059] The set of constant parameters is the set of constant parameter values possessed by the parameter placeholders, such as a list of constant values. Combining the position where the constant parameter is located, the position where the parameter placeholder is located can be determined, and thus the instance constant parameters with the same position can be determined. Subsequently, data aggregation is performed according to the position where the instance constant parameters are located to obtain the set of constant parameters under each parameter placeholder.
[0060] In some alternative embodiments, the above step S2022 includes:
[0061] Step a1, replacing the constant parameter with a parameter placeholder to generate an initial query template.
[0062] Step a2, performing data aggregation using the initial query template to determine the set of constant parameters corresponding to the parameter placeholder.
[0063] Identify all data query parameters for specific domain language (DSL) programming, determine the constant parameters among the data query parameters, and replace the constant parameters with parameter placeholders to obtain an initial query template, as Figure 6 shown.
[0064] Aggregate the initial query templates with the same structure according to the data query detail instances to obtain the set of constant parameters under each parameter placeholder, that is, the set of constant parameters corresponding to the parameter placeholder, as Figure 6 shown.
[0065] In the above embodiments, by replacing the constant parameter with a parameter placeholder, generating the corresponding initial query template, and performing data aggregation using the initial query template, the fetching logics that represent the same can be normalized, which is convenient for improving the coverage of data acquisition and the quality of data acquisition.
[0066] Step S2023, replacing the parameter placeholder with the set of constant parameters to generate a data query template.
[0067] Based on the set of constant parameters aggregated according to the detail query instances, identify whether the set of constant parameters corresponding to each parameter placeholder is a variable, and combine the variable detection results to perform parameter replacement on the parameter placeholder using the corresponding parameter replacement rule to obtain the final data query template.
[0068] In some alternative embodiments, the above step S2023 includes:
[0069] Step b1, identifying whether the set of constant parameters corresponding to the parameter placeholder is a variable parameter.
[0070] Step b2, if the set of constant parameters corresponding to the parameter placeholder is not a variable parameter, then backfill the non-variable parameter to the parameter placeholder to obtain a data query template.
[0071] Step b3, if the constant parameter corresponding to the parameter placeholder is a variable parameter, backfill the variable parameter to the parameter placeholder according to the preset parameter backfill rule to obtain a data query template.
[0072] A variable parameter represents a parameter that will change during the data query process. Identify the parameters in the constant parameter set corresponding to each parameter placeholder to determine whether the parameters are the same. If the parameters change in an increasing or decreasing order, etc., it means that the constant parameter set corresponding to the parameter placeholder is a variable parameter; if the parameters are fixed, it means that the constant parameter set corresponding to the parameter placeholder is a constant parameter (i.e., a non-variable parameter).
[0073] Specifically, if the constant parameter set corresponding to the parameter placeholder is a non-variable parameter, replace the parameter placeholder corresponding to the non-variable parameter in the initialization query template with the constant corresponding to the non-variable parameter to obtain the corresponding data query template. As Figure 6 shown, backfill the parameter placeholder ${2} with the specific constant parameter value "customer service" to obtain the corresponding data query template.
[0074] The preset parameter backfill rule is a rule for replacing parameter placeholders with variable parameters set in advance. For example, for a variable parameter, keep the parameter placeholder unchanged. As Figure 6 shown, keep the parameter placeholder ${1} unchanged. For example, for a variable parameter that increases by date, the parameter placeholder ${1} can be sequentially replaced in the order of date increase. As Figure 6 shown, the parameter placeholder ${1} is identified as a date variable. Of course, other rules can also be used, and it can be determined according to actual needs here.
[0075] If the constant parameter set corresponding to the parameter placeholder is a variable parameter, the parameter placeholder corresponding to the variable parameter in the initialization query template can be replaced with the constant corresponding to the variable parameter according to the preset parameter backfill rule to obtain the corresponding data query template.
[0076] In the above embodiments, by identifying the constant parameters in the constant parameter set corresponding to the parameter placeholders, and combining non-variable parameters and variable parameters for corresponding parameter backfilling to obtain the corresponding data query template, the templatization of query information is realized, which is convenient for adapting to data queries of various storage objects and improves the application scenarios of data queries.
[0077] Step S203, generate prompt description information based on the template information of the data query template. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, and will not be elaborated here.
[0078] Step S204: Analyze the data query characteristics of the query information to determine the availability of the data queried by the data query template. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments and will not be elaborated here.
[0079] Step S205: Based on the availability and the prompt description information, perform data acquisition to obtain the target data that matches the query information. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments and will not be elaborated here.
[0080] The data acquisition method provided in this embodiment performs data cleaning on the log data generated by data queries to extract the query information of the data, thereby enabling targeted extraction of the query information and ensuring that the query information can cover the core information of the data query. By aggregating data according to the constant parameters in the data query parameters to determine the constant parameter set of the parameter placeholders corresponding to the constant parameters, the normalization of the same logic is achieved, and the coverage of the data query is improved. By using the constant parameters in the constant parameter set to replace the corresponding parameter placeholders to generate the final data query template, the structuring and templating of the data query parameters are realized, and the application scenarios of the data query are expanded.
[0081] In this embodiment, a data acquisition method is provided, which can be used in computer devices such as computers and servers. Figure 7 is a flowchart of the data acquisition method according to an embodiment of the present disclosure, as Figure 7 shown, and this process includes the following steps:
[0082] Step S301: Obtain the query information of the data. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments and will not be elaborated here.
[0083] Step S302: Structurally process the data query parameters according to the query information to generate a data query template. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments and will not be elaborated here.
[0084] Step S303: Generate prompt description information based on the template information of the data query template. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments and will not be elaborated here.
[0085] Step S304: Analyze the data query characteristics of the query information to determine the availability of the data queried by the data query template.
[0086] Specifically, the above step S304 includes:
[0087] Step S3041: Parse the query information to determine the data query characteristics of the query information in multiple dimensions.
[0088] As described above, the query information contains the core information required for data query, such as query identification ID, query object identification ID, storage object identification ID, query structured language DSL, and query time. By parsing the query information, the query object identification for initiating the data query can be determined, and the relevant attributes of the query object can be determined. Combining the relevant attributes of the query object, the data query characteristics generated by the query information in multiple dimensions can be determined. Specifically, the query dimensions may include the department to which the query object belongs, query object attributes (such as data analysis objects), and global queries, etc. Different query dimensions have different data query characteristics. By parsing the query information, the query data can be statistically analyzed from each dimension to determine the data query characteristics generated by the query information in multiple dimensions.
[0089] Step S3042, perform feature fusion on the data query characteristics according to each dimension, and determine the availability of the data queried by the data query template based on the feature fusion result.
[0090] The data query characteristics of each dimension are weighted and fused according to the corresponding weight parameters to obtain the feature fusion result of the data query characteristics of multiple dimensions. The relevance between the data queried by the data query template and the query information is evaluated through the feature fusion result, and the availability of the data queried by the data query template is determined in combination with this relevance.
[0091] Specifically, the above step S3042 includes:
[0092] Step c1, obtain the heat weight of the data query characteristics under each dimension.
[0093] Step c2, perform weighted processing on the data query characteristics of each dimension according to the heat weight to obtain the availability of the query data.
[0094] The heat weight is used to represent the importance degree of the data query characteristics. This heat weight is manually specified in the initial stage and can be adaptively adjusted based on the supervised learning method in the subsequent stage to make the heat weight match the data query characteristics.
[0095] Perform weighted processing on the data query characteristics under each dimension according to the heat weight corresponding to each data query characteristic to obtain the weighted result, and use this weighted result to represent the availability of the query data. In a specific example, if there are data query characteristics in 3 dimensions: department heat h d , analyst heat h a and global heat h g , where the department heat h dRefers to the access popularity of a department to the DSL data query template, reflecting the frequency of access to the data acquisition path within the same department. The reason is that there is a high degree of correlation and convergence in the business and data acquisition paths of query objects within the same department; analyst popularity h a Refers to the frequency of access to the data acquisition path by the analysis object group. The reason is that the analysis object has authority as the management subject of the business data acquisition path; global popularity h g Reflects the access situation of the general public to this data acquisition path. Weighted processing is performed based on the data query characteristics of these three dimensions, as follows:
[0096] score = m 1 *h d +m 2 *h a +m 3 *h g
[0097] Among them, m 1 、m 2 and m 3 are the popularity weights of department popularity h d 、analyst popularity h a and global popularity h g respectively; score is the result of weighted processing, used to evaluate the availability of the query data.
[0098] In the above implementation, by fusing the data query characteristics of each dimension, it is convenient to evaluate the availability of the query data by combining multiple dimensions, so as to consider the personalized characteristics in the data query process and achieve the correlation between the data query result and the data query information.
[0099] Step S305, perform data acquisition based on the availability and prompt description information to obtain the target data that matches the query information.
[0100] Specifically, the above step S305 includes:
[0101] Step S3051, sort the query data obtained by querying the data query template based on the availability to generate a data sorting result.
[0102] Perform data query according to the query information through the data query template, so as to obtain multiple pieces of query data corresponding to the query information. Sort the multiple pieces of query data output by the data query target in descending order or ascending order according to their corresponding availability to obtain the data sorting result corresponding to the query data.
[0103] Step S3052, use the data sorting result to screen out the target prompt description information from the prompt description information.
[0104] According to the data sorting result, the query data with the highest availability can be determined, and the query data with the highest availability is the data with the highest relevance to the query information. Since each query data is obtained according to the corresponding prompt description information, the target prompt description information corresponding to the query data with the highest availability can be determined from the prompt description information.
[0105] Step S3053, obtain the target data matching the query information from the data storage object according to the target prompt description information.
[0106] The data storage object is an object for storing data, such as a data table, a data set, etc. When performing a data query according to the query information, the process of guiding the data query from the data storage object according to the determined target prompt description information is used to obtain the target data matching the query information from the data storage object.
[0107] The data acquisition method provided in this embodiment determines the data query characteristics of the query information in multiple dimensions by parsing the query information, and fuses the multi-dimensional data query characteristics to determine the availability of the query data. Therefore, the personalized characteristics in the query information can be considered, and the phenomenon that the data query result is fixed due to only considering a single global popularity characteristic can be avoided, thereby improving the relevance between the query data and the query information. By sorting the query data according to the availability, the target prompt description information of the highly available data can be determined in combination with the data sorting result, and then the subsequent data query can be performed according to the target prompt description information, improving the data acquisition efficiency and data acquisition accuracy, and ensuring a high degree of matching between the query data and the query information.
[0108] In this embodiment, a data acquisition device is also provided. This device is used to implement the above embodiment and the preferred implementation manner, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0109] This embodiment provides a data acquisition device, as Figure 8 shown, including:
[0110] An acquisition module 401, configured to acquire the query information of the data.
[0111] A processing module 402, configured to perform a structured process on the data query parameters according to the query information to generate a data query template.
[0112] A generation module 403, configured to generate prompt description information based on the template information of the data query template.
[0113] An analysis module 404, configured to analyze the data query characteristics of the query information and determine the availability of the data queried by the data query template.
[0114] A data acquisition module 405, configured to perform data acquisition based on the availability and the prompt description information to obtain target data matching the query information.
[0115] In some alternative embodiments, the acquisition module 401 includes:
[0116] A log acquisition unit, configured to acquire log data generated for the data query.
[0117] A data cleaning unit, configured to clean the log data and extract the query information of the data from the log data.
[0118] In some alternative embodiments, the processing module 402 includes:
[0119] An identification unit, configured to identify the constant parameters in the data query parameters and determine the parameter placeholder corresponding to the constant parameters.
[0120] An aggregation unit, configured to perform data aggregation according to the constant parameters and determine the constant parameter set corresponding to the parameter placeholder.
[0121] A parameter replacement unit, configured to replace the parameter placeholder with the constant parameter set to generate a data query template.
[0122] In some alternative embodiments, the above-mentioned aggregation unit includes:
[0123] A first parameter replacement subunit, configured to replace the constant parameters with the parameter placeholder to generate an initialization query template.
[0124] A data aggregation subunit, configured to perform data aggregation using the initialization query template and determine the constant parameter set corresponding to the parameter placeholder.
[0125] In some alternative embodiments, the above-mentioned parameter replacement unit includes:
[0126] A variable identification subunit, configured to identify whether the constant parameter set corresponding to the parameter placeholder is a variable parameter.
[0127] A first backfill subunit, configured to, if the constant parameter set corresponding to the parameter placeholder is not a variable parameter, backfill the non-variable parameter to the parameter placeholder to obtain a data query template.
[0128] In some alternative embodiments, the above-mentioned parameter replacement unit further includes:
[0129] A second backfilling subunit, configured to backfill a variable parameter to a parameter placeholder according to a preset parameter backfilling rule to obtain a data query template if the constant parameter corresponding to the parameter placeholder is a variable parameter.
[0130] In some alternative embodiments, the analysis module 404 includes:
[0131] A query feature determination unit, configured to parse query information and determine data query features of the query information in multiple dimensions.
[0132] A feature fusion unit, configured to perform feature fusion on the data query features according to each dimension and determine the availability of the data queried by the data query template based on the feature fusion result.
[0133] In some alternative embodiments, the above feature fusion unit includes:
[0134] A weight acquisition subunit, configured to acquire the popularity weights of the data query features in each dimension.
[0135] A weighting subunit, configured to perform weighting processing on the data query features in each dimension according to the popularity weights to obtain the availability of the query data.
[0136] In some alternative embodiments, the data acquisition module 405 includes:
[0137] A sorting unit, configured to sort the query data obtained by querying the data query template based on the availability to generate a data sorting result.
[0138] A target prompt determination unit, configured to use the data sorting result to screen out target prompt description information from the prompt description information.
[0139] A target data determination unit, configured to obtain target data matching the query information from the data storage object according to the target prompt description information.
[0140] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0141] The data acquisition device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0142] The data acquisition device provided in this embodiment generates a corresponding data query template by structurally processing data query parameters according to query information, thereby enabling the abstraction of data query information at the data structure level to support various storage objects such as data sets and data tables. Combining the template information of the data query template, it generates prompt description information for user data acquisition. By analyzing the data query characteristics of the query information, it determines the availability of the queried data, and combines the availability with the prompt description information to guide the data acquisition process to obtain target data that matches the query information. Thus, by using the prompt description information to guide the data acquisition, it reduces the user's understanding cost, is conducive to improving the quality and coverage of the acquired data. At the same time, incorporating the data query characteristics into the data acquisition process is conducive to enhancing the relevance between the queried data and the query information, effectively improving the data acquisition efficiency and data acquisition quality, and greatly enhancing the data acquisition effect.
[0143] The embodiment of the present disclosure also provides a computer device having the above-mentioned Figure 8 data acquisition device.
[0144] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure. As shown in Figure 9 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common main board or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 9 In
[0145] FIG. 16, one processor 10 is taken as an example.
[0146] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0147] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0148] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.
[0149] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 9 Taking the connection through the bus as an example.
[0150] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (for example, an LED), and a haptic feedback device (for example, a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0151] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.
[0152] Embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0153] A part of the present disclosure can be applied as a computer program product, such as computer program instructions, which when executed by a computer, can call or provide the methods and / or technical solutions according to the present disclosure through the operation of the computer. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0154] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A data acquisition method, characterized in that: The method comprises: Get query information of data; Performing structured processing on the data query parameters according to the query information to generate a data query template; Generate prompt description information based on the template information of the data query template; Analyze the data query characteristics of the query information to determine the availability of the data queried by the data query template; Data is acquired based on the availability and the prompt description information to obtain target data matching the query information.
2. The method according to claim 1, characterized in that: The query information for obtaining data includes: Get the log data generated by the data query; The log data is cleaned and query information of the data is extracted from the log data.
3. The method according to claim 1, characterized in that: The step of performing structural processing on the data query parameters according to the query information to generate a data query template includes: Identify constant parameters in the data query parameters, and determine parameter placeholders corresponding to the constant parameters; Performing data aggregation according to the constant parameters to determine a constant parameter set corresponding to the parameter placeholder; The parameter placeholder is replaced by the constant parameter set to generate the data query template.
4. The method according to claim 3, characterized in that The aggregating data according to the constant parameters to determine the constant parameter set corresponding to the parameter placeholder includes: Replacing the constant parameter with the parameter placeholder to generate an initialization query template; The initialization query template is used to perform data aggregation to determine a constant parameter set corresponding to the parameter placeholder.
5. The method according to claim 3 or 4, characterized in that: The step of replacing the parameter placeholder with the constant parameter set to generate the data query template includes: Identify whether the constant parameter set corresponding to the parameter placeholder is a variable parameter; If the constant parameter set corresponding to the parameter placeholder is a non-variable parameter, the non-variable parameter is used to backfill the parameter placeholder to obtain the data query template.
6. The method according to claim 5, characterized in that Also includes: If the constant parameter set corresponding to the parameter placeholder is a variable parameter, the variable parameter is backfilled into the parameter placeholder according to a preset parameter backfill rule to obtain the data query template.
7. The method according to claim 1, characterized in that The analyzing the data query characteristics of the query information to determine the availability of the data queried by the data query template includes: Parsing the query information to determine data query features of the query information in multiple dimensions; The data query features are fused according to various dimensions, and the availability of the data queried by the data query template is determined based on the feature fusion results.
8. The method according to claim 7, characterized in that The step of fusing the data query features according to various dimensions and determining the availability of the data queried by the data query template based on the feature fusion result includes: Obtain the heat weight of the data query feature under each dimension; The data query features of each dimension are weighted according to the heat weight to obtain the availability of the query data.
9. The method according to claim 1, characterized in that: The acquiring of data based on the availability and the prompt description information to obtain target data matching the query information includes: Sorting the query data obtained by querying the data query template based on the availability to generate a data sorting result; Filtering target prompt description information from the prompt description information using the data sorting result; The target data matching the query information is acquired from the data storage object according to the target prompt description information.
10. A data acquisition device, characterized in that: The device comprises: The acquisition module is used to obtain query information of data; A processing module, used for performing structured processing on data query parameters according to the query information to generate a data query template; A generating module, used for generating prompt description information based on the template information of the data query template; An analysis module, used to analyze the data query characteristics of the query information and determine the availability of the data queried by the data query template; A data acquisition module is used to acquire data based on the availability and the prompt description information to obtain target data matching the query information.
11. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data acquisition method according to any one of claims 1 to 9 by executing the computer instructions.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data acquisition method according to any one of claims 1 to 9.
13. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the data acquisition method according to any one of claims 1 to 9.