Method for optimizing large-model intelligent data statistical query

Through the methods of dynamic granularity generation and incremental update, the problems of storage waste and computing burden in data statistics query in the prior art are solved, and efficient and flexible data query and storage optimization are achieved.

CN120216538APending Publication Date: 2025-06-27SHANGHAI AMARSOFT INFORMATION & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510261133.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art relies on preset query granularity and data tables in data statistics query, and cannot dynamically generate non-minimum granularity data, resulting in waste of storage and computational burden.

Method used

Through the methods of dynamic granularity generation and incremental update, the query request is parsed, the query granularity is judged, the required granularity is generated in real time, and the data is extracted from the original data table according to the preset period, and the merge and update are performed to avoid storing all granular data.

Benefits of technology

Reduces storage space and computing costs, improves query response time and flexibility, and avoids duplicate computing and storage redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216538A_ABST
    Figure CN120216538A_ABST
Patent Text Reader

Abstract

The invention discloses a method for optimizing large-model intelligent data statistical query, and particularly relates to the field of large-model-based data statistical query, which comprises the following steps of: analyzing a query request: analyzing the query request according to a preset query dimension to obtain a plurality of query conditions which respectively correspond to different dimensions; judging whether each query condition is the minimum query granularity or not; judging query granularity: if the query condition is judged to be the minimum query granularity, querying a corresponding result from the first data table; and if the query condition is not the minimum query granularity, dynamically determining a query range based on the current load and the data distribution condition, and querying a corresponding result from the second data table. Through dynamic granularity generation and incremental updating, storage of all granularity data is avoided, the storage space and the calculation cost are reduced, and during each query, data are only generated under the actually needed granularity, incremental data are processed in real time, and the whole data set is not repeatedly calculated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data statistical query based on large models. More specifically, the present invention relates to a method for optimizing intelligent data statistical query of large models. Background Art

[0002] Data statistical query is a process of analyzing, summarizing, and calculating data to extract useful information. It usually involves operations such as filtering, grouping, aggregating, and calculating on large datasets to support decision-making, reporting, and trend analysis. Common statistical queries include calculations such as summation, average, count, maximum value, and minimum value, and are widely used in fields such as business analysis and data mining;

[0003] Publication No. CN109828993A discloses a method and device for querying statistical data. The problem with the prior art is that it only relies on preset query granularity and data tables, and does not provide a mechanism for dynamically generating non-minimum granularity data, resulting in the need to calculate and store all granularity data in advance, which will cause unnecessary storage waste and computational burden. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for optimizing intelligent data statistical query of large models. Through dynamic granularity generation and incremental update, storing all granularity data is avoided, and storage space and computational cost are reduced to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: A method for optimizing intelligent data statistical query of large models, including:

[0006] Parsing the query request: Parse the query request according to the preset query dimension to obtain multiple query conditions, each corresponding to a different dimension; for each query condition, determine whether it is the minimum query granularity;

[0007] Judging the query granularity: If the query condition is determined to be the minimum query granularity, query the corresponding result from the first data table; if the query condition is not the minimum query granularity, dynamically determine the query range based on the current load and data distribution, and query the corresponding result from the second data table;

[0008] Dynamic granularity generation and incremental update: When the granularity of the query condition is not the minimum query granularity, generate the required granularity in real time; extract data from the original data table according to the preset period, and merge and update the newly added or changed data with the existing data;

[0009] Data query and merging: Include the first query and the second query;

[0010] First query: If the query condition is already at the minimum granularity, query the results from the first data table; the first data table records the data that has been merged at the minimum granularity, and the results are directly returned during the query.

[0011] Second query: If the query condition is not at the minimum granularity, query the corresponding range of data records from the second data table; the second data table records the data extracted according to the minimum query granularity, which is used to dynamically generate data at other granularities.

[0012] Result generation and response: Merge the first query result and / or the second query result to form the third query result; if multiple dimensions of the query condition involve data at different granularities, then according to the preset merging rules, merge the query results from different sources and store them in the result data table.

[0013] In a preferred embodiment, it further includes a hierarchical query mechanism: use the first data table for queries at common granularities, pre-merge and store the data at common granularities; for queries at uncommon granularities, use the second data table for queries and dynamically merge the required results according to the query range.

[0014] Based on the granularity generation and screening scheme of the generative adversarial network, through the adversarial training of the generator and the discriminator, analyze and distinguish the commonness of the query granularity; the commonness of the query granularity includes common granularity and uncommon granularity.

[0015] The generator of the generative adversarial network generates query data at different granularities according to the query history and model feedback, and the discriminator of the generative adversarial network evaluates the generated query granularity to judge the commonness of the query granularity.

[0016] In a preferred embodiment, it further includes optimizing storage and query: generate different levels of the first data table and the second data table according to the query granularity and query dimensions, and optimize storage through dynamic merging and incremental updates.

[0017] In a preferred embodiment, establish a dynamic granularity generation model, and the input variables of the dynamic granularity generation model include the query request Q query , the original data set R orig , the query granularity P granularity , the time span T of the query span , the distribution characteristics D of the data dist , the generation model G model .

[0018] Based on the dynamic granularity generation model, establish a granularity generation function; define the function G model , which is used to generate R gen , and this function depends on the query conditions and the original data;

[0019] R gen= G model (R orig , Q query , P granularity , D dist )

[0020] where R gen is the generated data set;

[0021] Establish a multi-level granularity generation mechanism based on the dynamic granularity generation model;

[0022]

[0023] where is the granularity data generated at the k-th layer; f k is the generation function at the k-th layer; C k is the input feature at the k-th layer; W k is the weight matrix, which is used to control the weighted sum processing of the features at this layer; K is the total number of layers.

[0024] In a preferred embodiment, during the generation process, to adapt to the query granularity, time span, and data distribution characteristics, an adaptive adjustment factor α is introduced into the dynamic granularity generation model, and α dynamically adjusts the generation strategy according to D dist and T span ;

[0025]

[0026] where D dist (i) is the data distribution characteristic of the i-th dimension; T span (i) is the time span of the i-th condition in the query request; N is the number of data dimensions;

[0027] The generated data will be optimized by α to make the generation process adapt to the actual needs of the query, and the generated target granularity data will then be multiplied by this adjustment factor for optimization;

[0028]

[0029] where represents the generated data optimized by α;

[0030] An attention mechanism is introduced into the dynamic granularity generation model, and the attention mechanism enables the dynamic granularity generation model to assign different weights to different parts of the data when generating data;

[0031]

[0032] where Attention represents the i-th attention weight function; Q i is the query condition; R iIs the original data.

[0033] In a preferred embodiment, an incremental update model is constructed. The incremental update model is used to avoid full-scale calculation and only update the changed parts of the data.

[0034] The input variables of the incremental update model include the original data set R orig , the newly added incremental data ΔR new , the incremental data ΔR after the last update prev , the data update period T cycle ;

[0035] Perform incremental update calculation through the incremental update model; the incremental update calculation is used to calculate the difference between ΔR and the data of the last update to ensure that only the changed parts are processed.

[0036] ΔR = ΔR new -ΔR prev where ΔR new = Extract(R orig , T cycle )

[0037] where ΔR new is extracted from R through T cycle ; ΔR is the current incremental data; Extract means to extract incremental data from the original data set according to a preset period or query conditions. orig ; ΔR is the current incremental data; Extract means to extract incremental data from the original data set according to a preset period or query conditions.

[0038] In a preferred embodiment, the incremental update model includes incremental update based on a sparse matrix. By using the sparse matrix, only the changed parts are updated, avoiding unnecessary repeated calculation of the unchanged parts.

[0039] ΔR final = SparseUpdate(R orig , ΔR new , S)

[0040] where S is the sparse matrix; ΔR final represents the incremental data after sparse update;

[0041] The final incremental data will be combined with the original data to form the updated data. Based on this, a dynamically adjusted weight matrix W is introduced to control the influence of different incremental data.

[0042]

[0043] where W i is the dynamic weight matrix; R updated is the finally updated data set;

[0044] An incremental update batch processing strategy is introduced for the incremental update model to batch process multiple incremental data. The batch update formula is as follows:

[0045] R batchupdated = BatchUpdate(R orig , ΔR final )

[0046] where BatchUpdate indicates that this operation performs incremental updates through batch calculations; R batchupdated represents the data set after incremental updates are batch processed.

[0047] In a preferred embodiment, a hierarchical query mechanism is constructed. The purpose of the hierarchical query mechanism is to dynamically select appropriate data tables according to the commonness of query granularity during query to avoid storage and calculation redundancy. Data with common granularity is pre-merged and stored, while data with uncommon granularity is dynamically generated during query;

[0048] The input variables of the hierarchical query mechanism include R common , R rare , P granularity , Q query , M merge ; where R common represents storing data with common granularity; R rare represents storing data with uncommon granularity; P granularity is the query granularity; Q query is the query request; M merge is the merge function.

[0049] Judgment of the query granularity in constructing the hierarchical query mechanism; according to the query conditions, judge whether the query granularity is a common granularity. If so, directly query from the data table of common granularity, otherwise query from the data table of uncommon granularity; the formula for judging the granularity is as follows:

[0050]

[0051] In the above formula, a return value of 1 indicates a common granularity, and a return value of 0 indicates an uncommon granularity; if the query condition is a common granularity, directly query data from the first data table R common ; if it is an uncommon granularity, then query from the second data table R rare and perform dynamic merging. The query merge formula is as follows:

[0052] R final = M merge (R common , R rare )

[0053] where M mergeExecute different merging strategies according to different query granularities; R final Represents the final query result.

[0054] In a preferred embodiment, a generative adversarial network is used to analyze the commonness of query granularities. The generator of the generative adversarial network generates query data of different granularities based on query history and model feedback. The discriminator of the generative adversarial network evaluates the commonness of query granularities and finally filters out queries with common and uncommon granularities;

[0055] The input variables of the generative adversarial network include G generator , D discriminator , Q real , Q gen , P common ; where G generator is the generator; D discriminator is the discriminator; Q real is the real query granularity data; Q gen is the generated query granularity data; P common is the common granularity data;

[0056] Based on the generative adversarial network, a generator loss function is established; the loss function of the generator is defined as:

[0057]

[0058] where P data is the distribution of query data; Q is the query data; G generator (Q) is the query granularity data generated by the generator; is the expected value; L gen is the loss function of the generator;

[0059] Based on the generative adversarial network, a discriminator loss function is established; the discriminator is used to distinguish between real query granularities and generated query granularities, and the discriminator loss function is defined as:

[0060]

[0061] where L disc is the loss function of the discriminator; P gen is the distribution of generator data, which is the probability distribution of the query granularity data generated by G generator ;

[0062] Based on the generative adversarial network, generative adversarial training is performed; the generator and the discriminator are gradually optimized through the adversarial training process; the formula for adversarial training is:

[0063] The technical effects and advantages of the present invention:

[0064] 1. Through dynamic granularity generation and incremental update, storing all granularity data is avoided, reducing storage space and computational costs. During each query, data is generated only at the granularity actually required, and incremental data is processed in real time instead of recomputing the entire dataset;

[0065] 2. The solution supports dynamically generating or querying data at the required granularity according to different query granularities; the data structure generated on demand supports different granularity query requirements, improving query response time and flexibility;

[0066] 3. By introducing a hierarchical query mechanism, appropriate data tables are selected based on the commonness of query granularities, avoiding excessive storage space occupied by uncommon granularity queries and effectively reducing the computational burden of full-scale calculation;

[0067] 4. Based on the training mechanism of the generative adversarial network, the commonness of query granularities is classified by the generator and discriminator, ensuring that preprocessed data is preferentially used for common granularity queries and reducing redundant calculations during generation and query. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0070] Refer to the attached Figure 1 description. A method for optimizing intelligent data statistical query of a large model according to an embodiment of the present invention includes:

[0071] Parsing the query request: Parse the query request according to the preset query dimensions to obtain multiple query conditions, each corresponding to a different dimension; for each query condition, determine whether it is the minimum query granularity;

[0072] Judging the query granularity: If the query condition is determined to be the minimum query granularity (for example, the time dimension is "day"), query the corresponding result from the first data table; if the query condition is not the minimum query granularity (for example, the time dimension is "one week"), dynamically determine the query range based on the current load and data distribution, and query the corresponding result from the second data table;

[0073] Dynamic Granularity Generation and Incremental Update: When the granularity of the query condition is not the minimum query granularity, the required granularity is generated in real time; for example, when querying data for a week, the system automatically generates data with the corresponding granularity from the original data, without the need to pre-store pre-computed data for all granularities; data is extracted from the original data table according to a preset period, and the newly added or changed data is merged and updated with the existing data, achieving the effect of only processing incremental data and avoiding repeated calculations;

[0074] Data Query and Merge: including the first query and the second query;

[0075] First Query: If the query condition is already the minimum granularity (for example, the time is a certain day), query the result from the first data table; the first data table records the data that has been merged according to the minimum granularity, and the result is directly returned during the query;

[0076] Second Query: If the query condition is not the minimum granularity (for example, the time is a week), query the data records within the corresponding range from the second data table; the second data table records the data extracted according to the minimum query granularity, which is used to dynamically generate data of other granularities;

[0077] Result Generation and Response: Merge the first query result and / or the second query result to form the third query result; if multiple dimensions of the query condition involve data of different granularities, then according to the preset merge rule, merge the query results from different sources and store them in the result data table.

[0078] It also includes a hierarchical query mechanism: use the first data table for queries of common granularities, pre-merge and store data of common granularities; for queries of uncommon granularities, use the second data table for queries and dynamically merge the required results according to the query range; this method ensures that queries of uncommon granularities do not occupy too much storage and avoids the computational cost of full-scale calculations;

[0079] Granularity Generation and Screening Scheme Based on Generative Adversarial Networks, through the adversarial training of the generator and discriminator, analyze and distinguish the commonness of query granularities; the commonness of query granularities includes common granularities and uncommon granularities;

[0080] The generator of the generative adversarial network generates query data of different granularities according to the query history and model feedback, and the discriminator of the generative adversarial network evaluates the generated query granularities to judge the commonness of the query granularities.

[0081] It also includes optimized storage and query: generate different levels of the first data table and the second data table according to the query granularity and query dimensions, and optimize storage through dynamic merge and incremental update; only generate appropriate data dynamically during the query, avoiding the high storage overhead caused by storing multi-granularity data.

[0082] In large-scale data queries, the goal of the dynamic granularity generation model is to dynamically generate data according to the query granularity requirements, rather than pre-storing data of all granularities. When the query granularity is greater than the minimum granularity (for example, querying "weekly" data while the minimum granularity is "daily"), target granularity data is generated in real time from the original data, which not only involves granularity conversion but also optimization according to the specific query conditions.

[0083] Establish a dynamic granularity generation model. The input variables of the dynamic granularity generation model include the query request Q query , the original data set R orig , the query granularity P granularity , the query time span T span , the distribution characteristics D of the data dist , and the generation model G model ; where Q query contains information such as the dimensions, time range, and query granularity of the query; R orig includes high-frequency, low-granularity data; P granularity includes days, weeks, months, etc., and P granularity determines how the data needs to be aggregated or transformed; D dist characterizes the distribution of the data in different dimensions; G model generates target granularity data based on the original data, and G model includes models based on deep learning such as Transformer, LSTM, or statistical modeling methods.

[0084] Establish a granularity generation function based on the dynamic granularity generation model; define the function G model , used to generate R gen , and this function depends on the query conditions and the original data;

[0085] R gen = G model (R orig , Q query , P granularity , D dist )

[0086] where R gen is the generated data set;

[0087] Establish a multi-level granularity generation mechanism based on the dynamic granularity generation model; the multi-level granularity generation mechanism divides the generation process into multiple levels, and each level is responsible for generating data of different granularities from the original data and considering the influence of different dimensions. The generation process of each level is represented by the following formula:

[0088]

[0089] where is the granularity data generated at the k-th level; fk is the generating function for the k-th layer and can be part of a neural network; C k is the input feature for the k-th layer, C k includes time, space or dimensional features; W k is the weight matrix, which is used to control the weighted sum processing of the features of this layer; K is the total number of layers; the generating function f for each layer k can include different deep learning structures, such as convolutional layers, fully connected layers, etc., and generate data through multi-layer abstraction to capture more complex granularity conversion relationships.

[0090] During the generation process, to better adapt to the query granularity, time span, and distribution characteristics of the data, an adaptive adjustment factor α is introduced into the dynamic granularity generation model, and α is based on D dist and T span dynamically adjusts the generation strategy;

[0091]

[0092] where D dist (i) is the data distribution characteristic of the i-th dimension; T span (i) is the time span of the i-th condition in the query request, which represents the time range of the query, such as one week; N is the number of dimensions of the data, which is used to represent the complexity of the data;

[0093] The generated data will be optimized by α to make the generation process adapt to the actual needs of the query, and the generated target granularity data will then be multiplied by this adjustment factor for optimization;

[0094]

[0095] where represents the generated data optimized by α;

[0096] To enhance the effect of generating granularity, an attention mechanism is introduced into the dynamic granularity generation model. The attention mechanism enables the dynamic granularity generation model to assign different weights to different parts of the data when generating data, so as to better capture the important features in the data;

[0097]

[0098] where Attention represents the i-th attention weight function, which is used to evaluate the importance of the data; Q i is the query condition, which is used to generate the corresponding granularity data; R i is the original data;

[0099] By introducing the attention mechanism, the dynamic granularity generation model can relatively flexibly adjust the generation strategy, and based on the requirements of the query granularity and the distribution of the data, ensure that the generated data has better accuracy and consistency.

[0100] Construct an incremental update model, which is used to avoid full-scale calculation and only update the changed part of the data; this is especially important when the data volume is relatively large, as it can reduce the overhead of repeated calculations and ensure the efficiency of the system;

[0101] The input variables of the incremental update model include the original dataset R orig , the newly added incremental data ΔR new , the incremental data ΔR after the last update prev , the data update period T cycle ; where ΔR new includes the data that has changed within the latest period; T cycle is used to determine the time range for extracting incremental data from the original data;

[0102] Perform incremental update calculation through the incremental update model; the incremental update calculation is used to calculate the difference between ΔR and the data of the last update to ensure that only the changed part is processed;

[0103] ΔR = ΔR new - ΔR prev where ΔR new = Extract(R orig , T cycle )

[0104] where ΔR new is extracted from R cycle through T orig ; ΔR is the current incremental data, and ΔR only includes the newly added or changed part within this period; Extract represents extracting incremental data from the original dataset according to a preset period or query conditions.

[0105] The incremental update model includes incremental update based on a sparse matrix, which is implemented by using a sparse matrix to only update the changed part and avoid unnecessary repeated calculations for the unchanged part;

[0106] ΔR final = SparseUpdate(R orig , ΔR new , S)

[0107] where S is a sparse matrix, and the sparse matrix represents the changed part of the incremental data. The sparse matrix has a low computational complexity and only updates the corresponding part when the data changes; ΔR finalRepresents the incremental data after sparse update, that is, the processed data used for the final update calculation; SparseUpdate means only updating the changed parts in the incremental data, and using the sparse matrix S to efficiently process the changed parts of the data to avoid unnecessary calculations;

[0108] The final incremental data will be combined with the original data to form the updated data. To ensure that the contribution of the incremental data does not overly bias towards certain parts, a dynamically adjusted weight matrix W is introduced based on this to control the influence of different incremental data;

[0109]

[0110] Among them, W i is the dynamic weight matrix, and the dynamic weight matrix is used to control the influence of the incremental data on the final update result; N is used to represent the complexity of the data structure; i represents the i-th dimension or feature of the incremental data; R updated is the dataset after the final update;

[0111] An incremental update batch processing strategy is introduced for the incremental update model, which is used to batch process multiple incremental data. The batch update formula is as follows:

[0112] R batchupdated = BatchUpdate(R orig , ΔR final )

[0113] Among them, BatchUpdate means that this operation performs incremental updates through batch calculations, reducing the processing time and improving the efficiency; R batchupdated represents the dataset after batch processing of incremental updates. Multiple incremental data are updated together to improve the processing efficiency and reduce the calculation time.

[0114] Construct a hierarchical query mechanism. The purpose of the hierarchical query mechanism is to dynamically select appropriate data tables according to the commonness of the query granularity during query to avoid storage and calculation redundancy. Data with common granularity is pre-merged and stored in advance, while data with uncommon granularity is dynamically generated during query;

[0115] The input variables of the hierarchical query mechanism include R common , R rare , P granularity , Q query , M merge ; Among them, R common represents storing data with common granularity, such as pre-merging and storing by "day"; R rare represents storing data with uncommon granularity, such as pre-merging and storing by "week"; P granularity is the query granularity, such as "day", "week", "month"; Qquery is a query request, Q query which contains the dimensions and granularity requirements of the query; M merge is a merging function, which is used to merge data of different granularities.

[0116] Construct the judgment of the query granularity in the hierarchical query mechanism; according to the query conditions, judge whether the query granularity is a common granularity. If so, directly query from the data table of the common granularity; otherwise, query from the data table of the uncommon granularity. The formula for judging the granularity is as follows:

[0117]

[0118] In the above formula, a return value of 1 indicates a common granularity, and a return value of 0 indicates an uncommon granularity; if the query condition is a common granularity, directly query data from the first data table R common ; if it is an uncommon granularity, then query from the second data table R rare and perform dynamic merging. The merging formula for the query is as follows:

[0119] R final = M merge (R common , R rare )

[0120] where M merge is used to execute different merging strategies according to different query granularities, including summation, average value, maximum value, etc.; R final represents the final query result.

[0121] The generative adversarial network is used to analyze the commonness of the query granularity. The generator of the generative adversarial network generates query data of different granularities according to the query history and model feedback. The discriminator of the generative adversarial network evaluates the commonness of the query granularity, and finally filters out the queries of common and uncommon granularities;

[0122] The input variables of the generative adversarial network include G generator , D discriminator , Q real , Q gen , P common ; where G generator is the generator, which is used to generate query data of different granularities; D discriminator is the discriminator, which is used to evaluate whether the generated query granularity is a common granularity; Q real is the real query granularity data, which is obtained from historical data; Q gen is the generated query granularity data, which is generated by the generator; P common is the common granularity data, which is the target category judged by the discriminator;

[0123] Based on the generative adversarial network, a loss function for the generator is established; the goal of the generator is to generate data as close as possible to the true query granularity, and the loss function of the generator is defined as:

[0124]

[0125] where P data is the distribution of the query data; Q is the query data; G generator (Q) is the query granularity data generated by the generator; is the expected value; L gen is the loss function of the generator, which aims to measure the gap between the generated query granularity data and the true query granularity data. The smaller the loss function, the closer the generated query granularity is to the true data;

[0126] Based on the generative adversarial network, a loss function for the discriminator is established; the discriminator is used to distinguish between the true query granularity and the generated query granularity, and the loss function of the discriminator is defined as:

[0127]

[0128] where L disc is the loss function of the discriminator, and its purpose is to enable the discriminator to correctly distinguish between Q real and Q gen . The discriminator is trained by maximizing the correct determination of the true data and the incorrect determination of the generated data; P gen is the distribution of the generator data, which is the probability distribution of the query granularity data generated by G generator ;

[0129] Based on the generative adversarial network, generative adversarial training is performed; the generator and the discriminator are gradually optimized through the adversarial training process, where the generator continuously improves the quality of the generated data, and the discriminator continuously improves the ability to distinguish between the true data and the generated data, and finally intelligently filters the query granularity; the formula for adversarial training is:

[0130]

[0131] By minimizing the generator loss L gen and maximizing the discriminator loss L disc , the generator and the discriminator are continuously optimized.

[0132] Regarding the above solution, it should be noted that the present invention is based on the dynamic generation and incremental update of query granularity to ensure the efficiency of data query at different granularities and storage optimization; in the query request, it will be judged whether the query granularity is the minimum granularity. If so, the result will be directly queried from the first data table storing common granularities; if not, according to the actual query granularity, the second data table will be used for query and the required granularity data will be dynamically generated; this can avoid storing and processing all granularity data in advance in the scenario of extremely large data volume, thereby reducing storage space and computing costs; at the same time, dynamic merging is performed according to the query conditions and data distribution, avoiding repeated calculations. Especially in the case of incremental update, only the newly added data part is updated instead of recalculating the entire data set; in this way, the requirements of different granularity queries can be responded to in real time while effectively reducing the complexity of storage.

[0133] Among them, the dynamic granularity generation model calculates and generates the required granularity data in real time based on the query request and the data distribution, while the incremental update model avoids the recalculation of the entire data volume; each time the data is updated, only the changed part is processed, and the data is merged through incremental update, avoiding repeated calculations of historical data; in addition, the generative adversarial network GAN is introduced to analyze and screen the queries of common granularities and uncommon granularities. Through the adversarial training of the generator and the discriminator, the commonness of the query granularity is identified and classified; the generator is responsible for generating possible query granularities, and the discriminator judges whether it conforms to the characteristics of common granularities, thereby improving the accuracy and efficiency of the query.

[0134] While maintaining the efficiency of the query, this solution avoids storing too much data of uncommon granularities; the entire process adopts a hierarchical query mechanism, using the first data table to store data of common granularities, and using the second data table and the dynamic generation mechanism to process queries of uncommon granularities; in practical applications, this design is suitable for processing large-scale data sets that need to query according to different granularities. Especially in the case of continuous increase in data volume, it can effectively reduce the burden of storage and computing and avoid performance bottlenecks.

[0135] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing large model intelligent data statistical query, comprising: Parsing query requests: Parsing query requests according to preset query dimensions to obtain multiple query conditions corresponding to different dimensions; For each query condition, determine whether it is the minimum query granularity; Determine the query granularity: if the query condition is determined to be the minimum query granularity, query the corresponding result from the first data table; If the query condition is not the minimum query granularity, the query range is dynamically determined based on the current load and data distribution, and the corresponding results are queried from the second data table; Features: Dynamic granularity generation and incremental update: When the granularity of the query condition is not the minimum query granularity, the required granularity is generated in real time; data is extracted from the original data table according to the preset cycle, and the newly added or changed data is merged and updated with the existing data; Data query and merge: including the first query and the second query; First query: if the query condition is the minimum granularity, query the result from the first data table; the first data table records the data that has been merged according to the minimum granularity, and the result is directly returned when querying; Second query: if the query condition is not the minimum granularity, query the data records in the corresponding range from the second data table; the second data table records the data extracted according to the minimum query granularity, which is used to dynamically generate data of other granularities; Result generation and response: merge the first query result and / or the second query result to form a third query result; if multiple dimensions of the query conditions involve data of different granularities, merge the query results from different sources according to the preset merging rules and store them in the result data table.

2. The method for optimizing large-model intelligent data statistical query according to claim 1, characterized in that: It also includes a hierarchical query mechanism: a first data table is used for queries of common granularity, and data of common granularity is pre-merged and stored; For queries of uncommon granularity, use the second data table to query and dynamically merge the required results based on the query scope; A granularity generation and screening scheme based on a generative adversarial network analyzes and distinguishes the commonality of query granularity through adversarial training of the generator and the discriminator. The commonness of query granularity includes common granularity and uncommon granularity; The generator of the generative adversarial network generates query data of different granularities based on the query history and model feedback, and the discriminator of the generative adversarial network evaluates the generated query granularity to determine the commonality of the query granularity.

3. The method for optimizing large-model intelligent data statistical query according to claim 2, characterized in that: It also includes optimizing storage and query: generating first data tables and second data tables at different levels according to query granularity and query dimension, and optimizing storage through dynamic merging and incremental updating.

4. The method for optimizing large-model intelligent data statistical query according to claim 3 is characterized by: Establish a dynamic granularity generation model. The input variables of the dynamic granularity generation model include the query request Q query , original data set R orig , query granularity P granularity , query time span T span , data distribution characteristics D dist , Generate model G model ; Based on the dynamic granularity generation model, a granularity generation function is established; the function G is defined model , used to generate R gen , the function depends on the query conditions and the original data; R gen =G model (R orig ,Q query ,P granularity ,D dist ) Where R gen The generated dataset; Establish a multi-level granularity generation mechanism based on a dynamic granularity generation model; in Granular data generated for the kth layer; f k is the generating function of the kth layer; C k is the input feature of the kth layer; W k is a weight matrix, which is used to control the weighting and processing of the features of this layer; K is the total number of layers.

5. The method for optimizing large-model intelligent data statistical query according to claim 4, characterized in that: In the generation process, in order to adapt to the query granularity, time span and data distribution characteristics, an adaptive adjustment factor α is introduced into the dynamic granularity generation model. α is based on D dist and T span Dynamically adjust generation strategies; Where D dist (i) is the data distribution feature of the i-th dimension; T span (i) is the time span of the i-th condition in the query request; N is the number of dimensions of the data; The generated data will be optimized by α to adapt the generation process to the actual needs of the query. The generated target granularity data will then be multiplied by this adjustment factor for optimization; in represents the generated data after optimization by α; The attention mechanism is introduced into the dynamic granularity generation model. The attention mechanism enables the dynamic granularity generation model to assign different weights to different data parts when generating data. Where Attention represents the i-th attention weight function; Q i is the query condition; R i is the original data.

6. The method for optimizing large-model intelligent data statistical query according to claim 5, characterized in that: Build an incremental update model. The incremental update model is used to avoid full calculation and only update the changed parts of the data. The input variables of the incremental update model include the original dataset R orig , New incremental data ΔR new , the incremental data after the last update ΔR prev , Data update cycle T cycle ; Incremental update calculation is performed through the incremental update model; the incremental update calculation is used to calculate the difference between ΔR and the last updated data to ensure that only the changed part is processed; ΔR=ΔR new -ΔR prev whereΔR new =Extract(R orig ,T cycle ) Where ΔR new By T cycle Extracted from R orig ; ΔR is the current incremental data; Extract means extracting incremental data from the original data set according to the preset period or query condition.

7. The method for optimizing large-model intelligent data statistical query according to claim 6, characterized in that: The incremental update model includes incremental updates based on sparse matrices, which only updates the changed parts by using sparse matrices to avoid unnecessary repeated calculations of the unchanged parts. ΔR final =SparseUpdate(R orig ,ΔR new ,S) Where S is a sparse matrix; ΔR final Represents incremental data after sparse update; The final incremental data will be combined with the original data to form updated data, based on which a dynamically adjusted weight matrix W is introduced to control the influence of different incremental data; Where W i is the dynamic weight matrix; R updated This is the final updated dataset; The incremental update model introduces a batch processing strategy for incremental updates to process multiple incremental data in batches. The batch update formula is as follows: R batchupdated =BatchUpdate(R orig ,ΔR final ) BatchUpdate indicates that the operation is updated by batch computing increments; R batchupdated Represents a dataset that has been incrementally updated through batch processing.

8. The method for optimizing large-model intelligent data statistical query according to claim 7, characterized in that: Construct a hierarchical query mechanism. The purpose of the hierarchical query mechanism is to dynamically select appropriate data tables based on the commonness of the query granularity to avoid storage and computing redundancy. Data of common granularity is merged and stored in advance, while data of uncommon granularity is dynamically generated during query. The hierarchical query mechanism is that the input variables include R common , R rare , P granularity , Q query 、M merge ; where R common Indicates the storage of data of common granularity; R rare Indicates the storage of data of uncommon granularity; P granularity is the query granularity; Q query For query request; M merge is the merge function; Determine the query granularity in constructing a hierarchical query mechanism; According to the query conditions, determine whether the query granularity is a common granularity. If so, query directly from the data table of common granularity, otherwise query from the data table of uncommon granularity; the formula for determining granularity is as follows: In the above formula, a return value of 1 indicates a common granularity, and a return value of 0 indicates an uncommon granularity; If the query condition is a common granularity, directly from the first data table R common Query data in If it is an uncommon particle size, then from the second data table R rare Query and perform dynamic merging. The query merging formula is as follows: R final =M merge (R common ,R rare ) Among them, M merge Used to execute different merging strategies according to different query granularities; R final Represents the final query result.

9. The method for optimizing large-model intelligent data statistical query according to claim 8, characterized in that: Generative adversarial networks are used to analyze the commonality of query granularity. The generator of the generative adversarial network generates query data of different granularities based on query history and model feedback. The discriminator of the generative adversarial network evaluates the commonality of query granularity and finally screens out queries of common and uncommon granularities. The input variables of the generative adversarial network include G generator , D discriminator , Q real , Q gen , P common ; where G generator is the generator; D discriminator is the discriminator; Q real is the actual query granularity data; Q gen is the query granularity data generated; common It is data of common granularity; The generator loss function is established based on the generative adversarial network; the generator loss function is defined as: Where P data is the distribution of query data; Q is the query data; G generator (Q) is the query granularity data generated by the generator; is the expected value; L gen is the loss function of the generator; The discriminator loss function is established based on the generative adversarial network; the discriminator is used to distinguish the real query granularity from the generated query granularity, and the discriminator loss function is defined as: Where L disc is the loss function of the discriminator; P gen is the distribution of the generator data, which is represented by G generator The probability distribution of the generated query-granular data; Generate adversarial training based on generative adversarial networks; The generator and discriminator are gradually optimized through the adversarial training process; the formula for adversarial training is:

Citation Information

Patent Citations

  • A statistical data query method and device

    CN109828993A