Method and system for realizing quick retrieval of hundred million-level data

By constructing query sub-intervals and updating progress in real time, the problems of inefficiency and system crashes in billions of data queries are solved, and fast and stable data retrieval is achieved.

CN120596717AInactive Publication Date: 2025-09-05BEIJING ZHONGAN NEBULA SOFTWARE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510703261.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing databases are inefficient when querying billions of data points. They are prone to querying unnecessary data blocks, resulting in slow query speeds and long response times. In addition, the efficiency of conventional retrieval with multiple conditions and large time ranges decreases, which may cause the system to crash.

Method used

By determining whether a query result set exists in the memory, constructing query subintervals, querying data in batches, and updating the query progress in real time, repeated queries and waste of system resources are avoided.

Benefits of technology

It achieves rapid retrieval of billions of data, avoids repeated queries and waste of system resources, reduces the query of unnecessary data blocks, and improves query speed and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596717A_ABST
    Figure CN120596717A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for realizing quick retrieval of hundred million-level data, and relates to the field of data search. The method comprises the following steps: judging whether a corresponding query result set exists in a memory or not, and if so, directly returning the query result set; otherwise, constructing a query sub-interval; calculating the total amount of data in the query range, selecting a query sub-interval, and presetting the number of query targets; querying data meeting query conditions in the selected query subinterval, storing the queried target data into a query result set, and repeating the step until all the queried data meeting the query conditions or the data meeting the query conditions reach a preset query target number; calculating the query progress, judging whether the query is completed or not, and if yes, ending the query; and otherwise, executing a circulation process until a preset condition is met. According to the method, the retrieval efficiency can be improved, system crash caused by overlarge data is prevented, the response time is shortened, and the query speed is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data search technology, and in particular to a method and system for realizing rapid retrieval of billions of data. Background Art

[0002] With the advent of the big data era, data volumes are rapidly increasing, with billions or even more becoming commonplace. Rapid data retrieval and analysis are becoming increasingly important across all industries. However, this massive amount of data presents challenges in rapidly retrieving and processing it. When faced with large datasets, traditional data retrieval methods can be inefficient, time-consuming, and resource-intensive, failing to meet practical needs.

[0003] When querying billions of data, the query speed of commonly used databases easily reaches a bottleneck, and sometimes the system crashes due to the large amount of data. Although the ClickHouse database can achieve millisecond-level query speeds when querying billions of data, the storage structure of ClickHouse can only use indexes to accurately find the required data blocks when querying large amounts of data, which will cause unnecessary data block scanning. In conventional searches, it is impossible to generate indexes for all query conditions. Therefore, in conventional searches, directly querying ClickHouse data will result in a decrease in efficiency when querying multiple conditions and large time ranges. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for realizing rapid retrieval of billions of data, which can solve the problems existing in the prior art of low retrieval efficiency, easy query of unnecessary data blocks, slow query speed and long response time.

[0005] The present invention is achieved in that:

[0006] In a first aspect, the present application provides a method for quickly searching billions of data, comprising the following steps:

[0007] Determine whether a corresponding query result set exists in the memory according to the query condition. If so, directly return the query result set; otherwise, construct a query sub-interval based on the query range in the query condition according to the set rules; wherein the query sub-interval includes one or more;

[0008] Calculate the total amount of data within the query range, select the query sub-interval, and preset the target number of queries;

[0009] Search for data that meets the query conditions in the selected query sub-interval, and store the found target data in the query result set. Repeat this step until all the data that meets the query conditions are found or the number of data that meets the query conditions reaches the preset query target number.

[0010] Calculate the query progress based on the target number of data items in the query result set, and determine whether the query is completed based on the total amount of data and the query progress. If so, terminate the query; otherwise, execute the loop process until the preset conditions are met.

[0011] The cycle process includes:

[0012] A scheduled query task is set for the unqueried range, the queried target data is also stored in the query result set, and the number of target data in the query result set is updated in real time to calculate the query progress; wherein, the preset condition is that the query progress reaches 100%.

[0013] Furthermore, constructing the query subinterval includes the following substeps:

[0014] Determine whether the query time exceeds the set value based on the query range; if so, split the query range according to the set conditions to obtain query sub-intervals; otherwise, use the query range as the query sub-interval.

[0015] In a second aspect, the present application provides a system for realizing rapid retrieval of billions of data, which comprises a preprocessing module, a first query module and a second query module;

[0016] The preprocessing module determines whether there is a corresponding query result set in the memory according to the query condition. If so, the query result set is directly returned; otherwise, a query sub-interval is constructed according to the query range in the query condition; wherein the query sub-interval includes one or more

[0017] The first query module calculates the total amount of data within the query range, selects the query sub-interval, and presets the target number of queries;

[0018] Search for data that meets the query conditions in the selected query sub-interval, and store the found target data in the query result set. Repeat this step until all the data that meets the query conditions are found or the number of data that meets the query conditions reaches the preset query target number.

[0019] The second query module calculates the query progress based on the number of target data items in the query result set, and determines whether the query is completed based on the total amount of data and the query progress. If so, the query is terminated; otherwise, a loop process is executed until a preset condition is met; the loop process includes: setting a scheduled query task for the unqueried range, storing the queried target data in the query result set, and updating the number of target data items in the query result set in real time, and calculating the query progress; wherein, the preset condition is that the query progress reaches 100%.

[0020] In a third aspect, the present application provides an electronic device comprising a memory for storing one or more programs; a processor; and when the one or more programs are executed by the processor, the method as described in any one of the first aspects is implemented.

[0021] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of the above-mentioned first aspects.

[0022] Compared with the prior art, the present invention has at least the following advantages or beneficial effects:

[0023] The present invention proposes a method and system for realizing rapid retrieval of billions of data. The method first determines whether there is a query result set that meets the query conditions in the memory, which can effectively avoid repeated queries and reduce the waste of system resources. The query range in the query conditions is constructed into query sub-intervals according to the set rules, and the queries are performed in batches, which can effectively avoid the situation where the storage crashes due to excessive short-term data volume. The query data is marked, which can facilitate users to understand the subsequent data volume and thus better plan subsequent work. When querying the data in each query sub-interval, the query progress is judged, which can avoid repeated queries. At the same time, if all data has been obtained before all query sub-intervals are queried, it can avoid querying unnecessary data blocks, thereby reducing the waste of system resources and speeding up the query progress. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 This is a flow chart of an embodiment of a method for realizing rapid retrieval of billions of data in the present invention;

[0026] Figure 2 This is a structural block diagram of the present invention for realizing rapid retrieval of billions of data;

[0027] Figure 3 This is a structural block diagram of an electronic device provided by an embodiment of the present invention.

[0028] Icon: 101, memory; 102, processor; 103, communication interface. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. Some implementation methods of the present application are described in detail below in conjunction with the drawings. In the absence of conflict, the following embodiments and the various features in the embodiments can be combined with each other.

[0030] Example

[0031] See also Figure 1 The method for realizing rapid retrieval of billions of data includes the following steps:

[0032] Determine whether a corresponding query result set exists in the memory according to the query condition. If so, directly return the query result set; otherwise, construct a query sub-interval based on the query range in the query condition according to the set rules; wherein the query sub-interval includes one or more;

[0033] Determine whether the query time exceeds the set value based on the query range; if so, split the query range according to the set conditions to obtain query sub-intervals; otherwise, use the query range as the query sub-interval.

[0034] For example, if the query condition determines that the query time is less than the set value, the query can be performed directly without splitting the query range in the query condition; for example, the query condition is 2022-12-01 to 2022-12-01, the number of pages is 1, and the number of items displayed per page is 15. Querying data in the background and displaying it directly according to the display rules will not cause the query time to be too long or the cache to crash due to excessive query data volume.

[0035] Set the query conditions to time 2022-12-01 to 2022-12-05, page number 1, and the number of items displayed per page is 15.

[0036] The query range in the query condition is determined by time, and the sub-interval of the query based on time is

[0037] [{2022-12-01 00:00:00~2022-12-01 23:59:59},

[0038] {2022-12-02 00:00:00~2022-12-02 23:59:59},

[0039] {2022-12-03 00:00:00~2022-12-03 23:59:59},

[0040] {2022-12-04 00:00:00~2022-12-04 23:59:59},

[0041] {2022-12-05 00:00:00~2022-12-05 23:59:59}].

[0042] Divide the query range into five parts and query each part separately. This can reduce system usage and query time, and avoid storage paralysis caused by excessive data volume in a short period of time. Query the split query sub-intervals one by one. If the corresponding data has been queried in the previous intervals, redundant data can be avoided, reducing unnecessary system consumption. Calculate the total amount of data in the query range, select the query sub-interval, and preset the target number of query items.

[0043] Calculate the total amount of data within the query range, select the query sub-interval, and preset the target number of queries;

[0044] Search for data that meets the query conditions in the selected query sub-interval, and store the found target data in the query result set. Repeat this step until all the data that meets the query conditions are found or the number of data that meets the query conditions reaches the preset query target number.

[0045] For example, selecting a query sub-interval and querying the preset query target number according to the query conditions can avoid the problem of long system response time caused by displaying after all data has been queried. When displaying data to the user, first extract a part of the data. After the user has viewed this part of the data, determine whether all the data has been viewed. If not, set a scheduled task to continue querying subsequent data.

[0046] The query conditions are: 2022-12-01 to 2022-12-05, page count 1, and number of items displayed per page 15. The front-end passes the current page as 1, and queries 15 items per page. The back-end actually queries 15 pages of data by default, returns the first page, and caches the remaining 14 pages in SEARCH_CACHE_DATA. The query result is set as key. If the user clicks on a page that is still within the range of 15, the cached data is directly retrieved; otherwise, the program queries the database again. If there are only 5 pages of data that meet the conditions in the current query sub-range, 5 pages of data are directly returned, and the subsequent operation proceeds to the next query sub-range.

[0047] Calculate the query progress based on the target number of data items in the query result set, and determine whether the query is completed based on the total amount of data and the query progress. If so, terminate the query; otherwise, execute the loop process until the preset conditions are met.

[0048] The cycle process includes:

[0049] A scheduled query task is set for the unqueried range, the queried target data is also stored in the query result set, and the number of target data in the query result set is updated in real time to calculate the query progress; wherein, the preset condition is that the query progress reaches 100%.

[0050] For example, the current query subinterval is traversed to find the total number of items that meet the criteria. Since there are still subsequent query subintervals that have not been queried, the progress percentage of the query progress found first is stored in the cache SEARCH_CACHE_DATA. The cache SEARCH_CACHE_DATA also stores the total amount of data within the query range and the current time. The array type of the query progress value can be set to value, and the query condition is added to the thread pool. The thread pool is the smallest unit for completing tasks in the computer. By placing the query condition in the thread pool, a query task can be set, allowing the computer to complete the query autonomously without affecting the main task (display). When the query condition is time, the query progress includes the progress percentage, the total number, and the current time.

[0051] Real-time query progress assessment avoids repeated queries and wastes system resources. For example, within the current query subinterval, 225 matching data items are found, stored in the SEARCH_RESULT_CACHE cache, and the result set is returned. If the returned result set contains the SEARCH_CACHE_DATA flag and progress is not 100%, the front-end adds a scheduled page task to continue the query. The thread pool continues to query the total number of items based on the passed query criteria until the SEARCH_CACHE_DATA flag reaches 100%. This process is then returned to the front-end, which terminates the scheduled task and concludes the entire query process.

[0052] Based on the same inventive concept, please refer to Figure 2 ,The present invention also proposes a system for realizing rapid retrieval of 100 million-level data, comprising a pre-processing module, a first query module and a second query module;

[0053] The preprocessing module determines whether there is a corresponding query result set in the memory according to the query condition. If so, the query result set is directly returned; otherwise, a query sub-interval is constructed according to the query range in the query condition; wherein the query sub-interval includes one or more

[0054] The first query module calculates the total amount of data within the query range, selects the query sub-interval, and presets the target number of queries;

[0055] Search for data that meets the query conditions in the selected query sub-interval, and store the found target data in the query result set. Repeat this step until all the data that meets the query conditions are found or the number of data that meets the query conditions reaches the preset query target number.

[0056] The second query module calculates the query progress based on the number of target data items in the query result set, and determines whether the query is completed based on the total amount of data and the query progress. If so, the query is terminated; otherwise, a loop process is executed until a preset condition is met; the loop process includes: setting a scheduled query task for the unqueried range, storing the queried target data in the query result set, and updating the number of target data items in the query result set in real time, and calculating the query progress; wherein, the preset condition is that the query progress reaches 100%.

[0057] For the specific implementation process of the above system, please refer to the method for realizing rapid retrieval of billions of data provided in the embodiment of this application, which will not be repeated here.

[0058] See also Figure 3 , Figure 3 A structural block diagram of an electronic device provided in an embodiment of the present invention. The electronic device includes a memory 101, a processor 102 and a communication interface 103, and the memory 101, the processor 102 and the communication interface 103 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, such as program instructions / modules corresponding to a system for realizing rapid retrieval of billions of data provided in an embodiment of the present application, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.

[0059] Among them, the memory 101 can be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0060] The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0061] I understand. Figure 3 The structure shown is only for illustration, and the electronic device may also include Figure 3 More or fewer components than shown, or with Figure 3 Different configurations shown. Figure 3 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0062] In summary, the embodiments of the present application provide a method and system for realizing rapid retrieval of billions of data. First, it is determined whether there is a query result set that meets the query conditions in the memory, which can effectively avoid repeated queries and reduce waste of system resources; the query range in the query conditions is constructed into query sub-intervals according to the set rules, and the queries are performed in batches, which can effectively avoid the situation where the storage crashes due to excessive short-term data volume; the query data is marked, which can facilitate users to understand the subsequent data volume and thus better plan subsequent work; when querying the data in each query sub-interval, the query progress is judged to avoid repeated queries. At the same time, if all data has been obtained before all query sub-intervals are queried, unnecessary data blocks can be avoided, thereby reducing waste of system resources and speeding up query progress.

[0063] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for realizing rapid retrieval of billions of data, characterized in that: The following steps are involved: Determine whether a corresponding query result set exists in the memory according to the query condition. If so, directly return the query result set; otherwise, construct a query sub-interval based on the query range in the query condition according to the set rules; wherein the query sub-interval includes one or more; Calculate the total amount of data within the query range, select the query sub-interval, and preset the target number of queries; Search for data that meets the query conditions in the selected query sub-interval, and store the found target data in the query result set. Repeat this step until all the data that meets the query conditions are found or the number of data that meets the query conditions reaches the preset query target number. Calculate the query progress based on the target number of data items in the query result set, and determine whether the query is completed based on the total amount of data and the query progress. If so, terminate the query; otherwise, execute the loop process until the preset conditions are met. The cycle process includes: A scheduled query task is set for the unqueried range, the queried target data is also stored in the query result set, and the number of target data in the query result set is updated in real time to calculate the query progress; wherein, the preset condition is that the query progress reaches 100%.

2. A method for realizing rapid retrieval of billions of data as claimed in claim 1, characterized in that: Constructing a query subinterval includes the following substeps: Determine whether the query time exceeds the set value based on the query range; if so, split the query range according to the set conditions to obtain query sub-intervals; otherwise, use the query range as the query sub-interval.

3. A system for realizing rapid retrieval of billions of data, characterized in that: It includes a pre-processing module, a first query module and a second query module; The preprocessing module determines whether there is a corresponding query result set in the memory according to the query condition. If so, the query result set is directly returned; otherwise, a query sub-interval is constructed according to the query range in the query condition; wherein the query sub-interval includes one or more The first query module calculates the total amount of data within the query range, selects the query sub-interval, and presets the target number of queries; Search for data that meets the query conditions in the selected query sub-interval, and store the found target data in the query result set. Repeat this step until all the data that meets the query conditions are found or the number of data that meets the query conditions reaches the preset query target number. The second query module calculates the query progress based on the number of target data items in the query result set, and determines whether the query is completed based on the total amount of data and the query progress. If so, the query is terminated; otherwise, a loop process is executed until a preset condition is met; the loop process includes: setting a scheduled query task for the unqueried range, storing the queried target data in the query result set, and updating the number of target data items in the query result set in real time, and calculating the query progress; wherein, the preset condition is that the query progress reaches 100%.

4. An electronic device, characterized in that: include: a memory for storing one or more programs; processor; When the one or more programs are executed by the processor, the method according to any one of claims 1 to 4 is implemented.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.