Data query method, device, electronic device and storage medium

By receiving data query requests in the data query method and querying matching results in the offline data storage area, the problem of low data query efficiency in the prior art is solved, and fast response and efficient query are achieved.

CN114090631BActive Publication Date: 2025-05-16HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111289007.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-02
Publication Date
2025-05-16
Estimated Expiration
2041-11-02

AI Technical Summary

Technical Problem

In the prior art, data query efficiency is low and query results cannot be obtained quickly, especially when the time range corresponding to the query conditions is large.

Method used

A data query method is proposed, by receiving a data query request, determining a target query area, and querying the query results matching the target query conditions and time information in the first offline data storage area. The method includes preset processing rules, performing query operations in the historical data based on configured query conditions to count query results.

Benefits of technology

By directly storing query results and pre-executing query operations, storage needs are reduced, query efficiency and controllability of query time are improved, and query needs can be quickly responded to query needs and improved user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114090631B_ABST
    Figure CN114090631B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of data processing technology, and in particular to a data query method, device, electronic device and storage medium, which are used to solve the problems of low query efficiency and inability to quickly query and obtain query results. When receiving a data query request triggered by a target object and determining that the target query area targeted by the data query request is a first offline data storage area, query and obtain the target query result in the first offline data storage area, which stores corresponding time information. Based on the configured query conditions, query operations are performed in the corresponding historical data to obtain the query results. In this way, the query efficiency of data can be improved, the controllability of data query time can be improved, the query needs of related objects can be quickly responded to, and the user experience of related objects can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data query method, device, electronic device and storage medium. Background Art

[0002] In order to better understand the operating status of the business, you can usually query the quantity changes of each type of data involved in the business. For example, for the music playback business, you can determine the operating status of the music playback business by querying the data changes of the number of songs that meet the set conditions within the historical time period.

[0003] Under the related technology, in order to count the changes in the quantity of each type of data in the business, it is usually necessary to archive a copy of the full data on a global scale every day, and partition the archived data every day by time and store it in a distributed search engine. In the query process, according to the query conditions triggered by the user, the data restricted by the query conditions are aggregated to obtain the statistical results corresponding to the query conditions, wherein the full data records the full data on a global scale including the latest released data.

[0004] However, in order to analyze data changes, it is necessary to save the full amount of archived data for a long time. Therefore, the total amount of data involved is very high. The accumulated archived storage of the full amount of data will inevitably take up a huge amount of storage space. When the time range corresponding to the query condition is large, it is necessary to call a distributed search engine to search and aggregate the full amount of data stored in different partitions. The query process involves a large amount of data, and the query efficiency is very low. The query time required for querying based on different query conditions is uncontrollable, and the query results cannot be obtained quickly. Summary of the invention

[0005] The embodiments of the present disclosure provide a data query method, device, electronic device and storage medium to solve the problems in the prior art of low query efficiency and inability to quickly query and obtain query results.

[0006] The specific technical solutions provided by the embodiments of the present invention are as follows:

[0007] In the first aspect, a data query method is proposed, comprising:

[0008] Receiving a data query request triggered by a target object in an operable page, and determining a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query;

[0009] When it is determined that the target query area is the first offline data storage area, searching the first offline data storage area for a target query result that matches the target query condition and the target time information;

[0010] Among them, the first offline data storage area stores query results corresponding to various time information and obtained by operating according to a preset first processing rule. The first processing rule includes executing query operations in corresponding historical data based on various configured query conditions to statistically obtain the query results. The target query condition is included in the various query conditions, and the historical data at least includes content description information of data in a global scope.

[0011] Optionally, when it is determined that the target storage area is the second offline data storage area, the following operations are performed:

[0012] In the second offline data storage area, searching for a target information query result that matches the target query condition and the target time information;

[0013] Among them, the second offline data storage area stores various storage data corresponding to various time information and obtained by operating according to a preset second processing rule. The second processing rule includes aggregating the operated information of each offline data based on the corresponding offline database to obtain the various storage data. The offline data includes various types of information of data in the global scope.

[0014] Optionally, the first processing rule further includes:

[0015] The second engine is periodically used to periodically obtain offline variable data generated within a specified historical period up to the current time information from the corresponding offline database, and process the offline variable data into data to be stored with compliant content;

[0016] Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to periodically perform query operations in the updated historical data based on each configured query condition, and statistically obtain query results corresponding to each query condition;

[0017] Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the query result under the current time information, and store each index keyword and the corresponding query result in the first offline data storage area.

[0018] Optionally, the first processing rule further includes:

[0019] Monitor changes in content description information of data in a global scope in real time, and when it is determined that there is target data with changed content description information, use a third engine to process the target data at the current time information into content-compliant data to be stored;

[0020] Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to perform a query operation in the updated historical data based on each configured query condition, and statistically obtain the real-time query results corresponding to each query condition;

[0021] Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the real-time query result under the current time information, and store each index keyword and the corresponding real-time query result in the first offline data storage area.

[0022] Optionally, the processing of the offline variable data into content-compliant data to be stored includes:

[0023] In each data table included in the offline variable data, at least the content description information stored in correspondence with each data identification information is aggregated into data content in a single table form;

[0024] The text content in the data content is replaced with the content in the corresponding coded form, and the content marked as invalid in the data content is deleted, and the processed data content is used as the data to be stored with compliant content.

[0025] Optionally, updating the historical data stored in the first engine includes:

[0026] Determine each newly added data associated with a newly released state in the offline variable data, and directly store each newly added data in a storage area where historical data is located in the first engine;

[0027] Determine each invalid data associated with a deletion state in the offline variable data, and delete each data corresponding to the each invalid data in the historical data stored in the first engine;

[0028] Determine each to-be-updated data of the associated information adjustment status in the offline variable data, and modify each data corresponding to the to-be-updated data in the historical data stored by the first engine based on the content description information of each to-be-updated data.

[0029] Optionally, generating corresponding index keywords for each query condition includes any one of the following operations:

[0030] Pre-configure corresponding digital value results for each data attribute information in each type, respectively, and pre-configure corresponding weight parameters for each type, and use the digital value results corresponding to each data attribute information constrained by each query condition and the weighted results determined by the weight parameters as index keywords corresponding to each query condition;

[0031] Pre-configure different digital value results for each data attribute information in each type, and respectively determine the digital value results corresponding to each data attribute information constrained by each query condition, and respectively accumulate the binary form of each digital value result constrained by the same query condition as the index keyword generated by the corresponding query condition;

[0032] Different digital value results are configured in advance for each type of data attribute information, and a hash algorithm is used to generate a hash value of a specified length corresponding to each digital value result, as well as each hash value corresponding to each data attribute information constrained by each query condition, to generate index keywords corresponding to each query condition.

[0033] Optionally, each query condition is restricted to a statistical indicator characterizing the query result, and when it is determined that the current time information is the end date of a preset time period, the following operations are performed:

[0034] Determine each time information included in the time period corresponding to the current time information, and establish the data change result of each query condition in the time period under the corresponding statistical indicator based on each query condition and the corresponding query result saved corresponding to each time information;

[0035] Establish a correspondence between the index keywords corresponding to each query condition and the data change results within the time period, and store the index keywords and data change results corresponding to each query condition in the first offline data storage area corresponding to the time period.

[0036] Optionally, the second processing rule further includes:

[0037] Based on the corresponding offline database, the second engine is periodically used to aggregate at least the operated information stored in the data tables corresponding to the offline data corresponding to each data identification information into single-table data with compliant content;

[0038] Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively;

[0039] The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

[0040] Optionally, the second processing rule further includes:

[0041] Monitor the changes of the operated information of the data in the global scope in real time, and when it is determined that there is target data with changed operated information, update the offline data based on the target data at the current time;

[0042] Using the second engine, in each data table corresponding to the updated offline data, at least the operated information stored in the corresponding data identification information is aggregated into a single table of data with compliant content;

[0043] Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively;

[0044] The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

[0045] Optionally, the querying of target query results matching the target query condition and the target time information includes:

[0046] When it is determined that the target object triggers the data query request in the first operable page, a target index keyword is generated based on the target query condition, and a target query result corresponding to the target index keyword under the target time information is searched in the first offline data storage area;

[0047] When it is determined that the target object triggers the data query request in the second operable page, a target identification keyword is generated based on the data identification information carried in the target query condition, and the target query result corresponding to the target identification keyword under the target time information is queried in the second offline data storage area.

[0048] In a second aspect, a data query device is provided, comprising:

[0049] A receiving unit receives a data query request triggered by a target object in an operable page, and determines a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query;

[0050] A query unit, when determining that the target query area is a first offline data storage area, searches the first offline data storage area for a target query result that matches the target query condition and the target time information;

[0051] Among them, the first offline data storage area stores query results corresponding to various time information and obtained by operating according to a preset first processing rule. The first processing rule includes executing query operations in corresponding historical data based on various configured query conditions to statistically obtain the query results. The target query condition is included in the various query conditions, and the historical data at least includes content description information of data in a global scope.

[0052] Optionally, when it is determined that the target storage area is the second offline data storage area, the query unit is configured to perform the following operations:

[0053] In the second offline data storage area, searching for a target information query result that matches the target query condition and the target time information;

[0054] Among them, the second offline data storage area stores various storage data corresponding to various time information and obtained by operating according to a preset second processing rule. The second processing rule includes aggregating the operated information of each offline data based on the corresponding offline database to obtain the various storage data. The offline data includes various types of information of data in the global scope.

[0055] Optionally, the device further includes a storage unit, and the storage unit is used to further perform the following operations using the first processing rule:

[0056] The second engine is periodically used to periodically obtain offline variable data generated within a specified historical period up to the current time information from the corresponding offline database, and process the offline variable data into data to be stored with compliant content;

[0057] Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to periodically perform query operations in the updated historical data based on each configured query condition, and statistically obtain query results corresponding to each query condition;

[0058] Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the query result under the current time information, and store each index keyword and the corresponding query result in the first offline data storage area.

[0059] Optionally, the device further includes a storage unit, and the storage unit is used to further perform the following operations using the first processing rule:

[0060] Monitor changes in content description information of data in a global scope in real time, and when it is determined that there is target data with changed content description information, use a third engine to process the target data at the current time information into content-compliant data to be stored;

[0061] Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to perform a query operation in the updated historical data based on each configured query condition, and statistically obtain the real-time query results corresponding to each query condition;

[0062] Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the real-time query result under the current time information, and store each index keyword and the corresponding real-time query result in the first offline data storage area.

[0063] Optionally, when the offline variable data is processed into content-compliant data to be stored, the storage unit is used to:

[0064] In each data table included in the offline variable data, at least the content description information stored in correspondence with each data identification information is aggregated into data content in a single table form;

[0065] The text content in the data content is replaced with the content in the corresponding coded form, and the content marked as invalid in the data content is deleted, and the processed data content is used as the data to be stored with compliant content.

[0066] Optionally, when updating the historical data stored in the first engine, the storage unit is used to:

[0067] Determine each newly added data associated with a newly released state in the offline variable data, and directly store each newly added data in a storage area where historical data is located in the first engine;

[0068] Determine each invalid data associated with a deletion state in the offline variable data, and delete each data corresponding to the each invalid data in the historical data stored in the first engine;

[0069] Determine each to-be-updated data of the associated information adjustment status in the offline variable data, and modify each data corresponding to the to-be-updated data in the historical data stored by the first engine based on the content description information of each to-be-updated data.

[0070] Optionally, when generating corresponding index keywords for each query condition respectively, the storage unit is used to perform any one of the following operations:

[0071] Pre-configure corresponding digital value results for each data attribute information in each type, respectively, and pre-configure corresponding weight parameters for each type, and use the digital value results corresponding to each data attribute information constrained by each query condition and the weighted results determined by the weight parameters as index keywords corresponding to each query condition;

[0072] Pre-configure different digital value results for each data attribute information in each type, and respectively determine the digital value results corresponding to each data attribute information constrained by each query condition, and respectively accumulate the binary form of each digital value result constrained by the same query condition as the index keyword generated by the corresponding query condition;

[0073] Different digital value results are configured in advance for each type of data attribute information, and a hash algorithm is used to generate a hash value of a specified length corresponding to each digital value result, as well as each hash value corresponding to each data attribute information constrained by each query condition, to generate index keywords corresponding to each query condition.

[0074] Optionally, each query condition is restricted to a statistical indicator characterizing the query result, and when it is determined that the current time information is an end date of a preset time period, the storage unit is used to perform the following operations:

[0075] Determine each time information included in the time period corresponding to the current time information, and establish the data change result of each query condition in the time period under the corresponding statistical indicator based on each query condition and the corresponding query result saved corresponding to each time information;

[0076] Establish a correspondence between the index keywords corresponding to each query condition and the data change results within the time period, and store the index keywords and data change results corresponding to each query condition in the first offline data storage area corresponding to the time period.

[0077] Optionally, the query unit is configured to further perform the following operations using the second processing rule:

[0078] Based on the corresponding offline database, the second engine is periodically used to aggregate at least the operated information stored in the data tables corresponding to the offline data corresponding to each data identification information into single-table data with compliant content;

[0079] Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively;

[0080] The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

[0081] Optionally, the query unit is configured to further perform the following operations using the second processing rule:

[0082] Monitor the changes of the operated information of the data in the global scope in real time, and when it is determined that there is target data with changed operated information, update the offline data based on the target data at the current time;

[0083] Using the second engine, in each data table corresponding to the updated offline data, at least the operated information stored in the corresponding data identification information is aggregated into a single table of data with compliant content;

[0084] Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively;

[0085] The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

[0086] Optionally, when searching for a target query result that matches the target query condition and the target time information, the query unit is used to:

[0087] When it is determined that the target object triggers the data query request in the first operable page, a target index keyword is generated based on the target query condition, and a target query result corresponding to the target index keyword under the target time information is searched in the first offline data storage area;

[0088] When it is determined that the target object triggers the data query request in the second operable page, a target identification keyword is generated based on the data identification information carried in the target query condition, and the target query result corresponding to the target identification keyword under the target time information is queried in the second offline data storage area.

[0089] In a third aspect, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any one of the methods described in the first aspect are implemented.

[0090] In a fourth aspect, a computer-readable storage medium stores a computer program thereon, so that when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.

[0091] The beneficial effects of the present invention are as follows:

[0092] In an embodiment of the present disclosure, a data query method, device, electronic device and storage medium are proposed, which receive a data query request triggered by a target object in an operable page, determine a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query; when the target query area is determined to be a first offline data storage area, query a target query result matching the target query condition and the target time information in the first offline data storage area, wherein the first offline data storage area stores query results obtained by operating according to a preset first processing rule and corresponding to each time information, wherein the first processing rule includes executing a query operation in corresponding historical data based on each configured query condition to obtain the query result by statistics, wherein the target query condition is included in each query condition, and the historical data at least includes content description information of data in a global scope.

[0093] In this way, the full amount of data is no longer archived by date, but the query results of the data are directly stored. After executing the query according to various possible query conditions in advance, it is only necessary to store the query results corresponding to all query conditions to meet the subsequent query needs. At the same time, considering that the query conditions in the actual query process are controllable, the method of storing query results can greatly reduce the amount of data that needs to be stored, which can not only improve the query efficiency of the data, but also improve the controllability of the query time of the data, and can quickly respond to the query needs of related objects and improve the user experience of related objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 It is a query interaction diagram based on periodic storage operations in an embodiment of the present disclosure;

[0095] Figure 2a This is a schematic diagram of a hierarchical structure for storing data in a first offline data storage area according to a first processing rule in an embodiment of the present disclosure;

[0096] Figure 2b A schematic diagram of an implementation process of storing data in a first offline data storage area in an embodiment of the present disclosure;

[0097] Figure 2c It is a flowchart of processing offline variable data into content-compliant data to be stored in an embodiment of the present disclosure;

[0098] Figure 3a It is a schematic diagram of another implementation process of storing data in the first offline data storage area in the real-time example of the present disclosure;

[0099] Figure 3b A schematic diagram of the operation process of obtaining data to be stored with compliant content in an embodiment of the present disclosure;

[0100] Figure 4 A schematic diagram of an implementation flow of storing data in a second offline data storage area in an embodiment of the present disclosure;

[0101] Figure 5 It is a schematic diagram of another implementation process of storing data in the second offline data storage area in an embodiment of the present disclosure;

[0102] Figure 6a A schematic diagram of a data query process in an embodiment of the present disclosure;

[0103] Figure 6b A schematic diagram showing the display of query results in the implementation of the present disclosure;

[0104] Figure 6c A schematic diagram of query results displayed based on query conditions in an embodiment of the present disclosure;

[0105] Figure 6d A schematic diagram of an operable page for triggering a query in a second offline data storage area in an embodiment of the present disclosure;

[0106] Figure 6e A result schematic diagram of the query result presented based on the query result obtained in the second offline data storage area according to an embodiment of the present disclosure;

[0107] Figure 6f This is a schematic diagram of a query result obtained based on monitoring of data in an online database in an embodiment of the present disclosure;

[0108] Figure 7 Schematic diagram of the logical structure of the data query device in the embodiment of the present disclosure;

[0109] Figure 8Schematic diagram of the physical structure of the data query device in the embodiment of the present disclosure. DETAILED DESCRIPTION

[0110] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the technical solution of the present disclosure, rather than all the embodiments. Based on the embodiments recorded in the present disclosure document, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present disclosure.

[0111] The terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.

[0112] For ease of understanding, the terms involved in the embodiments of the present disclosure are explained below:

[0113] Offline calculation: A method of calculating a large amount of data. The real-time requirements for the input data used in the calculation process are not high. All input data are known before the calculation begins. The input data used in the calculation process will not change. The calculation is performed on the premise that the result must be obtained immediately after the calculation is performed according to the set conditions. The offline calculation process can be performed periodically according to actual processing needs. In the embodiment of the present disclosure, offline operations can be performed periodically on offline data. For example, under the processing method with a timeliness of T+1, for the data generated on the same day, it can be used as input data for offline calculations on the next day.

[0114] Offline database: A storage area for storing offline data, which stores all data-related content in the global scope, including content description information of the data, information about the data being operated, and information about the source of the data. In the disclosed embodiment, the offline database can be a data warehouse tool (hive) based on the distributed system infrastructure (Hadoopde). Hive provides functions such as data extraction, transformation, and loading. It is a mechanism that can store, query and analyze large-scale data stored in Hadoop. The offline database can be understood as a subset of the online database, which is equivalent to a backup of the data generated in the online database within a certain period of time. For example, under the T+1 processing mode, the offline database stores all data stored in the online database before the time node under the T+1 processing mode.

[0115] Data content description information: is information used to describe data, and the content description information of data may include different contents for different types of data. For example, in the scenario of processing audio data, the online data is audio data, such as songs, and the data content description information includes basic information of the audio data, the style of the audio data, the artist information involved in the audio data, the lyrics format and other information. The basic information includes song name, song status, song alias, song language and the like.

[0116] Data operation information: is information used to characterize the operation status of data by users. For example, in the scenario of processing audio data, the operation information of audio data may include the number of likes, collections, and reposts corresponding to the audio data.

[0117] The first engine: specifically a search engine, which can provide data retrieval services, perform query operations based on the provided query conditions, and realize real-time search and result aggregation. In other words, a search engine can be understood as a retrieval technology that can use configured strategies to retrieve results from the index and provide feedback based on user needs and preset algorithm logic.

[0118] The second engine is specifically an engine with computing functions. In the embodiment of the present disclosure, the second engine can provide a data channel between the offline database and the first engine, process the data obtained from the offline database into data that meets the needs, and then provide the processed data to the first engine. For example, the second engine can specifically be a spark engine, or a self-developed engine that can realize the above functions.

[0119] The third engine: an engine with computing functions. In the embodiment of the present disclosure, the third engine can provide a data channel between the online database and the first engine, process the data in the online database with content description information changes or operation conditions changes into data that meets the needs, and provide the processed data to the first engine. For example, the third engine can specifically be a Flink engine, or a self-developed engine that can realize the above functions.

[0120] The index keyword, also called index, in the embodiment of the present disclosure, specifically refers to a storage structure unit of data in the first engine.

[0121] Persistence: refers to permanently storing data in a specified storage area, which may be a storage device or medium, such as a disk. In the disclosed embodiment, it mainly refers to storing objects in memory in a database, wherein all data in the first engine are stored in memory.

[0122] Processing equipment: It can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.

[0123] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.

[0124] Under relevant technologies, in order to meet the query needs of relevant personnel on the operational status of the business, a data display platform capable of displaying the current status of the data is usually provided to relevant personnel, so as to realize a centralized display of the data situation within the business scope and assist relevant personnel in obtaining the data performance under the set query conditions.

[0125] In the specific implementation process of the relevant technology, in order to meet the query requirements under various conditions and cope with various complex query scenarios, distributed search engines such as Elasticsearch are usually used as storage media, and a full set of data is archived every day and partitioned and stored by day. Subsequently, partitioned queries and aggregation are performed in the full data stored in the partitions according to the configured query conditions.

[0126] For example, assuming that there are 50 million songs in a music player application, then under the relevant technology, it is necessary to archive and store at least 50 million songs every day. The total amount of data stored in the past year is nearly 20 billion songs. When it is determined that the growth of the number of songs in a year needs to be counted based on the configured query conditions, the number of songs per day in the past year is queried in the archived data, and the final query results are obtained by aggregation.

[0127] In this way, the implementation scheme under the relevant technology will inevitably have two problems. First, since the full amount of data needs to be archived every day, the amount of stored data is extremely large, which requires a huge storage space. Secondly, under the storage method of the relevant technology, when querying the data performance within a period of time based on complex query conditions, it is necessary to perform query operations separately in the storage data of the corresponding days, and aggregate the query results, which makes the query efficiency very low, and the query time is greatly affected by the query data range and query conditions. The larger the amount of data within the query data range and the more constraints included in the query conditions, the longer the response time to obtain the query results, which makes the query efficiency very low and the query time unstable, which greatly affects the query experience of related objects.

[0128] Application Scenario Overview

[0129] The data query method proposed in the embodiment of the present disclosure can be applied to the scenario of querying data according to the configured query conditions. The present disclosure can be applied to the scenario of querying various types of data, and the data type targeted may be at least one or a combination of audio type data, video type data, or text type data. In addition, according to the actual query needs, the query results that can be obtained by the relevant objects may be the statistical results of the data volume, or the statistical results of the data operation status.

[0130] For example, in a scenario where the operating status of a music and video playback application is queried, when the relevant object wishes to know the changes in the amount of music data and / or video data, or when the relevant object wishes to count the amount of data that meets the query conditions, or when the relevant object wishes to know the operation status of certain data, the corresponding query results can be obtained based on the technical solution proposed in the present disclosure.

[0131] In some optional embodiments of the present disclosure, considering that when the query result targeted by the query condition is a statistical result of the data volume, it is impossible to obtain the result directly based on the information of the data itself, and there are many query situations involved, various possible query conditions can be pre-enumerated according to the data type and query needs, and then based on the configured query conditions, offline query operations are performed on the full amount of historical data respectively, and the query results that meet the constraints of the query conditions are statistically obtained, wherein the query conditions may be restricted to statistical indicators corresponding to the query results, and the statistical indicators are used to determine the expression form of the query results obtained by statistics. The possible statistical indicators in the present disclosure include the total amount at a specified time granularity, the increment at a specified time granularity, and the proportion at a specified time granularity.

[0132] In this way, the solution proposed in the present disclosure changes the original data storage method. It no longer archives the full amount of data by date, but directly stores the query results of the data. After executing the query in advance according to various possible query conditions, it is only necessary to store the query results corresponding to all the query conditions to meet the subsequent query needs. At the same time, considering that the query conditions in the actual query process are controllable, the method of storing the query results can greatly reduce the amount of data that needs to be stored, which can not only improve the query efficiency of the data, but also improve the controllability of the query time of the data, and can quickly respond to the query needs of related objects and enhance the user experience of related objects.

[0133] In some other optional embodiments of the present disclosure, when the query result targeted by the query condition is a statistical result of the data being operated, the query result under the specified time information can be obtained directly based on the queried data, so that in the storage process, the entire amount of data is directly stored after being cleaned, processed, etc.

[0134] In this way, compared with the form of directly storing the full amount of data under related technologies, directly storing part of the data information in the form of key-value in this scenario can also reduce the storage space occupied by data storage and improve data query efficiency.

[0135] See also Figure 1 As shown, it is a query interaction diagram based on periodic storage operations in an embodiment of the present disclosure. Figure 1 The content shown is that in the scenario of periodically performing storage operations, the processing device uses the second engine to obtain offline data from the offline database to perform offline calculations, wherein the second engine is equivalent to providing an offline computing system. The processing device calls the second engine to successively perform data query, data cleaning, and data conversion and processing operations based on the offline data in the offline database to obtain data with compliant content. Then, according to actual processing needs, the data with compliant content is input into the first engine, so that the first engine, based on the scheduled task, stores the query results obtained based on the query conditions in the first offline data storage area, or, according to actual processing needs, the data with compliant content can be processed and directly stored in the second offline data storage area. After completing the storage of the data, the processing device selectively obtains the query results from the first offline storage area or the second offline storage area according to the query needs, wherein the processing device can provide a data display platform for related objects and support data query requests initiated by related objects.

[0136] In actual application scenarios, in order to support the query needs of related objects, scheduled tasks are usually executed periodically in the early morning based on offline data to ensure that related objects can obtain stored data processed based on the latest updated offline data before going to work.

[0137] In the embodiment of the present disclosure, considering that the data query process depends on the data storage process, the storage process in different application scenarios is first described below in conjunction with the accompanying drawings:

[0138] It should be noted that there are two types of full data in the embodiments of the present disclosure. One is the full data in the online database, which is used to support online data display, so the real-time and latest data is always stored. The other is the full data in the offline database. The full data in the offline database is equivalent to a subset of the full data in the online database. Usually, when the offline incremental data is periodically obtained, the data in the online database is stored in the offline database. Among them, in the embodiments of the present disclosure, the online database can be a storage area under the relational database management system (RDBMS) architecture. The offline database usually uses the hive database, and the online database is specifically a mysql database under the RDBMS.

[0139] Scenario 1: Store the query results at the data volume level under each query condition.

[0140] In the embodiment of the present disclosure, when the application process does not focus on the specific content description information of the data, but only focuses on the statistical results of the data under the content description information of a specified type, the processing method shown in scenario 1 can be adopted. For each of the configured limited number of query conditions, the query results are pre-queried to obtain the query results, and then only the query results under each query condition are stored during the storage process, without storing the full amount of data, so as to save storage costs and improve query efficiency.

[0141] It should be noted that, in the storage process defined in scenario 1, the processing device can set a quartz timer task to regularly schedule the query and storage operations of the executed data. In the specific processing process, the processing device adopts a first processing rule, and obtains query results according to the configured query conditions in the historical data up to the configured storage time, and stores each query result corresponding to the storage time in the first offline data storage area, wherein the first processing rule includes at least the configured query conditions and the storage time that triggers the execution of the query and storage operations. The first offline data storage area can specifically be a persistent storage area different from the memory, such as a MySQL database.

[0142] In the embodiment of the present disclosure, when the processing device performs specific query and storage operations according to the storage timing of the query and storage operations triggered in the first processing rule, there may be the following two possible situations depending on the query timing: one is to periodically perform data query and storage operations, and the other is to monitor data changes in real time and trigger the execution of data query and storage operations when the data changes.

[0143] See also Figure 2aAs shown, it is a schematic diagram of the hierarchical structure for storing data in the first offline data storage area according to the first processing rule in an embodiment of the present disclosure. In the hierarchical structure constructed based on scenario 1, it includes a data processing layer, a data service layer, a storage layer, and a business service layer.

[0144] The data processing layer includes data acquisition operations performed on the online database and the offline database respectively, and is used to use the engine to query the data that needs to be processed from the database respectively. The operations implemented by the data processing layer include collecting the data that needs to be processed, cleaning and processing the data, and other operations.

[0145] The data service layer is used to call the first engine based on various query conditions configured, perform query operations in the updated historical data stored in the first engine, and obtain query results.

[0146] The storage layer is used to map the query results obtained based on the query conditions into the form of index keywords and corresponding query results, and store the corresponding relationship between the index keywords and the query results in the first offline data storage area.

[0147] The business service layer is used to support related objects to initiate query requests and provide data query services for related objects.

[0148] In the following description, the corresponding storage processes will be described respectively for two possible storage opportunities that trigger query and storage operations.

[0149] Scenario 1.1: Periodically trigger the execution of data query and storage operations.

[0150] See also Figure 2b As shown in FIG. 1 , it is a schematic diagram of an implementation process of storing data in the first offline data storage area in an embodiment of the present disclosure. Figure 2b , the process of storing data in the first offline data storage area by the processing device in the embodiment of the present disclosure is described:

[0151] Step 201: The processing device periodically uses the second engine to regularly obtain offline variable data generated within a specified historical period up to the current time information from the corresponding offline database, and processes the offline variable data into content-compliant data to be stored.

[0152] Specifically, the processing device can periodically use the second engine according to the first processing rule to regularly obtain offline variable data generated within a specified historical period up to the current time information from the corresponding offline database, and continue to use the second engine to process the offline variable data into content-compliant data to be stored, wherein, in the case of periodically obtaining offline variable data, the specified historical period refers to the time interval between two adjacent periods, and the current time information refers to the time information corresponding to the current period.

[0153] In the embodiment of the present disclosure, when the processing device uses the second engine to process the offline variable data into content-compliant data to be stored, the following operations may be performed:

[0154] See also Figure 2c As shown in the figure, it is a schematic diagram of the process of processing offline variable data into content-compliant data to be stored in the embodiment of the present disclosure. Figure 2c , the processing flow of offline variable data is explained:

[0155] Step 201a: The processing device aggregates at least the content description information stored corresponding to each data identification information in each data table included in the offline variable data into data content in a single table form.

[0156] Specifically, considering that data is currently stored in data tables, and different content description information of the same data is stored in different data tables, when the processing device processes offline variable data, it is usually necessary to aggregate the content description information of the same data from different data tables to obtain the data content in a single table, wherein the offline variable data is the offline data generated in the offline database between the current cycle and the previous cycle.

[0157] It should be noted that the content aggregated from multiple data tables is set according to actual processing needs, and the present disclosure does not impose too many restrictions on this.

[0158] For example, for stored data, it is usually stored in multiple tables. Taking song data as an example, there is a song basic information table (including storing basic information such as song name, status, alias, language, etc.), song style table (recording song style), song artist table, song lyrics table, etc. Considering that the data stored in the subsequent first engine has only one index structure, it is necessary to aggregate the contents corresponding to the same data storage in different data tables into data content in the form of a single table. In other words, aggregate offline variable data from multiple tables into a single table.

[0159] In this way, the storage format of the data can be sorted according to actual processing needs, and different contents of the same data of concern in multiple data tables can be aggregated to obtain the data content in a single table.

[0160] Step 201b: The processing device replaces the text content in the single-table data content with the corresponding coded content, deletes the content marked as invalid in the data content, and uses the processed data content as compliant data to be stored.

[0161] In the disclosed embodiment, after obtaining the data content in the form of a single table, the processing device needs to further process and clean the data content in the form of a single table in order to improve the query performance in the subsequent query process for the data.

[0162] Specifically, the processing device can convert text content in the data content that is difficult to read into encoded content, and delete content marked as invalid in the data content based on the status of the data, and then use the processed data content as compliant content to be stored.

[0163] For example, for song data, the language of the song data is stored in the data table in text form. In order to ensure accurate identification of the language, the text content can be converted into encoded content in data form, such as representing "language 1" with the number 1, and "language 2" with the number 2, etc.

[0164] For another example, considering that there are various forms of data content information stored in the data table for data, but it may not be necessary to pay attention to the actual query process, you can set filtering rules according to actual processing needs, such as filtering out deleted and uncopyrighted data from the data content, deleting the publishing source information of the data, etc.

[0165] In this way, by processing the data content, the data content can be processed into a content form that can be normally recognized, which provides convenience for subsequent query operations performed based on the data content.

[0166] It should be noted that in the implementation of the present disclosure, for the processing operations limited by steps 201a and 201b, the present disclosure does not specifically limit the execution order of the two operations. In the actual application process, according to the actual processing needs, it is also possible to selectively execute the data cleaning and processing operations limited by step 201b first, and then execute the data aggregation operations limited by step 201a. The present disclosure does not make specific restrictions here.

[0167] Step 202: The processing device updates the historical data stored in the first engine based on the data to be stored, and calls the first engine to periodically perform query operations in the updated historical data based on the configured query conditions, and obtains statistical query results corresponding to each query condition.

[0168] Specifically, after the processing device obtains the data to be stored, it updates the historical data stored in the first engine based on the data to be stored, wherein the first engine stores the full amount of historical data, and the historical data can be understood as partial information of the full amount of data in the offline data under the same time span. All data in the first engine can be understood as stored in the memory, and the first engine can be understood as having its own storage medium.

[0169] For example, assuming that the business function of audio data begins to be used on September 16, and there are 5 million audio data on September 16, then on September 17, the first engine will store at least 5 million audio data content description information. Assuming that 200,000 audio data are added on September 17, on September 18, the first engine will have 5.2 million audio data content description information. At the same time, if 1 million audio data out of the 5 million audio data have their content description information updated on September 17, we will also update these 1 million audio data to the latest status on t+1 on the 18th. In other words, the first engine always stores historical data processed based on the latest full data on t+1. In a special and extreme case, if the audio data has not changed for a month, it can be understood that the same data is stored in the first engine for the past month.

[0170] When the processing device updates the historical data stored in the first engine, the following three possible update operations may be stored in the actual update process, taking into account the changes and adjustments made to the data to be stored relative to the stored historical data:

[0171] Operation 1: Based on the data to be stored, data is added to the historical data stored in the first engine.

[0172] The processing device determines each newly added data associated with a newly released state in the offline variable data, and directly stores the each newly added data in a storage area where the historical data is located in the first engine.

[0173] Specifically, when the processing device determines that each newly added data with a newly released status is associated with the offline variable data, the corresponding newly added data is added to the historical data stored in the first engine. In other words, the newly added data is stored in the storage area where the historical data is located.

[0174] For example, assuming that the query and storage operations of data are periodically performed at 0:00 a.m. every day, the offline variable data obtained by the scheduled query at 0:00 a.m. every day is specifically: offline data generated between 0:00 a.m. of the previous day (the start of the previous cycle) and 0:00 a.m. of the current day (the start of the current cycle). After obtaining the data to be stored based on the offline variable data, if it is determined that there are new data in the data to be stored, the data addition operation is performed in the historical data stored by the first engine based on each new data, wherein the historical data stored by the first engine is obtained based on the offline data processing up to 0:00 a.m. of the previous day, and the historical data may be composed of content description information of each offline data.

[0175] Operation 2: based on the data to be stored, delete the data in the historical data stored in the first engine.

[0176] The processing device determines each invalid data associated with a deletion state in the offline variable data, and deletes each data corresponding to the each invalid data in the historical data stored by the first engine.

[0177] For example, assuming that the query and storage operations of data are periodically performed at 0:00 a.m. every day, the offline variable data after 0:00 a.m. every day is specifically: offline data generated between 0:00 a.m. of the previous day (the start of the previous cycle) and 0:00 a.m. of the current day (the start of the current cycle). After obtaining the data to be stored according to the offline variable data, if it is determined that there is invalid data to be deleted in the data to be stored, a data deletion operation is performed in the historical data stored by the first engine based on each invalid data, wherein the historical data stored by the first engine is obtained based on the offline data processing up to 0:00 a.m. of the previous day.

[0178] Operation three: based on the data to be stored, modify the data in the historical data stored in the first engine.

[0179] The processing device determines each to-be-updated data of the associated information adjustment status in the offline variable data, and modifies each data corresponding to the to-be-updated data in the historical data stored by the first engine based on the content description information of the each to-be-updated data.

[0180] For example, assuming that the query and storage operations of data are periodically performed at 0:00 a.m. every day, the offline variable data after 0:00 a.m. every day is specifically: offline data generated between 0:00 a.m. of the previous day (the start of the previous cycle) and 0:00 a.m. of the current day (the start of the current cycle). After obtaining the data to be stored according to the offline variable data, if it is determined that there are modified data to be updated in the data to be stored, the data modification operation is performed in the historical data stored by the first engine based on each data to be updated, wherein the historical data stored by the first engine is obtained based on the offline data processing up to 0:00 a.m. of the previous day.

[0181] In this way, according to the status of the data included in the data to be stored, any one or a combination of operations of adding, deleting, and modifying the data is performed in the historical data stored in the first engine, so that the historical data is updated to the latest status in the current cycle, so that the first engine always saves the historical data obtained based on the latest offline data processing, thereby ensuring the reliability and validity of the data based on which subsequent query operations are based.

[0182] Furthermore, after the processing device updates the historical data stored in the first engine, the processing device calls the first engine to periodically perform query operations in the updated historical data based on the configured query conditions, and statistically obtains query results corresponding to each query condition, wherein the query conditions are obtained by exhaustively combining the data attribute information on which the query is based.

[0183] It should be noted that the content description information of the data includes various types of data attribute information. Considering that in the actual query process, some or all types of data attribute information are usually used as elements for generating query conditions based on actual query needs, the query conditions that can be generated are necessarily limited. The processing device can provide the selected types of data attribute information to the relevant objects according to the actual query process, and exhaustively generate various query conditions.

[0184] For example, it is assumed that the data attribute information types that the processing device configures for song data in the actual query process and generates the target query conditions include: the language of the song data, such as "language 1", "language 2", "language 3", "language 4", and "all" and other language restrictions; the style of the song data, such as "pop", "rock", "folk", and "all" and other style restrictions; the region of the song data, such as "region 1", "region 2", "region 3", and "all" and other region restrictions; the release status of the song data, such as "released", "unreleased", and "all" and other status restrictions; the lyrics form of the song data, such as "text lyrics", "line by line lyrics", "word by word lyrics", and "all" and other lyrics form restrictions. Based on this, the processing device can perform a combined query based on multiple types of data attribute information, or can perform a query based on a single data attribute information. For example, when querying all songs in language 1, the language of the song data can be selected as "language 1", and the other data attribute information can be selected as "all".

[0185] For another example, continuing with the above-mentioned assumed attribute information type, in the process of exhaustively enumerating all possible combinations of data attribute information to obtain various query conditions, assuming that there are 10 commonly used languages ​​of song data, 20 styles of song data, 5 regions of song data, 8 release statuses of songs, and 4 lyrics, all of their combinations are 10*20*5*8*4=32,000, so 32,000 query conditions can be generated.

[0186] After generating each query condition, the processing device calls the first engine, performs a query operation in the updated historical data based on each query condition, and obtains query results corresponding to each query condition by statistics.

[0187] In this way, when the number of query conditions is limited, by enumerating all possible query conditions and executing query operations according to each query condition, the corresponding query results can be obtained in advance based on the historical data updated in the current period. This is equivalent to having queried the query results corresponding to various possible query conditions with the help of a scheduled task before the actual query operation occurs.

[0188] Step 203: The processing device generates corresponding index keywords for each query condition, establishes a corresponding relationship between each index keyword and the query result under the current time information, and stores each index keyword and the corresponding query result in the first offline data storage area.

[0189] In the disclosed embodiment, the processing device calls the first engine to perform query operations for each query condition in the updated historical data stored in the first engine to obtain corresponding query results, and generates corresponding index keywords for each query condition, and establishes a corresponding relationship between each index keyword and the query result under the current time information, and stores each index keyword and the corresponding query result in the first offline data storage area.

[0190] It should be noted that, in the embodiment of the present disclosure, an optional operation mode is that each time the processing device calls the first engine to perform a query operation based on a query condition, a corresponding storage operation is performed, so that the generated query results can be stored in time to avoid data loss due to emergencies. In addition, when a storage operation is performed based on the obtained query results, the corresponding current time information stored can refer to the generation time of the currently obtained offline variable data, and at the same time, the current time information can also refer to the deadline corresponding to the updated historical data.

[0191] For example, assuming that data query and storage operations are performed periodically, and offline variable data is obtained at 0:00 a.m. every day, and assuming that the current date is March 26, the offline variable data obtained at 0:00 a.m. on March 26 was generated between 0:00 a.m. on March 25 and 0:00 a.m. on March 26. Therefore, when the query and storage are completed based on the offline variable data, since the generation time of the offline variable data is March 25, and after the historical data stored in the first engine is updated based on the offline variable data, the latest time of the updated historical data is the end time of March 25. Therefore, after performing a query operation in the updated historical data based on each query condition, the current time information corresponding to the storage is March 25.

[0192] In the embodiment of the present disclosure, when the processing device generates corresponding index keywords for each query condition, there may be the following possible generation methods according to different generation forms of the index keywords:

[0193] Generation method 1: The processing device generates an index keyword through bit operation.

[0194] Specifically, the processing device pre-configures corresponding digital value results for each data attribute information in each type, and pre-configures corresponding weight parameters for each type, and uses each data attribute information constrained by each query condition, each corresponding digital value result and a weighted result determined by the weight parameter as the index keyword corresponding to each query condition.

[0195] For example, assuming that the data attribute information types that can be selected for song data are: language, style, region, and popularity of the song data, then when generating query conditions, you can select n (n>=0) types of data attribute information from these four types and combine them to obtain query conditions. Among them, for the unselected data attribute information types, the default setting can be to select "all" in the corresponding type of data attribute information.

[0196] Assume that the language of the song data includes 5 selectable languages, namely language 1-5; the style of the song data includes 20 selectable styles; the popularity of the song data includes 4 selectable hotnesses; and the region of the song data includes 10 selectable regions. Then, the weight parameter can be configured as 1 for the language of the song data in advance, and the data value results can be configured to occupy 1 bit for various selectable languages, such as 1 for language 1, 2 for language 2, 3 for language 3, etc.; the weight parameter can be configured as 10 for the style of the song data, and the data value results of the style of the song data can be configured to occupy 2 bits, such as 12 for popularity, etc.; the data value result of the popularity of the song data can be configured to occupy 1 bit, and the weight parameter can be configured to 1000, such as 2 for S-level popularity; the data value result of the region of the video data can be configured to occupy 1 bit, and the weight parameter can be configured to 10000, such as 3 for region 1. Then, for the query conditions constrained by the data attribute information of "region 1, language 1, popularity, S-level popularity", the index keyword key generated by the corresponding query condition is: 1*1+12*10+2*1000+3*10000=32121.

[0197] Generation method 2: The processing device generates index keywords through binary operation.

[0198] Specifically, the processing device pre-configures different digital value results for each data attribute information in each type, and determines each data attribute information constrained by each query condition, the corresponding digital value results, and accumulates the binary form of each digital value result constrained by the same query condition as the index keyword generated by the corresponding query condition.

[0199] For example, it is assumed that the language of the song data includes 5 selectable languages, namely language 1-5; the style of the song data includes 20 selectable styles; the popularity of the song data includes 4 selectable hotnesses; and the region of the song data includes 10 selectable regions. Then, in the actual processing process, the digital value result can be set for the language of the song data, such as setting the digital value result corresponding to language 1 to 1, and the corresponding binary form is 0000 0001; for the style of the song data, for example, setting the digital value result corresponding to popularity to 13, and the corresponding binary form is 0000 1101; for the popularity of the song data, for example, setting the data value result corresponding to S-level popularity to 27, then the corresponding binary form is 0001 1011; for the region of the song data, for example, setting the digital value result corresponding to region 1 to: 36, then the corresponding binary form is 0010 0100, then for the query conditions constrained by the data attribute information of "region 1 language 1 popularity S-level popularity", the superposition result of the binary form of each data attribute information is: 0000 0001+0000 1101+0001 1011+0010 0100=01001101, that is, the index keyword key generated corresponding to the query condition is: 0100 1101.

[0200] Generation method three: The processing device generates an index keyword through a hash operation.

[0201] Specifically, the processing device pre-configures different digital value results for each type of data attribute information, and uses a hash algorithm to generate a hash value of a specified length corresponding to each digital value result, as well as each hash value corresponding to each data attribute information constrained by each query condition, to generate index keywords corresponding to each query condition.

[0202] In the disclosed embodiment, when generating an index keyword corresponding to a query condition, the data attribute information included in the query condition is first determined, and then a hash algorithm is used to generate a corresponding hash value for the data attribute information, and the concatenated hash value can be selectively used as the index keyword for the corresponding query condition, or the superposition result of the hash value corresponding to the data attribute information can be selectively used as the index keyword for the corresponding query condition.

[0203] In this way, various possible generation methods can be used to generate index keywords that uniquely correspond to each query condition, which are used to distinguish each query condition. This is equivalent to mapping unique values ​​to query conditions that combine complex data attribute information. In the subsequent query process based on the stored query results, the mapped index keywords are used instead of the query conditions to perform matching operations, which can further improve the query processing performance.

[0204] Furthermore, after establishing a correspondence between each index keyword and a query result under current time information, the processing device stores each index keyword and the corresponding query result in the first offline data storage area.

[0205] Specifically, the processing device calls the first engine based on each query condition constructed, executes query operations respectively in the updated historical data stored in the first engine, obtains query results statistically, and then obtains the index keyword uniquely corresponding to each query condition through index keyword generation methods such as binary, bit operations or hash operations, and then persistently stores the query results, that is, the index keywords and query results are stored correspondingly in the first offline data storage area.

[0206] In addition, considering that the data in the offline database comes from the online database, the storage media in the storage process illustrated in scenario 1.1 of the present disclosure are sequentially: first transmitted from the online database to the offline database, then transmitted from the offline database to the memory of the first engine to store data, and then transmitted from the memory of the first engine to store data to the first offline data storage area.

[0207] In summary, in the embodiment of the present disclosure, the offline database and the first offline data storage area belong to different databases. The offline database may specifically be hive, and the first offline data storage area may specifically be a mysql database under RDBMS; the online database may also be a mysql database under RDBMS. In other words, the content stored in the online database and the content stored in the first offline data storage area may be stored in different data tables of the mysql database.

[0208] For example, in the scenario of querying song data, for the query condition of "region 1, language 1, popular S-level popularity", the corresponding index keyword is: key = 32121, and the first engine is called to query the total number of songs 1800w, date 2021-03-25, so key can be used as the primary key, and the query results and the current date can be used as the corresponding values ​​and stored in the first offline storage area. Alternatively, the index keyword key can be used as the primary key, and the query results can be used as the value corresponding to the primary key and stored in the storage area set for the corresponding date 2021-03-25.

[0209] Under the storage method corresponding to the current scenario 1.1 of the present disclosure, the total amount that the processing device needs to store is: the total number of query conditions*the number of days+a full set of historical data.

[0210] In this way, when the total number of possible query conditions is limited, each time a data query and storage operation is performed, only the query results equal to the total number of query conditions need to be persistently stored in the first offline data storage area, so that all query scenarios can be covered. Compared with the storage method of storing full data every day in the related art, the storage cost can be greatly reduced.

[0211] In particular, in the embodiments of the present disclosure, it is considered that in various query scenarios, the statistical indicators that characterize the query results may be different. For example, in some application scenarios, it may be desired to obtain the total amount of data corresponding to the query conditions, in other application scenarios, it may be desired to obtain the proportion of the total amount of data under the query conditions to the total amount of data, and in some application scenarios, it may be desired to obtain the data increment under the query conditions. Therefore, when storing data, the processing device can organize the query results into a suitable data format based on the statistical indicators specified in the query conditions. In particular, when the statistical indicators are not specified in the query results, the total amount of data under the corresponding query conditions is queried by default.

[0212] On this basis, in the embodiments of the present disclosure, in order to take into account the application scenarios of obtaining query results at different time granularities, it is necessary to determine the data change results at the time granularity under the corresponding statistical indicators according to the preset time granularity, and store the data change results obtained after processing. Among them, the time granularity involved in the embodiments of the present disclosure is represented by means of time periods set according to time units such as weeks, months, quarters, and years.

[0213] Specifically, when the processing device determines that each query condition is restricted with a statistical indicator characterizing the query result and determines that the current time information is the end date of a preset time period, the processing device determines each time information included in the time period corresponding to the current time information, and establishes the data change results of each query condition within the time period under the corresponding statistical indicator based on each query condition and the corresponding query result saved corresponding to each time information, and then establishes a correspondence between the index keyword corresponding to each query condition and the data change results within the time period, and stores the index keyword and the corresponding data change results corresponding to each query condition in the first offline data storage area corresponding to the time period.

[0214] For example, assuming that the statistical indicator constrained in the query condition is the data increment, the periodically stored query results are specifically the difference between the total amount of data obtained by the current query condition query and the total amount of data obtained by the previous period query, that is, the data increment under the corresponding query condition. Assuming that it is necessary to determine the new audio data every month, every week, and every quarter, the daily data increment can be obtained through the task executed on a daily basis, and a week, a month, and a year can be used as the time period respectively. Taking a week as an example, assuming that the current time is the last day of the week, the data increment of each day in the week can be accumulated to obtain the data change result within a week, that is, the data increment within a week, and in the first offline data storage area, the index keywords and the corresponding data change results corresponding to each query condition are stored respectively corresponding to the current week.

[0215] In this way, additional conditions for data query can be generated with the help of time granularities such as day, week, month, quarter, and year, so that when querying, data can be directly returned according to the time granularity selected by the relevant objects, thereby further helping to improve the query performance of the data.

[0216] At the same time, with the help of the data query and storage process proposed in the disclosed scenario 1.1, after the historical data stored in the first engine is updated based on the latest processed offline variable data, the query results under various query conditions are calculated through a scheduled task, and the query results and query conditions are bound to generate a unique index keyword and the corresponding value form, which are persistently stored in the first offline data storage area, so that all subsequent query operations only need to convert the query conditions into corresponding index keywords, and then obtain the corresponding query results through simple data query, which saves storage costs and improves data query efficiency.

[0217] Scenario 1.2: Detect data changes in real time and trigger data query and storage operations when data changes.

[0218] See also Figure 3a As shown, it is another implementation flow diagram of storing data in the first offline data storage area in the real-time example of the present disclosure. Figure 3a , the process of storing data in the first offline data storage area by the processing device in the embodiment of the present disclosure is described:

[0219] Step 301: The processing device monitors the changes in content description information of data in the global scope in real time, and when it is determined that there is target data with changed content description information, a third engine is used to process the target data under the current time information into content-compliant data to be stored.

[0220] For details, see Figure 3b As shown, it is a schematic diagram of the operation process of obtaining content-compliant data to be stored in an embodiment of the present disclosure. The processing device can use a data acquisition system to monitor in real time the changes in the content description information of the global data stored in the online database. Among them, every time the data in the online database of the RDBMS architecture changes, a change record will be generated in real time and recorded in the binary log (binlog) file.

[0221] After the processing device uses the data acquisition system to automatically monitor the changes in the binlog file, it performs a structured conversion on the changed data in the binlog file into a form that can be transmitted using the message queue transmission method, and then transmits the target data with changed content description information to the third engine through the message queue, and calls the third engine to process and convert the target data with changed content description information in the message queue into content-compliant data to be stored, and then transmits the data to be stored to the first engine through the message queue, wherein the message queue is used for communication between different processing components, converts data into messages, and transmits them in the message queue through messages. Kafka is a commonly used message queue.

[0222] It should be noted that the process of processing the target data into content-compliant data to be stored involved in step 301 of the present disclosure can adaptively adopt the processing form indicated in steps 201a-201b, and the present disclosure will not further explain it in detail.

[0223] Step 302: The processing device updates the historical data stored in the first engine based on the data to be stored, and calls the first engine to perform query operations in the updated historical data based on the configured query conditions, and obtains statistical real-time query results corresponding to each of the query conditions.

[0224] Specifically, after the processing device obtains the data to be stored, it updates the historical data stored in the first engine based on the data to be stored, calls the first engine, and performs query operations in the updated historical data based on the configured query conditions, and statistically obtains the real-time query results corresponding to each query condition. Among them, the operation of updating the historical data involved in step 302 and the execution of query operations based on the query conditions are the same as the implementation method described in scenario 1.1, and will not be repeated here.

[0225] Step 303: The processing device generates corresponding index keywords for each query condition, establishes a corresponding relationship between each index keyword and the real-time query result under the current time information, and stores each index keyword and the corresponding real-time query result in the first offline data storage area.

[0226] Specifically, the processing device can continue to use the index keyword generation method shown in step 203 in scenario 1.1 to generate corresponding index keywords for each query condition, establish a corresponding relationship between each index keyword and the real-time query result under the current time information, and store each index keyword and the corresponding real-time query result in the first offline data storage area.

[0227] It should be noted that, in the disclosed embodiment, the first offline storage area for storing query results in scenario 1.1 and the first offline data storage area for storing query results in scenario 1.2 may respectively correspond to different data tables in a database, such as different data tables in a MySQL database.

[0228] In this way, in the processing method disclosed in scenario 1.2, the processing device monitors the data changes in the online database, and directly updates the historical data in the first engine based on the acquired data changes in the online database, and calls the first engine to perform query operations based on various query conditions to obtain more real-time query results compared to periodic queries and storage, so that the stored query results can represent the status of the online data and provide a query basis for related objects to query the query results of real-time changes.

[0229] Scenario 2: Storing query results that represent the data being operated on.

[0230] In the embodiment of the present disclosure, when the operation status of the data is concerned during the application process, the query results illustrated in the scenario 2 of the present disclosure can be used to obtain the operation status of the data within the set time range, such as the changes in the operation status of the data such as the number of views, the number of likes, the number of favorites, and the number of comments. In the query process based on the storage content obtained in scenario 2, the query condition can be the identification information of the data, which is used to obtain the operation status corresponding to the data. The identification information of the data and the information representing the operation status of the data can be stored in the form of key-value, so you only need to pay attention to the storage process of the data. In addition, in the embodiment of the present disclosure, the implementation process of scenario 1 is independent of the implementation process of scenario 2, and the present disclosure does not specifically limit the order of execution of scenario 1 and scenario 2.

[0231] It should be noted that, considering that scenario 2 of the present disclosure focuses on reflecting the operation of data, the amount of data that needs to be stored in a targeted manner is larger than that of scenario 1. The data stored in scenario 2 only needs to undergo simple data aggregation, processing, and conversion operations, and can be directly stored in the second offline data storage area. Therefore, considering that the amount of data that needs to be saved in scenario 2 is large, the second offline data storage area can be different from the first offline data storage area. For example, the first offline data storage area can use RDBMS as its storage medium, such as MySQL database; the second offline data storage area can use HBase as its storage medium, wherein HBase has better support for massive data storage, and only supports key-value queries, and does not support multi-condition combination queries.

[0232] In the embodiment of the present disclosure, when the processing device performs specific processing and storage operations according to the storage timing of the data processing and storage operations triggered in the second processing rule, there may be the following two possible situations according to the different timings of performing data processing and storage: one is to periodically perform data processing and storage operations, and the other is to detect data changes in real time and trigger the execution of data processing and storage operations when the data being operated changes. In the following description, the corresponding processing and storage processes will be described for the two possible situations respectively:

[0233] Scenario 2.1: Periodically trigger the execution of data processing and storage operations.

[0234] participate Figure 4 As shown in FIG. 1 , it is a schematic diagram of an implementation process of storing data in the second offline data storage area in an embodiment of the present disclosure. Figure 4 , the process of storing data in the second offline data storage area by the processing device in the embodiment of the present disclosure is described:

[0235] Step 401: The processing device periodically uses the second engine based on the corresponding offline database to aggregate at least the operated information stored in the corresponding data identification information in each data table corresponding to the offline data into a single table data with compliant content.

[0236] Specifically, the processing device can periodically use the second engine according to the second processing rule to perform data aggregation, processing, and conversion operations on the full amount of data included in the corresponding offline database to obtain single-table data with compliant content, wherein the data content aggregated by the second engine is set according to actual processing needs.

[0237] For example, you can configure the system to transfer data from the online database to the offline database at 0:00 a.m. every day, and perform aggregation, processing, and conversion operations on the data stored in multiple data tables in the offline database to obtain single-table data with compliant content.

[0238] Step 402: The processing device generates corresponding identification keywords based on the data identification information of each data in the single table data, and establishes key values ​​corresponding to each identification keyword based on the operated information stored in the single table data corresponding to each data.

[0239] Specifically, after the processing device aggregates the data in the offline database to obtain single-table data with compliant content, it performs the following operations for each data in the single-table data: generates a corresponding identification keyword based on the identification information (Identity, ID) of a data, and uses at least the operated information of the data as the key value corresponding to the identification keyword, wherein the operated information in the key value refers to the total amount of operations on various data recorded in the offline database up to the current cycle.

[0240] It should be noted that in the embodiments of the present disclosure, the content included in the key value corresponding to the configuration keyword can be configured according to actual processing needs. In some optional embodiments, the operated information of the data can be selected as the corresponding key value. In other optional embodiments, in order to adapt to the business needs and functional expansion of the data system, the content description information of the data can be selected as the corresponding key value based on the selected data's operated information.

[0241] For example, the relationship between the identification keyword and the corresponding key value generated by the processing device for an offline data X is as follows: ID1-{the total number of likes of offline data X; the total number of reposts of offline data X; the total number of plays of offline data X; the total number of collections of offline data X; …}.

[0242] For another example, in order to meet business needs, the relationship between the identification keyword and the corresponding key value generated by the processing device for an offline data Y is as follows: ID2-{the total number of likes of offline data Y; the total number of reposts of offline data Y; the total number of plays of offline data Y; the total number of collections of offline data Y; the performing object information included in offline data Y; the popularity information of offline data Y...}.

[0243] Step 403: The processing device stores the corresponding relationship between each identification keyword and key value established at the current time in the second offline data storage area.

[0244] Specifically, after the processing device generates a correspondence between an identification keyword and a corresponding key value for each offline data in the offline database at the current time, the correspondence between each identification keyword and the key value established at the current time can be stored in the second offline data storage area, wherein the current time refers to the deadline of the offline data in the offline database.

[0245] In addition, considering that the data in the offline database comes from the online database, the storage media in the storage process illustrated in Scenario 2.1 of the present disclosure are first transmitted from the online database to the offline database, and then transmitted from the offline database to the second offline data storage area.

[0246] For example, the processing device may be triggered once a day to write the corresponding relationship between the identification keyword and the key value obtained from the offline data into HBase. The time day is a basic attribute of HBase and is used to distinguish data from multiple days.

[0247] In this way, after processing the offline data, the desired content can be extracted and then stored. Compared with the existing technology of directly storing the full amount of data, the required storage space can be reduced to a certain extent.

[0248] Scenario 2.2: Detect data changes in real time and trigger data processing and storage operations when data changes.

[0249] See also Figure 5 As shown in FIG. 1 , it is another schematic diagram of a process for implementing data storage in the second offline data storage area in the embodiment of the present disclosure. Figure 5 , the process of storing data in the second offline data storage area by the processing device in the embodiment of the present disclosure is described:

[0250] Step 501: The processing device monitors changes in the operated information of data in the global scope in real time, and when it is determined that there is target data whose operated information has changed, updates the offline data based on the target data at the current time.

[0251] Specifically, the processing device monitors the changes of the operated information of the global data stored in the online database in real time, wherein the process of monitoring the online database is the same as the process indicated in the above step 301, and the present disclosure will not explain it in detail here.

[0252] Furthermore, when the processing device determines that there is target data whose operated information has changed, the offline data is updated based on the target data at the current time. Since the offline data in the offline database is periodically obtained from the online database, when the target data is used to update the offline data, the data in the offline database can be updated based on the target data according to the addition, deletion, and modification operations performed on the data. The specific update logic is the same as the update logic in the above step 202.

[0253] Step 502: The processing device uses the second engine to aggregate at least the operated information stored in the corresponding data identification information in each data table in the updated offline data into a single table of data with compliant content.

[0254] Specifically, when executing step 502, the processing device may use the second engine to execute the operation indicated in step 401 based on the updated offline data to obtain single-table data with compliant content, which will not be described in detail in this disclosure.

[0255] Step 503: The processing device generates corresponding identification keywords based on the data identification information of each data in the single table data, and establishes key values ​​corresponding to each identification keyword based on the operated information stored in the single table data corresponding to each data.

[0256] Specifically, when executing step 503, the processing device may adopt the processing method of step 402 above to establish key values ​​corresponding to the identification keywords of each data based on the single table data, and the present disclosure does not elaborate on the relevant processing process.

[0257] Step 504: The processing device stores the corresponding relationship between each identification keyword and key value established at the current time in the second offline data storage area.

[0258] Specifically, when the processing device executes step 504, the processed data may be stored in the second offline data storage area according to the operation restricted by the above step 403, and the present disclosure does not elaborate on the specific processing process.

[0259] In this way, the corresponding data information can be saved in real time according to the changes in the data operation conditions, providing a query basis for the subsequent real-time query process.

[0260] The following is combined with Figure 6a , which is a schematic diagram of the data query process in the embodiment of the present disclosure, and the following is combined with the attached Figure 6a , explain the specific data query process:

[0261] Step 601: The processing device receives a data query request triggered by a target object in an operable page, and determines a target query area targeted by the data query request.

[0262] Specifically, the processing device receives a data query request triggered by a target object in an operable page, and determines a target query area targeted by the data query request based on the operable page operated by the target object, wherein the data query request includes target query conditions and target time information for data query, and the target query area includes a first offline data storage area and a second offline data storage area.

[0263] In the embodiment of the present disclosure, if the target object is determined, the data query request is triggered in the first operable page for displaying the data statistics obtained according to the content description information, and the target query area targeted by the data query request is determined to be the first offline data storage area. In addition, if the target object is determined, the data query request is triggered in the second operable page for displaying the data change trend obtained according to the operated information, and the target query area targeted by the data query request is determined to be the second offline data storage area.

[0264] Step 602: When the processing device determines that the target query area is the first offline data storage area, the first offline data storage area is searched for a target query result that matches the target query condition and the target time information.

[0265] Specifically, when the processing device determines that the target object triggers the data query request in the first operable page, it generates a target index keyword based on the target query condition, and queries the first offline data storage area for a target query result corresponding to the target index keyword under the target time information.

[0266] It should be noted that in some optional embodiments of the present disclosure, when the first offline data storage area stores query results corresponding to various time information and obtained by operating according to a preset first processing rule, the processing device may first determine the storage area corresponding to the target time information setting, or the processing device may first determine the various query results stored corresponding to the target time information, and then search for the target query result corresponding to the target index keyword in the storage area or the various query results stored corresponding to the target time information based on the target index keyword.

[0267] In some other optional embodiments of the present disclosure, when the storage format of the query results in the first offline data storage area is that the index keyword is the primary key and the time information and the query results are the corresponding values, the processing device can first determine the various values ​​stored corresponding to the index keyword based on the generated index keyword, and obtain the target value of the target time information with the time information as the target time information, and use the query results within the target value as the target query result.

[0268] In the disclosed embodiment, the first processing rule includes executing a query operation in corresponding historical data based on various query conditions configured to statistically obtain the query result, the target query condition is included in the various query conditions, and the historical data includes at least content description information of data in a global scope, wherein the process of storing data in the first offline data storage area has been described in detail in the process illustrated in the above scenario 1, and will not be repeated here.

[0269] In this way, with the help of the aforementioned data storage process, the storage method of full storage under the relevant technology is no longer adopted, but the query results corresponding to all query conditions are directly stored. When the query conditions are controllable, the storage space cost can be greatly saved, so that no matter how complex the query conditions are, high-efficiency queries can be achieved, and the query time of the data can be controlled. In the query process, it is only necessary to generate the corresponding index keywords according to the processing needs to obtain the stored query results, which greatly reduces the query time of the data, improves the query efficiency, can quickly respond to the query needs of related objects, and improves the user experience of related objects.

[0270] It should be noted that, in the embodiment of the present disclosure, when the processing device determines that the target storage area is the second offline data storage area, the target information query result matching the target query condition and the target time information is searched in the second offline data storage area.

[0271] Specifically, when the processing device determines that the target object triggers the data query request in the second operable page, it generates a target identification keyword based on the data identification information carried in the target query condition, and queries the target query result corresponding to the target identification keyword under the target time information in the second offline data storage area.

[0272] It should be noted that, in the embodiment of the present disclosure, the second offline data storage area stores various storage data corresponding to various time information and obtained by operating according to a preset second processing rule. The second processing rule includes aggregating the operated information of each offline data based on the corresponding offline database to obtain the various storage data. The offline data includes various types of information of data in a global scope. The process of storing data in the second offline data storage area has been described in detail in the description of the aforementioned scenario 2 and will not be repeated here.

[0273] In this way, the processing device can query the data operation status in the second offline data storage area according to the content stored in the second offline data storage area to obtain the corresponding query result.

[0274] Furthermore, in an embodiment of the present disclosure, after obtaining the query results corresponding to the target query conditions, the query results can be displayed to relevant objects in a preset display format, wherein the display format may be a pie chart, a line chart, a bar chart or the like.

[0275] For example, see Figure 6b As shown, it is a schematic diagram of displaying the query results in the implementation of the present disclosure. Figure 6b The display content constructed based on the query result is schematically shown in FIG. Figure 6b The query conditions shown are presented, and the query results are presented in the form of a pie chart to assist relevant objects in determining the song release status and the health of song data.

[0276] For example, see Figure 6c , which is a schematic diagram of query results displayed based on query conditions in an embodiment of the present disclosure, according to Figure 6c The processing device can present corresponding query results according to the query conditions selected by the relevant objects, such as Figure 6c As shown, in the scenario of querying song data, when the relevant object chooses to display songs with a release time between 2021-09-01 and 2021-09-09, the release time is used as the query condition to query the query results stored for each day from 2021-9-1 to 2021-9-9, and based on the query results obtained, an intuitive curve chart is drawn.

[0277] For example, see Figure 6d and 6e As shown, Figure 6d This is a schematic diagram of an operable page for triggering a query in a second offline data storage area in an embodiment of the present disclosure. Figure 6eFIG. 1 is a schematic diagram of a result presented based on a query result obtained in a second offline data storage area according to an embodiment of the present disclosure. In a scenario where music data is queried, the processing device responds to the related object in Figure 6d In the interface shown, the query request is initiated based on the song ID: 1875246245, time: 2021-09-03 to 2021-09-15, and is presented to the relevant objects in a targeted manner. Figure 6e The query results shown are for relevant personnel to know the operation status of the song data with song ID 1875246245.

[0278] For example, see Figure 6f As shown, it is a schematic diagram of the query result obtained based on monitoring the data in the online database in the embodiment of the present disclosure. Figure 6f It can be seen from the contents shown that the technical solution proposed in the present disclosure can monitor the changes of data in real time, analyze the coverage of data according to the actual processing needs, and present the changes of various data report indicators of concern, among which, Figure 6f The data report indicators presented in the data report are conventional indicators in this field and will not be described in detail here. For example, the GAP indicator is used for gap analysis.

[0279] In this way, query efficiency is improved. No matter how complex the query conditions are, high-efficiency queries can be achieved, which makes data query time controllable, improves data query performance, can quickly respond to query requests for related objects and quickly generate charts, and improves the user experience of related objects.

[0280] Based on the same inventive concept, refer to Figure 7 As shown, it is a schematic diagram of the logical structure of the data query device in the embodiment of the present disclosure. The data query device 700 includes: a receiving unit 701, a query unit 702, and a storage unit 703, wherein:

[0281] The receiving unit 701 receives a data query request triggered by a target object in an operable page, and determines a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query;

[0282] The query unit 702, when determining that the target query area is a first offline data storage area, searches the first offline data storage area for a target query result that matches the target query condition and the target time information;

[0283] Among them, the first offline data storage area stores query results corresponding to various time information and obtained by operating according to a preset first processing rule. The first processing rule includes executing query operations in corresponding historical data based on various configured query conditions to statistically obtain the query results. The target query condition is included in the various query conditions, and the historical data at least includes content description information of data in a global scope.

[0284] Optionally, when it is determined that the target storage area is the second offline data storage area, the query unit 702 is configured to perform the following operations:

[0285] In the second offline data storage area, searching for a target information query result that matches the target query condition and the target time information;

[0286] Among them, the second offline data storage area stores various storage data corresponding to various time information and obtained by operating according to a preset second processing rule. The second processing rule includes aggregating the operated information of each offline data based on the corresponding offline database to obtain the various storage data. The offline data includes various types of information of data in the global scope.

[0287] Optionally, the device further includes a storage unit 703, and the storage unit 703 is used to further perform the following operations using the first processing rule:

[0288] The second engine is periodically used to periodically obtain offline variable data generated within a specified historical period up to the current time information from the corresponding offline database, and process the offline variable data into data to be stored with compliant content;

[0289] Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to periodically perform query operations in the updated historical data based on each configured query condition, and statistically obtain query results corresponding to each query condition;

[0290] Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the query result under the current time information, and store each index keyword and the corresponding query result in the first offline data storage area.

[0291] Optionally, the device further includes a storage unit 703, and the storage unit 703 is used to further perform the following operations using the first processing rule:

[0292] Monitor changes in content description information of data in a global scope in real time, and when it is determined that there is target data with changed content description information, use a third engine to process the target data at the current time information into content-compliant data to be stored;

[0293] Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to perform a query operation in the updated historical data based on each configured query condition, and statistically obtain the real-time query results corresponding to each query condition;

[0294] Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the real-time query result under the current time information, and store each index keyword and the corresponding real-time query result in the first offline data storage area.

[0295] Optionally, when the offline variable data is processed into content-compliant data to be stored, the storage unit 703 is used to:

[0296] In each data table included in the offline variable data, at least the content description information stored in correspondence with each data identification information is aggregated into data content in a single table form;

[0297] The text content in the data content is replaced with the content in the corresponding coded form, and the content marked as invalid in the data content is deleted, and the processed data content is used as the data to be stored with compliant content.

[0298] Optionally, when updating the historical data stored in the first engine, the storage unit 703 is used to:

[0299] Determine each newly added data associated with a newly released state in the offline variable data, and directly store each newly added data in a storage area where historical data is located in the first engine;

[0300] Determine each invalid data associated with a deletion state in the offline variable data, and delete each data corresponding to the each invalid data in the historical data stored in the first engine;

[0301] Determine each to-be-updated data of the associated information adjustment status in the offline variable data, and modify each data corresponding to the to-be-updated data in the historical data stored by the first engine based on the content description information of each to-be-updated data.

[0302] Optionally, when generating corresponding index keywords for each query condition respectively, the storage unit 703 is used to perform any one of the following operations:

[0303] Pre-configure corresponding digital value results for each data attribute information in each type, respectively, and pre-configure corresponding weight parameters for each type, and use the digital value results corresponding to each data attribute information constrained by each query condition and the weighted results determined by the weight parameters as index keywords corresponding to each query condition;

[0304] Pre-configure different digital value results for each data attribute information in each type, and respectively determine the digital value results corresponding to each data attribute information constrained by each query condition, and respectively accumulate the binary form of each digital value result constrained by the same query condition as the index keyword generated by the corresponding query condition;

[0305] Different digital value results are configured in advance for each type of data attribute information, and a hash algorithm is used to generate a hash value of a specified length corresponding to each digital value result, as well as each hash value corresponding to each data attribute information constrained by each query condition, to generate index keywords corresponding to each query condition.

[0306] Optionally, each query condition is restricted to a statistical indicator characterizing the query result, and when it is determined that the current time information is an end date of a preset time period, the storage unit 703 is configured to perform the following operations:

[0307] Determine each time information included in the time period corresponding to the current time information, and establish the data change result of each query condition in the time period under the corresponding statistical indicator based on each query condition and the corresponding query result saved corresponding to each time information;

[0308] Establish a correspondence between the index keywords corresponding to each query condition and the data change results within the time period, and store the index keywords and data change results corresponding to each query condition in the first offline data storage area corresponding to the time period.

[0309] Optionally, the query unit 702 is configured to further perform the following operations using the second processing rule:

[0310] Based on the corresponding offline database, the second engine is periodically used to aggregate at least the operated information stored in the data tables corresponding to the offline data corresponding to each data identification information into single-table data with compliant content;

[0311] Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively;

[0312] The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

[0313] Optionally, the query unit 702 is configured to further perform the following operations using the second processing rule:

[0314] Monitor the changes of the operated information of the data in the global scope in real time, and when it is determined that there is target data with changed operated information, update the offline data based on the target data at the current time;

[0315] Using the second engine, in each data table corresponding to the updated offline data, at least the operated information stored in the corresponding data identification information is aggregated into single-table data with compliant content;

[0316] Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively;

[0317] The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

[0318] Optionally, when searching for a target query result that matches the target query condition and the target time information, the query unit 702 is used to:

[0319] When it is determined that the target object triggers the data query request in the first operable page, a target index keyword is generated based on the target query condition, and a target query result corresponding to the target index keyword under the target time information is searched in the first offline data storage area;

[0320] When it is determined that the target object triggers the data query request in the second operable page, a target identification keyword is generated based on the data identification information carried in the target query condition, and the target query result corresponding to the target identification keyword under the target time information is queried in the second offline data storage area.

[0321] See also Figure 8 As shown, it is a schematic diagram of the physical structure of the data query device in the embodiment of the present disclosure. Based on the same inventive concept, it can include a memory 801 and a processor 802.

[0322] The memory 801 is used to store computer programs executed by the processor 802. The memory 801 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc. The processor 802 may be a central processing unit (CPU) or a digital processing unit, etc. The specific connection medium between the above-mentioned memory 801 and the processor 802 is not limited in the embodiments of the present disclosure. The embodiments of the present disclosure are Figure 8 In the embodiment, the memory 801 and the processor 802 are connected via a bus 803. The bus 803 is connected to the processor 802 via a bus 803. Figure 8 The connections between other components are shown in bold lines, and are not intended to be limiting. Bus 803 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0323] The memory 801 may be a volatile memory, such as a random-access memory (RAM); the memory 801 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or the memory 801 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 801 may be a combination of the above memories.

[0324] Processor 802, used to execute the following when calling the computer program stored in memory 801: Figure 6a The embodiment shown in provides a data query method.

[0325] Based on the same inventive concept, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the data query method in any of the above method embodiments is implemented.

[0326] In summary, in an embodiment of the present disclosure, a data query method, device, electronic device and storage medium are proposed, which receive a data query request triggered by a target object in an operable page, determine a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query; when the target query area is determined to be a first offline data storage area, query the target query result matching the target query condition and target time information in the first offline data storage area, wherein the first offline data storage area stores query results corresponding to each time information and obtained by operating according to a preset first processing rule, wherein the first processing rule includes executing a query operation in corresponding historical data based on each configured query condition to statistically obtain the query result, wherein the target query condition is included in each query condition, and the historical data at least includes content description information of data in a global scope.

[0327] In this way, the full amount of data is no longer archived by date, but the query results of the data are directly stored. After executing the query according to various possible query conditions in advance, it is only necessary to store the query results corresponding to all query conditions to meet the subsequent query needs. At the same time, considering that the query conditions in the actual query process are controllable, the method of storing query results can greatly reduce the amount of data that needs to be stored, which can not only improve the query efficiency of the data, but also improve the controllability of the query time of the data, and can quickly respond to the query needs of related objects and improve the user experience of related objects.

[0328] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0329] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0330] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0331] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0332] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0333] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A data query method, characterized in that: include: Receiving a data query request triggered by a target object in an operable page, and determining a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query; When it is determined that the target query area is the first offline data storage area, searching the first offline data storage area for a target query result that matches the target query condition and the target time information; The first offline data storage area stores query results corresponding to various time information and obtained by operating according to a preset first processing rule, wherein the first processing rule includes executing a query operation in corresponding historical data based on various configured query conditions to obtain the query results by statistics, the target query condition is included in the various query conditions, and the historical data at least includes content description information of data in a global scope; the first offline data storage area is a persistent storage area different from a memory; The first processing rule further includes: periodically using the second engine to periodically obtain offline variable data generated within a specified historical period up to the current time information from the corresponding offline database, and process the offline variable data into content-compliant data to be stored; based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to periodically perform query operations in the updated historical data based on the configured query conditions, and statistically obtain the query results corresponding to each query condition; respectively generate corresponding index keywords for each query condition, and establish a corresponding relationship between each index keyword and the query result under the current time information, and store each index keyword and the corresponding query result in the first offline data storage area.

2. The method according to claim 1, characterized in that When the target storage area is determined to be the second offline data storage area, perform the following operations: In the second offline data storage area, searching for a target information query result that matches the target query condition and the target time information; Among them, the second offline data storage area stores various storage data corresponding to various time information and obtained by operating according to a preset second processing rule. The second processing rule includes aggregating the operated information of each offline data based on the corresponding offline database to obtain the various storage data. The offline data includes various types of information of data in the global scope.

3. The method according to claim 1, characterized in that The first processing rule further includes: Monitor changes in content description information of data in a global scope in real time, and when it is determined that there is target data with changed content description information, use a third engine to process the target data at the current time information into content-compliant data to be stored; Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to perform a query operation in the updated historical data based on each configured query condition, and statistically obtain the real-time query results corresponding to each query condition; Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the real-time query result under the current time information, and store each index keyword and the corresponding real-time query result in the first offline data storage area.

4. The method according to claim 1, characterized in that The processing of the offline variable data into content-compliant data to be stored includes: In each data table included in the offline variable data, at least the content description information stored in correspondence with each data identification information is aggregated into data content in a single table form; The text content in the data content is replaced with the content in the corresponding coded form, and the content marked as invalid in the data content is deleted, and the processed data content is used as the data to be stored with compliant content.

5. The method according to claim 1, characterized in that The updating of the historical data stored in the first engine includes: Determine each newly added data associated with a newly released state in the offline variable data, and directly store each newly added data in a storage area where historical data is located in the first engine; Determine each invalid data associated with a deletion state in the offline variable data, and delete each data corresponding to the each invalid data in the historical data stored in the first engine; Determine each to-be-updated data of the associated information adjustment status in the offline variable data, and modify each data corresponding to the to-be-updated data in the historical data stored by the first engine based on the content description information of each to-be-updated data.

6. The method according to any one of claims 1 to 5, characterized in that: Generating corresponding index keywords for each query condition respectively includes any one of the following operations: Pre-configure corresponding digital value results for each data attribute information in each type, respectively, and pre-configure corresponding weight parameters for each type, and use the digital value results corresponding to each data attribute information constrained by each query condition and the weighted results determined by the weight parameters as index keywords corresponding to each query condition; Pre-configure different digital value results for each data attribute information in each type, and respectively determine the digital value results corresponding to each data attribute information constrained by each query condition, and respectively accumulate the binary form of each digital value result constrained by the same query condition as the index keyword generated by the corresponding query condition; Different digital value results are configured in advance for each type of data attribute information, and a hash algorithm is used to generate a hash value of a specified length corresponding to each digital value result, as well as each hash value corresponding to each data attribute information constrained by each query condition, to generate index keywords corresponding to each query condition.

7. The method according to any one of claims 1 to 5, characterized in that: Each query condition is limited to a statistical indicator characterizing the query result. When it is determined that the current time information is the end date of a preset time period, the following operations are performed: Determine each time information included in the time period corresponding to the current time information, and establish the data change result of each query condition in the time period under the corresponding statistical indicator based on each query condition and the corresponding query result saved corresponding to each time information; Establish a correspondence between the index keywords corresponding to each query condition and the data change results within the time period, and store the index keywords and data change results corresponding to each query condition in the first offline data storage area corresponding to the time period.

8. The method according to claim 2, characterized in that The second processing rule further includes: Based on the corresponding offline database, the second engine is periodically used to aggregate at least the operated information stored in the data tables corresponding to the offline data corresponding to each data identification information into single-table data with compliant content; Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively; The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

9. The method according to claim 2 or 8, characterized in that The second processing rule further includes: Monitor the changes of the operated information of the data in the global scope in real time, and when it is determined that there is target data with changed operated information, update the offline data based on the target data at the current time; Using the second engine, in each data table corresponding to the updated offline data, at least the operated information stored in the corresponding data identification information is aggregated into single-table data with compliant content; Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively; The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

10. The method according to claim 2 or 8, characterized in that The target query result that matches the target query condition and the target time information includes: When it is determined that the target object triggers the data query request in the first operable page, a target index keyword is generated based on the target query condition, and a target query result corresponding to the target index keyword under the target time information is searched in the first offline data storage area; When it is determined that the target object triggers the data query request in the second operable page, a target identification keyword is generated based on the data identification information carried in the target query condition, and the target query result corresponding to the target identification keyword under the target time information is queried in the second offline data storage area.

11. A data query device, characterized in that: include: A receiving unit receives a data query request triggered by a target object in an operable page, and determines a target query area targeted by the data query request, wherein the data query request includes a target query condition and target time information for performing data query; A query unit, when determining that the target query area is a first offline data storage area, searches the first offline data storage area for a target query result that matches the target query condition and the target time information; The first offline data storage area stores query results corresponding to various time information and obtained by operating according to a preset first processing rule, wherein the first processing rule includes executing a query operation in corresponding historical data based on various configured query conditions to obtain the query results by statistics, the target query condition is included in the various query conditions, and the historical data at least includes content description information of data in a global scope; the first offline data storage area is a persistent storage area different from a memory; The storage unit is used to further perform the following operations using the first processing rule: periodically using the second engine to regularly obtain offline variable data generated within a specified historical time period up to the current time information from the corresponding offline database, and process the offline variable data into content-compliant data to be stored; based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to periodically perform query operations in the updated historical data based on the configured query conditions, and statistically obtain the query results corresponding to each query condition; corresponding index keywords are generated for each query condition, and a corresponding relationship between each index keyword and the query result under the current time information is established, and each index keyword and the corresponding query result are stored in the first offline data storage area.

12. The device according to claim 11, characterized in that When the target storage area is determined to be the second offline data storage area, the query unit is configured to perform the following operations: In the second offline data storage area, searching for a target information query result that matches the target query condition and the target time information; Among them, the second offline data storage area stores various storage data corresponding to various time information and obtained by operating according to a preset second processing rule. The second processing rule includes aggregating the operated information of each offline data based on the corresponding offline database to obtain the various storage data. The offline data includes various types of information of data in the global scope.

13. The device according to claim 11 or 12, characterized in that The device further includes a storage unit, and the storage unit is used to further perform the following operations using the first processing rule: Monitor changes in content description information of data in a global scope in real time, and when it is determined that there is target data with changed content description information, use a third engine to process the target data at the current time information into content-compliant data to be stored; Based on the data to be stored, the historical data stored in the first engine is updated, and the first engine is called to perform a query operation in the updated historical data based on each configured query condition, and statistically obtain the real-time query results corresponding to each query condition; Generate corresponding index keywords for each query condition respectively, establish corresponding relationships between each index keyword and the real-time query result under the current time information, and store each index keyword and the corresponding real-time query result in the first offline data storage area.

14. The device according to claim 11, characterized in that When the offline variable data is processed into content-compliant data to be stored, the storage unit is used to: In each data table included in the offline variable data, at least the content description information stored in correspondence with each data identification information is aggregated into data content in a single table form; The text content in the data content is replaced with the content in the corresponding coded form, and the content marked as invalid in the data content is deleted, and the processed data content is used as the data to be stored with compliant content.

15. The device according to claim 11, characterized in that When the historical data stored in the first engine is updated, the storage unit is used to: Determine each newly added data associated with a newly released state in the offline variable data, and directly store each newly added data in a storage area where historical data is located in the first engine; Determine each invalid data associated with a deletion state in the offline variable data, and delete each data corresponding to the each invalid data in the historical data stored in the first engine; Determine each to-be-updated data of the associated information adjustment status in the offline variable data, and modify each data corresponding to the to-be-updated data in the historical data stored by the first engine based on the content description information of each to-be-updated data.

16. The device according to any one of claims 11 to 15, characterized in that When generating corresponding index keywords for each query condition respectively, the storage unit is used to perform any one of the following operations: Pre-configure corresponding digital value results for each data attribute information in each type, respectively, and pre-configure corresponding weight parameters for each type, and use the digital value results corresponding to each data attribute information constrained by each query condition and the weighted results determined by the weight parameters as index keywords corresponding to each query condition; Pre-configure different digital value results for each data attribute information in each type, and respectively determine the digital value results corresponding to each data attribute information constrained by each query condition, and respectively accumulate the binary form of each digital value result constrained by the same query condition as the index keyword generated by the corresponding query condition; Different digital value results are configured in advance for each type of data attribute information, and a hash algorithm is used to generate a hash value of a specified length corresponding to each digital value result, as well as each hash value corresponding to each data attribute information constrained by each query condition, to generate index keywords corresponding to each query condition.

17. The device according to any one of claims 11 to 15, characterized in that Each query condition is limited to a statistical indicator characterizing the query result. When it is determined that the current time information is the end date of a preset time period, the storage unit is used to perform the following operations: Determine each time information included in the time period corresponding to the current time information, and establish the data change result of each query condition in the time period under the corresponding statistical indicator based on each query condition and the corresponding query result saved corresponding to each time information; Establish a correspondence between the index keywords corresponding to each query condition and the data change results within the time period, and store the index keywords and data change results corresponding to each query condition in the first offline data storage area corresponding to the time period.

18. The device according to claim 12, characterized in that The query unit is used to further perform the following operations using the second processing rule: Based on the corresponding offline database, the second engine is periodically used to aggregate at least the operated information stored in the data tables corresponding to the offline data corresponding to each data identification information into single-table data with compliant content; Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively; The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

19. The device according to claim 12 or 18, characterized in that The query unit is used to further perform the following operations using the second processing rule: Monitor the changes of the operated information of the data in the global scope in real time, and when it is determined that there is target data with changed operated information, update the offline data based on the target data at the current time; Using the second engine, in each data table corresponding to the updated offline data, at least the operated information stored in the corresponding data identification information is aggregated into single-table data with compliant content; Based on the data identification information of each data in the single table data, corresponding identification keywords are generated respectively, and based on the operated information respectively stored in the single table data corresponding to each data, key values ​​corresponding to each identification keyword are established respectively; The corresponding relationship between each identification keyword and key value established at the current time is stored in the second offline data storage area.

20. The device according to claim 11 or 18, characterized in that When searching for a target query result that matches the target query condition and the target time information, the query unit is used to: When it is determined that the target object triggers the data query request in the first operable page, a target index keyword is generated based on the target query condition, and a target query result corresponding to the target index keyword under the target time information is searched in the first offline data storage area; When it is determined that the target object triggers the data query request in the second operable page, a target identification keyword is generated based on the data identification information carried in the target query condition, and the target query result corresponding to the target identification keyword under the target time information is queried in the second offline data storage area.

21. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 10 are implemented.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Big data online analysis method and system

    CN113342843A

  • Data loading method and device, computer program product and storage medium

    CN113377777A