Data storage method and data query method

By using storage time information and sharding time index in data storage for data sharding storage, the problems of low cache hit rate and large overhead of hot and cold data merging in the prior art are solved, and efficient data hot and cold separation and stream computing system performance improvement are achieved.

CN120029521APending Publication Date: 2025-05-23HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311560993.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the storage module capacity is saturated, existing data storage solutions can easily lead to a decrease in cache hit rate and an increase in hot and cold data merging overhead, and cannot effectively realize hot and cold data separation.

Method used

By obtaining the storage time information of the data to be stored, and using the shard time index to determine the target storage shard, the data is stored in pieces according to the time period, and the data is separated by hot and cold.

Benefits of technology

It improves data locality, improves data hit rate, reduces the overhead of hot and cold data, and improves the overall performance of the stream computing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029521A_ABST
    Figure CN120029521A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data storage method and a data query method, and the data storage method comprises the steps: obtaining to-be-stored data, and determining a target storage unit corresponding to the to-be-stored data; extracting at least one piece of storage time information from the to-be-stored data; at least one target storage fragment is determined according to the at least one piece of storage time information and fragment time indexes corresponding to a plurality of storage fragments in a target storage unit, and the storage time information is in one-to-one correspondence with the target storage fragments; and storing the to-be-stored data to the at least one target storage fragment according to the at least one piece of storage time information. The to-be-stored data is stored in the at least one target storage fragment in the target storage unit according to the storage time information, fragment storage of the data in different time periods is achieved, the data locality is better, the hit rate is higher, and due to the fact that the data popularity is strongly related to the time periods, data cold and hot separation is achieved, and the data storage efficiency is improved. And the combination overhead of cold and hot data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification relate to the field of computer technology, and in particular to a data storage method and a data query method. Background Art

[0002] With the development of computer technology, stream computing has been increasingly widely used. Stream computing is a data processing technology that reads data from a continuous data stream, quickly filters, converts, analyzes or trains data in real time, and then passes the processing results to applications, data storage or other processing engines, so as to obtain valuable information from massive data in real time.

[0003] At present, in the data storage process, data is usually stored in the same storage module. When the storage module capacity is saturated, the data will be stored in other idle storage modules. However, the above solution is prone to problems such as reduced cache hit rate and increased overhead of merging hot and cold data. Therefore, a data storage solution with high cache hit rate and separation of hot and cold data is urgently needed. Summary of the invention

[0004] In view of this, an embodiment of this specification provides a data storage method. One or more embodiments of this specification also relate to a data query method, a data storage device, a data query device, a computing device, a computer-readable storage medium and a computer program to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a data storage method is provided, including:

[0006] Acquire the data to be stored, and determine the target storage unit corresponding to the data to be stored;

[0007] Extracting at least one storage time information from the data to be stored;

[0008] Determine at least one target storage shard according to at least one storage time information and shard time indexes respectively corresponding to a plurality of storage shards in the target storage unit, wherein the storage time information corresponds to the target storage shard in a one-to-one manner;

[0009] According to at least one storage time information, the data to be stored is stored in at least one target storage slice.

[0010] According to a second aspect of an embodiment of this specification, a data query method is provided, including:

[0011] receiving a data query request, wherein the data query request carries data query information;

[0012] Parsing the data query information and determining at least one query time information;

[0013] According to at least one query time information, the target storage data corresponding to the data query information is searched from at least one target storage shard, wherein the at least one target storage shard is obtained based on the at least one storage time information and the shard time index corresponding to the multiple storage shards in the target storage unit, the storage time information corresponds one-to-one to the target storage shard, the at least one target storage shard includes multiple storage data, and the storage data is stored in the at least one target storage shard based on the at least one storage time information.

[0014] According to a third aspect of an embodiment of this specification, there is provided a data storage device, including:

[0015] An acquisition module is configured to acquire the data to be stored and determine a target storage unit corresponding to the data to be stored;

[0016] An extraction module, configured to extract at least one storage time information from the data to be stored;

[0017] A determination module is configured to determine at least one target storage shard according to at least one storage time information and shard time indexes respectively corresponding to a plurality of storage shards in the target storage unit, wherein the storage time information corresponds to the target storage shard in a one-to-one manner;

[0018] The first storage module is configured to store the data to be stored in at least one target storage slice according to at least one storage time information.

[0019] According to a fourth aspect of the embodiments of this specification, a data query device is provided, including:

[0020] A receiving module is configured to receive a data query request, wherein the data query request carries data query information;

[0021] A parsing module, configured to parse the data query information and determine at least one query time information;

[0022] The first search module is configured to search for target storage data corresponding to data query information from at least one target storage shard according to at least one query time information, wherein the at least one target storage shard is obtained based on at least one storage time information and a shard time index corresponding to a plurality of storage shards in a target storage unit, the storage time information corresponds one-to-one to the target storage shard, the at least one target storage shard includes a plurality of storage data, and the storage data is stored in the at least one target storage shard based on the at least one storage time information.

[0023] According to a fifth aspect of an embodiment of this specification, a computing device is provided, including:

[0024] Memory and processor;

[0025] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect or the second aspect are implemented.

[0026] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the method provided in the first aspect or the second aspect are implemented.

[0027] According to a seventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect or the second aspect above.

[0028] A data storage method provided by an embodiment of the present specification obtains data to be stored and determines a target storage unit corresponding to the data to be stored; extracts at least one storage time information from the data to be stored; determines at least one target storage shard according to at least one storage time information and a shard time index corresponding to a plurality of storage shards in the target storage unit, wherein the storage time information corresponds to the target storage shard one-to-one; stores the data to be stored in at least one target storage shard according to at least one storage time information. By storing the data to be stored in at least one target storage shard in the target storage unit according to the storage time information, shard storage of data in different time periods is achieved, making data locality better and data hit rate higher, and because the data heat is strongly related to the time period, the separation of hot and cold data is achieved, reducing the overhead of merging hot and cold data. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is an architectural diagram of a first data storage system provided by an embodiment of this specification;

[0030] Figure 2 is an architectural diagram of a second data storage system provided by an embodiment of this specification;

[0031] Figure 3 is a flow chart of a data storage method provided by one embodiment of this specification;

[0032] Figure 4 This is a schematic diagram of extracting storage time information in a data storage method provided by an embodiment of this specification;

[0033] Figure 5 This is a schematic diagram of data deletion in a data storage method provided by an embodiment of this specification;

[0034] Figure 6 is a flow chart of a data query method provided by an embodiment of this specification;

[0035] Figure 7 is a process flow chart of a data storage method provided by one embodiment of the present specification;

[0036] Figure 8 is a processing flow chart of a data query method provided by an embodiment of this specification;

[0037] Fig. 9 is an architectural diagram of a third data storage system provided by an embodiment of this specification;

[0038] Fig.10 is a schematic diagram of the structure of a data storage device provided by an embodiment of this specification;

[0039] Fig.11 It is a structural schematic diagram of a data query device provided by an embodiment of this specification;

[0040] Fig.12 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0041] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0042] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0043] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0044] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0045] First, the terms involved in one or more embodiments of this specification are explained.

[0046] Log Structured Merge Tree: Log Structured Merge Tree (LSM-tree) is a hierarchical, ordered, disk-oriented data structure in the storage field. Its core idea is to make full use of the fact that the performance of disk batch sequential writes is much higher than that of random writes, and store data modification and deletion operations in an append manner, thereby improving the overall performance of the storage system.

[0047] Compaction: Compaction refers to the process of merging LSM-tree data files. In the LSM-tree storage structure, data at different levels can be merged regularly to ensure that the read link of the LSM-tree does not become infinitely long. During the merging process, redundant data can also be cleaned up to free up disk space.

[0048] State storage: State storage refers to a storage module in a stream computing system that is used to temporarily store intermediate results of calculations or save data or events read over a period of time.

[0049] Hot and cold separation: In a storage system, frequently accessed data is called hot data, and rarely accessed data is called cold data. Storing hot and cold data in different files, locations, and media is called hot and cold separation.

[0050] Sharded storage: In a storage system, the full amount of data is divided into multiple data subsets according to certain rules. Multiple data subsets are physically isolated from each other, and data from different subsets are stored in different files. This storage method is called sharded storage. Each data subset is called a shard.

[0051] Structured Query Language Hint: Structured Query Language Hint (SQL Hint) is a mechanism provided by standard SQL that allows users to add some artificial hints to SQL statements to optimize the execution plan or operation mode of SQL.

[0052] Business Intelligence (BI) analysis scenarios are one of the most common scenarios in stream computing. Such scenarios usually aggregate (group by) and analyze data at the hour, day, and month granularity at the SQL level to obtain valuable data information. Such jobs usually contain project time information (such as hours, days, and months) in the key of the stream computing state storage, which makes the read and write access of the state storage have obvious time locality: data in the same project time period will be written continuously; data in the same project time period will also be accessed continuously and intensively within a certain period of time. The current state storage engine will store data from different project time periods in an LSM-tree structure by default. This storage method obviously cannot utilize valuable time information. In addition, when the scale of state data is large, the state storage engine is prone to problems such as reduced cache hit rate and increased overhead of merging cold and hot data, which leads to increased overhead of the input and output (IO) and core processor (CPU) of the stream computing system, and thus leads to lower overall throughput of the stream computing system. Therefore, data is currently stored mainly in the following two ways:

[0053] The first method is to use the tiered storage function of the persistent key-value storage engine to store the most recently written data in the upper layer of the LSM-tree as much as possible. Specifically, when each piece of data is written, an auto-increment identifier is generated to indicate the early or late writing of the data. At the same time, an identifier-time mapping table is generated to roughly identify the writing time of the data. When the last layer of files in the LSM-tree is merged, the identifier-time mapping table is used to place the data written in the most recent period in the upper layer of the LSM-tree, and only the data written earlier is placed in the bottom layer of the LSM-tree. In this way, in the scenario where the latest written data is accessed more frequently, the system as a whole can obtain better read performance. However, in the above solution, there is no necessary correlation between the hot and cold data and the writing time of the data. The frequently accessed hot data may be written earlier. Therefore, the above solution cannot accurately divide the hot and cold data according to the project requirements, and cannot play an effective role in the stream computing BI analysis scenario. In addition, the above solution can only place the old data uniformly in the bottom layer of the LSM-tree and the new data in the upper layer of the LSM-tree. However, in the BI analysis scenario of stream computing, data in different time periods will have different degrees of hotness and coldness. The above solution is difficult to make a fine-grained distinction between data with multiple degrees of hotness and coldness, and the effect of hot and cold separation is relatively limited. At the same time, the merge strategy of the above solution will place the hot data file in the second-to-last layer of the LSM-tree, making the merge frequency of the second-to-last layer higher, consuming a lot of additional CPU and IO resources.

[0054] The second method is to select several keys for each level of the LSM-tree, and the data files in each level will be physically cut into multiple fragments by the keys, and each fragment consists of multiple files. However, the above solution does not utilize the time attribute of the project data. When applied to the stream computing BI scenario, it can only partially reduce the data merging overhead. In addition, since the fragments do not distinguish between hot and cold data, the merging process will still merge the hot and cold data, which will lead to poor locality of data access and reduced performance on the one hand, and consume a lot of CPU and IO resources on the other hand.

[0055] Therefore, in the embodiments of this specification, according to the characteristics of typical aggregation analysis jobs in stream computing scenarios, a solution is proposed for physically sharding the data in the state storage according to time based on the time attribute of the project data. This solution can better utilize the time attribute of the project data to achieve the effect of hot and cold separation. At the same time, the sharded storage format is more friendly to access locality, so that only one shard will be accessed in a time period, thereby obtaining better access performance and greatly reducing the CPU and IO overhead of data merging, thereby improving the overall performance of the stream computing system storage.

[0056] Specifically, the data to be stored is obtained, and a target storage unit corresponding to the data to be stored is determined; at least one storage time information is extracted from the data to be stored; at least one target storage shard is determined according to the at least one storage time information and the shard time indexes corresponding to the multiple storage shards in the target storage unit, wherein the storage time information corresponds one-to-one to the target storage shard; according to the at least one storage time information, the data to be stored is stored in the at least one target storage shard.

[0057] In this specification, a data storage method is provided. This specification also relates to a data query method, a data storage device, a data query device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0058] See also Figure 1 , Figure 1 The first data storage system provided by an embodiment of the present specification is shown in the architecture diagram. The data storage system may include a client 100 and a server 200;

[0059] The client 100 is used to send data to be stored to the server 200;

[0060] The server 200 is used to obtain the data to be stored and determine the target storage unit corresponding to the data to be stored; extract at least one storage time information from the data to be stored; determine at least one target storage shard according to the at least one storage time information and the shard time indexes corresponding to the multiple storage shards in the target storage unit, wherein the storage time information corresponds to the target storage shard one-to-one; store the data to be stored in at least one target storage shard according to the at least one storage time information.

[0061] By applying the solution of the embodiments of this specification, the data to be stored is stored in at least one target storage shard in the target storage unit according to the storage time information, thereby realizing sharded storage of data in different time periods, so that data locality is better and the data hit rate is higher. In addition, since the data heat is strongly related to the time period, the separation of hot and cold data is realized, reducing the merging overhead of cold and hot data.

[0062] See also Figure 2 , Figure 2The architecture diagram of the second data storage system provided by an embodiment of the present specification is shown. The data storage system may include multiple clients 100 and a server 200, wherein the client 100 may include a terminal-side device and the server 200 may include a cloud-side device. A communication connection may be established between multiple clients 100 through the server 200. In the data storage scenario, the server 200 is used to provide data storage services between multiple clients 100. Multiple clients 100 may serve as a sender or a receiver respectively, and realize communication through the server 200.

[0063] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the data storage scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 stores the data to be stored in at least one target storage shard according to the data stream, and pushes the storage result to other clients that have established communication.

[0064] The client 100 and the server 200 are connected via a network. The network provides a medium for a communication link between the client 100 and the server 200. The network may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, etc. before being released to the server 200.

[0065] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 100 can be based on the software development kit (SDK, Software Development Kit) of the corresponding service provided by the server 200, such as based on the real-time communication (RTC, Real Time Communication) SDK development and acquisition. The client 100 can be deployed in an electronic device, and needs to rely on the device to run or some APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0066] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that provide support for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0067] It is worth noting that the data storage method provided in the embodiments of this specification is generally executed by the server, but in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the data storage method provided in the embodiments of this specification. In other embodiments, the data storage method provided in the embodiments of this specification may also be jointly executed by the client and the server.

[0068] See also Figure 3 , Figure 3 A flowchart of a data storage method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0069] Step 302: Acquire data to be stored, and determine a target storage unit corresponding to the data to be stored.

[0070] In one or more embodiments of this specification, the data to be stored can be obtained, and the target storage unit corresponding to the data to be stored can be determined, so that based on the data storage method proposed in the embodiments of this specification, the data to be stored can be stored in the target storage unit to complete the data storage.

[0071] Specifically, the data to be stored refers to a data storage object. The data to be stored can be data in different formats, such as table data, text data, code data, etc. The data to be stored can also be data in different scenarios, such as data in transaction scenarios, data in social statistics scenarios, data in education scenarios, etc. The target storage unit can be a separate database or a table (state) in a database. The target storage unit is specifically set according to the actual situation, and the embodiments of this specification do not impose any restrictions on this.

[0072] In practical applications, there are many ways to obtain the data to be stored, which are selected according to the actual situation, and the embodiments of this specification do not limit this. In one possible implementation of this specification, the data to be stored sent by the user through the client can be received. In another possible implementation of this specification, the data to be stored can be read from other data acquisition devices or databases.

[0073] It should be noted that there are multiple ways to determine the target storage unit corresponding to the data to be stored, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the unit identifier of the target storage unit corresponding to the data to be stored sent by the user through the client can be received, so as to determine the target storage unit based on the unit identifier. In another possible implementation of this specification, any idle storage unit can be selected as the target storage unit.

[0074] In an optional embodiment of the present specification, in order to facilitate information reading or deletion, the data to be stored may be first stored in a data storage cache. When the amount of data in the data storage cache is greater than a preset cache threshold, the data to be stored in the data storage cache is stored in the target storage unit. That is, the above-mentioned acquisition of the data to be stored may include the following steps:

[0075] When the amount of data in the data storage cache is greater than a preset cache threshold, the data to be stored is obtained from the data storage cache.

[0076] Specifically, the data storage cache refers to a storage space for temporarily storing data, through which the stored data can be quickly accessed. The data storage cache includes but is not limited to random access memory (RAM) and solid state drive (SSD). The data volume refers to the amount of data stored in the data storage cache. The preset cache threshold is set according to the actual situation, and the embodiments of this specification do not make any restrictions on this.

[0077] It should be noted that if the amount of data in the data storage cache is less than or equal to the preset cache threshold, it means that the preset storage cache can store the data to be stored, and the data to be stored at this time does not need to be stored in the target storage unit, but can be stored in the data storage cache; if the amount of data in the data storage cache is greater than the preset cache threshold, it means that the cache space of the preset storage cache is full and no new data to be stored can be stored. At this time, the data to be stored in the preset storage cache can be read out and stored in the target storage unit to release the space of the data storage cache. Therefore, when the amount of data in the data storage cache is greater than the preset cache, the data to be stored can be obtained from the data storage cache, and further based on the data storage method proposed in the embodiment of this specification, the data to be stored can be stored in the target storage unit.

[0078] In actual applications, when obtaining data to be stored from the data storage cache, all the data to be stored in the data storage cache can be obtained, or only part of the data to be stored can be obtained. The amount of data to be stored obtained from the data storage cache is selected according to the actual situation, and the embodiments of this specification do not impose any limitations on this.

[0079] By applying the solution of the embodiments of the present specification, when the amount of data in the data storage cache is greater than a preset cache threshold, the data to be stored is obtained from the data storage cache, avoiding all data query requests from accessing the slow permanent storage unit, thereby reducing the data access time and improving the overall data processing performance.

[0080] In an optional embodiment of the present specification, after determining the target storage unit corresponding to the data to be stored, the unit type of the target storage unit may be determined, wherein the unit type may be a storage unit with shard storage enabled or a storage unit with shard storage not enabled. Further, after obtaining the data to be stored and determining the target storage unit corresponding to the data to be stored, the following steps may be further included:

[0081] When the target storage unit does not enable shard storage, the data to be stored is stored in a preset sub-unit in the target storage unit.

[0082] It should be noted that if the target storage unit does not enable shard storage, it means that it does not support the data to be stored being divided and stored in multiple storage shards. In this case, the data to be stored can be stored in a preset sub-unit in the target storage unit. The preset sub-unit is a separate storage shard (default-shard) in the target storage unit. If the target storage unit enables shard storage, it means that the data to be stored can be divided and stored in multiple storage shards. In this case, at least one storage time information can be extracted from the data to be stored, and the corresponding target storage shard can be determined based on the at least one storage time information. Among them, the storage allocation function of the target storage unit can be pre-configured to be enabled, or it can be enabled by the user through SQLHint.

[0083] By applying the solution of the embodiment of this specification, when the target storage unit does not enable shard storage, the data to be stored is stored in a preset sub-unit in the target storage unit, thereby realizing flexible storage of the data to be stored.

[0084] Step 304: extract at least one storage time information from the data to be stored.

[0085] In one or more embodiments of the present specification, after acquiring the data to be stored and determining the target storage unit corresponding to the data to be stored, further, at least one storage time information may be extracted from the data to be stored.

[0086] Specifically, the storage time information refers to the time attribute information contained in the data to be stored. The storage time information includes but is not limited to the storage time field and the storage time tag. For example, if the data to be stored is a table, the fields such as "hour, day, month" in the table are the storage time information in the data to be stored.

[0087] In practical applications, there are many ways to extract at least one storage time information from the data to be stored, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, at least one storage time information can be extracted from the data to be stored by information matching. In another possible implementation of this specification, the user can receive the specified time information sent by SQL Hint, and generate a time information extractor based on the specified time information, so as to use the time information extractor to extract at least one storage time information from the data to be stored.

[0088] In an optional embodiment of the present specification, the extracting of at least one storage time information from the data to be stored may include the following steps:

[0089] Obtaining data prompt information, wherein the data prompt information includes specified time information;

[0090] Based on the specified time information, determining at least one location information storing the time information;

[0091] At least one storage time information is extracted from the data to be stored according to the location information.

[0092] Specifically, data hint information refers to hint information sent by the user through SQL Hint, and the data hint information includes the specified time information for sharding. The specified time information includes but is not limited to the specified time field and the specified time label. The location information refers to the location information of at least one storage time information in the user key (userkey).

[0093] It should be noted that in the stream computing aggregation analysis operator, the aggregation field corresponds to the user key in the state storage. The user can specify the specified time information for sharding through SQL Hint. When the user SQL is converted into a stream computing job, an information extractor (Partition Key Extractor) can be generated based on the specified time information, where the information extractor is used to record the location information of the storage time information; the information extractor is used to extract at least one storage time information from the data to be stored.

[0094] See also Figure 4 , Figure 4 FIG. 1 shows a schematic diagram of extracting storage time information in a data storage method provided by an embodiment of the present specification, such as Figure 4 As shown in the figure, the user key consists of four fields: biz+date+hour+Cate_id. The user specifies date+hour as the specified time field through SQLHint. Further, a field extractor can be generated based on the specified time field. Then the field extractor can extract the date+hour field from each received data according to the rules as the storage time field (partition key) of the data.

[0095] The scheme of the embodiment of this specification is applied to obtain data hint information, wherein the data hint information includes specified time information; based on the specified time information, determine the location information of at least one storage time information; and extract at least one storage time information from the data to be stored according to the location information. The SQL Hint method enables the state storage layer to perceive the specified time information specified by the user, thereby better optimizing the physical storage format of the data, enhancing the locality of data access, and improving the overall performance of the system.

[0096] Step 306: Determine at least one target storage shard according to at least one storage time information and shard time indexes corresponding to a plurality of storage shards in the target storage unit, wherein the storage time information corresponds one-to-one to the target storage shard.

[0097] In one or more embodiments of the present specification, data to be stored is obtained, and a target storage unit corresponding to the data to be stored is determined; after extracting at least one storage time information from the data to be stored, at least one target storage shard can be further determined based on the at least one storage time information and a shard time index corresponding to each of multiple storage shards in the target storage unit.

[0098] Specifically, the target storage unit may include multiple storage shards, and the storage shards may be an LSM-tree hierarchical storage structure, and each storage shard corresponds to a shard time index. The shard time index can be understood as a binary search index, and the shard time index is used to index the storage shard corresponding to the time information. The shard time index is obtained based on the storage time information of the data stored in the storage shard. The target storage shard is used to store the data to be stored.

[0099] In actual applications, when determining at least one target storage shard based on at least one storage time information and the shard time indexes corresponding to multiple storage shards in the target storage unit, the time information corresponding to the at least one storage time information can be searched from the shard time index, and the storage shard corresponding to the time information can be further used as the target storage shard corresponding to the at least one storage time information.

[0100] It should be noted that if the target storage shard corresponding to the storage time information is not found in the target storage unit based on the shard time index, the target storage shard may be constructed in the target storage unit based on the storage time information.

[0101] Step 308: According to at least one storage time information, store the data to be stored in at least one target storage slice.

[0102] In one or more embodiments of the present specification, data to be stored is obtained, and a target storage unit corresponding to the data to be stored is determined; at least one storage time information is extracted from the data to be stored; after at least one target storage shard is determined based on at least one storage time information and a shard time index corresponding to a plurality of storage shards in the target storage unit, further, the data to be stored can be stored in at least one target storage shard based on the at least one storage time information.

[0103] In practical applications, when storing the data to be stored in at least one target storage shard according to at least one storage time information, the data to be stored can be directly stored in the at least one target storage shard according to at least one storage time information. The data to be stored can also be stored in the LSM-tree storage structure corresponding to the at least one target storage shard according to at least one storage time information.

[0104] By applying the solution of the embodiments of this specification, the data to be stored is stored in at least one target storage shard in the target storage unit according to the storage time information, thereby realizing sharded storage of data in different time periods, so that data locality is better and the data hit rate is higher. In addition, since the data heat is strongly related to the time period, the separation of hot and cold data is realized, reducing the merging overhead of cold and hot data.

[0105] In an optional embodiment of the present specification, storing the data to be stored in at least one target storage slice according to at least one storage time information may include the following steps:

[0106] Dividing the data to be stored according to at least one storage time information, and determining the sub-data to be stored corresponding to the at least one storage time information;

[0107] According to the correspondence between the storage time information and the target storage slice, the to-be-stored sub-data respectively corresponding to at least one storage time information is stored in at least one target storage slice.

[0108] It should be noted that the data to be stored corresponding to different storage time information can be stored in different storage shards. Therefore, after determining at least one storage time information and the target storage shard corresponding to each storage time information, the data to be processed can be divided into multiple sub-data to be stored according to the storage time information, and each sub-data to be stored can be stored in the target storage shard corresponding to the storage time information.

[0109] By applying the solution of the embodiments of this specification, the data to be stored is stored in at least one target storage slice in the target storage unit according to the storage time information, thereby realizing the shard storage of data in different time periods, making the data locality better and the data hit rate higher.

[0110] In an optional embodiment of the present specification, after storing the data to be stored in at least one target storage slice according to at least one storage time information, the following steps may also be included:

[0111] Determine a data deletion time of at least one target storage shard, wherein the data deletion time is obtained based on an expiration time of each stored data in the at least one target storage shard;

[0112] When the current time is greater than or equal to the data deletion time, the stored data in at least one target storage slice is deleted.

[0113] It should be noted that under the shard storage architecture based on time information, data in different time periods will be written to different storage shards, and ordinary data merging will not mix data from different time periods together. Furthermore, since the life cycle (data writing time and expiration time) of data in the same storage shard is relatively similar, the embodiments of this specification introduce a time to live (TTL, Time To Live) expired data cleanup strategy based on shard granularity to efficiently clean up expired shard data.

[0114] In actual applications, there are many ways to determine the data deletion time of at least one target storage shard, and the specific selection is based on the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the expiration time of each stored data in the target storage shard can be obtained, and the latest expiration time can be used as the data deletion time of the target storage shard. In another possible implementation of this specification, a preset time threshold can be added to the latest expiration time to obtain the data deletion time.

[0115] Furthermore, if the current time is greater than or equal to the data deletion time, it means that all stored data in the target storage slice has expired, and at this time the data in the target storage slice can be deleted as a whole.

[0116] By applying the solution of the embodiments of this specification, the data deletion time of at least one target storage shard is determined; when the current time is greater than or equal to the data deletion time, the stored data in at least one target storage shard is deleted, thereby achieving no IO overhead and very little CPU overhead when cleaning invalid data, thereby improving data cleaning efficiency.

[0117] See also Figure 5 , Figure 5 A schematic diagram of data deletion in a data storage method provided by an embodiment of this specification is shown. Figure 5 As shown, assuming that the current time (Current Time) is 2000, the data deletion time (Expire Time) of storage shard 1-1 is 1000, the data deletion time of storage shard 1-2 is 2500, and the data deletion time of storage shard 1-3 is 4000, since the current time is greater than the data deletion time of storage shard 1-1, the stored data in storage shard 1-1 can be deleted at this time.

[0118] In an optional embodiment of the present specification, after storing the data to be stored in at least one target storage slice according to at least one storage time information, the following steps may also be included:

[0119] Get the data information of the target storage shard;

[0120] Based on the data information, the storage data in the target storage shard is updated.

[0121] Specifically, the data information of the target storage shard includes, but is not limited to, data storage capacity and data access information, wherein the data access information includes, but is not limited to, the number of data accesses and data access popularity.

[0122] It should be noted that when a small amount of disordered data flows into the stream computing system, small files are easily generated in the first storage layer of the LSM-tree structure of the target storage shard, affecting the read and write performance of the storage system. Therefore, in the embodiments of this specification, an additional data merging mechanism can be introduced to update the data stored in the target storage shard to solve the problem of small files under the shard storage architecture.

[0123] In actual applications, there are many ways to obtain the data information of the target storage shard, which can be selected according to the actual situation, and the embodiments of this specification do not limit this. In one possible implementation of this specification, the data information of the target storage shard can be read from the storage log of the target storage shard. In another possible implementation of this specification, the data information of the target storage shard sent by the user through the client can be received.

[0124] By applying the solution of the embodiment of this specification, data information of the target storage shard is obtained; based on the data information, the storage data in the target storage shard is updated, which solves the problem of small files under the shard storage architecture and improves the read and write performance of the storage system.

[0125] In actual applications, there are multiple ways to update the stored data in the target storage slice based on the data information. The specific method is selected according to the actual situation, and the embodiments of this specification do not impose any limitation on this.

[0126] In a possible implementation of the present specification, the target storage shard includes a first storage layer, and the data information includes a first data storage capacity of the first storage layer; and the updating of the storage data in the target storage shard based on the data information may include the following steps:

[0127] When the first data storage amount is less than or equal to the storage amount threshold of the first storage layer and greater than a preset merging threshold, the storage data in the first storage layer is merged.

[0128] Specifically, the first storage layer refers to level-0 in the LSM-tree, and the storage capacity threshold and the preset merge threshold are set according to the actual situation, and the embodiments of this specification do not impose any restrictions on this. For example, the storage capacity threshold is half of the total storage capacity of the first storage layer.

[0129] It should be noted that the data in the LSM-tree is stored in layers. After the amount of data stored in the first layer reaches a certain size, the data in the first storage layer will be merged with the data in the second storage layer. As a result, the data in the second storage layer will be merged frequently and rewritten repeatedly, consuming system resources and thus affecting the read and write performance. Therefore, in the embodiments of this specification, when the amount of the first data stored in the first storage layer reaches the preset merge threshold but is less than or equal to half of the total storage capacity of the first storage layer, the files within the first storage layer can be merged, and the merge result is still placed in the first storage layer without being merged with the lower-layer files.

[0130] Applying the solution of the embodiments of this specification, when the amount of the first data is less than or equal to the storage capacity threshold of the first storage layer and greater than the preset merge threshold, the stored data in the first storage layer is merged. This solves the problem of small files in the sharded storage architecture and improves the read and write performance of the storage system.

[0131] In another possible implementation manner of this specification, the target storage shard includes a first storage layer and a second storage layer, the storage priority of the first storage layer is higher than that of the second storage layer, and the data information includes data access information; the above-mentioned updating of the stored data in the target storage shard based on the data information may include the following steps:

[0132] When the data access information is less than or equal to the preset access threshold, the stored data in the first storage layer is merged into the second storage layer, and the stored data in the first storage layer is deleted.

[0133] Specifically, the second storage layer refers to level-1 in the LSM-tree, and the preset access threshold is specifically set according to the actual situation, and the embodiments of this specification do not make any limitations in this regard.

[0134] It should be noted that if there is no new data written into the LSM-tree, the data in the first storage layer can be merged into the second storage layer. Since redundant data is merged, the space occupancy will be less, and the read link of the LSM-tree will be shorter, thus achieving an improvement in read performance. In the stream computing scenario, since the data writing time in the storage shard is relatively close, and the access time is also relatively close. When a certain storage shard has passed the period of concentrated writing or reading, the probability of accessing the data in this shard will become smaller. Therefore, when the data access information is less than or equal to the preset access threshold, all the stored data in the first storage layer can be merged into the second storage layer, and the stored data in the first storage layer is deleted.

[0135] By applying the solution of the embodiment of this specification, when the data access information is less than or equal to the preset access threshold, the storage data in the first storage layer is merged into the second storage layer, and the storage data in the first storage layer is deleted. This realizes the clearing of small files in the first storage layer and improves the read performance of the target storage shard.

[0136] In an optional embodiment of the present specification, in order to support the smooth migration of user data from the normal state to the shard storage state, and to restore from the shard storage state to the default state, the embodiment of the present specification can also support seamless migration between the shard architecture and the normal architecture. Specifically, when migrating from the default state to the shard storage state, the data can be traversed out from the original LSM-tree storage in sequence, and then written to the new shard storage structure; when the shard storage state is restored to the default state, the data can be traversed out from all shard storages in sequence, and then written to the general LSM-tree storage structure.

[0137] See also Figure 6 , Figure 6 A flowchart of a data query method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0138] Step 602: Receive a data query request, wherein the data query request carries data query information.

[0139] Step 604: parse the data query information and determine at least one query time information.

[0140] Step 606: According to at least one query time information, search for target storage data corresponding to the data query information from at least one target storage shard, wherein the at least one target storage shard is obtained based on at least one storage time information and a shard time index corresponding to each of multiple storage shards in the target storage unit, the storage time information corresponds one-to-one to the target storage shard, the at least one target storage shard includes multiple storage data, and the storage data is stored in the at least one target storage shard based on the at least one storage time information.

[0141] Specifically, the data query information represents the query demand, and the data query information includes but is not limited to the target storage unit to be queried and at least one query time information. The query time information includes but is not limited to the query time field and the query time tag. The target storage data is the final data query result.

[0142] It should be noted that there are multiple ways to parse data query information and determine at least one query time information, which can be selected based on actual conditions, and the embodiments of this specification do not impose any restrictions on this. In an optional embodiment of this specification, an information extractor can be used to extract at least one query time information from the data query information. In another possible implementation of this specification, a data extraction model can be used to parse at least one query time information from the data query information.

[0143] By applying the solution of the embodiments of this specification, since the stored data is stored in the target storage shards in the target storage unit, sharded storage of data in different time periods is achieved, which makes data locality better and the data hit rate higher. In addition, since the data popularity is strongly related to the time period, the separation of hot and cold data is achieved, reducing the merging overhead of cold and hot data.

[0144] In an optional embodiment of the present specification, after receiving the data query request, the following steps may also be included:

[0145] Searching the data storage cache for target storage data corresponding to the data query information;

[0146] Parsing the data query information and determining at least one query time information, including:

[0147] In a case where the target storage data is not included in the data storage cache, the data query information is parsed to determine at least one query time information.

[0148] It should be noted that after receiving a data query request, the target storage data requested by the data query request may be stored in the data storage cache, but not yet stored in the target storage unit. Therefore, in order to save data query resources, after receiving the data query request, the target storage data corresponding to the data query information can be searched from the data storage cache first. If the target storage data is included in the data storage cache, the data query is completed; if the target storage data is not included in the data storage cache, the data query information can be parsed to determine at least one query time information, and the target storage data can be searched from at least one target storage shard based on the at least one query time information.

[0149] Applying the solution of the embodiment of this specification, searching the data storage cache for target storage data corresponding to the data query information;

[0150] In the case that the target storage data is not included in the data storage cache, the data query information is parsed and at least one query time information is determined, thereby saving data query time and improving data query efficiency.

[0151] In an optional embodiment of the present specification, searching for target storage data corresponding to the data query information from at least one target storage shard according to at least one query time information may include the following steps:

[0152] Based on the shard time index, search for at least one target storage shard corresponding to the query time information;

[0153] In a case where the target storage unit does not include the target storage shard, determining that the data search fails;

[0154] In the case where the target storage unit includes a target storage slice, the target storage data corresponding to the data query information is extracted from the target storage slice.

[0155] It should be noted that the implementation method of "searching for at least one target storage shard corresponding to query time information based on the shard time index" is the same as the implementation method of "determining at least one target storage shard based on at least one storage time information and the shard time indexes corresponding to multiple storage shards in the target storage unit", and the embodiments of this specification do not impose any limitations on this.

[0156] In actual applications, if the target storage unit does not include the target storage shard, it means that the target storage unit does not include the target storage data and the data search fails; if the target storage unit includes the target storage shard, the target storage data can be read from the LSM-tree corresponding to the target storage shard.

[0157] By applying the solution of the embodiments of the present specification, based on the shard time index, at least one target storage shard corresponding to the query time information is searched; when the target storage unit does not include the target storage shard, it is determined that the data search has failed; when the target storage unit includes the target storage shard, the target storage data corresponding to the data query information is extracted from the target storage shard, thereby improving data query efficiency and reducing query merging overhead.

[0158] See also Figure 7 , Figure 7 A flowchart of a data storage method according to an embodiment of the present specification is shown, which specifically includes:

[0159] The user specifies the specified time information in the job SQL through SQL Hint and other methods, and then the storage time information in the data to be stored can be transparently transmitted to the storage layer; in the storage layer, the data belonging to different time periods are stored in shards using the above data storage method, and the data in different shards are physically isolated. In other words, the job structured query statement is converted into a stream computing job, and the time information hint is converted into a shard storage architecture.

[0160] It should be noted that since data access popularity is strongly correlated with time information, sharded storage allows for better data locality, a higher storage system hit rate, and can also achieve hot and cold data separation, reducing the overhead of merging hot and cold data.

[0161] See also Figure 8 , Figure 8 A process flow chart of a data query method provided by an embodiment of the present specification is shown, which specifically includes:

[0162] When the data query starts, the target storage data corresponding to the data query information can be searched from the data storage cache to determine whether the target storage data is included in the data storage cache: if it is included, the target storage data is returned and the data query is terminated; if it is not included, the data query information is parsed to determine at least one query time information, and a binary search is performed based on the at least one query time information, that is, based on the shard time index, the target storage shard corresponding to at least one query time information is searched to determine whether the target storage shard exists: if it does not exist, invalid (NULL) is returned, it is determined that the data search fails, and the data query is terminated; if it exists, the target storage data is extracted from the LSM-tree corresponding to the target storage shard, and it is determined whether the target storage data is included in the LSM-tree: if it is included, the target storage data is returned and the data query is terminated; if it is not included, invalid (NULL) is returned, it is determined that the data search fails, and the data query is terminated.

[0163] See also Fig. 9 , Fig. 9 An architecture diagram of a third data storage system provided by an embodiment of the present specification is shown, wherein the data storage system includes a plurality of target storage units with shard storage enabled, a plurality of target storage units with shard storage disabled, information extractors corresponding to the target storage units, and a data storage cache (Write Buffer) and a data query cache (Block Cache) shared by all storage shards;

[0164] like Fig. 9As shown, when the amount of data in the data storage cache is greater than a preset cache threshold, the data to be stored is obtained from the data storage cache, and the target storage unit corresponding to the data to be stored is determined; when the target storage unit does not enable shard storage, the data to be stored is stored in a preset sub-unit in the target storage unit; when the target storage unit enables shard storage, an information extractor is used to extract at least one storage time information from the data to be stored; at least one target storage shard is determined based on at least one storage time information and shard time indexes corresponding to multiple storage shards in the target storage unit, wherein the storage time information corresponds one-to-one to the target storage shard; the data to be stored is divided based on at least one storage time information, and the sub-data to be stored corresponding to at least one storage time information is determined; based on the correspondence between the storage time information and the target storage shard, the sub-data to be stored corresponding to at least one storage time information is stored in the LSM-tree storage structure corresponding to at least one target storage shard.

[0165] For example, taking the stream computing scenario of e-commerce as an example, the data to be stored is the transaction volume of all stores every hour. Using the above data storage solution, the project data at time 1 can be stored in an LSM-tree, the project data at time 2 can be stored in an LSM-tree, and so on. Under the time-slicing architecture, the file where the project data at time 1 is located will not be merged with the file where the project data at time 2 is located, which can effectively reduce the merging overhead of cold and hot data.

[0166] By applying the solution of the embodiments of this specification, the state storage layer can identify the time information in the user data through the SQL Hint method, perceive the time attribute of the stored data, and optimize the storage structure by using the time characteristics; by identifying the time information in the data to be stored, the data belonging to different time periods are stored in shards to achieve cold and hot separation, thereby effectively improving data locality, reducing the IO and CPU overhead caused by the merging of cold and hot data, and improving the overall performance of the system in large data volume scenarios; under the sharded storage architecture, a shard merging strategy is introduced to greatly improve the efficiency of cleaning up expired and invalid data; a small file merging mechanism is introduced to solve the small file problem under the sharded architecture; and seamless migration between the sharded architecture and the ordinary architecture is supported to further improve the usability of the sharded storage architecture.

[0167] Corresponding to the above data storage method embodiment, this specification also provides a data storage device embodiment. Fig.10 FIG. 1 is a schematic diagram showing a structure of a data storage device provided by an embodiment of the present specification. Fig.10 As shown, the device comprises:

[0168] The acquisition module 1002 is configured to acquire the data to be stored and determine the target storage unit corresponding to the data to be stored;

[0169] The extraction module 1004 is configured to extract at least one storage time information from the data to be stored;

[0170] The determination module 1006 is configured to determine at least one target storage shard according to at least one storage time information and shard time indexes respectively corresponding to a plurality of storage shards in the target storage unit, wherein the storage time information corresponds to the target storage shard in a one-to-one manner;

[0171] The first storage module 1008 is configured to store the data to be stored in at least one target storage slice according to at least one storage time information.

[0172] Optionally, the device further includes: a second storage module configured to store the data to be stored in a preset sub-unit in the target storage unit when the target storage unit does not enable shard storage.

[0173] Optionally, the determination module 1006 is further configured to divide the data to be stored according to at least one storage time information, and determine the sub-data to be stored corresponding to at least one storage time information; according to the correspondence between the storage time information and the target storage shard, the sub-data to be stored corresponding to at least one storage time information is stored in at least one target storage shard.

[0174] Optionally, the device also includes: a deletion module, configured to determine a data deletion time of at least one target storage shard, wherein the data deletion time is obtained based on an expiration time of each stored data in at least one target storage shard; and when the current time is greater than or equal to the data deletion time, deleting the stored data in at least one target storage shard.

[0175] Optionally, the device further includes: an update module configured to obtain data information of the target storage slice; and update the storage data in the target storage slice based on the data information.

[0176] Optionally, the target storage shard includes a first storage layer, and the data information includes a first data storage capacity of the first storage layer; the update module is further configured to merge the storage data in the first storage layer when the first data storage capacity is less than or equal to a storage capacity threshold of the first storage layer and greater than a preset merge threshold.

[0177] Optionally, the target storage slice includes a first storage layer and a second storage layer, the storage priority of the first storage layer is higher than the storage priority of the second storage layer, and the data information includes data access information; the update module is further configured to merge the storage data in the first storage layer into the second storage layer and delete the storage data in the first storage layer when the data access information is less than or equal to a preset access threshold.

[0178] Optionally, the extraction module 1004 is further configured to obtain data prompt information, wherein the data prompt information includes specified time information; based on the specified time information, determine location information of at least one storage time information; and extract at least one storage time information from the data to be stored according to the location information.

[0179] Optionally, the acquisition module 1002 is further configured to acquire the data to be stored from the data storage cache when the amount of data in the data storage cache is greater than a preset cache threshold.

[0180] By applying the solution of the embodiments of this specification, the data to be stored is stored in at least one target storage shard in the target storage unit according to the storage time information, thereby realizing sharded storage of data in different time periods, so that data locality is better and the data hit rate is higher. In addition, since the data heat is strongly related to the time period, the separation of hot and cold data is realized, reducing the merging overhead of cold and hot data.

[0181] The above is a schematic scheme of a data storage device of this embodiment. It should be noted that the technical scheme of the data storage device and the technical scheme of the data storage method described above are of the same concept, and the details of the technical scheme of the data storage device that are not described in detail can be found in the description of the technical scheme of the data storage method described above.

[0182] Corresponding to the above data query method embodiment, this specification also provides a data query device embodiment, Fig.11 FIG. 1 shows a schematic diagram of a data query device provided by an embodiment of the present specification. Fig.11 As shown, the device comprises:

[0183] The receiving module 1102 is configured to receive a data query request, wherein the data query request carries data query information;

[0184] The parsing module 1104 is configured to parse the data query information and determine at least one query time information;

[0185] The first search module 1106 is configured to search for target storage data corresponding to data query information from at least one target storage shard according to at least one query time information, wherein the at least one target storage shard is obtained based on at least one storage time information and a shard time index corresponding to a plurality of storage shards in a target storage unit, the storage time information corresponds one-to-one to the target storage shard, the at least one target storage shard includes a plurality of storage data, and the storage data is stored in the at least one target storage shard based on the at least one storage time information.

[0186] Optionally, the device also includes: a second search module, configured to search for target storage data corresponding to the data query information from the data storage cache; a parsing module 1104, further configured to parse the data query information and determine at least one query time information when the target storage data is not included in the data storage cache.

[0187] Optionally, the first search module 1106 is further configured to search for target storage shards corresponding to at least one query time information based on the shard time index; determine that the data search fails when the target storage shard is not included in the target storage unit; and extract the target storage data corresponding to the data query information from the target storage shard when the target storage unit includes the target storage shard.

[0188] By applying the solution of the embodiments of this specification, since the stored data is stored in the target storage shards in the target storage unit, sharded storage of data in different time periods is achieved, which makes data locality better and the data hit rate higher. In addition, since the data popularity is strongly related to the time period, the separation of hot and cold data is achieved, reducing the merging overhead of cold and hot data.

[0189] The above is a schematic scheme of a data query device of this embodiment. It should be noted that the technical scheme of the data query device and the technical scheme of the above data query method belong to the same concept, and the details not described in detail in the technical scheme of the data query device can be referred to the description of the technical scheme of the above data query method.

[0190] Fig.12 The block diagram of a computing device provided by an embodiment of the present specification is shown. The components of the computing device 1200 include but are not limited to a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and the database 1250 is used to store data.

[0191] The computing device 1200 also includes an access device 1240 that enables the computing device 1200 to communicate via one or more networks 1260. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1240 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0192] In one embodiment of the present specification, the above components of the computing device 1200 and Fig.12 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.12 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0193] The computing device 1200 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1200 may also be a mobile or stationary server.

[0194] The processor 1220 is used to execute the following computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned data storage method or data query method.

[0195] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned data storage method and data query method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned data storage method or data query method.

[0196] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned data storage method or data query method.

[0197] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned data storage method and data query method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned data storage method or data query method.

[0198] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data storage method or data query method.

[0199] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the above-mentioned data storage method and data query method belong to the same concept, and the details not described in detail in the technical scheme of the computer program can be referred to the description of the technical scheme of the above-mentioned data storage method or data query method.

[0200] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0201] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0202] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0203] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0204] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A data storage method, include: Acquire data to be stored, and determine a target storage unit corresponding to the data to be stored; Extracting at least one storage time information from the data to be stored; Determine at least one target storage shard according to the at least one storage time information and the shard time indexes respectively corresponding to the plurality of storage shards in the target storage unit, wherein the storage time information corresponds to the target storage shard in a one-to-one manner; According to the at least one storage time information, the data to be stored is stored in the at least one target storage slice.

2. The method according to claim 1, after obtaining the data to be stored and determining the target storage unit corresponding to the data to be stored, further comprising: include: When the target storage unit does not enable shard storage, the data to be stored is stored in a preset sub-unit in the target storage unit.

3. The method according to claim 1, storing the data to be stored in the at least one target storage shard according to the at least one storage time information, include: Dividing the data to be stored according to the at least one storage time information, and determining the sub-data to be stored corresponding to the at least one storage time information respectively; According to the correspondence between the storage time information and the target storage slice, the to-be-stored sub-data respectively corresponding to the at least one storage time information are stored in the at least one target storage slice.

4. The method according to claim 1, further comprising storing the data to be stored in the at least one target storage shard according to the at least one storage time information. include: Determining a data deletion time of the at least one target storage shard, wherein the data deletion time is obtained based on an expiration time of each stored data in the at least one target storage shard; When the current time is greater than or equal to the data deletion time, the stored data in the at least one target storage slice is deleted.

5. The method according to claim 1, further comprising storing the data to be stored in the at least one target storage shard according to the at least one storage time information. include: Obtaining data information of the target storage shard; Based on the data information, the storage data in the target storage slice is updated.

6. The method according to claim 5, wherein the target storage slice comprises a first storage layer, and the data information comprises a first data storage capacity of the first storage layer; Based on the data information, the stored data in the target storage slice is updated, include: When the first data storage amount is less than or equal to a storage amount threshold of the first storage layer and greater than a preset merging threshold, the storage data in the first storage layer is merged.

7. The method according to claim 5, wherein the target storage slice comprises a first storage layer and a second storage layer, the storage priority of the first storage layer is higher than the storage priority of the second storage layer, and the data information comprises data access information; Based on the data information, the stored data in the target storage slice is updated, include: In a case where the data access information is less than or equal to a preset access threshold, the storage data in the first storage layer is merged into the second storage layer, and the storage data in the first storage layer is deleted.

8. The method according to claim 1, wherein at least one storage time information is extracted from the data to be stored. include: Acquire data prompt information, wherein the data prompt information includes specified time information; Based on the specified time information, determining at least one location information for storing the time information; At least one storage time information is extracted from the data to be stored according to the location information.

9. The method according to claim 1, wherein obtaining the data to be stored, include: When the amount of data in the data storage cache is greater than a preset cache threshold, the data to be stored is obtained from the data storage cache.

10. A data query method, include: receiving a data query request, wherein the data query request carries data query information; Parsing the data query information to determine at least one query time information; According to the at least one query time information, the target storage data corresponding to the data query information is searched from at least one target storage shard, wherein the at least one target storage shard is obtained based on the at least one storage time information and the shard time indexes corresponding to the multiple storage shards in the target storage unit, the storage time information corresponds one-to-one to the target storage shard, the at least one target storage shard includes multiple storage data, and the storage data is stored in the at least one target storage shard based on the at least one storage time information.

11. The method according to claim 10, after receiving the data query request, further include: Searching the data storage cache for target storage data corresponding to the data query information; The step of parsing the data query information and determining at least one query time information includes: In a case where the target storage data is not included in the data storage cache, the data query information is parsed to determine at least one query time information.

12. The method according to claim 10, wherein the target storage data corresponding to the data query information is searched from at least one target storage shard according to the at least one query time information, include: Based on the shard time index, searching for the target storage shards corresponding to the at least one query time information respectively; If the target storage unit does not include the target storage slice, determining that the data search fails; In a case where the target storage unit includes the target storage slice, the target storage data corresponding to the data query information is extracted from the target storage slice.

13. A computing device, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 9 or any one of claims 10 to 12 are implemented.

14. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 9 or any one of claims 10 to 12.

Citation Information

Cited By

  • Operation and maintenance data processing method and system for AI intelligent agent

    CN121187518A

  • An operation and maintenance data processing method and system for an AI agent

    CN121187518B