Message index processing method, device, and storage medium

By converting the index file format in the local disk into a continuous index file with address and storing it to a distributed storage system, the problems of high storage costs and low query efficiency are solved, and low-cost and efficient index information storage and query are achieved.

WO2025149862A1PCT designated stage expired Publication Date: 2025-07-17CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Patent Information

Application Number
PCT/IB2025/050066
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2025-01-03
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the prior art, persistent storage of message index information on local disk leads to high storage costs and limited capacity, and multiple random accesses are required during querying, resulting in low query performance.

Method used

Convert the index file format in the local disk into a continuous index file and store it in an external distributed storage system to reduce the number of reads and improve query efficiency.

Benefits of technology

It reduces storage costs, expands storage capacity, and obtains multiple index information through a read operation at a time, improves query efficiency and solves the problem of read and write amplification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050066_17072025_PF_FP_ABST
    Figure IB2025050066_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a message index processing method, a device, and a storage medium. The method comprises: acquiring a first index file from a local disk, wherein the first index file comprises a plurality of index slots and a plurality of index entries used for storing index information of different messages, index information corresponding to each of at least two messages mapped to a first index slot is stored in at least two index entries at non-contiguous addresses, and the first index slot stores address information of an index entry among the at least two index entries, said index entry storing index information of a message last mapped to the first index slot; performing format conversion on the first index file to obtain a second index file; and storing the second index file in an external distributed storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field of Message Index Processing Method, Device and Storage Medium

[0001] This application relates to the field of cloud computing technology, and in particular, to a message index processing method, device, and storage medium. Background Art

[0002] In a distributed system, messages generated by a message producer will be stored in a log file, and a message index will be established for each message so that users can query messages based on the message index.

[0003] In a traditional message index processing solution, after a message generated by a message producer is stored in a log file, index information corresponding to the message is generated and the index information corresponding to the message is persistently stored on a local disk. However, persistently storing the index information corresponding to each message on a local disk not only has a high storage cost, but also the disk capacity is limited and cannot store a large amount of index information. Moreover, when querying a certain target message, it is necessary to randomly access the local disk multiple times to query the index information corresponding to the target message, resulting in a read amplification problem, that is, the entire query process has a large query overhead and a long query time, thus affecting the query performance of messages. Summary of the Invention

[0004] Embodiments of this application provide a message index processing method, device, and storage medium to reduce the storage cost of message indexes and improve the query efficiency at the same time.

[0005] In a first aspect, an embodiment of this application provides a message index processing method, the method includes: obtaining a first index file for storing index information of multiple messages from a local disk, the first index file has been written, and the first index file includes multiple index slots and multiple index entries for storing index information of different messages; wherein, index information corresponding to at least two messages mapped to a first index slot is stored in at least two index entries with discontinuous addresses, and address information of a target index entry among the at least two index entries is stored in the first index slot, the target index entry refers to an index entry storing index information of the message finally mapped to the first index slot; performing format conversion on the first index file to obtain a second index file, wherein, in the second index file, the index information of the at least two messages is stored in at least two index entries with continuous addresses, and address range information corresponding to the at least two index entries is stored in the first index slot; storing the second index file in an external distributed storage system, and deleting the first index file.

[0006] Second aspect, an embodiment of the present application provides a message index processing device, the device includes: an acquisition module, configured to acquire, from a local disk, a first index file for storing index information of multiple messages, the first index file has been written, and the first index file includes multiple index slots and multiple index entries for storing index information of different messages; wherein, index information corresponding to at least two messages mapped to a first index slot is stored in at least two index entries with discontinuous addresses, and address information of a target index entry among the at least two index entries is stored in the first index slot, the target index entry refers to an index entry storing index information of the message finally mapped to the first index slot; a conversion module, configured to perform format conversion on the first index file to obtain a second index file, wherein, in the second index file, the index information of the at least two messages is stored in at least two index entries with continuous addresses, and address range information corresponding to the at least two index entries is stored in the first index slot; a storage module, configured to store the second index file in an external distributed storage system and delete the first index file. The index information corresponding to at least two messages mapped to the first index slot is stored in at least two index entries with discontinuous addresses, and address information of a target index entry among the at least two index entries is stored in the first index slot, the target index entry refers to an index entry storing index information of the message finally mapped to the first index slot; a conversion module, configured to perform format conversion on the first index file to obtain a second index file, wherein, in the second index file, the index information of the at least two messages is stored in at least two index entries with continuous addresses, and address range information corresponding to the at least two index entries is stored in the first index slot; a storage module, configured to store the second index file in an external distributed storage system and delete the first index file.

[0007] Third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the message index processing method as described in the first aspect.

[0008] Fourth aspect, an embodiment of the present application provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the message index processing method as described in the first aspect.

[0009] Fifth aspect, an embodiment of the present application provides a computer program product, including: a computer program, and when the computer program is executed by a processor of an electronic device, the processor executes the message index processing method as described in the first aspect.

[0010] In the message index processing solution provided by the embodiments of the present application, the first index file stored on the local disk can be subjected to format conversion processing, and the obtained second index file after conversion can be stored in an external distributed storage system to expand the storage space. Since the storage cost of the distributed storage system is often lower than that of the local disk and the capacity can be elastically scaled, the storage of index files with a large amount of data can be realized. The main purpose of the above file format conversion is to improve the query efficiency of the index information of messages by reducing the number of reads. Specifically, the first index file for storing the index information of multiple messages includes multiple index slots and multiple index entries for storing the index information of different messages, and the multiple index entries in the first index file are written with index information in sequence. When at least two messages are mapped to the first index slot, the index information corresponding to each of the at least two messages will be stored in at least two index entries with discontinuous addresses, and the address information of the target index entry in the at least two index entries will be stored in the first index slot, where the target index entry refers to the index entry storing the index information of the message finally mapped to the first index slot. It can be seen that since the index information of multiple messages related to the same index slot is discretely stored in index entries with non-continuous addresses, separate reads of multiple index entries are involved during querying. The first index file is subjected to format conversion to store the index information of at least two messages mapped to the first index slot in at least two index entries with continuous addresses, and store the address range information corresponding to the at least two index entries in the first index slot to obtain the second index file after format conversion. The second index file is stored in an external distributed storage system, and the first index file on the local disk is deleted. In this way, in the second index file, the index information of multiple messages related to the same index slot is stored in multiple index entries with continuous addresses, and the index information in these multiple index entries can be read at one time during querying, reducing the number of reads and thus improving the query efficiency, so that the external distributed storage system has better random read and write performance. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] FIG. 1 is a schematic diagram of the format of an original index file provided by the embodiments of the present application;

[0012]

[0013] ​Figure 2 is a schematic diagram of the composition of a message system provided by an embodiment of the present application;

[0014] Figure 3 is a flowchart of a method for writing an original index file provided by an embodiment of the present application;

[0015] Figure 4 is a schematic diagram of the writing process of an original index file provided by an embodiment of the present application;

[0016] Figure 5 is a flowchart of a method for querying an original index file provided by an embodiment of the present application;

[0017] Figure 6 is a flowchart of a message index processing method provided by an embodiment of the present application;

[0018] Figure 7 is a schematic diagram of the format of a new index file provided by an embodiment of the present application;

[0019] Figure 8 is a schematic diagram of the index file format conversion process provided by an embodiment of the present application;

[0020] Figure 9 is a schematic diagram of an index file management method provided by an embodiment of the present application;

[0021] Figure 10 is a flowchart of a message index processing method provided by an embodiment of the present application;

[0022] Figure 11 is a schematic diagram of the hierarchical architecture of an index service subsystem provided by an embodiment of the present application;

[0023] Figure 12 is a schematic diagram of the structure of a message index processing device provided by an embodiment of the present application;

[0024] Figure 13 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data The management needs to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or reject.

[0027] The following will elaborate on some embodiments of the present application in conjunction with the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. Additionally, the step timings in the following method embodiments are only examples and not strictly limited.

[0028] First, the terms or concepts involved in the embodiments of the present application will be explained:

[0029] CommitLog: A typical write-ahead log file used to store the messages sent by message producers. After receiving new messages sent by message producers, they are appended to the log file in near real-time according to a specific format.

[0030] Compaction: An asynchronous file format conversion method that deletes redundant data in the file, compresses the data in the file to solve the read amplification problem, reduces the number of disk I / O reads, and thereby improves data query performance.

[0031] Object storage: A data storage service that stores data as an independent object. Object storage can use technologies such as asynchronous cold transfer to reduce storage costs. Compared with traditional storage media, it has lower storage costs and larger storage capacity, but weaker random read and write performance.

[0032] The current message index processing solution mainly generates the index information corresponding to a message after the message generated by the message producer is stored in the log file, and persists the index information corresponding to the message on the local disk. However, persisting the index information corresponding to each message on the local disk not only has a relatively high storage cost, but also the disk capacity is limited and cannot store a large amount of index information. Moreover, when querying a certain target message, it is necessary to randomly access the local disk multiple times to query the index information corresponding to the target message, resulting in the read amplification problem and making the entire query process take a long time.

[0033] To solve the above technical problems, an embodiment of the present application provides a message index processing method. This method combines local storage with an external distributed storage system to store an index file for storing multiple message indexes. The uncompleted or newly completed index file is stored in the local disk. After the index file is completed, it is format-converted and then stored in the external distributed storage system. This can not only reduce the storage cost but also expand the storage capacity. In addition, the address information of multiple consecutive index entries can be recorded simultaneously in one index slot of the format-converted index file. By reading once, the index information of multiple index entries can be obtained simultaneously, reducing the number of reads and improving the query efficiency, thus solving the read / write amplification problem.

[0034] The following introduces the message index processing solution provided by the embodiment of the present application.

[0035] First, introduce the format of a traditional index file, and first refer to this index file as the original index file.

[0036] As shown in FIG. 1, the original index file includes an index header (Index Header), index slots (Slots), and index entries (which can also be referred to as index items Index Items). Among them, the index header is mainly used to store the metadata information of the original index file, which specifically includes the following 5 fields: the magic number (Magic Code) for determining the starting position of the index file, the start time stamp (start Time Stamp), the end time stamp (end Time Stamp), the number of used index slots (hash Slot Count), and the number of stored indexes (index Count). The byte length occupied by each field is shown in FIG. 1. The magic number (Magic Code), start time stamp (start Time Stamp), end time stamp (end Time Stamp), number of used index slots (hash Slot Count), and number of stored indexes (index Count). The byte length occupied by each field is shown in FIG. 1.

[0037] Among them, the magic number used to determine the starting position of the index file is mainly used to identify that the file is an index file and can be regarded as a kind of identifier for the file type. The start timestamp and end timestamp refer to the earliest and latest storage times of the messages corresponding to the index information stored in the original index file. This storage time refers to the time when the message is stored in the log file. For example, if the original index file stores the index information of the messages stored in the log file from the message stored at time T1 to the message written at time T2, then the start timestamp in the index header of the original index file is T1, and the end timestamp is T2. The number of used index slots refers to the number of index slots that have been used among all the index slots included in the original index file. As the index slots are continuously used, this number will be updated dynamically. The number of stored indexes refers to the number of index entries that have been used in the original index file. As the index entries are continuously used, this number will be updated dynamically.

[0038] Among them, the index slot is used to store the address information of the index item corresponding to the index information of the message mapped to this index slot. The original index file may include a fixed number of multiple index slots, such as index slot 1, index slot 2, etc. shown in Figure 1. The information used for storage in the index slot can be simply referred to as the index position. The index position refers to the position information corresponding to the target index item of the message mapped to this index slot. The target index item refers to the index item where the index information of this message is stored. As shown in Figure 1, each index slot occupies 4 bytes.

[0039] Among them, the original index file includes multiple index entries, such as index entry 1, index entry 2, etc. shown in Figure 1. Each index entry is used to store the index information of the message. That is to say, what is stored in the index entry is not the actual specific message, but the index information of the message. This index information is used to indicate the storage position of this message in the log file. In practical applications, the storage position of the message in the log file can be located through each piece of information included in the index information stored in the index entry, so as to read the corresponding information.

[0040] As shown in Figure 1, each index entry in the original index file may include the following 7 fields: hash code, topic identifier (topicld), queue identifier (Queueld) field, physical position (Offset), message length (Size), time difference (timeDiff), index slot pointer (slotValue).

[0041] Among them, the hash value is obtained by performing a hash calculation on the identifier of the message. It is possible to determine whether it is the message to be queried by comparing the hash value stored in the index entry with the hash value corresponding to the information to be queried. The physical location refers to the sequential storage location of the message in the log file. The message length is used to describe the length of the message corresponding to the message in the log file. The time difference refers to the time difference value between the storage timestamp of the message in the log file and the start timestamp in the original index file. The index slot pointer is used to describe the address information of the previous index entry pointed to by the current index entry. Here, the previous index entry means that the previous index entry and the current index entry respectively correspond to messages mapped to the same index slot, and the index information in this index entry is later than the index information in the previous index entry. Then, the address information of the previous index entry can be obtained through the pointer information stored in the index information.

[0042] The message index processing method provided by the embodiments of the present application can be applied to the message system shown in Figure 2. As shown in Figure 2 shown, macroscopically, the message system includes a log service subsystem and an index service subsystem. The log service subsystem is used to receive messages sent by message producers and store the messages in a log file. The index service subsystem is used to create index information for the messages stored in the log file and store it in the original index file. Message producers can be, for example, application / content provider servers such as e-commerce servers and video servers.

[0043] Based on the format of the original index file shown in Figure 1 above, the process of generating an index file is introduced through the embodiment shown in Figure 3.

[0044] Figure 3 is a flowchart of a method for writing an original index file provided by the embodiments of the present application. As shown in Figure 3, the method may include the following steps:

[0045] 301. Write the currently received message into the log file.

[0046] 302. Determine the identifier information and index information of the message. The index information is used to indicate the storage location of the message in the log file.

[0047] 303. According to the identifier information and index information of the message, determine the second index slot and the second index entry corresponding to the message in the third index file.

[0048] 304. Store the third index file in the local disk.

[0049] In specific implementation, after the message service subsystem receives a message sent by a message producer, it can write the currently received message into a log file. The index service subsystem creates index information for the message and writes the index information into a third index file, where the third index file is an index file having the format of the above-mentioned original index file. In practical applications, the index service subsystem may not create index information for a message immediately after the message is written into the log file and can create it later.

[0050] In the process of writing the index information of the above message into the third index file, first, determine the identification information and index information of the message, and the index information is used to indicate the storage location of the message in the log file.

[0051] Among them, the identification of the message can be extracted from the message according to the type of message identification defined in advance. In this embodiment, the identification of the message is represented as key. In fact, the identification of a message is not necessarily unique, that is, different messages may have a certain same identification. Moreover, a message can have at least one identification. For example, taking an order in an e-commerce scenario as an example, an order message can include an order number and an order status, and these two pieces of information can be used as the two identifications of the order message. Thus, it can be understood that two order messages can correspond to the same order number as the identification.

[0052] Among them, the index information can include the hash value of the message, the topic identification queue identification corresponding to the message, the message length information, the time difference information, etc., as shown in the fields included in the index item shown in FIG. 1. Among them, the hash value of the message in an index item corresponds to an identification of the message. That is to say, if a message corresponds to n identifications, then n index items will be stored corresponding to it.

[0053] For example, message 1 corresponds to 3 key values, and its key values are A, B, and C respectively. Message 2 corresponds to 2 key values, and its key values are A and B respectively. Message 3 corresponds to 1 key value, and its key value is A. These messages correspond to 6 index items in the index file. Then, if the message identification of the message to be queried is B, the index information of message 1 and message 2 can be returned from the index file.

[0054] After determining the identification information and index information of the message, then, according to the identification information and index information of the message, determine the second index slot and the second index item corresponding to the message in the third index file.

[0055] According to the identification information and index information of the message, the specific implementation process of determining the second index slot and the second index entry corresponding to the message in the third index file may include: determining the second index slot to which the message is mapped in the third index file according to the hash calculation result of the identification information of the message; writing the index information of the message into the second index entry according to the sorting of multiple index entries in the third index file, writing the address information of the second index entry into the second index slot, and adding the address information of the previous index entry to the second index entry, where the previous index entry refers to the index entry into which the index information of the previous message mapped to the second index slot is written.

[0056] The following takes Figure 4 as an example to illustrate the process of writing the index information of the message into the third index file. For the convenience of description, only the case where one message has one identification information is exemplified. If one message has multiple message identifications, the same processing process is performed for each message identification respectively.

[0057] For example, assume that the message identification of message A is key1, and assume that the hash value obtained by performing a hash calculation on it matches index slot 1 in Figure 4. Here, the hash value obtained by the hash calculation indicates the identification (i.e., number) of the index slot. Assume that at this time, all index entries in the third index file have not been used yet, then the index information of message A is written into index entry 1. At this time, the address information of index entry 1 is written into index slot 1. In addition, it is determined that slotValue in index entry 1 is empty because no other message has been mapped to index slot 1 before.

[0058] After that, the index information of message B is stored. Assume that the message identification of message B is key2, and assume that the hash value obtained by performing a hash calculation on it matches index slot 2 in Figure 4. According to the sorting of each index entry in the third index file, the index information is written into the index entry in sequence. Then, at this time, the index information of message B is written into index entry 2. At this time, the address information of index entry 2 is written into index slot 2. In addition, it is determined that slotValue in index entry 2 is empty because no other message has been mapped to index slot 2 before.

[0059] After that, the index information of message C is stored. Assume that the message identifier of message C is also key3. After performing a hash calculation on it, the obtained hash value matches index slot 1. Then, the index information is written into the index entries in sequence according to the sorting of each index entry in the third index file. At this time, the index information of message C will be written into index entry 3. At this time, the address information of index entry 3 is written into index slot 1, that is, the address information of index entry 3 replaces the previously written address information of index entry 1. In addition, it is determined that the slotValue in index entry 3 is the address information of index entry 1 corresponding to the previous message A mapped to index slot 1. The arrows in the figure indicate the address pointer pointing relationship corresponding to slotValue.

[0060] It can be seen that each index entry in the above index file contains the address information of the previous index entry. Among them, the first index entry (such as index entry 3) and the previous index entry of the first index entry (such as index entry 1) map their respective corresponding messages to the same index slot (such as index slot 1), and the index information in the first index entry (such as the index information of message C) is stored in the index file later than the index information in the previous index entry (such as the index information of message A).

[0061] According to the above index information writing process, the index information of multiple messages is written into the third index file. The above writing process can be executed in memory. During the writing process of the third index file, the already written information can be batched and written to the local disk. After the third index file is full, the next index file will be generated to continue writing the index information.

[0062] Next, in combination with the embodiment shown in FIG. 5, the query process of the above third index file is introduced.

[0063] FIG. 5 is a flowchart of a query method for an original index file provided by an embodiment of the present application. As shown in FIG. 5, the method may include the following steps:

[0064] 501. Determine the fourth index slot mapped by the identification information of the message to be queried in the third index file.

[0065] 502. According to the address information of the third index entry stored in the fourth index slot, read the index information in the third index entry.

[0066] 503. According to the address information of the previous fourth index entry stored in the third index entry, read the index information in the fourth index entry.

[0067] 504. If the address information of the previous index entry in the fourth index entry is empty, determine the index information of the message to be queried from the index information in the third index entry and the index information in the fourth index entry.

[0068] The process of querying index information is still illustrated with the example in FIG. 4. Suppose the message to be queried is message C, and its identification information is key3. In practical applications, the user will carry the identification information of the message to be queried in the triggered query request. The hash value obtained by performing a hash calculation on key3 matches index slot 1, and the address information of the index entry stored therein is read from index slot 1: the read address information is that of index entry 3, and thus the index information stored in index entry 3 is read. Since the slotValue in index entry 3 is the address information of index entry 1, the index information stored in index entry 1 is continued to be read. Since the slotValue in index entry 1 is empty, the continued reading is stopped. At this time, determine which of the index information stored in index entry 1 and the index information stored in index entry 3 is the index information of the message to be queried. Specifically, since both of these index information contain a hash code field, determine the index information whose hash value stored in the hash code field is the same as the hash value of key3 as the index information of the message to be queried. Since the hash value of key3 is stored in index entry 3, and the hash value of keyl is stored in index entry 1, it is determined that the index information stored in index entry 3 is the index information of the message to be queried.

[0069] The above introduces the process of writing and querying index information in the format of the original index file. It can be seen from the above embodiments that in the format of the original index file, for example, the process of querying the index information of the above message C involves multiple reads of index entries (index entry 3, index entry 1), and the query efficiency is low. Moreover, the generated index files are all stored on the local disk, and the storage capacity of the local disk is limited, while the cost of increasing the local disk is relatively large. Based on this, the embodiments of the present application provide a new method for processing message indexes.

[0070] FIG. 6 is a flowchart of a method for processing message indexes provided by an embodiment of the present application. As shown in FIG. 6, the method for processing message indexes may include the following steps:

[0071] 601. Obtain a first index file for storing index information of multiple messages from the local disk. The first index file The piece has been written. The first index file includes multiple index slots and multiple index entries for storing index information of different messages. Among them, the index information corresponding to at least two messages mapped to the first index slot is stored in at least two index entries with discontinuous addresses. The address information of the target index entry in at least two index entries is stored in the first index slot. The target index entry refers to the index entry storing the index information of the message finally mapped to the first index slot.

[0072] 602. Perform format conversion on the first index file to obtain a second index file. Among them, in the second index file, the index information of at least two messages is stored in at least two index entries with continuous addresses, and the address range information corresponding to at least two index entries is stored in the first index slot.

[0073] 603. Store the second index file in an external distributed storage system and delete the first index file.

[0074] In this embodiment, the first index file adopts the format of the original index file described above, and the first index file is an index file that has been written and stored on the local disk.

[0075] In order to improve the storage capacity of message indexes and reduce the storage cost of message indexes, in this embodiment, the index files stored on the local disk can be processed by separating hot and cold. The index files that have not been written or have just been written are stored on the local disk. After that, the index files that have been written can be format-converted at a certain time, and the format-converted index files are stored in an external distributed storage system. Among them, the external distributed storage system can be a distributed storage system such as object storage or database files. The type of the external distributed storage system is not limited in the embodiments of the present application. Among them, the index file being written or the index file that has just been written not long ago is a "hot index file", and the index file that has been written for a long time is called a "cold index file". The separation of hot and cold means that the hot index file is stored on the local disk, while the cold index file is stored in the distributed storage system.

[0076] In practical applications, since the external distributed storage system not only often has a lower storage cost than the local disk, but also has a larger storage capacity than the local disk and can be elastically scaled, it is possible to achieve low-cost storage of a large number of index files by means of the distributed storage system.

[0077] However, when using an external distributed storage system to store index files, the query time will increase significantly compared to local disks. Because when reading from local disks, even for the same 10 times, the time consumption will be shorter. However, the distributed storage system involves multiple networks, and the network time consumption will be more. Therefore, in order to improve the query efficiency when querying message index information from the distributed storage system, it is necessary to convert the format of the index files stored in the distributed storage system.

[0078] In practical applications, a timed task can be set to periodically transfer the index files that have been written to the local disk to an external distributed storage system.

[0079] Take the first index file in the format corresponding to the original index file that has been written to the local disk as an example. First, obtain the first index file that has been written and stores the index information of multiple messages from the local disk, and then perform format conversion on it to obtain the second index file in the new format, and store the second index file in the distributed storage system.

[0080] Among them, the first index file includes multiple index slots and multiple index entries for storing the index information of different messages. Among them, the index information corresponding to at least two messages mapped to the first index slot is stored in at least two index entries with discontinuous addresses. The address information of the target index entry among the at least two index entries is stored in the first index slot. The target index entry refers to the index entry that stores the index information of the message finally mapped to the first index slot. The format and writing process of the first index file can be understood by referring to the foregoing embodiments. Among them, for example, for message A and message C mentioned above, although they are mapped to the same index slot 1, under the characteristic that the index entries are written sequentially, the index information of the two messages is respectively stored in two discontinuous index entries, index entry 1 and index entry 3. When querying, it is necessary to perform chained reading of index entries based on the slotValue field. In the second index file after format conversion, the index information of at least two messages mapped to the first index slot is stored in at least two index entries with continuous addresses, and the address range information corresponding to the at least two index entries is stored in the first index slot. In addition, each index entry in the second index file does not contain the address information of the previous index entry, that is, there is no longer the slotValue field.

[0081]

[0082] ​Optionally, the address range information corresponding to the at least two index entries stored in the first index slot may be: the start address information of the at least two index entries and the number of index entries.

[0083] Specifically, the implementation process of converting the first index file into the second index file may include: writing the index information corresponding to at least two messages mapped to the same index slot in the first index file into at least two consecutively addressed index entries, deleting the address information of the previous index entry included in each index entry, and writing the address range information corresponding to the at least two index entries into this index slot.

[0084] The format of the second index file and the conversion process are schematically shown below with reference to FIGS. 7 and 8.

[0085] As shown in FIG. 7, compared with the format of the first index file shown in FIG. 1, the index header of the second index file remains unchanged, and the number of index slots also remains unchanged. Only the byte length occupied by each index slot is doubled (The index slot occupies 4 bytes in the first index file and 8 bytes in the second index file), and the number of index entries also remains unchanged. Only the field slotValue that occupies 4 bytes no longer exists in each index entry.

[0086] Among them, in the format of the second index file, the index slot may include the following two fields: index position and index total size, each occupying 4 bytes. The index position represents the start address of multiple index entries corresponding to the same index slot, and the index total size represents the total number of these multiple index entries.

[0087] Since in an index file, the number of index slots is much smaller than the number of index entries, the data volume of an index file can also be reduced through the above conversion of the index file format.

[0088] In Figure 8, the storage results of the index information of Message A, Message B, and Message C in the first index file are schematically shown. The writing process can refer to the foregoing embodiments: The message identifier of Message A is key1, which is mapped to index slot 1. The index information of Message A is written into index entry 1, and the slotValue in index entry 1 is empty; the message identifier of Message B is key2, which is mapped to index slot 2. The index information of Message B is written into index entry 2. The slotValue in index entry 2 is empty; the message identifier of Message C is key3, which is mapped to index slot 1. The index information of Message C is written into index entry 3, and the slotValue in index entry 3 is the address information of index entry 1; the address information currently stored in index slot 1 is the address information of index entry 3, and the address information currently stored in index slot 2 is the address information of index entry 2. The conversion process is as follows: Traverse each index slot in the first index file in sequence. For the currently traversed index slot 1, read the index information of Message C stored in index entry 3 according to the address information of index entry 3 stored therein, and then, according to the slotValue in index entry 3, read the index information of Message A stored in index entry 1. Since the slotValue in index entry 1 is empty, the reading stops. After that, store the index information of Message C and the index information of Message A read from the first index file into two consecutive index entries in the second index file. Assuming that all index entries in the second index file are empty at this time, then in order, store them into index entry 1 and index entry 2 of the second index file. And, write the address range information corresponding to index entry 1 and index entry 2 in the same index slot 1 of the second index file. For example: In index slot 1 of the second index file, set the index position to the address information of index entry 1, and set the index quantity to 2, which means that starting from index entry 1, there are a total of 2 consecutive index entries (i.e., index entry 1 and index entry 2) corresponding to index slot 1.

[0089] Similarly, the index information of Message B is stored in index entry 3 of the second index file. In index slot 2 of the second index file, set the index position to the address information of index entry 3, and set the index quantity to 1.

[0090]

[0091] ​Traverse each index slot in the first index file one by one based on the above conversion process, and perform the above conversion process to obtain the second index file after format conversion corresponding to the first index file. Then store the second index file in the distributed storage system and delete the first index file on the local disk.

[0092] The external distributed storage system has relatively poor random read and write performance compared with the local disk. A single read or write operation not only consumes more time but also has the problem of read and write amplification. Therefore, in the embodiments of the present application, in order to improve the read and write performance of the external distributed storage system, before storing the already written first index file in the external distributed storage system, the first index file can be subjected to format conversion processing so that the address information of the index items that can write the index information of multiple messages simultaneously in one index slot in the converted second index file can be obtained through a single read operation. Moreover, since the index information of multiple messages corresponding to the same index slot in the second index file is continuously written into consecutive index items, write merging can be performed at this time to reduce the number of writes.

[0093] In summary, by introducing a distributed storage system and based on the elastic scaling characteristics of the distributed storage system, massive storage of index files can be achieved with lower storage costs. In addition, in order to improve the query efficiency of the distributed storage system, the index file is subjected to the above format conversion processing so that the index items corresponding to the same index slot are address-continuous, so as to facilitate reading the index information in multiple index items at one time, solve the problem of read amplification (i.e., multiple read times), and improve the query efficiency.

[0094] As described above, in practical applications, a part of the index files (hot index files) are stored in the local disk in the format of the original index file, and another part of the index files (cold index files) are stored in the distributed storage system in the format of the new index file.

[0095] To facilitate the management of the numerous generated index files, the embodiments of the present application provide a method for managing index files in the form of a file management queue: generate a file management queue, and the file management queue contains the description information of multiple generated index files. The description information includes the time range and storage space corresponding to the respective index files. Among them, the time range indicates the time range when the multiple messages corresponding to the respective index files are written into the log file, and the storage space includes the local disk and the external distributed storage system.

[0096] ​It can be understood that in the above-mentioned "queue", the description information is written into the queue in sequence according to the order of the index file generation time (start time).

[0097] Based on the introduction of the foregoing embodiments, an index file may have the following different states: when it is first generated, it is in the state of "writing to file", that is, the state of writing the index information of the message into the index file, denoted as writing; after the writing is completed (i.e., full), the state of completing the format conversion, denoted as compact; the state of storing the index file with the format conversion completed into the distributed storage system, denoted as 7 T 7 OSS (i.e., object storage).

[0098] Therefore, in the file management queue, for each index file that has been stored in the queue, its corresponding description information can be adaptively modified based on the change of the index file state.

[0099] Among them, since the index files in the writing and compact states are stored on the local disk, and the index files in the OSS state are stored in the distributed storage system, the storage space of the index file can be determined based on the file state.

[0100] Illustrated with reference to Figure 9. Suppose the sequentially generated index files include index file 1 - index file 6. Among them, index files 1 - 3 have completed the format conversion and are stored in the distributed storage system, that is, in the OSS state; index files 4 - 5 have completed the format conversion and have not been stored in the distributed storage system yet, that is, in the compact state; index file 6 is in the writing state.

[0101] In addition, as shown in Figure 9, each index file is associated with corresponding time range information. Optionally, the time range information can be stored in a form after a certain encoding calculation. Denote the time range information as the key of the index file, then the corresponding description information of the index file is the value.

[0102] To achieve the feature of Non-Stop Write and improve the writing performance of index information, as shown in Figure 9, three different threads can be set in the index service subsystem to cooperate with each other: the index writing thread, the index query thread, and the timed task thread. The three threads are each responsible for different tasks, and the read-write lock is used to ensure the correctness under concurrent conditions.

[0103] Among them, the index writing thread is non-blocking, and its responsibility is to write index information into the index file in the writing state at the end of the queue. When an index file is full, this thread will automatically create a new index file at the end of the queue and switch to the next index file for writing. To improve the writing efficiency, this thread can also cache the index information in memory and write it in batches to the index file when the cache reaches a certain amount, so as to reduce the number of disk I / Os.

[0104] Among them, the index query thread can query index files in different states. The specific query strategy is as follows: for the index file in the writing state, this thread needs to wait for the index writing thread to finish writing the index information before it can query; for other index files that are already full, according to the storage space of the corresponding index file, query from the local disk or the distributed storage system.

[0105] Among them, the timed task thread is mainly responsible for performing format conversion operations on the files that are in the writing state and have been full. When performing the format conversion operation, this thread needs to first obtain the read-write lock of the corresponding index file to avoid concurrent access to the index file by other threads. After the format conversion is completed, the state of the index file is switched to the Compact state, and then the formatted index file needs to be uploaded to the distributed storage system. After the upload is completed, the state of the index file is switched to the uploaded state - OSS. During the upload process, this thread needs to release the read-write lock of the index file.

[0106] The following describes the index query process based on the above file management queue in conjunction with the embodiment shown in FIG. 10.

[0107] FIG. 10 is a flowchart of a message index processing method provided by an embodiment of the present application. As shown in FIG. 10, the method may include the following steps:

[0108] 1001. Receive a message query request, where the message query request includes the identification information of the message to be queried and the time information to be queried.

[0109] 1002. Determine the target index file to be queried according to the time information to be queried and the time ranges corresponding to the multiple index files.

[0110] 1003. Determine the target index file in the storage space of the target index file.

[0111] 1004. Obtain the index information of the message to be queried from the target index file according to the identification information of the message to be queried.

[0112] 1005. Query the message to be queried from the log file according to the index information of the message to be queried.

[0113] After receiving a message query request, the target index file to be queried can be determined according to the time information to be queried in the message query request and the time ranges respectively corresponding to multiple index files. Then, determine the target index file in the storage space of the target index file. Then, according to the identification information of the message to be queried in the message query request, obtain the index information of the message to be queried from the target index file. Finally, query the message to be queried from the log file according to the index information of the message to be queried.

[0114] Among them, as described above, the time ranges respectively corresponding to multiple index files in the file management queue can be represented by encoded key values. After encoding the time to be queried in the same way, based on the key values respectively corresponding to multiple index files and the key value corresponding to the time information to be queried in the query request, locate the target index file corresponding to the message to be queried.

[0115] In an optional embodiment, the specific implementation process of obtaining the index information of the message to be queried from the target index file according to the identification information of the message to be queried may include: if the target index file is a second index file stored in an external distributed storage system, determine the third index slot mapped by the identification information of the message to be queried in the second index file; according to the address range information corresponding to at least two index entries stored in the third index slot, read the corresponding at least two index information from the at least two index entries; determine the index information of the message to be queried from the at least two index information read.

[0116] Taking the situation of the embodiment shown in FIG. 8 above as an example for illustration. Suppose the identification information of the message to be queried is key3, perform a hash calculation on it, and the obtained hash value is mapped to index slot 1 in the second index file. Then, according to the address range information corresponding to index entry 1 and index entry 2 stored in index slot 1, read the index information from index entry 1 and index entry 2 at one time. Then, according to the respective hash code contents in index entry 1 and index entry 2, it is determined that the hash value corresponding to message C in index entry 2 is the same as the hash value of key3. Therefore, it is determined that the index information stored in index entry 2 is the index information of the message to be queried. Thus, it can be seen that multiple index entries corresponding to one index slot can be read at one time, reducing the number of reads.

[0117] If the target index file is the third index file stored on the local disk, determine the fourth index slot to which the identification information of the message to be queried is mapped in the third index file; according to the address information of the third index entry stored in the fourth index slot, read the index information in the third index entry; according to the address information of the previous fourth index entry stored in the third index entry, read the index information in the fourth index entry; if the address information of the previous index entry in the fourth index entry is empty, determine the index information of the message to be queried from the index information in the third index entry and the index information in the fourth index entry. The query process refers to the embodiment shown in FIG. 5 above and will not be elaborated here.

[0118] The above embodiments introduce the working process of the index service subsystem in the message system. In actual implementation, in order to improve the scalability of the subsystem and facilitate writing unit tests, the entire index service subsystem can adopt the idea of hierarchical design. As shown in FIG. 11, from top to bottom, an index service layer, an index file parsing layer, and a data storage layer are designed respectively. Different layers are responsible for handling different tasks, and the layers are decoupled from each other. The upper layer only depends on the services provided by the lower layer.

[0119] Among them, the index service layer provides a message index service for the log subsystem and is the entry for writing and reading the index information of messages. Specifically, when the index service layer obtains the message written to the log file, the index service layer generates the index information of the message and sends the generated index information to the index file parsing layer so that the index file parsing layer can write the generated index information into the corresponding index file. In addition, the index service layer can also be responsible for querying the message index. For example, in actual applications, when a user needs to query a specific message, a message query request can be sent to the index service layer. After receiving the message query request, the index service layer sends the message query request to the index file parsing layer so that the index file parsing layer can determine the target index file to be queried and the storage location of the target index file according to the time information to be queried in the message query request, and obtain the index information of the message to be queried from the target index file stored in the data storage layer according to the identification information of the message to be queried in the message query request. At the same time, the index service layer is also responsible for the life cycle management of the index file, such as creating an index file, performing format conversion processing on the index file that has been written, uploading the index file, destroying the index file, and other operations.

[0120] Among them, the index file parsing layer mainly parses the formats of individual index files in different states. Specifically, the index file parsing layer can read the target index file from the data storage layer and parse the target index file into a readable format for upper-layer calls. In addition, the index file parsing layer can also provide the KV query service for a single index file. For example, after receiving the message query request sent by the index service layer, the index file parsing layer can determine the target index file to be queried and the storage location of the target index file according to the time information to be queried in the message query request and the timestamp KEY values corresponding to multiple index files respectively, read the target index file from the data storage layer, and obtain the index information of the message to be queried from the target index file stored in the data storage layer according to the identification information of the message to be queried in the message query request.

[0121] Among them, the data storage layer is mainly used for storing the index information of messages and supports different types of storage methods, including object storage, local disk files, or database files, etc. Specifically, the data storage layer can store the index information of messages in the corresponding index files and store the index files in the local disk, object storage, or database files. When reading the index information of a certain message, the data storage layer is mainly used to obtain the index information of the message from the local disk, object storage, or database file and convert it into binary stream data and return it to the caller.

[0122] Based on this, the index service module divides the entire index service into three different layers by adopting the idea of hierarchical design, making the system have good scalability and maintainability, which is convenient for subsequent upgrades and maintenance. At the same time, the decoupling between layers is clear in responsibilities, which is convenient for unit testing and maintenance.

[0123] In addition, since the file management queue is stored in the memory, when a system crash occurs, the file management queue stored in the memory will disappear. In order to continue to use the file management queue to manage each index file, it is necessary to regenerate the file management queue after the system crash and restart. Specifically, in response to the restart operation after the system crash, obtain the index files that have been written from the distributed storage system and the local disk, and obtain the index files that have not been written from the local disk; regenerate the information contained in some index entries and index slots in the index files that have not been written, where some index entries and index slots refer to the index entries and index slots that were not stored in the local disk during the system crash; regenerate the file management queue according to the index files that have been written and the regenerated index files that have not been written.

[0124] The message index processing device of one or more embodiments of the present application will be described in detail below. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.

[0125] FIG. 12 is a schematic structural diagram of a message index processing device provided by an embodiment of the present application. As shown in FIG. 12, the device includes: an acquisition module 11, a conversion module 12, and a storage module 13.

[0126] The acquisition module 11 is used to obtain a first index file for storing index information of multiple messages from a local disk. The first index file has been written, and the first index file includes multiple index slots and multiple index entries for storing index information of different messages. Among them, the index information corresponding to at least two messages mapped to the first index slot is stored in at least two index entries with discontinuous addresses. The address information of the target index entry among the at least two index entries is stored in the first index slot. The target index entry refers to the index entry storing the index information of the message finally mapped to the first index slot.

[0127] The conversion module 12 is used to perform format conversion on the first index file to obtain a second index file. Among them, in the second index file, the index information of the at least two messages is stored in at least two index entries with continuous addresses, and the address range information corresponding to the at least two index entries is stored in the first index slot.

[0128] The storage module 13 is used to store the second index file in an external distributed storage system and delete the first index file.

[0129] Optionally, the address information of the previous index entry is included in each index entry in the first index file. Among them, the messages corresponding to the first index entry and the previous index entry of the first index entry are mapped to the same index slot, and the index information in the first index entry is stored in the first index file later than the index information in the previous index entry. The address information of the previous index entry is not included in each index entry in the second index file. Optionally, the address range information corresponding to the at least two index entries includes: the start address information of the at least two index entries and the number of index entries.

[0130] Optionally, the address range information corresponding to the at least two index entries includes: the start address information of the at least two index entries and the number of index entries.

[0131] Optionally, the obtaining module 11 is further configured to: write the currently received message into a log file; determine the identification information and index information of the message, where the index information is used to indicate the storage location of the message in the log file; determine, according to the identification information and index information of the message, a second index slot and a second index entry corresponding to the message in a third index file; and store the third index file in a local disk.

[0132] Optionally, the obtaining module 11 is further configured to: determine, according to the hash calculation result of the identification information of the message, a second index slot to which the message is mapped in a third index file; write the index information of the message into a second index entry according to the sorting of multiple index entries in the third index file, write the address information of the second index entry into the second index slot, and add the address information of the previous index entry to the second index entry, where the previous index entry refers to the index entry into which the index information of the previous message mapped to the second index slot is written.

[0133] Optionally, the obtaining module 11 is further configured to: generate a file management queue, where the file management queue contains description information of multiple generated index files, and the description information includes a time range and a storage space corresponding to the respective index files; where the time range indicates the time range in which multiple messages corresponding to the respective index files are written into the log file, and the storage space includes a local disk and an external distributed storage system.

[0134] Optionally, the obtaining module 11 is further configured to: receive a message query request, where the message query request includes the identification information of the message to be queried and the time information to be queried; determine a target index file to be queried according to the time information to be queried and the time ranges corresponding to the multiple index files; determine the target index file in the storage space of the target index file; obtain the index information of the message to be queried from the target index file according to the identification information of the message to be queried; and query the message to be queried from the log file according to the index information of the message to be queried.

[0135] Optionally, the obtaining module 11 is further configured to: if the target index file is the second index file stored in the external distributed storage system, determine a third index slot to which the identification information of the message to be queried is mapped in the second index file; read at least two corresponding index information from at least two index entries according to the address range information stored in the at least two index entries in the third index slot; and determine the index information of the message to be queried from the at least two read index information.

[0136] Optionally, the obtaining module 11 is further configured to: if the target index file is the third index file stored in the local disk, determine a fourth index slot mapped by the identification information of the message to be queried in the third index file; according to the address information of the third index entry stored in the fourth index slot, read the index information in the third index entry; according to the address information of the previous fourth index entry stored in the third index entry, read the index information in the fourth index entry; if the address information of the previous index entry in the fourth index entry is empty, then from the third index information in the index entry and the index information in the fourth index entry, determine the index information of the message to be queried.

[0137] Optionally, the obtaining module 11 is further configured to: in response to a restart operation after a system crash, obtain the index files that have been written from the distributed storage system and the local disk, and obtain the index files that have not been written from the local disk; regenerate the information included in some index entries and index slots in the index files that have not been written, where the some index entries and index slots refer to the index entries and index slots that were not stored in the local disk at the time of the system crash; according to the index files that have been written and the regenerated index files that have not been written, regenerate the file management queue.

[0138] The device shown in FIG. 12 may execute the steps provided in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, and details are not described herein again.

[0139] In a possible design, the structure of the message index processing device shown in FIG. 12 above may be implemented as an electronic device. As shown in FIG. 13, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. Among them, an executable code is stored on the memory 22. When the executable code is executed by the processor 21, the processor 21 can at least implement the message index processing method provided in the foregoing embodiments.

[0140] In addition, an embodiment of the present application provides a non-transitory machine-readable storage medium. An executable code is stored on the non-transitory machine-readable storage medium. When the executable code is executed by a processor of an electronic device, the processor can at least implement the message index processing method provided in the foregoing embodiments.

[0141] The device embodiments described above are merely illustrative. The network elements described as separate components may or may not be physically separated. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform. Of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the prior art can be embodied in the form of a computer product. This application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, not to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

Claims 1. A message index processing method, wherein, Including: Obtain a first index file for storing index information of multiple messages from a local disk. The first index file has been written. The first index file includes multiple index slots and multiple index entries for storing index information of different messages. Among them, the index information corresponding to at least two messages mapped to the first index slot is stored in at least two index entries with discontinuous addresses. The address information of the target index entry among the at least two index entries is stored in the first index slot. The target index entry refers to the index entry storing the index information of the message finally mapped to the first index slot. Perform format conversion on the first index file to obtain a second index file. Among them, in the second index file, the index information of the at least two messages is stored in at least two index entries with continuous addresses. The address range information corresponding to the at least two index entries is stored in the first index slot. Store the second index file in an external distributed storage system and delete the first index file.

2. The method according to claim 1, wherein Each index entry in the first index file contains the address information of the previous index entry. Among them, the messages corresponding to the first index entry and the previous index entry of the first index entry are mapped to the same index slot, and the index information in the first index entry is stored in the first index file later than the index information in the previous index entry. Each index entry in the second index file does not contain the address information of the previous index entry.

3. The method according to claim 1, wherein, The address range information corresponding to the at least two index entries includes: the start address information of the at least two index entries and the number of index entries.

4. The method according to any one of claims 1-3, wherein The method further includes: writing the currently received message into a log file; determining the identification information and index information of the message, where the index information is used to indicate the storage location of the message in the log file; determining the second index slot and the second index entry corresponding to the message in a third index file according to the identification information and index information of the message; storing the third index file in the local disk.

5. The method according to claim 4, wherein The determining the second index slot and the second index entry corresponding to the message in the third index file according to the identification information and index information of the message includes: determining the second index slot to which the message is mapped in the third index file according to the hash calculation result of the identification information of the message; writing the index information of the message into the second index entry according to the sorting of multiple index entries in the third index file, writing the address information of the second index entry in the second index slot, and adding the address information of the previous index entry in the second index entry. The previous index entry refers to the index entry where the index information of the previous message mapped to the second index slot is written.

6. The method according to claim 4, wherein The method further includes: generating a file management queue, where the file management queue contains the description information of multiple generated index files. The described description information includes the time range and storage space corresponding to the respective index files; wherein, the time range indicates the time range in which a plurality of messages corresponding to the respective index files are written into the log file, and the storage space includes a local disk and an external distributed storage system.

7. The method according to claim 6, wherein The method further includes: receiving a message query request, where the message query request includes identification information of the message to be queried and time information to be queried; determining a target index file to be queried according to the time information to be queried and the time ranges corresponding to the respective plurality of index files; determining the target index file in the storage space of the target index file; obtaining index information of the message to be queried from the target index file according to the identification information of the message to be queried; querying the message to be queried from the log file according to the index information of the message to be queried.

8. The method according to claim 7, wherein The obtaining index information of the message to be queried from the target index file according to the identification information of the message to be queried includes: if the target index file is the second index file stored in the external distributed storage system, determining a third index slot mapped by the identification information of the message to be queried in the second index file; reading at least two corresponding index information from at least two index entries according to the address range information corresponding to the at least two index entries stored in the third index slot; determining the index information of the message to be queried from the at least two index information read.

9. The method according to claim 7, wherein The obtaining index information of the message to be queried from the target index file according to the identification information of the message to be queried includes: if the target index file is the third index file stored in the local disk, determining a fourth index slot mapped by the identification information of the message to be queried in the third index file; reading the index information in the third index entry according to the address information of the third index entry stored in the fourth index slot; reading the index information in the fourth index entry according to the address information of the previous fourth index entry stored in the third index entry; if the address information of the previous index entry in the fourth index entry is empty, determining the index information of the message to be queried from the index information in the third index entry and the index information in the fourth index entry.

10. The method according to claim 6, wherein, The method further includes: in response to a restart operation after a system crash, obtaining the index files that have been written from the distributed storage system and the local disk, and obtaining the index files that have not been written from the local disk; regenerating the information included in some index entries and index slots in the index files that have not been written, where the some index entries and index slots refer to the index entries and index slots that were not stored in the local disk at the time of the system crash; regenerating the file management queue according to the index files that have been written and the regenerated index files that have not been written.

11. An electronic device, wherein, Includes: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor is caused to execute the message indexing processing method according to any one of claims 1 to 10.

12. A non-transitory machine-readable storage medium, wherein, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by the processor of the electronic device, the processor is caused to execute the message indexing processing method according to any one of claims 1 to 10.

13. A computer program product, wherein, Comprising: A computer program, which when executed by the processor of the electronic device, causes the processor to execute the message indexing processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Data operation method and system

    CN116028677A

  • Index construction method and device

    CN116414830A

  • Index creation method and device, data query method and device, equipment and storage medium

    CN117235069A

Cited By

  • Multi-modal intelligent identification-based cross-process deadlock detection method for swan gap platform

    CN121187814A