An index processing method and related device based on full-text retrieval
By adding document category identifiers to build index keywords and values in full-text search, the problem of document category index interleaving storage is solved, and a more fine-grained index sub-table function is realized, which improves retrieval efficiency.
Patent Information
- Application Number
- CN202110915363.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-08-10
AI Technical Summary
In the existing full-text search technology, the index information of document categories is stored interlaced, resulting in inefficient retrieval.
By adding document category identification based on the document's words to be updated and document identification identification, the index keywords and index values are constructed, index information in the key-value pair format is formed, and persisting to disk in the order of index keywords when the preset conditions are met, fine-grained document category storage is achieved.
It improves the efficiency of full-text retrieval, ensures that the index information of the same document category is stored sequentially, avoids interleaving, and achieves more convenient and fast retrieval processing.
Smart Images

Figure CN115705353B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of retrieval technology, and in particular to an index processing method and related devices based on full-text retrieval. Background Art
[0002] As full-text search is widely used in more and more business scenarios, the requirements for full-text search are also getting higher and higher. Full-text search refers to building an index for documents and storing them so that the required search results can be obtained through the index during subsequent searches.
[0003] In the related art, when building an index for a document, the index is only built based on the word segmentation and identification of the document, and the index is stored for subsequent retrieval. However, the index built in this way is not detailed enough in granularity, resulting in the interleaving storage of indexes corresponding to different document categories in business scenarios, which greatly reduces the retrieval efficiency. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides an index processing method and related devices based on full-text retrieval, so that the index information corresponding to the same document category is stored sequentially, avoiding the interleaved storage of index information corresponding to different document categories, and realizing a more fine-grained document category-based index sub-table function in full-text retrieval, thereby making the retrieval processing more convenient and faster during full-text retrieval, and greatly improving the retrieval efficiency.
[0005] The embodiments of the present application disclose the following technical solutions:
[0006] On the one hand, the present application provides an index processing method based on full-text retrieval, the method comprising:
[0007] Obtaining the to-be-updated words of the document to be written, the document identification mark and the document category mark to which the document to be written belongs; the document category mark is the mark of the document category set based on the business classification requirement, and the document category mark is used to construct the index keyword;
[0008] Based on the word to be updated, the document identification mark and the document category mark, construct the index keyword and the index value corresponding to the index keyword in the memory to obtain index information, wherein the index information is stored in a key-value pair format;
[0009] When a preset persistence trigger condition is met, the index information stored in the memory is persisted as a first index file and stored in a disk based on the numerical order of the index keywords.
[0010] On the other hand, the present application provides an index processing device based on full-text retrieval, the device comprising: an acquisition unit, a construction unit and a persistence unit;
[0011] The obtaining unit is configured to obtain the words to be updated in the document to be written, the document identification identifier, and the document category identifier to which the document to be written belongs; the document category identifier is an identifier of a document category set based on business classification requirements, and the document category identifier is used to construct an index keyword;
[0012] The constructing unit is configured to construct index information including the index keyword and the index value corresponding to the index keyword in a memory based on the words to be updated, the document identification identifier, and the document category identifier, and the index information is stored in a key-value pair format;
[0013] The persisting unit is configured to, when a preset persisting trigger condition is satisfied, persist the index information stored in the memory into a first index file and store it on a disk based on the numerical order of the index keywords.
[0014] On the other hand, the present application provides an index processing device for full-text retrieval, and the device includes a processor and a memory:
[0015] The memory is configured to store program code and transmit the program code to the processor;
[0016] The processor is configured to execute the method described in the above aspect according to the instructions in the program code.
[0017] On the other hand, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium is configured to store a computer program, and the computer program is configured to execute the method described in the above aspect.
[0018] As can be seen from the above technical solutions, different document category identifiers are set in advance according to the requirements of business classification. For the document to be written, based on the words to be updated in the document to be written and the document identification identifier, the document category identifier to which the document to be written belongs is added, and the index keyword and the index value corresponding to the index keyword are jointly constructed in the memory to obtain the index information in the key-value pair format. The document category identifier is used to construct the index keyword. When the preset persistence trigger condition is met, the index information stored in the memory is persisted in the numerical order of the index keyword, and the first index file is stored on the disk. Since the index keyword in the index information is constructed by the document category identifier, the constructed index information can represent the document category in a finer granularity, and thus can be sorted in order according to the document category in the first index file on the disk after the persistence trigger. Based on this, in the case of setting different document categories according to the requirements of business classification, the index information corresponding to the same document category is stored sequentially to avoid the interlaced storage of the index information corresponding to different document categories, and a finer granularity index sub-table function based on the document category is realized in the full-text retrieval, so that the retrieval process is more convenient and fast during the full-text retrieval, and the retrieval efficiency is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 FIG. is a schematic diagram of an application scenario of an index processing method based on full-text retrieval provided by an embodiment of the present application;
[0021] Figure 2 FIG. is a schematic flowchart of an index processing method based on full-text retrieval provided by an embodiment of the present application;
[0022] Figure 3 FIG. is a schematic diagram of the format of a forward index keyword provided by an embodiment of the present application;
[0023] Figure 4 FIG. is a schematic diagram of the format of a reverse index keyword provided by an embodiment of the present application;
[0024] Figure 5 FIG. is a schematic diagram of the format of an inverted list provided by an embodiment of the present application;
[0025] Figure 6 FIG. is a schematic diagram of the format of forward index information and reverse index information provided by an embodiment of the present application;
[0026] Figure 7 Schematic diagram of a file layer in a disk provided by an embodiment of the present application;
[0027] Figure 8 Schematic flowchart of a full-text retrieval method provided by an embodiment of the present application;
[0028] Figure 9 Schematic diagram of merging inverted lists of identification labels of multiple target documents provided by an embodiment of the present application;
[0029] Figure 10 Schematic diagram of an overall framework based on full-text retrieval provided by an embodiment of the present application;
[0030] Figure 11 Schematic diagram of an index processing device based on full-text retrieval provided by an embodiment of the present application;
[0031] Figure 12 Schematic diagram of the structure of a server provided by an embodiment of the present application;
[0032] Figure 13 Schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0033] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0034] In the related art, in the case of full-text retrieval, when constructing an index for a document, only the word segmentation and identification label of the document are used to construct the index. For example, the word segmentation and identification labels of documents 1, 2, 3, and 4 are respectively used to construct the index, and the constructed index is stored for subsequent retrieval. However, the index constructed in this way is not fine-grained enough. For example, when the documents need to be classified according to different users, in the case where document 1 and document 3 belong to user A, and document 2 and document 4 belong to user B, the index constructed in this way cannot represent user A or user B in a fine-grained manner, resulting in the indexes corresponding to user A and user B being stored in an interleaved manner, that is, the indexes corresponding to different document categories are stored in an interleaved manner, making the retrieval process relatively complex during full-text retrieval, thus greatly reducing the retrieval efficiency.
[0035] Based on this, the present application provides an index processing method and related device based on full-text retrieval, so that the index information corresponding to the same document category is stored in sequence to avoid the index information corresponding to different document categories being stored in an interleaved manner, and a more fine-grained index sub-table function based on document categories is realized in full-text retrieval, thereby making the retrieval process more convenient and fast during full-text retrieval and greatly improving the retrieval efficiency.
[0036] The index processing method based on full-text retrieval provided by this application can be applied to an index processing device based on full-text retrieval with data processing capabilities, such as a server, a terminal device, etc. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. In addition, the terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0037] In addition, the index processing device based on full-text retrieval provided by the embodiments of this application also has cloud storage capabilities. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technology, and distributed file systems, and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world.
[0038] Currently, the storage method of the storage system is as follows: Create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume, and this physical storage space may be composed of the disks of a certain storage device or several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object not only contains data but also contains additional information such as a data identifier (ID entity, ID), etc. The file system writes each object into the physical storage space of the logical volume respectively, and the file system will record the storage location information of each object. Thus, when the client requests to access the data, the file system can enable the client to access the data according to the storage location information of each object.
[0039] The process of the storage system allocating physical storage space for a logical volume is specifically as follows: According to the capacity estimation of the objects stored in the logical volume (this estimation often has a large margin relative to the actual capacity of the objects to be stored) and the group of a redundant array of independent disks (RAID), the physical storage space is pre-divided into stripes, and a logical volume can be understood as a stripe, thereby allocating physical storage space for the logical volume.
[0040] In addition, the index processing method based on full-text retrieval provided by the embodiments of the present application may also involve a blockchain. Among them, data such as index information and the first index file can be stored on the blockchain.
[0041] To facilitate understanding of the technical solution of the present application, the index processing method based on full-text retrieval provided by the embodiments of the present application will be introduced below in combination with an actual application scenario, taking a server as the index processing device based on full-text retrieval.
[0042] See Figure 1 , Figure 1 which is a schematic diagram of an application scenario of the index processing method based on full-text retrieval provided by the embodiments of the present application. In Figure 1 the application scenario shown, the application scenario includes a terminal device 101 and a server 102.
[0043] Among them, the terminal device 101 determines a document to be written to the disk as the document to be written, and sends a write request to the server 102. The write request needs to carry the document to be written and its related information. The related information may include, for example, a document identification identifier, a document category identifier, etc.; the server 102, as the aforementioned index processing device based on full-text retrieval, receives the write request sent by the terminal device 101 and performs relevant index processing on the document to be written carried by it.
[0044] First, the server 102 can obtain the document identification identifier of the document to be written and the document category identifier to which the document to be written belongs. The document category identifier is an identifier of a document category set based on business classification requirements, and the document category identifier is used to construct an index keyword; the server 102 also needs to obtain the words to be updated of the document to be written. For example, by performing operations such as word segmentation on the document to be written, the words to be updated of the document to be written are obtained.
[0045] Then, based on the above words to be updated and the document identification identifier, the server 102 adds the above document category identifier, and jointly constructs an index keyword and an index value corresponding to the index keyword in the memory to obtain index information in the form of key-value pairs. Since the index keyword in the index information is constructed by the document category identifier, the constructed index information can represent the document category at a finer granularity.
[0046] Finally, when the server 102 determines that a preset persistence trigger condition is met, it persists the index information stored in the memory in numerical order of the index keyword and stores it as a first index file on the disk. Based on the fact that the constructed index information can represent the document category at a finer granularity, after the persistence trigger, the index information can be arranged in an orderly manner according to the document category in the first index file on the disk.
[0047] Based on this, in the case of different document categories set according to business classification requirements, the index information corresponding to the same document category is stored sequentially to avoid the interleaved storage of the index information corresponding to different document categories, and to achieve a finer-grained index sub-table function based on document categories. Thus, during full-text retrieval, the retrieval process becomes more convenient and fast, greatly improving the retrieval efficiency.
[0048] Taking the business classification requirement including user classification requirement as an example, that is, the documents are classified in a fine-grained manner according to different users, the user to which the document belongs is determined, and a unique user identifier is assigned to each user. Then the document category identifier is the user identifier. Since the index keyword in the index information in this method is constructed by the user identifier, the constructed index information can represent the situation of the user to which the document belongs in a finer-grained manner. Furthermore, after the persistence trigger, it can be arranged in an orderly manner in the first index file on the disk according to the situation of the user to which the document belongs. In this way, the index information corresponding to the same user is stored sequentially, avoiding the interleaved storage of the index information corresponding to different users, and realizing a finer-grained index sub-table function based on users in full-text retrieval. Thus, during full-text retrieval, the retrieval process based on users becomes more convenient and fast, greatly improving the retrieval efficiency.
[0049] Next, in conjunction with the accompanying drawings, taking the server as the index processing device based on full-text retrieval, an index processing method based on full-text retrieval provided by an embodiment of the present application will be introduced.
[0050] See Figure 2 , Figure 2 which is a schematic flowchart of an index processing method based on full-text retrieval provided by an embodiment of the present application. As Figure 2 shown, the index processing method based on full-text retrieval includes the following steps:
[0051] S201. Obtain the words to be updated, the document identification identifier, and the document category identifier to which the document to be written belongs; the document category identifier is the identifier of the document category set according to the business classification requirement, and the document category identifier is used to construct the index keyword.
[0052] In the related art, an index is merely constructed based on the word segmentation and recognition identifiers of a document and stored for subsequent retrieval. The index constructed in this way is not detailed enough in terms of granularity, resulting in the indexes corresponding to different document categories in the business scenario being stored intertwined with each other, thus greatly reducing the retrieval efficiency. To solve the above problems, in the embodiments of the present application, different document categories are preset based on business classification requirements, and a unique identifier is assigned to each document category; any document to be written to the disk is used as a document to be written. For the document to be written, before constructing the index information, it is first necessary to obtain the data for constructing the index. It is not only necessary to obtain the words to be updated of the document to be written and the document recognition identifier of the document to be written, but also necessary to obtain the document category identifier to which the document to be written belongs. This document category identifier is used to construct the index keyword so that the index information constructed subsequently can represent the document category with a finer granularity.
[0053] Among them, the process of obtaining the words to be updated of the document to be written actually means: First, perform word segmentation processing on the document to be written, and multiple word segments of the document to be written can be obtained, forming a set of word segments to be written. Then, determine whether the document recognition identifier of the document to be written is already stored. The document recognition identifier refers to the identifier used to uniquely identify a document, such as a document number, etc., and can be represented by an unsigned 64-bit value; if so, it means that the document to be written is actually an updated version of the already written document corresponding to the document recognition identifier, and then the set of already written word segments of the already written document can be read from the memory and the disk through the document recognition identifier. Finally, by comparing the two sets of the set of word segments to be written and the set of already written word segments, the words to be updated of the document to be written can be obtained. Therefore, the present application provides a possible implementation. The step of obtaining the words to be updated of the document to be written in S201 may include, for example, the steps S2011 - S2013 in the following:
[0054] S2011: Perform word segmentation processing on the document to be written to obtain the set of word segments to be written of the document to be written.
[0055] S2012: If the document recognition identifier is already stored, obtain the set of already written word segments corresponding to the document recognition identifier from the memory and the disk.
[0056] In addition, if the document recognition identifier is not stored, it means that the set of already written word segments corresponding to the document recognition identifier is not included in the memory and the disk, and then the words to be updated of the document to be written can be directly obtained based on the set of word segments to be written of the document to be written, that is, the word segments to be written in the set of word segments to be written are used as the words to be updated of the document to be written.
[0057] S2013: Based on the set of word segments to be written and the set of already written word segments, obtain the words to be updated of the document to be written.
[0058] As an example, the document to be written is Document 1, and the document identification identifier of Document 1 is 000001. Perform word segmentation on Document 1 to obtain the set of words to be written for Document 1, which is Set A; when it is determined that the document identification identifier 000001 has been stored, read the set of words that have been written corresponding to the document identification identifier 000001 from the memory and disk, which is Set B; calculate the difference set between Set A and Set B to obtain the set of words to be inserted, A - B, and calculate the difference set between Set B and Set A to obtain the set of words to be deleted, B - A. The words to be inserted in the set of words to be inserted A - B and the words to be deleted in the set of words to be deleted B - A are used as the words to be updated for Document 1.
[0059] Among them, in the embodiments of the present application, the specific content of the business classification requirements is not specifically limited. The specific content of the business classification requirements can be determined according to the specific situation of the business scenario. Since different document categories are preset based on the business classification requirements and a unique identifier is assigned to each document category, therefore, through the document category to which the document to be written belongs, the document category identifier to which the document to be written belongs can be obtained.
[0060] For example, in the case of a large number of users in the user communication scenario, the business classification requirement can be the user classification requirement, that is, the documents are classified in fine granularity according to different users; at this time, a unique user identifier needs to be assigned to each user. Correspondingly, the document category identifier can be the user identifier, and the user identifier can be a personal identifier, or an enterprise identifier, or an account identifier, or any combination of the foregoing three. Therefore, the present application provides a possible implementation manner. When the business classification requirement includes the user classification requirement, the document category identifier includes the user identifier, and the user identifier includes any one or more combinations of the personal identifier, the enterprise identifier, or the account identifier.
[0061] As an example, the business classification requirement is the user classification requirement, a unique user identifier is assigned to each user, and the document category identifier is the user identifier; the document to be written is Document 1, the document category to which Document 1 belongs is User A, and the user identifier assigned to User A is 100001. Then the document category identifier to which Document 1 belongs is the user identifier of User A, that is, 100001.
[0062] For another example, in the case of a report with a very long time in the data report scenario, the business classification requirement can be a time classification requirement, that is, the documents are classified in fine granularity according to different times; at this time, a unique time identifier needs to be assigned to each period of time. Correspondingly, the document category identifier can be a time identifier, and the time identifier can be a date identifier, or a cycle identifier, or a month identifier, or a year identifier, or any combination of the foregoing four, etc. Therefore, the present application provides a possible implementation manner. When the business classification requirement includes a time classification requirement, the document category identifier includes a time identifier, and the time identifier includes any one or more combinations of a date identifier, a cycle identifier, a month identifier, or a year identifier.
[0063] As an example, the business classification requirement is a time classification requirement. A unique time identifier is assigned to each period of time, and the document category identifier is a time identifier; the document to be written is Document 1, the document category to which Document 1 belongs is XX / XX / XXXX, and the time identifier obtained by assigning XX / XX / XXXX is 100111. Then, the document category identifier to which Document 1 belongs is the time identifier of XX / XX / XXXX, that is, 100111.
[0064] S202. Based on the word to be updated, the document identification identifier, and the document category identifier, construct index keywords and index values corresponding to the index keywords in the memory to obtain index information, and the index information is stored in a key-value pair format.
[0065] In the embodiment of the present application, the word to be updated, the document identification identifier, and the document category identifier of the document to be written obtained in S201 are all data used to construct the index. Then, based on the above data, index keywords and index values corresponding to the index keywords can be constructed to obtain index information, and the index information is stored in the memory in a key-value pair format through a Log-Structured Merge Tree (LSM-tree). Since the index keywords in the index information are constructed through the document category identifier, the constructed index information can characterize the document category in a finer granularity.
[0066] Among them, the LSM-tree is composed of two or more data storage structures. The simplest LSM-tree consists of two components. One component resides in the memory and can be any data structure convenient for key-value lookup. The other component resides on the disk and can also be any data structure convenient for key-value lookup.
[0067] In fact, through the words to be updated, document identification identifiers, and document category identifiers, forward index keywords and corresponding forward index values of the forward index keywords can be constructed to obtain forward index information; or reverse index keywords and corresponding reverse index values of the reverse index keywords can be constructed to obtain reverse index information; that is, the index information can include forward index information and reverse index information.
[0068] Among them, the above-mentioned forward index information refers to using the document identification identifier of the document to be written as the forward index keyword, recording the words to be updated of the document to be written in the forward index value. On the basis that the document category identifier is used to construct the index keyword, the forward index keyword can be constructed through the document identification identifier and the document category identifier, the forward index value can be constructed by serializing the words to be updated, and the forward index keyword and the forward index value are stored in a key-value pair format in the memory, that is, the forward index information is obtained.
[0069] The above-mentioned reverse index information refers to using the words to be updated of the document to be written as the reverse index keyword, recording the document identification identifier of the document to be written in the reverse index value. On the basis that the document category identifier is used to construct the index keyword, the reverse index keyword can be constructed through the words to be updated and the document category identifier, the reverse index value can be constructed through the document identification identifier, and the reverse index keyword and the reverse index value are stored in a key-value pair format in the memory, that is, the reverse index information is obtained.
[0070] Therefore, the present application provides a possible implementation manner. For example, S202 may include the following steps S2021-S2026:
[0071] S2021. Based on the document identification identifier and the document category identifier, construct a forward index keyword.
[0072] The forward index keyword includes two parts: a forward prefix and a document identification identifier. Among them, the forward prefix is used to represent the document category to which the document belongs to support a finer-grained index sub-table function based on the document category, and the forward prefix needs to have a forward mark; then the forward prefix can be determined first through the document category identifier and the forward identifier, and then the forward index keyword can be constructed in combination with the document identification identifier. Therefore, the present application provides a possible implementation manner. For example, S2021 may include the following steps: Based on the document category identifier and the forward identifier, determine the forward prefix in the forward index keyword; based on the forward prefix and the document identification identifier, construct the forward index keyword.
[0073] As an example, such as Figure 3Schematic diagram of the format of a forward index keyword. The forward index keyword includes a forward prefix and a document identification identifier. The forward prefix includes a document category identifier and a forward identifier. Among them, the document category identifier can occupy 8 bytes, an unsigned 64-bit value, such as the user identification 100001 of user A; the forward identifier can occupy 1 byte, such as a fixed value 1, and the document identification identifier can also occupy 8 bytes, an unsigned 64-bit value, such as the document identification identifier 000001 of document 1, or the document identification identifier 000002 of document 2, or the document identification identifier 000003 of document 3, etc.
[0074] S2022. Perform serialization processing based on the word to be updated to construct a forward index value corresponding to the forward index keyword.
[0075] S2023. Store the forward index keyword and the forward index value in a key-value pair format in memory to obtain the forward index information in the index information.
[0076] S2024. Construct a reverse index keyword based on the word to be updated and the document category identifier.
[0077] The reverse index keyword includes two parts: a reverse prefix and the word to be updated. Among them, the reverse prefix is used to represent the document category to which the document belongs to support a more fine-grained index sub-table function based on the document category, and the reverse prefix needs to have a reverse mark; then the reverse prefix can be determined first through the document category identifier and the reverse identifier, and then combined with the word to be updated to construct the reverse index keyword. Therefore, the present application provides a possible implementation manner. S2024 can include the following steps, for example: determine the reverse prefix in the reverse index keyword based on the document category identifier and the reverse identifier; construct the reverse index keyword based on the reverse prefix and the word to be updated.
[0078] As an example, as Figure 4 Schematic diagram of the format of a reverse index keyword. The reverse index keyword includes a reverse prefix and the word to be updated. The reverse prefix includes a document category identifier and a reverse identifier. Among them, the document category identifier can occupy 8 bytes, an unsigned 64-bit value, such as the user identification 100001 of user A; the reverse identifier can occupy 1 byte, such as a fixed value 0, and the word to be updated can be, for example, the word to be updated 1, or the word to be updated 2, or the word to be updated 3, etc.
[0079] S2025. Construct a reverse index value corresponding to the reverse index keyword based on the document identification identifier.
[0080] The inverted index value actually stores the inverted list of the word to be updated. The inverted list includes three parts: a header, a deletion list, and an inverted list of document identification identifiers. Among them, the header is used to record the coding version corresponding to the deletion list and the inverted list of document identification identifiers. The deletion list is used to record the positions of the document identification identifiers of the documents that need to delete the word to be updated in the inverted list of document identification identifiers. The inverted list of document identification identifiers is used to record the reverse order of the document identification identifiers of the documents including the word to be updated.
[0081] In the process of constructing the inverted index value corresponding to the inverted index keyword based on the document identification identifier, it is necessary to determine whether the inverted index keyword has a corresponding historical inverted list. If not, it means that the inverted list has not been encoded for this inverted index keyword, and it is necessary to encode the inverted list of the word to be updated based on the document identification identifier to obtain the inverted index value corresponding to this inverted index keyword. If there is, it means that the inverted list has been encoded for this inverted index keyword, that is, the above historical inverted list, and it is necessary to insert the document identification identifier into the historical inverted list of document identification identifiers in the historical inverted list in an orderly manner.
[0082] On the basis that the words to be updated for constructing the inverted index keyword are divided into words to be inserted for word segmentation and words to be deleted for word segmentation, for the words to be inserted for word segmentation, directly inserting the document identification identifier into the historical inverted list of document identification identifiers in an orderly manner can obtain the inverted index value corresponding to the inverted index keyword. For the words to be deleted for word segmentation, not only the document identification identifier needs to be inserted into the historical inverted list of document identification identifiers in an orderly manner, but also the deletion mark of the corresponding document identification identifier needs to be added to the historical deletion list in the historical inverted list to obtain the inverted index value corresponding to the inverted index keyword. Therefore, the present application provides a possible implementation manner. S2025 may include steps A-step C in the following steps:
[0083] Step A: If the inverted index keyword has a corresponding historical inverted list and the word to be updated is a word to be inserted for word segmentation, insert the document identification identifier into the historical inverted list of document identification identifiers in the historical inverted list in an orderly manner to obtain the inverted index value corresponding to the inverted index keyword.
[0084] Step B: If the inverted index keyword has a corresponding historical inverted list and the word to be updated is a word to be deleted for word segmentation, insert the document identification identifier into the historical inverted list of document identification identifiers in the historical inverted list in an orderly manner, and add the deletion mark of the corresponding document identification identifier to the historical deletion list in the historical inverted list to obtain the inverted index value corresponding to the inverted index keyword.
[0085] Step C: If the inverted index keyword does not have a corresponding historical inverted list, encode the inverted list of the word to be updated based on the document identification identifier to obtain the inverted index value corresponding to the inverted index keyword.
[0086] As an example, as Figure 5 shown in the schematic diagram of the format of an inverted list. Taking the document identification identifiers of the respective documents including the term to be updated as 99999, 99998, 99997, 99996, 99995 respectively, where the document identification identifiers of the respective documents that need to delete the term to be updated are 99998 and 99995 as an example; then the inverted list of document identification identifiers uses varint encoding. Varint is a method of serializing an integer using one or more bytes, and will encode the integer into variable-length bytes. Specifically, in reverse order of the document identification identifiers, the first one is the document identification identifier with the largest value as the basis, and each subsequent value records the difference from the previous document identification identifier, that is, 99999, 1, 1, 1, 1. The deletion list also uses varint encoding. Specifically, in ascending order of the positions of the document identification identifiers of the respective documents that need to delete the term to be updated in the inverted list of document identification identifiers, the positions of the document identification identifiers in the inverted list of document identification identifiers start from 0, the first one is the specific position, and each subsequent value records the difference from the previous position, that is, 1, 3.
[0087] S2026. Store the reverse index keyword and the reverse index value in a key-value pair format in the memory to obtain the reverse index information in the index information.
[0088] As an example, as Figure 6 shown in the schematic diagram of the format of a forward index information and a reverse index information. Among them, the forward index information includes a forward index keyword and a forward index value. The forward index keyword includes a forward prefix and a document identification identifier. The forward index value includes the serialized term to be updated. The reverse index information includes a reverse index keyword and a reverse index value. The reverse index keyword includes a reverse prefix and the term to be updated. The reverse index value includes the inverted list of the term to be updated.
[0089] S203. When a preset persistence trigger condition is met, based on the numerical order of the index keywords, persist the index information stored in the memory into a first index file and store it on the disk.
[0090] In the embodiment of the present application, the data stored in the memory needs to be persisted into a file and stored on the disk subsequently, and the persistence requires certain triggering conditions. Therefore, the triggering conditions for persistence are preset as the preset persistence triggering conditions. After the index information constructed in S202 is stored in the memory in the form of key-value pairs, it is necessary to first determine whether the preset persistence triggering conditions are met. If so, then in the numerical order of the index keywords in the index information, the index information needs to be persisted into the first index file, and the first index file is stored on the disk through the LSM-tree. On the basis that the index information constructed in S202 can represent the document category in a finer granularity, after the persistence is triggered, the index information can be arranged in an orderly manner in the first index file on the disk according to the document category.
[0091] Based on this, the implementation manners of S201 - S203, when setting different document categories according to the business classification requirements, enable the index information corresponding to the same document category to be stored sequentially, so as to avoid the index information corresponding to different document categories from being stored interleaved with each other, and implement a finer granularity of index sub-table function based on document categories in full-text retrieval, thereby making the retrieval processing more convenient and fast during full-text retrieval, and greatly improving the retrieval efficiency.
[0092] Taking the business classification requirements including user classification requirements as an example, that is, the documents are classified in a finer granularity according to different users, the user to which the document belongs is determined, and a unique user identifier is assigned to each user, then the document category identifier is the user identifier. Since the index keywords in the index information are constructed through the user identifier in this method, the constructed index information can represent the situation of the user to which the document belongs in a finer granularity, and thus can be arranged in an orderly manner in the first index file on the disk according to the situation of the user to which the document belongs after the persistence is triggered. In this way, the index information corresponding to the same user is stored sequentially, avoiding the index information corresponding to different users from being stored interleaved with each other, and implementing a finer granularity of index sub-table function based on users in full-text retrieval, thereby making the retrieval processing based on users more convenient and fast during full-text retrieval, and greatly improving the retrieval efficiency.
[0093] Generally, when the memory usage status reaches the preset memory status, for example, the memory size reaches the preset memory threshold, or the memory occupancy rate reaches the preset occupancy rate, etc.; when the statistical time after persistence reaches the preset time, for example, the statistical time since the last persistence reaches the preset persistence cycle time; when the system is restarted, etc., the index information is triggered to be persisted into the first index file. Therefore, the present application provides a possible implementation manner, and the preset persistence triggering conditions include one or more of the following: the memory usage status reaches the preset memory status, the statistical time after persistence reaches the preset time, or the system is restarted.
[0094] In addition, in the embodiments of the present application, the disk includes multiple layers of file layers. The levels of the multiple layers of file layers are different. Each layer of file layer stores multiple index files. Each index file stores multiple index information within a continuous index key range. The index information inside the index file and between the index files is arranged in ascending order of the numerical value of the index key. As the number of index files stored in the file layer with a lower level increases, a part of the index files can be selected and merged with the index files stored in the file layer of the higher level. After the merge, they are stored in the file layer of the higher level. This merge method reduces the redundant index files and index information in the disk, so that during subsequent full-text retrieval, the retrieval process can be made more convenient and fast, thereby further improving the retrieval efficiency. Therefore, the present application provides a possible implementation. The disk includes a first file layer and a second file layer. The level of the second file layer is higher than that of the first file layer. The first file layer stores a first index file, and the second file layer stores a second index file. Correspondingly, the method may further include, for example, S204: Merging the first index file in the first file layer with the second index file in the second file layer and storing them in the second file layer.
[0095] On the basis of the above description, the disk may further include a third file layer. The level of the third file layer is higher than that of the second file layer. The third file layer stores a third index file. The second index file in the second file layer is merged with the third index file in the third file layer and stored in the third file layer, and so on, which will not be elaborated here.
[0096] As an example, as Figure 7 shown in the schematic diagram of the file layers in a disk, the first file layer included in the disk is Level 0, the second file layer is Level 1, the second file layer is Level 2, and so on. There is an overlap in the index key range between the first index files stored in Level 0. For example, there is an overlap in the index key range between 007.sst and 008.sst. For file layers other than Level 0, such as the second index files stored in Level 1 and the third index files stored in Level 2, there is no overlap in the index key range. For example, there is no overlap in the index key range between 005.sst and 004.sst, and there is no overlap in the index key range between 002.sst and 001.sst.
[0097] The index processing method based on full-text retrieval provided by the above embodiments sets identifiers for different document categories in advance according to business classification requirements. For a document to be written, based on the words to be updated in the document to be written and the document identification identifier, by adding the document category identifier to which the document to be written belongs, an index keyword and an index value corresponding to the index keyword are jointly constructed in memory to obtain index information in the form of key-value pairs. The document category identifier is used to construct the index keyword. When the preset persistence trigger condition is met, the index information stored in memory is persisted in numerical order of the index keyword to obtain a first index file stored on the disk. Since the index keyword in the index information is constructed by the document category identifier, the constructed index information can represent the document category in a finer granularity, and thus can be sorted in order by document category in the first index file on the disk after the persistence trigger. Based on this, in the case of setting different document categories according to business classification requirements, the index information corresponding to the same document category is stored sequentially to avoid the interleaved storage of the index information corresponding to different document categories, and a finer-granularity index sub-table function based on document categories is realized in full-text retrieval, so that the retrieval processing is more convenient and fast during full-text retrieval, and the retrieval efficiency is greatly improved.
[0098] In addition, in the related art, Elasticsearch is used to store the index. The retrieval engine used during full-text retrieval is Lucene. Through Lucene, only the index stored on the disk can be retrieved, and the index stored in memory is not persistently stored on the disk in real time, so the complete and accurate retrieval result cannot be retrieved in real time. Among them, Elasticsearch is a search server based on Lucene, which provides a full-text retrieval service with distributed multi-user capabilities.
[0099] Therefore, based on the above embodiments of the index processing method based on full-text retrieval, when performing full-text retrieval on the text to be retrieved and the document category identifier to be retrieved input by the user through the terminal device, the terminal device sends them to the server. The server constructs a reverse index keyword to be retrieved through the word to be retrieved in the text to be retrieved and the document category identifier to be retrieved. Considering the characteristics of the LSM-tree, it is necessary to retrieve the reverse index keyword to be retrieved not only on the disk but also in memory to obtain a list of inverted permutations of multiple target document identification identifiers corresponding to the reverse index keyword to be retrieved. The list of inverted permutations of multiple target document identification identifiers has priorities, and the list of inverted permutations of multiple target document identification identifiers is merged according to the priorities to obtain a complete list of inverted permutations of target document identification identifiers.
[0100] The following introduces a full-text retrieval method provided by an embodiment of the present application with reference to the accompanying drawings. See Figure 8 , Figure 8A flowchart of a full-text retrieval method provided by an embodiment of the present application. As Figure 8 shown, the full-text retrieval method includes the following steps:
[0101] S801. Obtain the search terms and the search document category identifier of the text to be retrieved.
[0102] In the embodiment of the present application, the text to be retrieved is segmented to obtain multiple search terms of the text to be retrieved, forming a set of search terms. The search terms in the set of search terms are used as the search terms.
[0103] S802. Based on the search terms and the search document category identifier, construct the reverse index keywords to be retrieved.
[0104] In the embodiment of the present application, the implementation manner of S802 can refer to the implementation manner of S2024. Based on the search document category identifier and the reverse identifier, determine the reverse prefix in the reverse index keywords to be retrieved; based on the reverse prefix and the search terms, construct the reverse index keywords to be retrieved.
[0105] S803. Retrieve in memory and on disk based on the reverse index keywords to be retrieved, and obtain an inverted list of multiple target document identification identifiers corresponding to the reverse index keywords to be retrieved.
[0106] In the embodiment of the present application, when retrieving the reverse index keywords to be retrieved in memory and on disk, in fact, multiple target reverse index values corresponding to the reverse index keywords to be retrieved can be read. The target reverse index values store the target inverted list of the search terms. Referring to the description of the inverted list in the above embodiment, multiple target reverse index values can be parsed to obtain an inverted list of multiple target document identification identifiers. Therefore, the present application provides a possible implementation manner. For example, S803 may include S8031 - S8032 in the following steps:
[0107] S8031. Retrieve in memory and on disk based on the reverse index keywords to be retrieved, and obtain multiple target reverse index values corresponding to the reverse index keywords to be retrieved.
[0108] S8032. Parse the multiple target reverse index values to obtain an inverted list of multiple target document identification identifiers.
[0109] S804. Merge the inverted lists of multiple target document identification identifiers according to the priorities of the inverted lists of multiple target document identification identifiers, and obtain a complete inverted list of target document identification identifiers.
[0110] In the embodiments of the present application, the inverted list of target document identification identifiers obtained from the memory has the highest priority. Among the inverted lists of target document identification identifiers obtained from the disk, the lower the level of the file layer where the inverted list of target document identification identifiers is stored, the higher the priority of the inverted list of target document identification identifiers. If two or more inverted lists of target document identification identifiers are both stored in the first file layer, the higher the write time, the higher the priority.
[0111] In the embodiments of the present application, the specific implementation manner of S804 refers to: traversing the first target document identification identifier among the multiple inverted lists of target document identification identifiers according to the priorities of the multiple inverted lists of target document identification identifiers, taking the target document identification identifier with the highest priority as the standard. If the first target document identification identifier with the highest priority has a deletion mark, the first target document identification identifier needs to be removed. If the first target document identification identifier with the highest priority does not have a deletion mark, the first target document identification identifier needs to be retained; traversing the next target document identification identifier in a similar manner as above; until all target document identification identifiers are traversed to obtain a complete inverted list of target document identification identifiers. Therefore, the present application provides a possible implementation manner. For example, S804 may include S8041 - S8042 in the following steps:
[0112] S8041. Traverse the target document identification identifiers in the multiple inverted lists of target document identification identifiers according to the priorities of the multiple inverted lists of target document identification identifiers to determine the target document identification identifier with the highest priority.
[0113] S8042. If the target document identification identifier with the highest priority has a deletion mark, remove the target document identification identifier. If the target document identification identifier with the highest priority does not have a deletion mark, retain the target document identification identifier to obtain a complete inverted list of target document identification identifiers.
[0114] As an example, as Figure 9 shown in a schematic diagram of merging multiple inverted lists of target document identification identifiers. Among them, the priorities of 4 inverted lists of target document identification identifiers are arranged from high to low from top to bottom. Traverse the first target document identification identifier 9. The target document identification identifier with the highest priority is 9[delete], which has a deletion mark, and remove the target document identification identifier 9; traverse the next target document identification identifier 8. The target document identification identifier with the highest priority is 8, which does not have a deletion mark, and retain the target document identification identifier 8, and so on, until traversing the last target document identification identifier 4. The target document identification identifier with the highest priority is 4, which does not have a deletion mark, and retain the target document identification identifier 4 to obtain a complete inverted list of target document identification identifiers, that is, 8, 7, 6, 5, 4.
[0115] The full-text retrieval method provided by the above embodiments constructs a reverse index keyword to be retrieved through the keyword to be retrieved in the text to be retrieved and the document category identifier to be retrieved. It is necessary to retrieve the reverse index keyword to be retrieved not only in the disk but also in the memory, and multiple target document identification ID inverted lists corresponding to the reverse index keyword to be retrieved can be obtained. The multiple target document identification ID inverted lists have priorities, and the multiple target document identification ID inverted lists are merged according to the priorities to obtain a complete target document identification ID inverted list. Based on this, during full-text retrieval, the inverted list of document identification IDs newly stored in the memory and the inverted list of historical stored document identification IDs in the disk can be merged in real time to achieve real-time retrieval and obtain complete and accurate retrieval results.
[0116] To better understand the index processing method and full-text retrieval method based on full-text retrieval provided by the embodiments of the present application, the overall framework based on full-text retrieval will be introduced below. Refer to Figure 10 , Figure 10 which is a schematic diagram of an overall framework based on full-text retrieval provided by the embodiments of the present application.
[0117] The overall framework based on full-text retrieval may include an acquisition module, a word segmentation module, a construction module, a persistence module, a merging module, and a retrieval module. Among them, the input of the acquisition module can be a document to be written or a text to be retrieved.
[0118] When the input of the acquisition module is a document to be written, the word segmentation module performs word segmentation processing on the document to be written to obtain the words to be updated in the document to be written. The construction module constructs an index keyword and an index value corresponding to the index keyword based on the words to be updated, the document identification ID of the document to be written, and the document category identifier to which the document to be written belongs, obtains forward index information and reverse index information, and stores the forward index information and reverse index information in the memory in the form of key-value pairs. The persistence module, when meeting the preset persistence trigger condition, persists the forward index information and reverse index information stored in the memory into an index file in the numerical order of the index keywords, stores the index file in the disk. In addition, the merging and deletion module merges the index files in the low-level file layer stored in the disk with the index files in the high-level file layer and stores them in the high-level file layer, and deletes the document identification IDs with deletion marks in the reverse index values of the reverse index information.
[0119] When the input of the acquisition module is the text to be retrieved, the word segmentation module performs word segmentation on the text to be retrieved to obtain the words to be retrieved of the text to be retrieved; the retrieval module constructs the reverse index keywords to be retrieved based on the words to be retrieved and the document category identifier to be retrieved, retrieves in the memory and disk to obtain the inverted list of multiple target document identification identifiers corresponding to the reverse index keywords to be retrieved, and merges the inverted lists of multiple target document identification identifiers according to the priority of the inverted lists of multiple target document identification identifiers to obtain the complete inverted list of target document identification identifiers.
[0120] Based on full-text retrieval, it supports fine-grained index sub-table function based on document category, can also update index information in real time, and as the index information and index files gradually increase, the retrieval performance can still remain stable and efficient. This method is widely used in full-text retrieval in business scenarios such as communication scenarios. The document categories can include enterprise employee address books, approvals, daily reports, weekly reports, reports, enterprise material retrieval, enterprise mailboxes, etc. The largest business scenario can reach more than 30 billion records, the number of index information is in the trillions +, and the storage capacity is dozens of terabytes (Terabyte, TB).
[0121] For the index processing method based on full-text retrieval provided in the above embodiments, the embodiments of the present application also provide an index processing device based on full-text retrieval.
[0122] See Figure 11 This figure is a schematic diagram of an index processing device based on full-text retrieval provided by an embodiment of the present application. As Figure 11 shown, the index processing device 1100 based on full-text retrieval includes: an acquisition unit 1101, a construction unit 1102, and a persistence unit 1103;
[0123] The acquisition unit 1101 is configured to acquire the words to be updated, the document identification identifier of the document to be written, and the document category identifier to which the document to be written belongs; the document category identifier is the identifier of the document category set based on business classification requirements, and the document category identifier is used to construct index keywords;
[0124] The construction unit 1102 is configured to construct index keywords and index values corresponding to the index keywords in the memory based on the words to be updated, the document identification identifier, and the document category identifier to obtain index information, and the index information is stored in a key-value pair format;
[0125] The persistence unit 1103 is configured to, when a preset persistence trigger condition is met, persist the index information stored in the memory into a first index file and store it in the disk based on the numerical order of the index keywords.
[0126] As a possible implementation manner, the construction unit 1102 is configured to:
[0127] Construct a forward index keyword based on the document recognition identifier and the document category identifier;
[0128] Perform serialization processing based on the word to be updated, and construct a forward index value corresponding to the forward index keyword;
[0129] Store the forward index keyword and the forward index value in a key-value pair format in memory to obtain the forward index information in the index information;
[0130] Construct a reverse index keyword based on the word to be updated and the document category identifier;
[0131] Construct a reverse index value corresponding to the reverse index keyword based on the document recognition identifier;
[0132] Store the reverse index keyword and the reverse index value in a key-value pair format in memory to obtain the reverse index information in the index information.
[0133] As a possible implementation, construct unit 1102 for:
[0134] Determine the positive prefix in the forward index keyword based on the document category identifier and the positive identifier;
[0135] Construct a forward index keyword based on the positive prefix and the document recognition identifier;
[0136] Determine the negative prefix in the reverse index keyword based on the document category identifier and the negative identifier;
[0137] Construct a reverse index keyword based on the negative prefix and the word to be updated.
[0138] As a possible implementation, construct unit 1102 for:
[0139] If the reverse index keyword has a corresponding historical inverted list and the word to be updated is the token to be inserted, insert the document recognition identifier into the historical document recognition identifier inverted list in the historical inverted list in an orderly manner to obtain the reverse index value corresponding to the reverse index keyword;
[0140] If the reverse index keyword has a corresponding historical inverted list and the word to be updated is the token to be deleted, insert the document recognition identifier into the historical document recognition identifier inverted list in the historical inverted list in an orderly manner, and add a deletion mark for the corresponding document recognition identifier to the historical deletion list in the historical inverted list to obtain the reverse index value corresponding to the reverse index keyword;
[0141] If the reverse index keyword does not have a corresponding historical inverted list, encode the inverted list of the word to be updated based on the document recognition identifier to obtain the reverse index value corresponding to the reverse index keyword.
[0142] As a possible implementation, an obtaining unit 1101 is configured to:
[0143] Perform word segmentation on the document to be written to obtain a set of words to be written for the document to be written;
[0144] If the document identification flag is already stored, obtain the set of written words corresponding to the document identification flag from the memory and the disk;
[0145] Based on the set of words to be written and the set of written words, obtain the words to be updated for the document to be written.
[0146] As a possible implementation, the disk includes a first file layer and a second file layer, the level of the second file layer is higher than that of the first file layer, the first file layer stores a first index file, and the second file layer stores a second index file; the apparatus further includes a merging unit;
[0147] The merging unit is configured to merge the first index file in the first file layer with the second index file in the second file layer and store them in the second file layer.
[0148] As a possible implementation, the preset persistence trigger conditions include one or more of the following:
[0149] The memory usage status reaches a preset memory status, the statistical time after persistence reaches a preset time, or the system restarts.
[0150] As a possible implementation, when the service classification requirement includes a user classification requirement, the document category identifier includes a user identifier, and the user identifier includes any one or more combinations of a personal identifier, an enterprise identifier, or an account identifier; when the service classification requirement includes a time classification requirement, the document category identifier includes a time identifier, and the time identifier includes any one or more combinations of a date identifier, a cycle identifier, a month identifier, or a year identifier.
[0151] As a possible implementation, the apparatus further includes a merging unit;
[0152] The obtaining unit 1101 is further configured to obtain the words to be retrieved for the text to be retrieved and the document category identifier to be retrieved;
[0153] The constructing unit 1102 is further configured to construct a reverse index keyword to be retrieved based on the words to be retrieved and the document category identifier to be retrieved;
[0154] The obtaining unit 1101 is further configured to perform a search in the memory and the disk based on the reverse index keyword to be retrieved, and obtain an inverted list of multiple target document identification flags corresponding to the reverse index keyword to be retrieved;
[0155] A merging unit, configured to merge multiple inverted lists of target document identification identifiers according to the priorities of the multiple inverted lists of target document identification identifiers, so as to obtain a complete inverted list of target document identification identifiers.
[0156] As a possible implementation manner, the obtaining unit 1101 is further configured to:
[0157] Retrieve based on the reverse index keywords to be retrieved in the memory and the disk, so as to obtain multiple target reverse index values corresponding to the reverse index keywords to be retrieved;
[0158] Parse the multiple target reverse index values to obtain multiple inverted lists of target document identification identifiers.
[0159] The index processing device based on full-text retrieval provided in the foregoing embodiment pre-sets the identifiers of different document categories according to the service classification requirements. For the document to be written, on the basis of the words to be updated and the document identification identifier of the document to be written, the identifier of the document category to which the document to be written belongs is added, and the index keyword and the index value corresponding to the index keyword are jointly constructed in the memory to obtain the index information in the key-value pair format. The document category identifier is used to construct the index keyword. When the preset persistence trigger condition is met, the index information stored in the memory is persisted according to the numerical order of the index keywords, and the first index file is stored on the disk. Since the index keyword in the index information is constructed by the document category identifier, the constructed index information can represent the document category in a finer granularity, and then can be sorted in order according to the document category in the first index file on the disk after the persistence trigger. Based on this, in the case of setting different document categories according to the service classification requirements, the index information corresponding to the same document category is stored sequentially, so as to avoid the index information corresponding to different document categories from being stored in an interleaved manner, and realize a finer-grained index sub-table function based on the document category in full-text retrieval, so that the retrieval processing is more convenient and fast during full-text retrieval, and the retrieval efficiency is greatly improved.
[0160] An embodiment of the present application further provides an index processing device for full-text retrieval. The index processing device for full-text retrieval provided in the embodiment of the present application will be introduced from the perspective of hardware implementation below.
[0161] See Figure 12 , Figure 12FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1200 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1222 (e.g., one or more processors) and a memory 1232, and one or more storage media 1230 (e.g., one or more mass storage devices) for storing application programs 1242 or data 1244. Among them, the memory 1232 and the storage media 1230 may be transient storage or persistent storage. The program stored in the storage media 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1222 may be configured to communicate with the storage media 1230 and execute a series of instruction operations in the storage media 1230 on the server 1200.
[0162] The server 1200 may further include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, and / or one or more operating systems 1241, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0163] The steps performed by the server in the above embodiments may be based on the Figure 12 server structure shown.
[0164] Among them, the CPU 1222 is used to perform the following steps:
[0165] Obtain the words to be updated, the document identification identifier, and the document category identifier to which the document to be written belongs; the document category identifier is an identifier of the document category set based on business classification requirements, and the document category identifier is used to construct an index keyword;
[0166] Based on the words to be updated, the document identification identifier, and the document category identifier, construct an index keyword and an index value corresponding to the index keyword in the memory to obtain index information, and the index information is stored in a key-value pair format;
[0167] When a preset persistent trigger condition is met, based on the numerical order of the index keywords, persist the index information stored in the memory as a first index file and store it on the disk.
[0168] For the index processing method based on full-text retrieval described above, an embodiment of the present application further provides a terminal device for index processing based on full-text retrieval, so that the above-mentioned index processing method based on full-text retrieval can be implemented and applied in practice.
[0169] See Figure 13 , Figure 13 which is a schematic structural diagram of a terminal device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA for short), etc. Taking the terminal device as a mobile phone as an example:
[0170] Figure 13 The block diagram shows a part of the structure of the mobile phone related to the terminal device provided by the embodiment of the present application. Refer to Figure 13 , the mobile phone includes: a radio frequency (RF) circuit 1310, a memory 1320, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, a wireless fidelity (WiFi) module 1370, a processor 1380, and a power supply 1390 and other components. Those skilled in the art can understand that Figure 13 the structure of the mobile phone shown in
[0171] does not limit the mobile phone, and may include more or fewer components than shown in the figure, or combine some components, or different component arrangements. Figure 13 The following specifically introduces each component of the mobile phone:
[0172] The RF circuit 1310 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is sent to the processor 1380 for processing. Additionally, the data designed for uplink transmission is sent to the base station. Generally, the RF circuit 1310 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1310 can also communicate with the network and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0173] The memory 1320 can be used to store software programs and modules. The processor 1380 runs the software programs and modules stored in the memory 1320 to implement various functional applications and data processing of the mobile phone. The memory 1320 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 1320 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0174] The input unit 1330 can be used to receive input numeric or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1330 can include a touch panel 1331 and other input devices 1332. The touch panel 1331, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 1331), and drive corresponding connection devices according to a pre-set program. Optionally, the touch panel 1331 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1380, and can receive and execute the commands sent by the processor 1380. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1331. In addition to the touch panel 1331, the input unit 1330 can also include other input devices 1332. Specifically, the other input devices 1332 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0175] The display unit 1340 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1340 can include a display panel 1341. Optionally, the display panel 1341 can be configured in forms such as a liquid crystal display (LCD for short) or an organic light-emitting diode (OLED for short). Further, the touch panel 1331 can cover the display panel 1341. After the touch panel 1331 detects a touch operation thereon or nearby, it transmits it to the processor 1380 to determine the type of touch event. Subsequently, the processor 1380 provides corresponding visual output on the display panel 1341 according to the type of touch event. Although in Figure 13 the touch panel 1331 and the display panel 1341 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1331 and the display panel 1341 can be integrated to realize the input and output functions of the mobile phone.
[0176] The mobile phone may further include at least one sensor 1350, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1341 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1341 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.
[0177] The audio circuit 1360, the speaker 1361, and the microphone 1362 can provide an audio interface between the user and the mobile phone. The audio circuit 1360 can transmit the electrical signal converted from the received audio data to the speaker 1361, and the speaker 1361 converts it into a sound signal for output; on the other hand, the microphone 1362 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1360 and then converted into audio data. After the audio data is output to the processor 1380 for processing, it is sent through the RF circuit 1310 to, for example, another mobile phone, or the audio data is output to the memory 1320 for further processing.
[0178] WiFi belongs to short - range wireless transmission technology. Through the WiFi module 1370, the mobile phone can help users send and receive emails, browse the web, and access streaming media, etc., and it provides users with wireless broadband Internet access. Although Figure 13 the WiFi module 1370 is shown, it can be understood that it does not belong to the essential components of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention according to needs.
[0179] The processor 1380 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1320, and by calling the data stored in the memory 1320, it executes various functions of the mobile phone and processes data. Optionally, the processor 1380 may include one or more processing units; preferably, the processor 1380 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 1380 either.
[0180] The mobile phone further includes a power source 1390 (such as a battery) for powering each component. Preferably, the power source can be logically connected to the processor 1380 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0181] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0182] In the embodiment of the present application, the memory 1320 included in the mobile phone can store program codes and transmit the program codes to the processor.
[0183] The processor 1380 included in the mobile phone can execute the index processing method based on full-text retrieval provided in the above embodiment according to the instructions in the program codes.
[0184] The embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute the index processing method based on full-text retrieval provided in the above embodiment.
[0185] The embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the index processing device based on full-text retrieval reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the index processing device based on full-text retrieval executes the index processing method provided in various optional implementation manners in the above aspects.
[0186] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiment can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium, and when the program is executed, it executes the steps including the above method embodiment; and the foregoing storage medium can be at least one of the following media: Read-Only Memory (ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.
[0187] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0188] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An index processing method based on full-text retrieval, characterized in that, The method includes: Obtaining the words to be updated in the document to be written, the document identification identifier, and the document category identifier to which the document to be written belongs; the document category identifier is the identifier of the document category set based on the business classification requirements, and the document category identifier is used to construct an index keyword; Based on the word to be updated, the document identification identifier, and the document category identifier, constructing the index keyword and the index value corresponding to the index keyword in the memory to obtain index information, and storing the index information in a key-value pair format; When a preset persistence trigger condition is met, based on the numerical order of the index keyword, persisting the index information stored in the memory into a first index file and storing it on the disk; The constructing the index keyword and the index value corresponding to the index keyword in the memory based on the word to be updated, the document identification identifier, and the document category identifier includes: Constructing a forward index keyword based on the document identification identifier and the document category identifier; Performing serialization processing based on the word to be updated to construct a forward index value corresponding to the forward index keyword; Correspondingly storing the forward index keyword and the forward index value in the memory in the key-value pair format to obtain the forward index information in the index information; Constructing a reverse index keyword based on the word to be updated and the document category identifier; Constructing a reverse index value corresponding to the reverse index keyword based on the document identification identifier; Correspondingly storing the reverse index keyword and the reverse index value in the memory in the key-value pair format to obtain the reverse index information in the index information.
2. The method according to claim 1, wherein The constructing a forward index keyword based on the document identification identifier and the document category identifier includes: Determining a forward prefix in the forward index keyword based on the document category identifier and the forward identifier; Constructing the forward index keyword based on the forward prefix and the document identification identifier; The constructing a reverse index keyword based on the word to be updated and the document category identifier includes: Determining a reverse prefix in the reverse index keyword based on the document category identifier and the reverse identifier; Constructing the reverse index keyword based on the reverse prefix and the word to be updated.
3. The method according to claim 1, wherein The constructing a reverse index value corresponding to the reverse index keyword based on the document identification identifier includes: If the reverse index keyword has a corresponding historical inverted list and the word to be updated is a token to be inserted, inserting the document identification identifier into the historical document identification identifier inverted list in the historical inverted list in an orderly manner to obtain the reverse index value corresponding to the reverse index keyword; If the reverse index keyword has a corresponding historical inverted list and the word to be updated is a token to be deleted, inserting the document identification identifier into the historical document identification identifier inverted list in the historical inverted list in an orderly manner, and adding a deletion mark corresponding to the document identification identifier to the historical deletion list in the historical inverted list to obtain the reverse index value corresponding to the reverse index keyword; If the reverse index keyword does not have a corresponding historical inverted list, encode the inverted list of the word to be updated based on the document identification identifier to obtain the reverse index value corresponding to the reverse index keyword.
4. The method according to claim 1, characterized in that, The obtaining of the word to be updated in the document to be written includes: Performing word segmentation processing on the document to be written to obtain a set of words to be written in the document to be written; If the document identification identifier is already stored, obtain the set of written words corresponding to the document identification identifier from the memory and the disk; Based on the set of words to be written and the set of written words, obtain the word to be updated in the document to be written.
5. The method according to claim 1, characterized in that The disk includes a first file layer and a second file layer. The level of the second file layer is higher than that of the first file layer. The first file layer stores the first index file, and the second file layer stores the second index file; The method further includes: Merging the first index file in the first file layer with the second index file in the second file layer and storing them in the second file layer.
6. The method according to any one of claims 1 to 5, characterized in that, The preset persistent trigger condition includes one or more of the following: The memory usage status reaches the preset memory status, the statistical time after persistence reaches the preset time, or the system restarts.
7. The method according to any one of claims 1-5, characterized in that, When the business classification requirement includes a user classification requirement, the document category identifier includes a user identifier, and the user identifier includes any one or more combinations of a personal identifier, an enterprise identifier, or an account identifier; when the business classification requirement includes a time classification requirement, the document category identifier includes a time identifier, and the time identifier includes any one or more combinations of a date identifier, a cycle identifier, a month identifier, or a year identifier.
8. The method according to claim 3, wherein The method further includes: Obtaining the word to be retrieved in the text to be retrieved and the document category identifier of the document to be retrieved; Based on the word to be retrieved and the document category identifier of the document to be retrieved, constructing a reverse index keyword to be retrieved; Performing a search in the memory and the disk based on the reverse index keyword to be retrieved to obtain a plurality of inverted lists of target document identification identifiers corresponding to the reverse index keyword to be retrieved; According to the priorities of the plurality of inverted lists of target document identification identifiers, performing a merging process on the plurality of inverted lists of target document identification identifiers to obtain a complete inverted list of target document identification identifiers.
9. The method according to claim 8, wherein The performing a search in the memory and the disk based on the reverse index keyword to be retrieved to obtain a plurality of inverted lists of target document identification identifiers corresponding to the reverse index keyword to be retrieved includes: Performing a search in the memory and the disk based on the reverse index keyword to be retrieved to obtain a plurality of target reverse index values corresponding to the reverse index keyword to be retrieved; Analyzing the plurality of target reverse index values to obtain the plurality of inverted lists of target document identification identifiers.
10. An index processing device based on full-text retrieval, characterized in that, The device includes: an obtaining unit, a constructing unit, and a persistent unit; The obtaining unit is configured to obtain the word to be updated in the document to be written, the document identification identifier, and the document category identifier to which the document to be written belongs; the document category identifier is an identifier of a document category set based on a business classification requirement, and the document category identifier is used to construct an index keyword; The building unit is used to build index information by building an index keyword and an index value corresponding to the index keyword in memory based on the word to be updated, the document identification identifier, and the document category identifier, and the index information is stored in a key-value pair format; The persistence unit is used to, when a preset persistence trigger condition is met, persist the index information stored in the memory into a first index file and store it on the disk based on the numerical order of the index keywords; The building unit is used to: Build a forward index keyword based on the document identification identifier and the document category identifier; Perform serialization processing based on the word to be updated to build a forward index value corresponding to the forward index keyword; Correspondingly store the forward index keyword and the forward index value in the memory in the key-value pair format to obtain the forward index information in the index information; Build a reverse index keyword based on the word to be updated and the document category identifier; Build a reverse index value corresponding to the reverse index keyword based on the document identification identifier; Correspondingly store the reverse index keyword and the reverse index value in the memory in the key-value pair format to obtain the reverse index information in the index information.
11. The device according to claim 10, characterized in that, The building unit is used to: Determine a forward prefix in the forward index keyword based on the document category identifier and the forward identifier; Build the forward index keyword based on the forward prefix and the document identification identifier; Determine a reverse prefix in the reverse index keyword based on the document category identifier and the reverse identifier; Build the reverse index keyword based on the reverse prefix and the word to be updated.
12. The device according to claim 10, characterized in that, The building unit is used to: If the reverse index keyword has a corresponding historical inverted list and the word to be updated is a word to be inserted for segmentation, insert the document identification identifier into the historical document identification identifier inverted list in the historical inverted list in an orderly manner to obtain a reverse index value corresponding to the reverse index keyword; If the reverse index keyword has a corresponding historical inverted list and the word to be updated is a word to be deleted for segmentation, insert the document identification identifier into the historical document identification identifier inverted list in the historical inverted list in an orderly manner, and add a deletion mark corresponding to the document identification identifier to the historical deletion list in the historical inverted list to obtain a reverse index value corresponding to the reverse index keyword; If the reverse index keyword does not have a corresponding historical inverted list, encode an inverted list of the word to be updated based on the document identification identifier to obtain a reverse index value corresponding to the reverse index keyword.
13. The device according to claim 10, characterized in that, The obtaining unit is used to: Perform word segmentation processing on the document to be written to obtain a set of words to be written for the document to be written; If the document identification identifier has been stored, obtain a set of words that have been written corresponding to the document identification identifier from the memory and the disk; Obtain the word to be updated for the document to be written based on the set of words to be written and the set of words that have been written.
14. The device according to claim 10, characterized in that, The disk includes a first file layer and a second file layer. The level of the second file layer is higher than that of the first file layer. The first file layer stores the first index file, and the second file layer stores the second index file. The apparatus further includes a merging unit; The merging unit is configured to merge the first index file in the first file layer with the second index file in the second file layer and store the merged file in the second file layer.
15. The device according to any one of claims 10-14, characterized in that, The preset persistence trigger condition includes one or more of the following: The memory usage status reaches a preset memory status, the statistical time after persistence reaches a preset time, or the system is restarted.
16. The device according to any one of claims 10 - 14, characterized in that, When the service classification requirement includes a user classification requirement, the document category identifier includes a user identifier, and the user identifier includes any one or more combinations of a personal identifier, an enterprise identifier, or an account identifier; when the service classification requirement includes a time classification requirement, the document category identifier includes a time identifier, and the time identifier includes any one or more combinations of a date identifier, a cycle identifier, a month identifier, or a year identifier.
17. The device according to claim 12, characterized in that, The apparatus further includes a merging unit; The obtaining unit is further configured to obtain a search term of the text to be retrieved and a document category identifier to be retrieved; The constructing unit is further configured to construct a reverse index keyword to be retrieved based on the search term and the document category identifier to be retrieved; The obtaining unit is further configured to perform a search in the memory and the disk based on the reverse index keyword to be retrieved, and obtain an inverted list of multiple target document identification identifiers corresponding to the reverse index keyword to be retrieved; The merging unit is configured to perform a merging process on the inverted lists of the multiple target document identification identifiers according to the priorities of the inverted lists of the multiple target document identification identifiers, and obtain a complete inverted list of target document identification identifiers.
18. The device according to claim 17, characterized in that, The obtaining unit is further configured to: Perform a search in the memory and the disk based on the reverse index keyword to be retrieved, and obtain multiple target reverse index values corresponding to the reverse index keyword to be retrieved; Parse the multiple target reverse index values to obtain the inverted lists of the multiple target document identification identifiers.
19. A device for index processing based on full-text retrieval, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1-9 based on the instructions in the program code.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1-9.
21. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1-9.
Citation Information
Patent Citations
File searching method and file searching device
CN105279278A
File storage and retrieval method and file storage and retrieval device
CN106294595A