Index storage method and device, retrieval engine, electronic equipment and storage medium

By selecting the memory or disk storage location according to the data volume of the index file, the shortcomings of existing retrieval engines in query speed and cost are solved, and more efficient index query and storage optimization are achieved.

CN114218161BActive Publication Date: 2025-10-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111638246.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-10-17
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

Existing disk-based retrieval engines cannot meet the needs of high-efficiency reading and processing, especially in terms of query speed and performance.

Method used

According to the data volume of the index file, the storage location can be flexibly selected, and the memory or disk storage structure can be adopted. The appropriate storage location can be selected according to the data volume of the index file, taking into account both performance and cost.

Benefits of technology

It improves index query speed, reduces storage costs, and achieves performance and cost optimization in different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218161B_ABST
    Figure CN114218161B_ABST
Patent Text Reader

Abstract

The present disclosure provides an index storage method, device, retrieval engine, electronic device and storage medium, and relates to the technical field of big data processing. The method specifically implements the following scheme: processing a target file to obtain an index file of the target file; determining a storage location of the index file based on a data amount of the index file, the storage location of the index file at least including a memory; and storing the index file based on a data storage structure corresponding to the storage location of the index file, the data storage structure at least including a memory data storage structure. Through the above scheme, a suitable storage location can be flexibly selected according to the data amount of the index file, and the processing efficiency is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of big data processing, and particularly relates to an index storage method and device, a search engine, an electronic device and a storage medium. BACKGROUND

[0002] The search engine technology has been a leading technology in the information acquisition field of computer science, and a disk-based search engine, hereinafter referred to as a disk search engine, is usually adopted. However, the disk search engine is increasingly unable to adapt to higher efficiency requirements in read processing, and therefore, how to provide a more efficient search storage method becomes a problem to be solved. SUMMARY

[0003] Embodiments of the present disclosure provide an index storage method, device, search engine, electronic device and storage medium to solve one or more technical problems in the prior art.

[0004] In a first aspect, the present disclosure provides an index storage method, comprising:

[0005] processing a target file to obtain an index file of the target file;

[0006] determining a storage location of the index file based on a data amount of the index file, wherein the storage location of the index file at least includes a memory;

[0007] storing the index file based on a data storage structure corresponding to the storage location of the index file, wherein the data storage structure at least includes a memory data storage structure.

[0008] In a second aspect, the present disclosure provides an index storage area device, comprising:

[0009] an index creation module configured to process a target file to obtain an index file of the target file;

[0010] a storage mode determination module configured to determine a storage location of the index file based on a data amount of the index file, wherein the storage location of the index file at least includes a memory;

[0011] an index storage module configured to store the index file based on a data storage structure corresponding to a storage mode of the index file, wherein the data storage structure at least includes a memory data storage structure.

[0012] In a third aspect, an embodiment of the present disclosure provides a search engine comprising the index formation device provided by the present disclosure.

[0013] In a fourth aspect, an embodiment of the present disclosure provides an electronic device comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected with the at least one processor; wherein

[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by any one of the embodiments of the present disclosure.

[0017] In a fifth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, the computer instructions being used to cause a computer to execute the method provided by any one of the embodiments of the present disclosure.

[0018] Through the above scheme, a suitable storage location can be flexibly selected according to the data amount of the index file, processing efficiency is ensured, and performance and cost can also be taken into account.

[0019] Other effects of the above-mentioned optional mode will be described in the following in conjunction with specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:

[0021] Figure 1 is a flowchart of an index storage method according to an embodiment of the present disclosure;

[0022] Figure 2 is a schematic block diagram of an index storage device according to an embodiment of the present disclosure;

[0023] Figure 3 is a schematic block diagram of an index storage module according to an embodiment of the present disclosure;

[0024] Figure 4 is a schematic block diagram of a search engine according to an embodiment of the present disclosure;

[0025] Figure 5 is a block diagram of an electronic device used to implement the index storage method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help in understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, descriptions of well-known functions and structures are omitted in the following description.

[0027] The computing hardware and the memory have made great breakthroughs in the past decade. Compared with the spinning magnetic disk hard drives of a decade ago, modern servers are often equipped with more than 100G of super large memory and super fast SSD storage. The new hardware facilities make the software development ideas of software developers completely different from a decade ago. More abundant memory and faster read-write disks can help developers fully improve the performance of software in various situations. Of course, the strong new hardware also brings an increase in hardware costs. Developers need to balance the costs of different hardware and use cost-acceptable hardware solutions in appropriate scenarios to design a software system with reasonable performance and cost.

[0028] As mentioned above, Lucene is a disk-based retrieval engine that adopts the way of storing all indexes on disk. Although the cost is low, it cannot meet the needs of business scenarios with higher requirements for query speed or performance. Lucene supports the mmap (memory mapping file) technology to map files to memory to improve file reading efficiency. However, this approach still uses a disk data storage structure, and the reading efficiency is still not high enough compared with using a memory data storage structure.

[0029] Based on the above description, the present disclosure provides an index storage method, device, retrieval engine, electronic device and storage medium, which selects a suitable storage location according to the data amount of an index file in the process of forming an index, thereby balancing performance and cost and customizing for business scenarios.

[0030] Figure 1 is a flowchart of an index storage method according to an embodiment of the present disclosure. As shown in Figure 1 The present disclosure provides an index storage method, comprising the following steps:

[0031] S101, processing a target file to obtain an index file of the target file;

[0032] S102, determining a storage location of the index file based on the data amount of the index file, wherein the storage location of the index file at least includes memory;

[0033] S103, storing the index file based on the data storage structure corresponding to the storage location of the index file, wherein the data storage structure at least includes a memory data storage structure.

[0034] In the present disclosure, the target file is a file to be indexed. The file can be various types of files, such as documents, emails, web pages, etc. The number of target files can be multiple. In the embodiments of the present disclosure, one target file is taken as an example for illustration, but it should be understood that the processing method for each target file is the same, so it is not described one by one.

[0035] In the present disclosure, the target file can be processed in S101 to obtain an index file of the target file by using various known suitable index creation methods. For example, the processing of the target file to obtain the index file of the target file can include the following steps: performing syntax analysis and language processing on the target file to form a series of terms, and forming a dictionary and an inverted index table through index creation.

[0036] In an embodiment of the present disclosure, the storage location of the index file at least includes a memory; further, the storage location of the index file can also include a disk.

[0037] In the present disclosure, the determination of the storage location of the index file based on the data amount of the index file in S102 can include the following steps: determining the data amount of the index file, judging the size of the data amount of the index file and a preset data amount threshold, when the data amount of the index file is less than the preset data amount threshold, determining the storage location of the index file as the memory, and when the data amount of the index file is greater than or equal to the preset data amount threshold, further determining the storage location of the index structure according to the data amount of the index structure contained in the index file.

[0038] That is, in the case that the data amount of the index file is appropriate, the index file can be stored in the memory according to the business needs, so as to improve the index query speed and further improve the search efficiency. The index file can also be stored in the disk according to the business needs, so as to reduce the cost. Or part of the index file is stored in the memory and part of the index file is stored in the disk according to the business needs, so as to balance the performance and the cost.

[0039] As an example, the storage location of the index file can be determined according to whether the data amount of the index file can be placed in one index instance. If the data amount of the index file can be placed in one index instance, the index file is stored in the memory data storage structure as much as possible. If the index file cannot be placed in the memory, the index structure with a large proportion in the index file is placed in the disk to save the memory. For example, the data amount of the business needs to be placed in a single index instance of 100w. If the index file is stored in the memory data storage structure, in the case that the single machine memory can be placed, the index file is stored in the memory data storage structure to achieve the highest search efficiency. In the case that the single machine memory cannot be placed, the index structure such as the inverted zipline is placed in the disk to save the memory.

[0040] In an embodiment of the present disclosure, the determining the storage location of the index file based on the data amount of the index file comprises: in a case where the data amount of the index file is less than a first preset data amount threshold, determining the storage location of the index file as the memory. That is, in a case where the data amount of the index file is less than the first preset data amount threshold, the index file is stored in the memory in the form of the in-memory data storage structure, so that the query speed of the index file can be greatly improved, and the retrieval performance is improved.

[0041] The first preset data amount threshold can be determined according to the size of the single machine memory, for example, can be 10% of the size of the single machine memory.

[0042] The determination manner of the first preset data amount threshold can comprise: obtaining the size of the single machine memory of the retrieval engine and the index file, and determining the size of the first preset data amount threshold according to the size of the single machine memory and the size of the memory required in the index retrieval. That is, the first preset data amount threshold is determined as a criterion that does not affect the required memory in the index retrieval.

[0043] In an embodiment of the present disclosure, the determining the storage location of the index file based on the data amount of the index file can comprise: in a case where the data amount of the index file is greater than or equal to the first preset data amount threshold, determining the storage location corresponding to each index structure of at least two index structures contained in the index file based on the data amount of the at least two index structures. That is, in a case where the data amount of the index file is greater than or equal to the first preset data amount threshold, it is determined whether each index structure is stored in the memory or the disk according to the data amount of the at least two index structures contained in the index file.

[0044] For example, the index structure with a small data amount can be stored in the memory in the in-memory data storage structure, and the index structure with a large data amount can be stored in the disk in the disk data storage structure. In this way, the query efficiency of the index structure with a small data amount is improved, the storage cost of the index structure with a large data amount is reduced, and the memory is not excessively occupied, so that the performance and the cost are considered.

[0045] In an embodiment of the present disclosure, the index file comprises at least two index structures. As an example, the index structures comprise a term dictionary, a frequencies inverted list, and a fields inverted list. It should be understood that the index structures contained in the index file are determined based on the design needs of the retrieval engine, and are not limited to the index structures listed in this example.

[0046] In an embodiment of the present disclosure, in a case where the data amount of the index file is greater than or equal to the first preset data amount threshold, determining the storage location corresponding to each index structure of the at least two index structures included in the index file based on the data amount of the at least two index structures can comprise: in a case where the data amount of an i-th index structure of the at least two index structures included in the index file is less than a second preset data amount threshold, determining the storage location of the i-th index structure as the memory, i being an integer greater than or equal to 1; in a case where the data amount of the i-th index structure of the at least two index structures included in the index file is greater than or equal to the second preset data amount threshold, determining the storage location of the i-th index structure as the disk. That is, the index structure with a data amount less than the second preset data amount threshold in the index file is stored in the memory in the form of a memory data storage structure, and the index structure with a data amount greater than or equal to the second preset data amount threshold in the index file is stored in the disk in the form of a disk data storage structure. In this way, the query efficiency of the small data amount index structure can be improved, and the storage cost of the large data amount index structure can be reduced.

[0047] The i-th index structure is any one of the at least two index structures included in the index file, that is, the processing mode for each index structure can be the same, and thus will not be described one by one.

[0048] The second preset data amount threshold can be the same as or different from the first preset data amount threshold. In an embodiment of the present disclosure, the second preset data amount threshold is less than the first preset data amount threshold. The second preset data amount threshold can be determined based on the size of the single machine memory, for example, can be 5% of the size of the single machine memory. The second preset data amount threshold is subject to the memory requirement when the index is retrieved. The second preset data amount threshold can also be subject to the first preset data amount threshold and the number of index structures included in the index file. For example, the second data amount threshold can be determined based on the quotient of the first preset data amount threshold and the first preset data amount threshold.

[0049] In an embodiment of the present disclosure, in a case where the data amount of the index file is greater than or equal to the first preset data amount threshold, determining the storage location corresponding to each index structure of the at least two index structures included in the index file based on the data amount of the at least two index structures can comprise: in a case where the data amount of an i-th index structure of the at least two index structures included in the index file is one of the N index structures with the smallest data amount, determining the storage location of the i-th index structure as the memory, N being an integer greater than or equal to 1; in a case where the data amount of the i-th index structure of the at least two index structures included in the index file is not one of the N index structures with the smallest data amount, determining the storage location of the i-th index structure as the disk. That is, the storage location of an index structure is determined according to the ranking of the data amount of the index structure. If the ranking of the data amount of an index structure is the smallest, the storage location of the index structure is determined as the memory, and vice versa. In this way, the query efficiency of a small data amount index structure can be improved, and the storage cost of a large data amount index structure can be reduced.

[0050] It should be understood that the index structure with the smallest data amount can be one or multiple. When the index structure with the smallest data amount is multiple, the multiple index structures are all stored in the memory as the memory data storage structure. In other words, the i-th index structure is any one of the at least two index structures included in the index file, that is, the processing mode of each index structure can be the same, and therefore will not be described in detail.

[0051] As an example, a small-scale data business scenario is taken as an example. In this business scenario, the index field length is small, the number of terms cut out is small, but the number of other non-retrieval fields of the business is large, and each field can be a long string. In this scenario, the disk space required by the dictionary and the inverted list in the index file is much smaller than the disk space required by the document field, and therefore the two index data of the dictionary and the inverted list can be stored in the memory as the memory data storage structure to greatly improve the efficiency of the inverted index list merging, and the document field index data can be stored in the disk as the disk data storage structure to reduce the storage cost and the memory occupation.

[0052] It should be understood that when the storage location of the index file is determined according to the sorting of the data amount of the index structure, the data amount can also be sorted into a preset level of data structure for storage in a memory data storage structure. For example, the index file includes four index structures A, B, C and D, and the data amount of the four index structures is sorted from B, C, D, A. In an embodiment, the last two index structures D and A in the sorting can be stored in a memory data storage structure, and the first two index structures B and C in the sorting can be stored in a disk data storage structure.

[0053] In an embodiment of the present disclosure, the storage location of the index file corresponds to a data storage structure; for example, when the storage location of the index file includes memory, the data storage structure corresponding to the storage location includes at least a memory data storage structure; when the storage location of the index file includes a disk, the data storage structure corresponding to the storage location can include a disk data storage structure.

[0054] In S103, the index file is stored based on the data storage structure corresponding to the storage location of the index file, which can include: when it is determined that the storage location of the index file is memory, the index file is stored in memory in a memory data storage structure; when it is determined that the storage location of the index file is a disk, the index file is stored in the disk in a disk data storage structure.

[0055] In an embodiment of the present disclosure, the index file includes at least two index structures, and when the index file is stored based on the data storage structure corresponding to the storage location of the index file, each index structure is stored in a data storage structure corresponding to each index structure, thereby realizing the storage of the index file.

[0056] It should be understood that the type and number of index structures included in the index file are determined based on the search engine. That is, when the search engine is designed, the type and number of index structures included in the index file formed by the search engine have been determined.

[0057] As an example, when the data amount of the index file is less than a first preset data amount threshold, and it is determined that the storage location of the index file is memory, each index structure included in the index file is stored in memory according to a memory data storage structure of the index structure.

[0058] As another example, in a case where the data amount of the index file is greater than or equal to a first preset data amount threshold, part of the index structures in the index file (for example, the data amount is less than a second preset threshold or the data amount is ranked as the minimum) determine the storage location as the memory, and part of the index structures (for example, the data amount is greater than or equal to the second preset threshold or the data amount is not ranked as the minimum) determine the storage location as the disk, then for the part of the index structures that determine the storage location as the memory, the index structures are stored in the memory according to the memory data storage structure of the index structures; and for the part of the index structures that determine the storage location as the disk, the index structures are stored in the disk according to the disk data storage structure of the index structures.

[0059] In an embodiment of the present disclosure, in order to facilitate the storage of each index structure, a virtual base class of each index structure can be generated in advance, and the memory data storage structure and the disk data storage structure of the index structure can be generated based on the virtual base class. In this way, after the storage location of the index structure is determined, the data storage structure corresponding to the storage location of the index structure can be directly called to store the index structure, thereby improving the storage efficiency.

[0060] In an embodiment of the present disclosure, the memory data storage structure and the disk data storage structure of the index structure are generated based on the virtual base class of the index structure and the corresponding data storage structure. The virtual base class of the index structure represents the general data storage structure of the index structure after the index structure is abstracted. After the virtual base class of the index structure is obtained, the memory data storage structure of the index structure is generated based on the virtual base class of the index structure and the memory data structure, and the disk data storage structure of the index structure is generated based on the virtual base class of the index structure and the disk data structure. Taking the inverted zipline as an example, the memory data storage structure of the inverted zipline is generated based on the virtual base class of the inverted zipline and the memory data structure. Illustratively, the memory data storage structure of the inverted zipline is a dictionary type data structure, the key is term, and the value is the zipline array of document ids. The disk data storage structure of the inverted zipline is based on the kv storage engine, the key is term, and the value is the zipline array after serialization, which is deserialized into the actual zipline array after being read from the disk.

[0061] As an example, the index structures include a term store, an inverted list store, and a document field. In this example, virtual base classes for the term store, the inverted list store, and the document field can be pre-generated: a term store virtual base class (TermStore), an inverted list store virtual base class (ListStore), and a document field virtual base class (FieldStore), as well as in-memory data store structures and disk data store structures for each index structure, such as a RAM TermStore and a DISK TermStore for the term store, a RAM ListStore and a DISK ListStore for the inverted list store, and a FieldListStore and a DISFieldStore for the document field. As an example, in the present disclosure, the polymorphic capabilities of the c++ language can be utilized to abstract each index structure separately, with different index structures extending the specific implementation of the different index structures on different media through polymorphism.

[0062] An example code for the term store virtual base class (TermStore) is as follows:

[0063] class Termstore{

[0064] public:

[0065] TermStore();

[0066] Virtual ~TermStore();

[0067] Virtual void func()=0;

[0068] };

[0069] The term store virtual base class Termstore is defined by Class, the class is accessible by the outside by public, the virtual functions of Termstore are defined by Virtual, and the pure virtual functions of Termstore, which must be overridden in the derived class, are defined by Virtual void func()=0.

[0070] Similarly, example codes for the inverted list store virtual base class (ListStore) and the document field virtual base class (FieldStore) are as follows:

[0071] class ListStore{

[0072] public:

[0073] ListStore();

[0074] Virtual~ListStore();

[0075] Virtual void func() = 0;

[0076] };

[0077] class FieldStore{

[0078] public:

[0079] FieldStore();

[0080] Virtual~FieldStore();

[0081] Virtual void func() = 0;

[0082] };

[0083] The sample code of the dictionary's memory data storage structure RAMTermStore is as follows:

[0084] class RAMTermstore: public TermStore{

[0085] public:

[0086] RAMTermStore();

[0087] Virtual~RAMTermStore();

[0088] Virtual void func(){}

[0089] };

[0090] The memory data storage structure RAMTermStore of the dictionary is derived from the dictionary virtual base class Termstore through Class definition. The class is limited to be accessible from the outside through public, the virtual function of RAMTermstore is limited through Virtual, and the virtual function of RAMTermstore is limited through Virtual void func(), which can be not overridden in the inheritance of derived classes.

[0091] Similarly, the dictionary's disk data storage structure DISKTermStore, the inverted zipper's memory data storage structure RAMListStore and disk data storage structure DISKListStore, and the document field's memory data storage structure FieldListStore and disk data storage structure DISFieldStore can be implemented in the above manner.

[0092] When the storage location of the index structure is determined, the data storage structure corresponding to the storage location of the index structure is directly called to store the index structure. For example, in one example, the dictionary and the inverted list are determined to exist in the memory, and the document data is determined to exist in the disk. Then, the data storage structures of RAMTermStore, RAMListStore and DISKFieldStore are directly called to store the dictionary, the inverted list and the document data, respectively.

[0093] Through the above scheme, each index structure has different implementations on different media. For different services, different combinations can be adopted to achieve the optimal combination of performance and cost.

[0094] In an embodiment of the present disclosure, when the data amount of the index file is less than a first preset data amount threshold, the storage location of the index file is determined to be the memory. In this case, storing the index file based on the data storage structure corresponding to the storage location of the index file comprises: generating a virtual base class of the index structure contained in the index file, generating a memory data storage structure of the index structure based on the virtual base class; and calling the data storage structure corresponding to the memory to store the index structure contained in the index file.

[0095] In an embodiment of the present disclosure, the index file comprises at least two index structures, and when the data amount of the index file is greater than or equal to the first preset data amount threshold, the storage location corresponding to each index structure in the at least two index structures is determined based on the data amount of the at least two index structures contained in the index file. In this case, storing the index file based on the data storage structure corresponding to the storage location of the index file comprises:

[0096] When the storage location of the i-th index structure is the memory, a virtual base class of the i-th index structure is generated, a memory data storage structure of the i-th index structure is generated based on the virtual base class, and the i-th index structure is stored based on the memory data storage structure;

[0097] When the storage location of the i-th index structure is the disk, a virtual base class of the i-th index structure is generated, a disk data storage structure of the i-th index structure is generated based on the virtual base class, and the i-th index structure is stored based on the disk data storage structure.

[0098] By the above scheme, a suitable storage location can be flexibly selected according to the data amount of the index file, processing efficiency is ensured, and performance and cost can also be considered. By the above scheme, the search engine can flexibly customize whether to store in the memory or in the disk according to different index data structures, or select other more suitable hardware, so as to make fine-grained adjustment and more flexible compromise in the performance and cost dimensions, so as to better control the cost in different business scenarios.

[0099] Figure 2 is a schematic block diagram of an index forming apparatus according to an embodiment of the present disclosure. As shown in Figure 3 The index forming apparatus 200 provided by the present disclosure includes an index creating module 210, a storage location determining module 220, and an index storing module 230.

[0100] The index creating module 210 is configured to process a target file to obtain an index file of the target file.

[0101] The storage location determining module 220 is configured to determine a storage location of the index file based on a data amount of the index file, wherein the storage location of the index file at least includes a memory.

[0102] The index storing module 230 is configured to store the index file based on a data storage structure corresponding to a storage manner of the index file, wherein the data storage structure at least includes a memory data storage structure.

[0103] The storage location determining module 220 is configured to determine the storage location of the index file as the memory when the data amount of the index file is less than a first preset data amount threshold.

[0104] The index storing module 230 is configured to call a memory data storage structure of an index structure included in the index file, and store the index structure included in the index file into the memory.

[0105] The storage location determining module 220 is configured to determine a storage location corresponding to each index structure of at least two index structures included in the index file based on data amounts of the at least two index structures when the data amount of the index file is greater than or equal to the first preset data amount threshold.

[0106] The storage location determining module is configured to determine the storage location of the i-th index structure as the memory when a data amount of the i-th index structure of the at least two index structures included in the index file is less than a second preset data amount threshold, wherein i is an integer greater than or equal to 1.

[0107] In a case where a data amount of an i-th index structure among the at least two index structures included in the index file is greater than or equal to a second preset data amount threshold, the storage location of the i-th index structure is determined to be the disk.

[0108] In a case where a data amount of an i-th index structure among the at least two index structures included in the index file is ranked as one of the minimum N index structures, the storage location of the i-th index structure is determined to be the memory, N is an integer greater than or equal to 1, and i is an integer greater than or equal to 1.

[0109] In a case where a data amount of an i-th index structure among the at least two index structures included in the index file is not ranked as one of the minimum N index structures, the storage location of the i-th index structure is determined to be the disk.

[0110] The index storage module is configured to:

[0111] In a case where the storage location of the i-th index structure is the memory, the memory data storage structure of the i-th index structure is called to store the i-th index structure in the memory.

[0112] In a case where the storage location of the i-th index structure is the disk, the disk data storage structure of the i-th index structure is called to store the i-th index structure in the disk.

[0113] Figure 3 FIG. 2 is a schematic block diagram of an index storage module according to an embodiment of the present disclosure. As shown in FIG. 2, the index storage module 230 includes a storage structure generation unit 310 and a storage unit 320. Figure 3

[0114] The storage structure generation unit 310 is configured to pre-generate a virtual base class of an index structure included in an index file, and generate a data storage structure corresponding to a storage location of the index structure based on the virtual base class.

[0115] The storage unit 320 is configured to call the data storage structure corresponding to the storage location to store the index structure included in the index file.

[0116] The type and quantity of the index structure included in the index file are determined based on a search engine. The storage structure generation unit 310 is configured to pre-generate a data storage structure of each index structure.

[0117] ​In an embodiment of the present disclosure, in a case where the data amount of the index file is less than a first preset data amount threshold, and the storage location of the index file is determined to be the memory, the storage unit 320 is configured to invoke a preset memory data storage structure of the index structure included in the index file, and store the index structure included in the index file into the memory.

[0118] In an embodiment of the present disclosure, in a case where the data amount of the index file is greater than or equal to the first preset data amount threshold, in a case where the storage location of the ith index structure is the memory, the storage unit 320 is configured to invoke a preset memory data storage structure of the ith index structure to store the ith index structure into the memory; and in a case where the storage location of the ith index structure is the disk, the storage unit 320 is configured to invoke a preset disk data storage structure of the ith index structure to store the ith index structure into the disk.

[0119] Figure 4 is a schematic block diagram of a search engine according to an embodiment of the present disclosure. As shown in Figure 4 The search engine 400 provided by the present disclosure includes a file collection module 410, an index creation module 420, a storage location determination module 430, an index storage module 440, an index library 450, a query module 460, and an index search module 470.

[0120] The file collection module 410 is configured to collect target files to be indexed. The target files can be obtained from web pages, databases, file systems, or manually input.

[0121] The index creation module 420 is configured to process the target files to obtain index files of the target files.

[0122] The storage location determination module 430 is configured to determine a storage location of the index file based on a data amount of the index file, wherein the storage location of the index file at least includes a memory.

[0123] The index storage module 440 is configured to store the index file based on a data storage structure corresponding to the storage location of the index file, wherein the data storage structure at least includes a memory data storage structure.

[0124] The index library 450 is configured to store the index files of the target files.

[0125] The query module 460 is configured to receive a query input of a user and return a query result to the user.

[0126] The index searching module 470 is configured to search the index file of the index library according to the query input of the user, so as to obtain the file related to the query input of the user and return the file to the user in a relevance order.

[0127] In the disclosed embodiment, the search engine 400 is a full-text search engine.

[0128] The functions of each module in each device in the embodiments of the present disclosure can be referred to the corresponding description in the above method, which will not be repeated here.

[0129] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0130] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0131] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0132] As shown in Figure 5 The device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage unit 508. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0133] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through computer networks, such as the Internet, and / or various telecommunication networks.

[0134] The computing unit 501 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, such as the network segment information processing method. For example, in some embodiments, the network segment information processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded to the RAM 503 and executed by the computing unit 501, one or more steps of the network segment information processing method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the network segment information processing method by any other appropriate means, such as by means of firmware.

[0135] The various implementations of the systems and techniques described above herein can be realized in a digital electronic circuit system, an integrated circuit system, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip system (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0136] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.

[0137] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0138] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0139] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0140] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions typically happening over a communication network. This relationship between a client and a server is created by executing computer programs on the respective computers with the client and server programs interacting across a network. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0141] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without departing from the desired results of the technology disclosed in the present disclosure, and are not limited herein.

[0142] The specific embodiments described above are not intended to be limiting, but rather to illustrate a few possible forms in which the technology described in this disclosure can be implemented. Many modifications, combinations, sub-combinations and alternatives are possible. Any of the disclosed features can be used in combination with any other disclosed feature. Any of the disclosed features can be used without combination with any other disclosed feature. Any modifications, equivalents, or alternative embodiments replacing a feature or combination of features are within the spirit and scope of the disclosure.

Claims

1. An index storage method, comprising: Processing the target file to obtain an index file of the target file; Determining a storage location of the index file based on the data volume of the index file, where the storage location of the index file at least includes a memory; Storing the index file based on a data storage structure corresponding to a storage location of the index file, wherein the data storage structure at least includes a memory data storage structure; Wherein, determining the storage location of the index file based on the data volume of the index file includes: if the data volume of the index file is less than a first preset data volume threshold, determining the storage location of the index file as a memory; if the data volume of the index file is greater than or equal to the first preset data volume threshold, determining the storage location corresponding to each of the at least two index structures contained in the index file based on the data volume of the at least two index structures; The first preset data volume threshold is determined according to the memory size of a single machine and the memory size required for index retrieval.

2. The method according to claim 1, wherein The storing of the index file based on the data storage structure corresponding to the storage location of the index file includes: calling a preset memory data storage structure of the index structure included in the index file, and storing the index structure included in the index file into the memory.

3. The method according to claim 1, wherein The determining, based on the data amount of the at least two index structures included in the index file, a storage location corresponding to each of the at least two index structures includes: When the data amount of the i-th index structure among the at least two index structures included in the index file is less than a second preset data amount threshold, determining that the storage location of the i-th index structure is a memory, where i is an integer greater than or equal to 1; When the data volume of the i-th index structure among the at least two index structures included in the index file is greater than or equal to a second preset data volume threshold, the storage location of the i-th index structure is determined to be a disk.

4. The method according to claim 1, wherein The determining, based on the data amount of the at least two index structures included in the index file, a storage location corresponding to each of the at least two index structures includes: When the data amount of the i-th index structure among the at least two index structures included in the index file is ranked as one of the N smallest index structures, determining that the storage location of the i-th index structure is a memory, where N is an integer greater than or equal to 1, and i is an integer greater than or equal to 1; When the data volume ranking of the i-th index structure among the at least two index structures included in the index file is not one of the N index structures with the smallest data volume, the storage location of the i-th index structure is determined to be a disk.

5. The method according to claim 4, wherein The storing of the index file based on the data storage structure corresponding to the storage location of the index file includes: In a case where the storage location of the i-th index structure is a memory, calling a memory data storage structure that pre-sets the i-th index structure to store the i-th index structure in the memory; In a case where the storage location of the i-th index structure is a disk, a preset disk data storage structure of the i-th index structure is called to store the i-th index structure in the disk.

6. An index storage device, comprising: An index creation module, used to process a target file to obtain an index file of the target file; a storage location determining module, configured to determine a storage location of the index file based on the data volume of the index file, wherein the storage location of the index file at least includes a memory; An index storage module, configured to store the index file based on a data storage structure corresponding to a storage location of the index file, wherein the data storage structure at least includes a memory data storage structure; The storage location determination module is further configured to, if the data volume of the index file is less than a first preset data volume threshold, determine that the storage location of the index file is the memory; and, if the data volume of the index file is greater than or equal to the first preset data volume threshold, determine, based on the data volume of the at least two index structures contained in the index file, a storage location corresponding to each of the at least two index structures; The first preset data volume threshold is determined according to the memory size of a single machine and the memory size required for index retrieval.

7. The index storage device according to claim 6, wherein: The index storage module is used to call a preset memory data storage structure of the index structure included in the index file, and store the index structure included in the index file into the memory.

8. The device according to claim 6, wherein The storage location determination module is configured to, when the data volume of the i-th index structure among the at least two index structures included in the index file is less than a second preset data volume threshold, determine that the storage location of the i-th index structure is the memory, where i is an integer greater than or equal to 1; When the data volume of the i-th index structure among the at least two index structures included in the index file is greater than or equal to a second preset data volume threshold, the storage location of the i-th index structure is determined to be a disk.

9. The device according to claim 6, wherein The storage location determination module is configured to determine that the storage location of the i-th index structure is a memory when the data amount of the i-th index structure among the at least two index structures included in the index file is ranked as one of the N index structures with the smallest data amount, where N is an integer greater than or equal to 1 and i is an integer greater than or equal to 1; When the data volume ranking of the i-th index structure among the at least two index structures included in the index file is not one of the N index structures with the smallest data volume, the storage location of the i-th index structure is determined to be a disk.

10. The device according to claim 9, wherein The index storage module is used for: In a case where the storage location of the i-th index structure is a memory, calling a memory data storage structure that pre-sets the i-th index structure to store the i-th index structure in the memory; In a case where the storage location of the i-th index structure is a disk, a preset disk data storage structure of the i-th index structure is called to store the i-th index structure in the disk.

11. A search engine comprising the index storage device according to any one of claims 6 to 10.

12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for data disc settling and computer-readable storage medium

    CN109032517A

  • File management method and device, electronic equipment and storage medium

    CN110147203A

  • Distributed indexing method and system for efficiently querying streaming data based on LSM

    CN113312312A