Information retrieval method, device and equipment and computer readable storage medium

By vectorizing the search text data and building a hierarchical index structure, the problem of not being able to fully utilize the parallel computing power of graphics processing units in the prior art is solved, and more efficient information retrieval is achieved, and performance and efficiency are improved.

CN120386832AActive Publication Date: 2025-07-29INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510874242.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing information retrieval methods cannot fully utilize the parallel computing power of the graphics processing unit, resulting in limited performance improvement.

Method used

By vectorizing the text data to be retrieved, a hierarchical index structure is constructed, and the adjacency table of each vector node contains a fixed number of similar vectors, adapting to the parallel computing architecture of the graphics processing unit, transitioning layer by layer from the coarse-grained index end to fine search, and using the parallel computing power of the GPU.

Benefits of technology

It improves the efficiency and performance of information retrieval, makes full use of the parallel processing capabilities of the graphics processing unit, and reduces storage space and computing overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386832A_ABST
    Figure CN120386832A_ABST
Patent Text Reader

Abstract

The invention discloses an information retrieval method, device and equipment and a computer readable storage medium, and relates to the technical field of natural language process.The method includes the steps that after to-be-retrieved text data is received, vectorization processing is conducted on the to-be-retrieved text data, the storage space and the calculation overhead are greatly reduced, a hierarchical index structure is constructed in advance, and the retrieval efficiency is improved. An adjacency list corresponding to each vector node in a hierarchical index structure is set to contain a fixed preset number of similar vectors, a parallel computing architecture of the graphic processing unit is adapted, parallel information retrieval of the graphic processing unit is facilitated, rough information retrieval from a coarse-grained index end is transited to fine information retrieval layer by layer, and the information retrieval efficiency of the graphic processing unit is improved. The information retrieval process is accelerated, the technical problem that the parallel computing capacity of the image processing unit cannot be fully utilized is solved, and the technical effects that the parallel processing capacity of the image processing unit is fully utilized, the performance is greatly improved, and the information retrieval efficiency is improved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to an information retrieval method, apparatus, device, and computer-readable storage medium. Background Art

[0002] With the continuous progress of artificial intelligence and natural language processing technologies, Retrieval-Augmented Generation (RAG) systems have been widely used in various generation tasks.

[0003] Currently, the commonly used information retrieval method mainly retrieves information by searching for shorter paths. The selection and traversal of neighbor nodes rely on a dynamic graph structure, which is not suitable for the Single Instruction, Multiple Data (SIMD) parallel architecture of the Graphics Processing Unit (GPU). It cannot fully utilize the parallel computing power of the graphics processing unit, resulting in limited performance improvement. Summary of the Invention

[0004] This application provides an information retrieval method, apparatus, device, and computer-readable storage medium to at least solve the problem in the related art that the parallel computing power of the graphics processing unit cannot be fully utilized, resulting in limited performance improvement.

[0005] This application provides an information retrieval method, including: Performing vectorization processing on the received text data to be retrieved to obtain a target query vector; Retrieving the target query vector by using the first layer of indexes at the coarse-grained index end in the pre-constructed hierarchical index structure to obtain a first adjacency list containing a preset number of candidate vector nodes; Performing parallel information retrieval layer by layer starting from the first adjacency list according to the arrangement order of each index layer in the hierarchical index structure to obtain a target retrieval result.

[0006] This application also provides an information retrieval apparatus, including: This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above information retrieval methods when executing the computer program.

[0007] This application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above information retrieval methods.

[0008] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above information retrieval methods.

[0009] Through the present application, after receiving the text data to be retrieved, by performing vectorization processing on the text data to be retrieved, the storage space and computational overhead are greatly reduced. A hierarchical index structure is pre-constructed, and the adjacency list corresponding to each vector node in the hierarchical index structure contains a fixed preset number of similar vectors, which is adapted to the parallel computing architecture of the graphics processing unit, thereby facilitating the parallel information retrieval by the graphics processing unit. By gradually transitioning from a rough information retrieval at the coarse-grained index end to a fine-grained information search, the information retrieval process is accelerated. Therefore, the technical problem of the limited performance improvement caused by the inability to fully utilize the parallel computing power of the image processing unit can be solved, achieving the technical effect that the parallel processing ability of the graphics processing unit is fully utilized, the performance is greatly improved, and the information retrieval efficiency is increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a flowchart of an embodiment of an information retrieval method provided by an embodiment of the present application; Figure 2 It is a flowchart of another embodiment of an information retrieval method provided by an embodiment of the present application; Figure 3 It is a block diagram of the structure of an information retrieval device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0013] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the drawings and specific embodiments.

[0015] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the information retrieval method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0016] The embodiments of this application provide an information retrieval method, and the method is described in detail in combination with the execution process of the information retrieval method.

[0017] See Figure 1 , Figure 1 , which is the implementation flowchart of an information retrieval method provided by the embodiments of this application. The method may include the following steps.

[0018] S101: Perform vectorization processing on the received text data to be retrieved to obtain a target query vector.

[0019] After receiving the text data to be retrieved, since the text data to be retrieved is in text format and occupies a large amount of storage space and has a large computational overhead, the received text data to be retrieved is vectorized to obtain a target query vector, thereby greatly reducing the storage space and computational overhead.

[0020] S102: Use the first layer index of the coarse-grained index end in the pre-constructed hierarchical index structure to retrieve the target query vector to obtain a first adjacency list containing a preset number of candidate vector nodes.

[0021] Pre-construct a hierarchical index structure in which each index layer is sorted according to coarse and fine granularities. Each layer of the index corresponds to a different granularity. For example, the top-layer index can be set to provide fast and rough search, and the bottom-layer index can be set to provide fine search. And set that the adjacency list corresponding to each vector node in the hierarchical index structure contains a fixed preset number of similar vector nodes. For example, based on the similarity or distance between vector nodes, an approximate algorithm is used to select the most relevant vector nodes as neighbors to construct an efficient adjacency list. After vectorizing the text data to be retrieved to obtain a target query vector, use the first layer of the index at the coarse-grained index end in the pre-constructed hierarchical index structure to retrieve the target query vector, and obtain a first adjacency list containing a preset number of candidate vector nodes, thereby obtaining a preset number of coarse-grained candidate vector nodes.

[0022] It should be noted that the preset number can be set and adjusted according to the actual situation, and the embodiments of the present application do not limit this. For example, it can be set to 5.

[0023] S103: Perform parallel information retrieval layer by layer starting from the first adjacency list according to the arrangement order of each index layer in the hierarchical index structure to obtain the target retrieval result.

[0024] After obtaining the first adjacency list containing a preset number of candidate vector nodes, perform parallel information retrieval layer by layer starting from the first adjacency list according to the arrangement order of each index layer in the hierarchical index structure to obtain the target retrieval result. For example, the warp structure of the GPU can be used to perform parallel calculations on multiple candidate vector nodes of a single retrieval task, thereby accelerating the similarity calculation, realizing batch queries, and making full use of the computing power of the GPU. The Compute Unified Device Architecture (CUDA) programming model can also be used to distribute multiple computing tasks in the retrieval process to multiple threads, thereby realizing the parallel execution of multiple retrieval tasks.

[0025] Through the present application, since after receiving the text data to be retrieved, by vectorizing the text data to be retrieved, the storage space and computing overhead are greatly reduced. A hierarchical index structure is pre-constructed, and it is set that the adjacency list corresponding to each vector node in the hierarchical index structure contains a fixed preset number of similar vectors, which adapts to the parallel computing architecture of the graphics processing unit, so as to facilitate the graphics processing unit to perform parallel information retrieval. By gradually transitioning from rough information retrieval at the coarse-grained index end to fine information search, the information retrieval process is accelerated. Therefore, the technical problem of limited performance improvement caused by the inability to fully utilize the parallel computing power of the image processing unit can be solved, and the technical effect of fully utilizing the parallel processing ability of the graphics processing unit, greatly improving the performance, and improving the information retrieval efficiency is achieved.

[0026] See Figure 2 , Figure 2 which is the implementation flowchart of another information retrieval method provided by the embodiments of this application. This method may include the following steps.

[0027] S201: Use a pre-trained language model to convert the text data to be retrieved into an initial query vector.

[0028] A language model for converting text data into vectors can be pre-trained in advance, and the pre-trained language model is used to convert the text data to be retrieved into an initial query vector. The initial query vector is generally a vector with a relatively high dimension.

[0029] The pre-trained language model can be BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), etc.

[0030] Before converting the text data to be retrieved into an initial query vector, preprocessing operations such as text cleaning, word segmentation, stop word removal, and lemmatization can also be performed on the text data to be retrieved, so as to improve the text quality.

[0031] S202: Perform dimensionality reduction processing on the initial query vector to obtain a target query vector.

[0032] After converting the text data to be retrieved into an initial query vector, perform dimensionality reduction processing on the initial query vector. Use methods such as Principal Component Analysis (PCA) or random projection to reduce the high-dimensional initial query vector to a lower dimension to obtain a target query vector. By performing dimensionality reduction processing on the initial query vector, compression of the high-dimensional vector is achieved, thereby reducing the video memory occupancy, reducing the calculation and storage overhead, and improving the retrieval efficiency.

[0033] After performing dimensionality reduction processing on the initial query vector to obtain a target query vector, vector compression can also be performed on the target query vector. For example, convert the target query vector represented by floating-point numbers to half-precision floating-point numbers (FP16), so as to reduce the storage space of the vector, and further reduce the video memory occupancy.

[0034] The Product Quantization (PQ) method can also be used to divide the target query vector into multiple sub-vectors, cluster each sub-vector, and replace the original sub-vector with the index of the cluster center to achieve efficient compression.

[0035] Vector quantization and compression may affect the retrieval accuracy. In practical applications, it is necessary to adjust the quantization parameters according to specific requirements and balance the compression ratio and retrieval accuracy.

[0036] When the video memory is insufficient, video memory paging technology can also be adopted to temporarily store some data in a high-speed storage medium such as a solid state drive (SSD) and quickly load it when needed.

[0037] S203: Use the first layer of indexes at the coarse-grained index end in the pre-built hierarchical index structure to retrieve the target query vector, and obtain a first adjacency list containing a preset number of candidate vector nodes.

[0038] S204: Starting from the first adjacency list, perform parallel information retrieval layer by layer in the order of the arrangement of each index layer in the hierarchical index structure, and obtain second adjacency lists corresponding to other index layers except the first layer of indexes in the hierarchical index structure.

[0039] Among them, each second adjacency list contains a preset number of candidate vector nodes.

[0040] After retrieving the first adjacency list containing a preset number of candidate vector nodes, starting from the first adjacency list, perform parallel information retrieval layer by layer in the order of the arrangement of each index layer in the hierarchical index structure, and obtain second adjacency lists corresponding to other index layers except the first layer of indexes in the hierarchical index structure. By gradually transitioning from rough information retrieval at the coarse-grained index end to fine information search, the information retrieval process is accelerated.

[0041] S205: Determine the first layer of indexes at the coarse-grained index end in the hierarchical index structure as the top-level index.

[0042] The index layers in the hierarchical index structure are arranged according to the coarse and fine granularity, and the first layer of indexes at the coarse-grained index end in the hierarchical index structure is determined as the top-level index.

[0043] S206: Obtain the second adjacency list corresponding to the bottom-level index in the retrieved hierarchical index structure.

[0044] The index layers in the hierarchical index structure are arranged according to the coarse and fine granularity, with the top level being the coarse-grained index end and the bottom level being the fine-grained index end. Obtain the second adjacency list corresponding to the bottom-level index in the retrieved hierarchical index structure. The second adjacency list corresponding to the retrieved bottom-level index also contains a preset number of candidate vector nodes.

[0045] S207: Calculate the similarity between each candidate vector node in the second adjacency list corresponding to the bottom-level index and the target query vector respectively, and sort the calculated similarities in descending order.

[0046] After obtaining the second adjacency list corresponding to the bottom layer index in the retrieved hierarchical index structure, calculate the similarity between each candidate vector node in the second adjacency list corresponding to the bottom layer index and the target query vector, and sort the calculated similarities in terms of magnitude. For example, the candidate vector nodes can be sorted in descending order of similarity or in ascending order of similarity.

[0047] S208: Determine the candidate vector node corresponding to the largest similarity in the sorting result obtained by sorting the similarities in terms of magnitude as the target retrieval result.

[0048] After sorting the calculated similarities in terms of magnitude, determine the candidate vector node corresponding to the largest similarity in the sorting result obtained by sorting the similarities in terms of magnitude as the target retrieval result. By sorting the similarities between each candidate vector node and the target query vector in terms of magnitude and determining the candidate vector node corresponding to the largest similarity in the sorting result as the target retrieval result, fine-grained retrieval of information is achieved, improving the accuracy of the retrieved information.

[0049] S209: Combine the target retrieval result with the pre-trained generation model to generate the target response result.

[0050] Pre-train a generation model for information retrieval. After determining the candidate vector node corresponding to the largest similarity in the sorting result obtained by sorting the similarities in terms of magnitude as the target retrieval result, combine the target retrieval result with the pre-trained generation model to generate the target response result. By combining the target retrieval result with the pre-trained generation model, rich context is provided for the generation model, enabling the generation of high-quality answers or content and improving the relevance and accuracy of the generated content.

[0051] The generation model can also be fine-tuned or optimized to better utilize the retrieved information and generate high-quality content. The generation model can also be compressed and pruned to reduce the number of model parameters and the video memory occupancy. According to the needs of the generation process, the model parameters can be loaded in segments to avoid loading the entire model at once and reduce the video memory pressure.

[0052] The combination method of the target retrieval result and the generation model may include combining through input fusion, using the target retrieval result as the input of the generation model, and directly embedding it into the input layer of the generation model. For example, splicing the target retrieval result behind the target query vector as the input of the generation model. It may also include combining through feature fusion, converting the target retrieval result into a feature vector and fusing it with the internal features of the generation model. For example, encoding the target retrieval result into a vector and then adding or splicing it with the hidden layer features of the generation model. In addition, the target retrieval result and the generation model can be combined through conditional generation, multimodal fusion, etc., and the embodiments of the present application do not limit this.

[0053] The embodiments of the present application solve the problems of high video memory occupancy, slow index access speed in the GPU environment, and inability to meet the high-efficiency retrieval requirements of the Retrieval-Augmented Generation (RAG) system for large-scale knowledge bases, optimize the use of video memory, improve the index access speed, and meet the requirements of the RAG system for real-time and large-scale data processing.

[0054] The RAG system has relatively high requirements for the real-time and relevance of the index. The hierarchical index structure should support dynamic updates to facilitate the timely addition of new knowledge to the system and keep the knowledge base up-to-date.

[0055] It is also possible to set a task scheduling strategy, reasonably allocate computing resources according to the urgency and importance of the retrieval request to ensure the timely response of high-priority tasks. In the environment of multiple GPUs or multiple nodes, evenly distribute tasks to avoid over-concentration or idleness of resources and improve the overall efficiency of the system. It is also possible to monitor the resource usage of the system, including GPU video memory, memory, computing load, etc., and dynamically adjust resource allocation. It is also possible to use historical data and the current state to predict the resource usage trend, perform resource scheduling or expansion in advance, and prevent the system from overloading.

[0056] In a specific implementation manner of the present application, the method may further include the following steps: Step 1: Obtain the access frequencies corresponding to each vector node in the hierarchical index structure; Step 2: Divide each vector node into a hot node set and a cold node set according to each access frequency; Step 3: Store the hot node set in the video memory of the graphics processing unit, and store the cold node set in other storage media except the video memory of the graphics processing unit.

[0057] For the convenience of description, the above three steps can be combined for explanation.

[0058] The information retrieval method provided by the embodiments of the present application can also obtain the access frequencies corresponding to each vector node in the hierarchical index structure, divide each vector node into a hot node set and a cold node set according to the access frequencies, store the hot node set in the video memory of the graphics processing unit, and store the cold node set in other storage media except the video memory of the graphics processing unit. By storing the frequently accessed data in the GPU video memory and the infrequently accessed data in other storage media (such as the host memory) according to the access frequency of the data, hierarchical storage of data is achieved, the utilization rate of the video memory is improved, and thus the information retrieval efficiency is improved.

[0059] In a specific implementation manner of the present application, after storing the hot node set in the video memory of the graphics processing unit and storing the cold node set in other storage media except the graphics processing unit, the method may further include the following steps: Step 1: Update the hot node set and the cold node set according to the access frequencies corresponding to each vector node in the hierarchically indexed structure collected dynamically, to obtain an update result; Step 2: Perform migration adjustment on each vector node in the video memory of the graphics processing unit and other storage media according to the update result.

[0060] For ease of description, the above two steps can be combined for explanation.

[0061] After storing the hot node set in the video memory of the graphics processing unit and storing the cold node set in other storage media except the graphics processing unit, update the hot node set and the cold node set according to the access frequencies corresponding to each vector node in the hierarchically indexed structure collected dynamically, to obtain an update result, and perform migration adjustment on each vector node in the video memory of the graphics processing unit and other storage media according to the update result. By updating the hot node set and the cold node set according to the access frequencies corresponding to each vector node, it is possible to dynamically load the required data from other storage media into the video memory of the graphics processing unit or unload the idle data from the video memory of the graphics processing unit to other storage media according to the real-time load and query requirements of the system, thereby optimizing the use of the video memory, reducing the pressure on the video memory of the graphics processing unit, further improving the utilization rate of the video memory, and improving the information retrieval efficiency. The data transmission path between the graphics processing unit and the Central Processing Unit (CPU) is optimized, and the data transmission bottleneck is reduced.

[0062] In a specific implementation manner of the present application, performing migration adjustment on each vector node in the video memory of the graphics processing unit and other storage media according to the update result may include the following steps: The asynchronous data copy mechanism of the unified computing device architecture migrates and adjusts each vector node in the graphics processing unit video memory and other storage media according to the update result.

[0063] When migrating and adjusting each vector node in the graphics processing unit video memory and other storage media, the asynchronous data copy mechanism of the unified computing device architecture migrates and adjusts each vector node in the graphics processing unit video memory and other storage media according to the update result. By using the asynchronous data copy mechanism of the unified computing device architecture, the overlap of data transmission and calculation is achieved, and the blocking of calculation by data transmission is reduced.

[0064] It is also possible to prefetch the data that may be needed in advance according to the query prediction and access pattern, reducing the latency of data loading.

[0065] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0066] The embodiment of the present application also provides an information retrieval device.

[0067] See Figure 3 , Figure 3 which is a structural block diagram of an information retrieval device provided by an embodiment of the present application. The device may include: A target query vector obtaining module 31, configured to perform vectorization processing on the received text data to be retrieved to obtain a target query vector; An adjacency list obtaining module 32, configured to retrieve the target query vector by using the first layer index of the coarse-grained index end in the pre-constructed hierarchical index structure to obtain a first adjacency list including a preset number of candidate vector nodes; A retrieval result obtaining module 33, configured to perform parallel information retrieval layer by layer starting from the first adjacency list according to the arrangement order of each index layer in the hierarchical index structure to obtain a target retrieval result.

[0068] Through this application, after receiving the text data to be retrieved, by vectorizing the text data to be retrieved, the storage space and computational overhead are greatly reduced. A hierarchical index structure is pre-constructed, and each adjacency list corresponding to a vector node in the hierarchical index structure contains a fixed preset number of similar vectors, which is adapted to the parallel computing architecture of the graphics processing unit, thus facilitating the graphics processing unit to perform parallel information retrieval. By gradually transitioning from a rough information retrieval from the coarse-grained index end to a fine information search, the information retrieval process is accelerated. Therefore, the technical problem of limited performance improvement caused by the inability to fully utilize the parallel computing power of the image processing unit can be solved, achieving the technical effect of fully utilizing the parallel processing power of the graphics processing unit, greatly improving the performance, and enhancing the information retrieval efficiency.

[0069] In a specific embodiment of this application, the device may further include: An access frequency acquisition module, configured to acquire the access frequencies respectively corresponding to each vector node in the hierarchical index structure; A hot and cold node set partitioning module, configured to partition each vector node into a hot node set and a cold node set according to each access frequency; A node set storage module, configured to store the hot node set in the graphics processing unit video memory and store the cold node set in other storage media except the graphics processing unit video memory.

[0070] In a specific embodiment of this application, the device may further include: An update result obtaining module, configured to update the hot node set and the cold node set according to the access frequencies respectively corresponding to each vector node in the dynamically collected hierarchical index structure to obtain an update result; A vector node migration and adjustment module, configured to perform migration and adjustment on each vector node in the graphics processing unit video memory and other storage media according to the update result.

[0071] In a specific embodiment of this application, the vector node migration and adjustment module is specifically a module that uses the asynchronous data copy mechanism of the unified computing device architecture to perform migration and adjustment on each vector node in the graphics processing unit video memory and other storage media according to the update result.

[0072] In a specific embodiment of this application, the target query vector obtaining module 31 may include: An initial query vector obtaining sub-module, configured to convert the text data to be retrieved into an initial query vector by using a pre-trained language model; A target query vector obtaining sub-module, configured to perform dimensionality reduction processing on the initial query vector to obtain a target query vector.

[0073] In a specific embodiment of this application, the device may further include:

[0074] The response result generation module is configured to, after obtaining the target retrieval result, combine the target retrieval result with a pre-trained generation model to generate a target response result.

[0075] In a specific embodiment of the present application, the retrieval result obtaining module 33 may include: The second adjacency list obtaining sub-module is configured to perform parallel information retrieval layer by layer starting from the first adjacency list according to the arrangement order of each index layer in the hierarchical index structure, and obtain second adjacency lists respectively corresponding to other index layers except the first index layer in the hierarchical index structure; wherein, each second adjacency list contains a preset number of candidate vector nodes; The top-level index determining sub-module is configured to determine the first index at the coarse-grained index end in the hierarchical index structure as the top-level index; The second adjacency list acquisition sub-module is configured to acquire the second adjacency list corresponding to the bottom index in the retrieved hierarchical index structure; The similarity magnitude sorting sub-module is configured to calculate the similarity between each candidate vector node in the second adjacency list corresponding to the bottom index and the target query vector respectively, and sort the calculated similarities by magnitude; The retrieval result determining sub-module is configured to determine the candidate vector node corresponding to the maximum similarity in the sorting result obtained by sorting the similarities by magnitude as the target retrieval result.

[0076] For the description of the features in the corresponding embodiment of the information retrieval device, reference may be made to the relevant description in the corresponding embodiment of the information retrieval method, which will not be elaborated here one by one.

[0077] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned embodiments of the information retrieval method.

[0078] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above-mentioned embodiments of the information retrieval method when running.

[0079] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0080] Embodiments of the present application further provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned information retrieval method embodiments are implemented.

[0081] Embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned information retrieval method embodiments are implemented.

[0082] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0083] The above has introduced in detail an information retrieval method, apparatus, device, and computer-readable storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. An information retrieval method, characterized in that: include: Perform vectorization processing on the received text data to be retrieved to obtain the target query vector; Retrieving the target query vector using a first-level index at a coarse-grained index end in a pre-built hierarchical index structure to obtain a first adjacency list containing a preset number of candidate vector nodes; According to the arrangement order of each index layer in the hierarchical index structure, information is searched in parallel layer by layer starting from the first adjacency table to obtain a target search result.

2. The information retrieval method according to claim 1, wherein: Also includes: Obtaining the access frequency corresponding to each vector node in the hierarchical index structure; Divide each vector node into a hot node set and a cold node set according to each access frequency; The hot node set is stored in a graphics processing unit memory, and the cold node set is stored in a storage medium other than the graphics processing unit memory.

3. The information retrieval method according to claim 2, wherein: After storing the hot node set in the graphics processing unit memory and storing the cold node set in a storage medium other than the graphics processing unit memory, the method further includes: The hot node set and the cold node set are updated according to the dynamically collected access frequencies corresponding to the respective vector nodes in the hierarchical index structure to obtain an update result; Each vector node in the graphics processing unit video memory and other storage media is migrated and adjusted according to the update result.

4. The information retrieval method according to claim 3, wherein Migrating and adjusting each vector node in the graphics processing unit memory and other storage media according to the update result includes: The asynchronous data copy mechanism of the unified computing device architecture is utilized to migrate and adjust each vector node in the graphics processing unit memory and other storage media according to the update result.

5. The information retrieval method according to claim 1, characterized in that, Perform vectorization on the received text data to be retrieved to obtain the target query vector, including: Using a pre-trained language model to convert the text data to be retrieved into an initial query vector; Performing dimensionality reduction processing on the initial query vector to obtain the target query vector.

6. The information retrieval method according to claim 1, wherein After obtaining the target search results, it also includes: The target retrieval result is combined with the pre-trained generation model to generate a target response result.

7. The information retrieval method according to any one of claims 1 to 6, characterized in that The method includes performing parallel information retrieval layer by layer starting from the first adjacency table according to the arrangement order of each index layer in the hierarchical index structure to obtain a target retrieval result, including: Starting from the first adjacency table, and performing parallel information retrieval layer by layer according to the arrangement order of each index layer in the hierarchical index structure, a second adjacency table corresponding to each index layer in the hierarchical index structure except the first index layer is obtained; wherein each second adjacency table contains the preset number of candidate vector nodes; Determine the first layer index at the coarse-grained index end in the hierarchical index structure as the top layer index; Obtaining a second adjacency list corresponding to the bottom-level index in the retrieved hierarchical index structure; Calculating the similarity between each candidate vector node in the second adjacency table corresponding to the bottom-level index and the target query vector, and sorting the calculated similarities; The candidate vector node corresponding to the largest similarity in the sorted results obtained by sorting the similarities is determined as the target retrieval result.

8. An information retrieval device, characterized in that, include: A target query vector obtaining module, configured to perform vectorization processing on the received text data to be retrieved, and obtain a target query vector; An adjacency list obtaining module, configured to retrieve the target query vector by using the first layer index of the coarse-grained index end in the pre-constructed hierarchical index structure, and obtain a first adjacency list including a preset number of candidate vector nodes; A retrieval result obtaining module, configured to perform parallel information retrieval layer by layer starting from the first adjacency list according to the arrangement order of each index layer in the hierarchical index structure, and obtain a target retrieval result.

9. An electronic device, characterized in that: Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the information retrieval method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the information retrieval method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Cold data placement strategy based on frequency correlation under low-power-consumption clustering environment

    CN107728938A

  • Information retrieval method and device based on text extension, electronic equipment and medium

    CN117992573A

  • Log-based storage for different data types in non-volatile memory

    US20190294333A1

Cited By

  • System and method for directly retrieving data in storage medium and electronic equipment

    CN121051138A