Distributed vector storage method based on hierarchical architecture

By constructing vector search indexes in heap memory and compressing them, combining hot and cold value calculations and hot and cold separation mechanisms of data segments, the scalability and real-time query problems of vector storage systems are solved, and efficient and low-latency vector storage is achieved.

CN120336322APending Publication Date: 2025-07-18ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510409764.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing vector storage methods are not scalable when facing the continuous growth of vector scale, occupy a large amount of storage space and cannot meet the real-time query needs.

Method used

A distributed vector storage method based on a hierarchical architecture is adopted to construct vector search indexes in heap memory and compress them, combining hot and cold value calculations and hot and cold separation mechanisms of data segments to realize streaming storage and access of vectors.

Benefits of technology

It provides a vector storage solution with low latency, low storage overhead and high scalability, supporting real-time query and analysis of large-scale vector data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336322A_ABST
    Figure CN120336322A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed vector storage method based on a hierarchical architecture, which comprises the following steps of: firstly, designing a vector representation method and a data organization method oriented to a real-time scene, ensuring the expandability required by vector storage by utilizing the hierarchical architecture, and storing billions of vector data; secondly, a vector compression method is designed, and storage space occupation needed by mass vector data is reduced; in addition, a data cold and hot value calculation method and a cold and hot separation mechanism during vector access are designed, so that high-quality and efficient large-scale vector query is provided. Therefore, efficient, low-delay and high-expandability vector storage support can be provided for real-time query and analysis of large-scale vector data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data storage, and particularly relates to a distributed vector storage method based on a hierarchical architecture. Background Art

[0002] With the rapid development of big data technology and network applications, a large amount of unstructured data is being generated at an extremely high rate, including text, images, audio, etc. Since unstructured data has a high dimension, it is extremely difficult to directly process unstructured data; in recent years, embedding models based on deep learning have become the mainstream solutions for processing unstructured data. These solutions use embedding models to convert unstructured data into high-dimensional vectors, and transform the storage, query, and analysis of unstructured data into operations on high-dimensional vectors, which have a wide range of application scenarios in production practice. How to design a vector storage method to better support upper-layer processing and analysis tasks has become an urgent problem to be solved in the industrial community.

[0003] Currently, with the rise of businesses such as social networks and e-commerce platforms and the rapid development of large language models, vector storage technology is facing new requirements and challenges. First, as the complexity of deep learning models, especially large language models, increases, the scale of vector data is also expanding, which is manifested in the synchronous growth of the number of vectors and the vector dimension; therefore, the vector storage system should have good scalability to cope with the theoretically infinite growth of vector data. Second, the huge amount of data will lead to an increase in users' storage overhead. How to store vectors using less space and reduce users' storage costs is an urgent problem to be solved. Third, the generation speed of vectors is accelerating, and the performance requirements for vector query and analysis are getting higher and higher. Therefore, the vector storage system needs to have real-time processing capabilities to cope with the rapid update and efficient query requirements of massive data.

[0004] The Chinese patent application with the publication number CN118606363A provides a method for retrieving and storing vectors. When storing vectors, the vector data to be stored is first converted into a vector fingerprint to be stored and divided into M sub-fingerprints to be stored. Then, the M sub-fingerprints are respectively stored in the node bitmaps of the pre-established M-layer bitmap structure. When retrieving vectors, the vector to be retrieved is converted into a data fingerprint to be retrieved and divided into M sub-fingerprints to be retrieved. Similar sub-fingerprints corresponding to the respective sub-fingerprints to be retrieved are determined in the node bitmaps, and the similar sub-fingerprints that meet the preset conditions are selected and spliced to filter out the similar data fingerprints, and the vectors corresponding to the similar data fingerprints are determined as the similar vectors of the vector to be retrieved. This method can avoid comparing data fingerprints one by one, reduce the number of comparisons during retrieval, and improve the retrieval speed and efficiency. However, in the scenario of massive data, the vector fingerprints and node bitmaps cannot all be accommodated in the local memory and / or disk, resulting in low scalability of this method. At the same time, as the total number of vectors increases, the number of vectors corresponding to the same data fingerprint also increases, resulting in a decrease in the accuracy of vector queries.

[0005] The Chinese patent application with the publication number CN119226274A provides a data storage method, recovery method, device and storage medium for a vector database. When storing vectors, the warehousing interface of the vector database is called according to the attribute information of the vector data to be stored, and it is judged whether the warehousing interface exists. If the warehousing interface exists, the vector data to be stored is sequentially saved to the local cache database, vector database and persistent database in the way of batch vector addition. When the service is started, it is checked whether it is necessary to perform a recovery operation on the data in the vector database. If a recovery operation is required, data recovery is performed according to the local cache database and / or persistent database. By performing multi-level backups of vectors, this method helps to enhance the security of vector data. However, it is necessary to store multiple copies of vectors without performing space compression operations such as compression. When inserting massive vector data, it is not conducive to controlling the storage cost of users. In addition, this method inserts vectors in a batch addition manner and cannot immediately establish an index for the newly inserted data in the real-time scenario of rapid data update, which will have a greater impact on the throughput rate and latency of real-time queries.

[0006] It can be seen from this that the existing vector storage methods store vectors in the local memory or disk, cannot cope with the continuous growth of the vector scale, and have insufficient scalability. Secondly, the existing vector storage methods directly store the uncompressed original vector data, occupying a large amount of storage space. In addition, the existing vector storage methods are batch storage, that is, a certain amount of data is accumulated, and then indexed and stored in the disk in batches, without optimizing the storage for the characteristics of real-time queries, resulting in low efficiency of large-scale data real-time queries. Summary of the Invention

[0007] In view of the above, the present invention provides a distributed vector storage method based on a hierarchical architecture. The method is based on a distributed architecture of hierarchical storage and can provide streaming vector storage and access with low latency, low storage overhead and high scalability.

[0008] A distributed vector storage method based on a layered architecture comprises the following steps:

[0009] (1) When inserting a vector, a vector search index is built in the heap memory, the inserted vector is cached on the heap memory, and inserted into the vector search index in sequence;

[0010] (2) Each time a fixed number of vectors are inserted, all the vectors and vector search indexes cached on the heap memory are serialized into one data segment, and the cache and vector search indexes on the heap memory are cleared;

[0011] (3) When constructing a data segment, a vector compression algorithm is used to compress the vector, and when reading, a corresponding decompression algorithm is used to decompress the vector;

[0012] (4) All constructed data segments are stored in layers according to timestamps, and each data segment moves from the lowest level of storage medium to the highest level of storage medium over time;

[0013] (5) When performing a vector query operation, the hot and cold values of each data segment are counted, and each data segment is queried in sequence according to the order of the hot and cold values. When the query result reaches the standard specified in advance by the user, the query is terminated early.

[0014] Furthermore, the specific implementation of step (1) is as follows:

[0015] 1.1 For any inserted vector s i , which is expressed as a triplet, namely s i = <id i ,ts i ,v i >, where id i is the unique identifier of the vector, ts i is the timestamp of the vector, v i is a vector value of fixed dimension, and the subscript i represents the index number of the vector;

[0016] 1.2 Create an empty vector search index I in the heap memory, read the index type and creation parameters from the user configuration file, and initialize the vector search index I according to the index type and creation parameters;

[0017] 1.3 Create an empty array B in heap memory, whose type is the same as the inserted vector s i same;

[0018] 1.4 Insert v i into the vector search index I, and at the same time add s i to the end of the array B.

[0019] Furthermore, the specific implementation of the step (2) is as follows: First, create a counting variable count in the heap memory with an initial value of 0; for each inserted vector, increment count, and when count reaches the set threshold, serialize all the vectors cached on the heap memory and the vector search index together into a data segment, reset count to 0, and clear the vector cache and vector search index on the heap memory;

[0020] The data segment consists of a vector area, an index area, a metadata area, and a footnote area. The vector area consists of vector items, offset value items, and the number of vectors. The vector items store all the vectors in the data segment in ascending order of timestamp. The offset value items store the position of each vector in the vector items in ascending order of timestamp. The number of vectors records the number of vectors stored in the vector area. The index area stores the serialized vector search index, and the range of this vector search index is all the vectors in the data segment. The metadata area stores the meta-information of the vectors and indexes in the data segment, including the vector dimension, the size of the vector search index, the type of the vector search index, the parameters of the vector search index, and the timestamp when the data segment is created. The footnote area stores the offsets of the positions of the vector area, the index area, and the metadata area in the data segment.

[0021] Furthermore, the specific implementation of the step (3) is as follows: For all the vectors inserted in the data segment, store the vector value v1 of the first vector s1 in IEEE 754 format. For each dimension in the vector value v i of the subsequent vectors s i , use the Gorilla algorithm to compress and store the compressed representation based on the same dimension in v i-1 ; when reading the vector s i , read from the first vector s1 in the data segment sequentially to s k . If the read vector s k is a compressed representation, then for each dimension in its vector value v i , use the Gorilla algorithm to decompress s i based on the same dimension in v i-1 ; where the subscripts i and k represent the indexes of the vectors, and v i and v i are the vector values of the vectors s i-1 and s i and s i-1 respectively.

[0022] Further, the specific implementation of step (4) is as follows: Let the storage media used for storing data segments be, from the lower level to the higher level, local memory, local disk, and remote disk; when a certain level of storage media is full, move the data segment with the earliest timestamp in that level to the storage media of the higher level; the local memory is the memory on the computer where the data segment is located, the local disk is the disk mounted on the computer where the data segment is located, and the remote disk is the disk mounted on another computer connected to the computer where the data segment is located through a local area network and accessed using a distributed file system.

[0023] Further, the specific implementation of step (5) is as follows: Calculate the cold-hot value of each data segment, which is obtained by weighted calculation of the access frequency of the data segment, the search hit rate on the data segment, the contribution degree of the vector pair to the search result on the data segment, and the freshness of the vector on the data segment, and update it in real time each time the data segment is accessed; when performing a vector query operation, first search for the vectors cached in the heap memory, and then query in each data segment in order from hot to cold according to the cold-hot value; after searching each data segment, judge the maximum distance between the vector result searched on the data segment and the input query target vector. If the maximum distance is less than the distance threshold set by the user, it means that the quality of the currently searched vector result is already better than the quality required by the user, and the query is terminated in advance without accessing the remaining data segments.

[0024] A computer device includes a memory and a processor. A computer program is stored in the memory, and the processor is used to execute the computer program to implement the above-mentioned distributed vector storage method based on a hierarchical architecture.

[0025] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned distributed vector storage method based on a hierarchical architecture.

[0026] Regarding the problem of distributed vector storage in the field of data storage, first, the present invention designs a vector representation method and a data organization method for real-time scenarios, and uses hierarchical storage to ensure the scalability required for vector storage, which can store billions of vector data; secondly, the present invention designs a vector compression method to reduce the storage space occupied by massive vector data; in addition, the present invention also designs a calculation method for data cold-hot values and a mechanism for separating cold and hot during vector access to provide high-quality and efficient large-scale vector queries. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic flowchart of the distributed vector storage method based on a hierarchical architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] To describe the present invention more specifically, the technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0029] As Figure 1 shown, the distributed vector storage method of the present invention based on a hierarchical architecture is specifically implemented as follows:

[0030] Step S11: When inserting a vector, construct a vector search index in the heap memory, cache the inserted vector on the heap memory, and insert it into the vector search index in sequence.

[0031] In this embodiment, this step specifically includes the following sub-steps:

[0032] 1-1: The format of each inserted vector is s i =<id i , ts i , v i >, where id i is the unique identifier of the inserted vector, ts i is the timestamp of the inserted vector, and v i is the inserted vector value of a fixed dimension. For example, the inserted vector s i =<512, 1731207992000, <-14.00625, 3.25》, id i =512 represents the unique identifier of the vector, which is the unique flag to distinguish vectors; ts i =1731207992000 represents the timestamp of the vector. This timestamp can reflect the time when the vector is generated or the time when the vector-related event occurs, or it can also be the logical timestamp specified by the data source, which needs to meet the requirement of non-strict increment, that is, for all i < j, there is ts i ≤ts j ; v i =<-14.00625, 3.25> represents the value of the vector. The vector value is a floating-point array of a fixed dimension. In this embodiment, the dimension of the vector value is 2.

[0033] 1-2: Create an empty vector search index I in the heap memory, read the index type (such as HNSW (Hierarchical Navigable Small World) or NSG (Navigating Spreading-out Graph), etc.) and creation parameters (including the initialization parameters required for creating each index type and the index capacity) from the user configuration file, and initialize the vector search index I according to the index type and creation parameters.

[0034] 1-3: Create an empty array B in the heap memory. The type of the array B is the same as that of the inserted vector s iThe same, with the same capacity as that of the vector search index I, and is used to cache the inserted vector data in the heap memory.

[0035] 1 - 4: For each inserted vector s i = <id i , ts i , v i 】>, insert v i into the vector search index I, and at the same time add s i to the end of the array B.

[0036] Step S12: Every time a fixed number of vectors are inserted, serialize all the vectors cached on the heap memory and the vector search index together into a data segment, and clear the cache on the heap memory and the vector search index.

[0037] In this embodiment, this step specifically includes the following sub - steps:

[0038] 2 - 1: Create a count variable in the heap memory, with an initial value of 0.

[0039] 2 - 2: For each inserted vector s i , increment count. When count reaches a pre - set threshold, serialize all the vectors cached on the heap memory and the vector search index together into a data segment, reset count to 0, and clear the vector cache and the vector search index on the heap memory.

[0040] The data segment consists of a vector area, an index area, a metadata area, and a footer area; the vector area of the data segment is located at the starting position of the data segment and consists of vector items, offset value items, and the number of vectors. The vector items store all the vectors s i = <id i , ts i , v i 】>, where id i is the unique identifier of the inserted vector, ts i is the timestamp of the inserted vector, and v i is the inserted vector value with a fixed dimension. These vectors are sorted by the unique identifier and arranged in sequence. The offset value items store the position of each vector in the vector items, that is, in the order of the unique identifier, sequentially store the number of bytes between the position of each vector in the vector area and the starting position of the vector area. The number of vectors stores the number of vectors stored in this vector area.

[0041] The index area of the data segment is located after the vector area and includes a serialized vector search index, and the range of the vector search index is all vectors in the data segment; the metadata area of the data segment includes the meta-information of vectors and indexes in the data segment, including vector dimension, vector search index size, vector search index type, vector search index parameters, and the timestamp when the data segment is created; the footnote area of the data segment includes the offsets of the vector area, index area, and metadata area in the data segment, that is, the number of bytes between the start position of each area and the start position of the data segment.

[0042] Step S13: When constructing the data segment, use a vector compression algorithm to compress the vectors, and use the corresponding decompression algorithm to decompress the vectors when reading.

[0043] In this embodiment, this step specifically includes the following sub-steps:

[0044] 3-1: In each data segment, for all vectors s i =<id i ,ts i ,v i >, where id i is the unique identifier of the inserted vector, ts i is the timestamp of the inserted vector, and v i is the inserted vector value of a fixed dimension. Store the uncompressed vector value v1 of the first vector in IEEE 754 format. For each dimension in the vector value v i of the subsequent vectors, use the Gorilla algorithm to compress it based on the same dimension in v i-1 and store the compressed representation.

[0045] Specifically, for example, the vector values of the first two vectors in the data segment are v1 = <14.0625, 3.25> and v2 = <15.5, 3.25>. Then for v1, store each dimension completely in IEEE754 floating-point format; for v2, the 15.5 in the first dimension is compressed using the Gorilla algorithm based on the 14.0625 in the same position in v1. The 3.25 in the second dimension is compressed using the Gorilla algorithm based on the 3.25 in the same position in v1, and store the compressed binary representations respectively. For subsequent vectors, use the same method to compress based on the previous vector and store the compressed binary representations.

[0046] 3-2: When reading the vector s k , read it sequentially from the first vector s1 in the data segment to s k . If s i is the compressed representation, then for each dimension in the vector value v i of s, use v i as the reference for each dimension in vi-1 The same dimension is used as the benchmark and the Gorilla algorithm is used for decompression.

[0047] Specifically, for example, when reading the first vector s1 in the data segment, since s1 is not compressed, the IEEE754 format decoding is used to obtain the floating-point number of each dimension; when reading the second vector s2, since s2 stores the compressed representation, when reading each dimension, the uncompressed or decompressed data of the same dimension in the previous vector value is used as a reference, and the Gorilla algorithm is used to decompress it to obtain the original floating-point number.

[0048] Step S14: All constructed data segments are stored in layers according to timestamps. As time goes by, each data segment moves from the lowest level storage medium to the highest level storage medium.

[0049] Specifically, the storage media used to store data segments are local memory, local disk, and remote disk from low to high levels; when the storage medium of a certain level is full, the data segment with the earliest timestamp in that level is moved to the storage medium of a higher level. Local memory refers to the memory on the computer where the data segment is located; local disk refers to the disk mounted on the computer where the data segment is located; remote disk refers to the disk mounted on other computers connected to the computer where the data segment is located through a local area network and accessed using a distributed file system. The distributed file system provides elasticity and high availability, and theoretically has unlimited storage space.

[0050] Step S15: When performing a vector query operation, the hot and cold values of each data segment are counted, and each data segment is queried in sequence according to the order of the hot and cold values. When the query result reaches the standard pre-specified by the user, the query is terminated early.

[0051] In this embodiment, this step specifically includes the following sub-steps:

[0052] 5-1: The hot and cold value of a data segment is calculated by weighted calculation based on the access frequency of the data segment, the search hit rate on the data segment, the contribution of the vector on the data segment to the search results, and the freshness of the vector on the data segment, and is updated in real time each time the data segment is accessed.

[0053] Specifically, assuming that the system time when the hot and cold values are updated this time is t, the system time when the hot and cold values were last updated is t', and the system startup time is t0, then the access frequency AF(t) of the data segment between t' and t refers to the number of queries that accessed the data segment, and the contribution of AF(t) to the hot and cold values is H A (t) = (1-d t )(1-e -γAF(t) )+d t H A(t’), where d t = e -η(t-t’) , η and γ are pre-specified parameters, H A (t0) = 0; The search hit rate SH(t) of this data segment between t’ and t refers to the number of times the vectors in this data segment appear in the search results. The contribution of SH(t) to the hot-cold value is H S (t) = (1 - d t )(1 - e -γSH(t) ) + d t H S (t’), where d t = e -η(t-t’) , η and γ are pre-specified parameters, H S (t0) = 0; The contribution degree H C (t) of this data segment to the search results refers to the similarity between the search results from this data segment and the query vector during this vector query, where d t = e -η(t-t’) , R is the result set from this data segment, s.v represents the vector value of vector s, q represents the query vector, d represents the vector distance metric function, η and γ are pre-specified parameters, H C (t0) = 0; The freshness H F (t) of the vectors on this data segment refers to the difference between the vector timestamp and the current system time, where S represents the set of all vectors on this data segment, s.t represents the timestamp of vector s, H F (t) = 0; The hot-cold value H of this data segment = ω A H A + ω S H S + ω C H C + ω F H F , where ω A , ω S , ω C and ω F are pre-specified parameters.

[0054] 5-2: When performing a vector query operation, first search for the vectors cached in the heap memory, and then query in each data segment in the order from hot to cold according to the hot-cold value.

[0055] Specifically, maintain a priority queue in the heap memory, store the pointers of each data segment and the hot-cold value of this data segment in the order of decreasing hot-cold value; when updating the hot-cold value of a certain data segment, synchronously adjust the hot-cold value and the order of the data segments in the priority queue; when performing a query operation, traverse the data segments in the order of the hot-cold value in the priority queue.

[0056] 5-3: When performing a vector query operation, if the maximum distance between all result vectors in the current result set and the query vector is less than the threshold set by the user, the query is terminated prematurely.

[0057] Specifically, assume that the distance threshold set by the user is θ. During the process of traversing the data segment, after each data segment is searched, calculate the maximum distance θ between all result vectors in the current result set and the query vector. t , if θ t < θ, it means that the quality of the current results has already exceeded the quality required by the user. Then the query is terminated prematurely and the remaining data segments are not accessed.

[0058] The above description of the embodiments is to facilitate the understanding and application of the present invention by those of ordinary skill in the art. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A distributed vector storage method based on a layered architecture, comprising the following steps: (1) When inserting a vector, a vector search index is built in the heap memory, the inserted vector is cached on the heap memory, and inserted into the vector search index in sequence; (2) Each time a fixed number of vectors are inserted, all the vectors and vector search indexes cached on the heap memory are serialized into one data segment, and the cache and vector search indexes on the heap memory are cleared; (3) When constructing a data segment, a vector compression algorithm is used to compress the vector, and when reading, a corresponding decompression algorithm is used to decompress the vector; (4) All constructed data segments are stored in layers according to timestamps, and each data segment moves from the lowest level of storage medium to the highest level of storage medium over time; (5) When performing a vector query operation, the hot and cold values of each data segment are counted, and each data segment is queried in sequence according to the order of the hot and cold values. When the query result reaches the standard specified in advance by the user, the query is terminated early.

2. The distributed vector storage method based on a hierarchical architecture according to claim 1, wherein: The specific implementation of step (1) is as follows: 1.1 For any inserted vector s i , represent it in the form of a triple, i.e., s i = <id i , ts i , v i >, where id i is the unique identifier of the vector, ts i is the timestamp of the vector, v i is the vector value of a fixed dimension, and the subscript i represents the index number of the vector; 1.2 Create an empty vector search index I in the heap memory, read the index type and creation parameters from the user configuration file, and initialize the vector search index I according to the index type and creation parameters; 1.3 Create an empty array B in the heap memory, whose type is the same as the inserted vector s i ; 1.4 Insert v i into the vector search index I, and at the same time add s i to the end of the array B.

3. The distributed vector storage method based on a hierarchical architecture according to claim 1, wherein: The specific implementation method of step (2) is as follows: first, a counting variable count is created in the heap memory, with an initial value of 0; each time a vector is inserted, count is incremented, and when count reaches a set threshold, all vectors and vector search indexes cached on the heap memory are serialized into a data segment, count is reset to 0, and the vector cache and vector search index on the heap memory are cleared; The data segment is composed of a vector area, an index area, a metadata area and a footer area. The vector area is composed of vector items, offset value items and vector numbers. The vector items store all vectors in the data segment in ascending order of timestamps, the offset value items store the position of each vector in the vector item in ascending order of timestamps, and the vector number records the number of vectors stored in the vector area; the index area stores a serialized vector search index, and the scope of the vector search index is all vectors in the data segment; the metadata area stores meta-information of vectors and indexes in the data segment, including vector dimension, vector search index size, vector search index type, vector search index parameters and the timestamp of data segment creation; the footer area stores the offsets of the positions of the vector area, index area and metadata area in the data segment.

4. The distributed vector storage method based on a hierarchical architecture according to claim 1, wherein: The specific implementation method of step (3) is as follows: for all vectors inserted in the data segment, the vector value v1 of the first vector s1 is stored in the IEEE 754 format. For each dimension of the subsequent vector s i 's vector value v i , taking the same dimension in v i-1 as the reference, the Gorilla algorithm is used to compress s i and store the compressed representation; when reading the vector s k , starting from the first vector s1 in the data segment, read sequentially until s k . If the read vector s i is the compressed representation, then for each dimension of its vector value v i , taking the same dimension in v i-1 as the reference, the Gorilla algorithm is used to decompress s i ; where the subscripts i and k represent the index numbers of the vectors, and v i and v i-1 are the vector values of the vectors s i and s i-1 respectively.

5. The distributed vector storage method based on a hierarchical architecture according to claim 1, wherein: The specific implementation method of step (4) is as follows: let the storage media used to store data segments be local memory, local disk and remote disk from low level to high level respectively; when the storage medium of a certain level is full, move the data segment with the earliest timestamp in this level to the storage medium of a higher level; the local memory is the memory on the computer where the data segment is located, the local disk is the disk mounted on the computer where the data segment is located, and the remote disk is the disk mounted on other computers connected to the computer where the data segment is located through a local area network and accessed using a distributed file system.

6. The distributed vector storage method based on a hierarchical architecture according to claim 1, characterized in that: The specific implementation of step (5) is as follows: count the hot and cold values of each data segment, which are obtained by weighted calculation of the access frequency of the data segment, the search hit rate on the data segment, the contribution degree of the vector on the data segment to the search result, and the freshness of the vector on the data segment, and update them in real time every time the data segment is accessed; when performing a vector query operation, first search for the vectors cached in the heap memory, and then query in each data segment in the order from hot to cold according to the hot and cold values; after searching each data segment, judge the maximum distance between the vector result searched on the data segment and the input query target vector. If the maximum distance is less than the distance threshold set by the user, it means that the quality of the currently searched vector result has been better than the quality required by the user, and the query is terminated in advance without accessing the remaining data segments.

7. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and characterized in that: The processor is used to execute the computer program to implement the distributed vector storage method based on a hierarchical architecture according to any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it is used to implement the distributed vector storage method based on a hierarchical architecture according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Vector data retrieval and storage method and device

    CN118606363A

  • Data storage method and device of vector database, data recovery method and device and storage medium

    CN119226274A