High-dimensional data library falling method, device and equipment based on hybrid index and medium
By converting high-dimensional data into configuration formats and building hybrid indexes, the problems of high-dimensional data storage overhead and low retrieval efficiency are solved, and efficient data storage and retrieval is achieved.
Patent Information
- Application Number
- CN202510535561.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional database systems have low retrieval efficiency and high storage overhead when processing high-dimensional vector data, making it difficult to support the rapid storage, indexing and retrieval of large-scale high-dimensional vector data.
Using a high-dimensional data database deleting method based on hybrid index, multi-modal high-dimensional data is converted into high-dimensional vectors in configuration format through the configuration interface, compressing and dimensional reduction processing are performed, HNSW index and IVFFLAT index are constructed, mixed indexes are obtained, and mixed indexes are deposited into a centralized relational database.
It reduces the overhead of deleting the database, improves data retrieval efficiency, supports the rapid storage, indexing and retrieval of large-scale high-dimensional vector data, and has good scalability and resource utilization.
Smart Images

Figure CN120448385A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, equipment and medium for storing high-dimensional data in a database based on a hybrid index. Background Art
[0002] With the rapid development of artificial intelligence and big data technologies, the application scenarios of high-dimensional vector data (such as image feature vectors, text embedding vectors, and other multimodal high-dimensional data) are increasing. This high-dimensional data needs to be properly stored, so database systems need to be adapted accordingly. However, traditional database systems still have the following problems when processing high-dimensional vector data:
[0003] 1) Low retrieval efficiency: Data in high-dimensional space is sparsely distributed, and traditional indexing methods (such as B+ trees) are difficult to effectively support efficient similarity retrieval.
[0004] 2) High storage overhead: The dimensions of high-dimensional vector data are usually between hundreds and thousands, and direct storage will lead to a significant increase in storage overhead.
[0005] Therefore, an efficient high-dimensional data storage method is needed to support the rapid storage, indexing and retrieval of large-scale high-dimensional vector data, and have good scalability and resource utilization. Summary of the Invention
[0006] In view of the above, it is necessary to provide a method, device, equipment and medium for storing high-dimensional data based on hybrid indexing, aiming to solve the problems of high storage overhead of high-dimensional data and low retrieval efficiency after storage.
[0007] A high-dimensional data storage method based on a hybrid index, the high-dimensional data storage method based on a hybrid index comprising:
[0008] In response to an instruction to store the multimodal high-dimensional data in a centralized relational database, calling a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format;
[0009] Compressing and reducing the dimension of the high-dimensional vector to obtain a vector to be stored;
[0010] Constructing an HNSW index and an IVFFLAT index of the vector to be stored based on a hybrid index strategy to obtain a hybrid index of the vector to be stored;
[0011] The mixed index and the vector to be stored are stored in the centralized relational database.
[0012] A high-dimensional data storage device based on a hybrid index, the high-dimensional data storage device based on a hybrid index comprising:
[0013] a conversion unit, configured to, in response to a storage instruction for storing multimodal high-dimensional data in a centralized relational database, call a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format;
[0014] A processing unit, configured to compress and reduce the dimension of the high-dimensional vector to obtain a vector to be stored;
[0015] A construction unit, configured to construct an HNSW index and an IVFFLAT index of the vector to be stored based on a hybrid index strategy, to obtain a hybrid index of the vector to be stored;
[0016] A storage unit is used to store the mixed index and the vector to be stored in the centralized relational database.
[0017] A computer device, comprising:
[0018] a memory storing at least one instruction; and
[0019] The processor executes the instructions stored in the memory to implement the high-dimensional data storage method based on hybrid indexing.
[0020] A computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the high-dimensional data storage method based on hybrid indexing.
[0021] It can be seen from the above technical solutions that the present invention calls the configuration interface to convert multimodal high-dimensional data into high-dimensional vectors in a configuration format. Through a unified vectorization interface, efficient conversion and interaction between different modal data can be achieved; high-dimensional vectors are compressed and dimensionally reduced to obtain vectors to be stored in the database, which can effectively reduce storage overhead; based on the hybrid index strategy, HNSW indexes and IVFFLAT indexes of vectors to be stored in the database are constructed to obtain hybrid indexes of vectors to be stored in the database. The hybrid index structure can support the rapid indexing of large-scale high-dimensional vector data; the hybrid index and vectors to be stored in the database are stored in a centralized relational database. Since the high-dimensional data is compressed and dimensionally reduced before being stored in the database, not only the storage overhead is reduced, but also the storage efficiency is improved. At the same time, combined with the hybrid index structure, the data retrieval efficiency can also be improved after storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flow chart of a preferred embodiment of the high-dimensional data storage method based on hybrid indexing of the present invention.
[0023] Figure 2 It is a functional module diagram of a preferred embodiment of the high-dimensional data storage device based on hybrid indexing of the present invention.
[0024] Figure 3 It is a structural diagram of a computer device according to a preferred embodiment of the present invention for implementing a high-dimensional data storage method based on a hybrid index. DETAILED DESCRIPTION
[0025] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] like Figure 1 FIG. 1 is a flow chart of a preferred embodiment of the high-dimensional data storage method based on hybrid indexing according to the present invention. The order of the steps in the flow chart can be changed according to different requirements, and some steps can be omitted.
[0027] The high-dimensional data storage method based on hybrid indexing is applied to one or more computer devices, which are devices that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Their hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0028] The computer device can be any electronic product that can interact with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.
[0029] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0030] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0031] Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0032] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0033] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0034] S10 , in response to an instruction to store multimodal high-dimensional data in a centralized relational database, calling a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format.
[0035] In this embodiment, the multimodal high-dimensional data may include, but is not limited to, transaction data and investment portfolio data in the financial sector, and electronic medical records and medical imaging data in the healthcare sector. This data may be in multiple modalities, such as text, images, and audio.
[0036] In this embodiment, the centralized relational database refers to a system in which all data is stored in a centralized system, which simplifies data management and maintenance and provides good consistency and security. For example, the centralized relational database may be a RASESQL (Reliability Availability Stability Enterprise SQL, an online relational database service based on a cloud computing platform) database.
[0037] In this embodiment, the storage instruction can be automatically triggered when it is detected that the multimodal high-dimensional data is uploaded to a designated page.
[0038] In this embodiment, the configuration interface is a unified interface for converting multimodal high-dimensional data into vectors in a specified format.
[0039] Specifically, before calling the configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format, the method further includes:
[0040] Collect data from different modalities;
[0041] Preprocess the collected data to obtain data to be processed in multiple modes;
[0042] Call the corresponding feature extraction algorithm to extract features from the data to be processed of each modality to obtain feature data corresponding to each modality;
[0043] Obtaining a vectorization function required to convert characteristic data corresponding to each mode into the configuration format;
[0044] The obtained vectorized functions are encapsulated into the same interface class to obtain the configuration interface.
[0045] For example, data in various modalities such as text, images, and audio can be collected, and the collected data can be cleaned to remove noise, outliers, and duplicate data (for example, for text data, HTML (Hypertext Markup Language) tags, special characters, etc. can be removed; for image data, blurred or damaged images can be removed) to obtain the data to be processed.
[0046] Furthermore, the corresponding feature extraction algorithm is called to extract features from the data to be processed in each modality. For example, for text data, word embedding models (such as Word2Vec (Word to Vector, word vector), BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder model based on the Transformer architecture), etc.) can be used to convert text into vector representations; for image data, pre-trained convolutional neural networks (such as ResNet (Residual Neural Network, residual network), VGG (Visual Geometry Group, visual geometry group)), etc.) can be used to extract image feature vectors; for audio data, Mel (Mel Frequency Cepstrum Coefficient, Mel Frequency Cepstrum Coefficient, MFCC), deep neural networks (such as VGGish (Visual Geometry Groupish, audio feature extractor)), etc. can be used to extract audio features.
[0047] Furthermore, defining the configuration format can determine a unified vector data type and dimension, ensuring that data from different modalities have consistent representations after conversion to vectors. For example, vectors from all modalities can be converted to fixed-length floating-point vectors.
[0048] Furthermore, corresponding vectorization functions are implemented for data of different modalities to convert the raw data into vectors in a unified format. These functions can be encapsulated in a unified interface class for easy calling. For example, a Vectorizer class can be defined that contains functions such as text_to_vector, image_to_vector, and audio_to_vector.
[0049] Through the above embodiments, data in various modalities such as text, images, and audio can be uniformly converted into specified data types based on the configured unified interface, thereby assisting in achieving efficient interaction and retrieval between data in different modalities.
[0050] S11, compressing and reducing the dimension of the high-dimensional vector to obtain a vector to be stored.
[0051] Whether in the financial field or the medical and health field, the amount of data is huge.
[0052] For example, transaction data in the financial sector generates a large number of orders daily, especially during special events like June 18th and Singles' Day, which sees a dramatic increase in transaction data volume. Another example is imaging data and electronic reports in the healthcare sector, which have a high volume due to the increasing adoption of paperless processes and increasingly high-definition images.
[0053] However, due to limited storage and computing resources, reducing data storage space is a priority.
[0054] In this embodiment, compressing and reducing the dimensionality of the high-dimensional vector to obtain the vector to be stored includes:
[0055] Obtaining a pre-configured discrete value set; wherein the discrete value set is used to store a preset number of data expression templates;
[0056] Performing clustering processing on the high-dimensional vector according to the preset number to obtain the preset number of cluster centers;
[0057] Obtaining each dimension value of the high-dimensional vector;
[0058] Calculate the distance between each dimension value and each cluster center;
[0059] For each dimension value, the data expression template corresponding to the cluster center closest to the dimension value is determined as the target template corresponding to the dimension value;
[0060] Map each dimension value to a discrete value according to the target template corresponding to each dimension value;
[0061] The discrete values are processed by principal component analysis to reduce the dimension, and the vector to be stored is obtained.
[0062] The data expression template is used to define the data expression form of each dimension, such as the data expression form of the size and color of image data.
[0063] Through the above embodiment, high-dimensional data is first quantized and compressed to reduce the amount of data, and then principal component analysis is performed on the compressed and quantized data to achieve further dimensionality reduction processing of the data, thereby retaining valid information in the data and speeding up the processing speed in a resource-constrained environment.
[0064] S12: constructing an HNSW (Hierarchical Navigable Small World) index and an IVFFLAT (Inverted File with Flat Quantization) index of the vector to be stored based on a hybrid indexing strategy to obtain a hybrid index of the vector to be stored.
[0065] In this embodiment, the HNSW index and IVFFLAT index of the vector to be stored are constructed based on the hybrid index strategy to obtain the hybrid index of the vector to be stored, including:
[0066] Constructing the HNSW index and constructing the IVFFLAT index;
[0067] Obtaining a mapping relationship between the vector to be stored in the HNSW index and the IVFFLAT index;
[0068] According to the mapping relationship, pointers pointing to each other are added in the HNSW index and the IVFFLAT index, and a unified index access interface of the HNSW index and the IVFFLAT index is constructed to obtain a mixed index of the vector to be stored in the database.
[0069] Wherein, pointers pointing to each other are added in the HNSW index and the IVFFLAT index according to the mapping relationship, so that an association relationship can be established between the HNSW index and the IVFFLAT index.
[0070] Among them, building a unified index access interface for the HNSW index and the IVFFLAT index can shield the internal implementation details of the HNSW index and the IVFFLAT index from upper-layer applications and support consistent operations such as insertion, query, and deletion.
[0071] The parameters of the HNSW index and the IVFFLAT index can also be adjusted jointly based on the actual query load and data characteristics. For example, if data is frequently updated, the HNSW index construction parameters can be appropriately reduced to speed up updates, while the IVFFLAT index clustering parameters can be adjusted to improve search accuracy.
[0072] Through the above embodiments, efficient vector retrieval can be achieved based on the hybrid index architecture.
[0073] In this embodiment, constructing the HNSW index includes:
[0074] Randomly select one or more vectors from the vectors to be stored as top-level nodes;
[0075] Obtain the number and data attributes of the vectors to be placed in the library, and configure the number of layers and the number of nodes in each layer according to the number and data attributes of the vectors to be placed in the library; wherein each layer has a pyramid structure, and the layer closer to the top layer contains a lower number of nodes;
[0076] Calculating the distance between each vector in the vector to be stored and the top node;
[0077] In the order of the distance from near to far, each vector is sequentially assigned to the corresponding layer according to the number of layers and the number of nodes in each layer;
[0078] For each node in each layer, calculate the distance between every two nodes and connect the two nodes with the closest distance in sequence;
[0079] For each node between adjacent layers, calculate the distance between each two nodes and connect each two nodes with the closest distance in sequence;
[0080] The currently obtained multi-layer index structure is obtained as the HNSW index.
[0081] The distance between each vector in the vector to be stored and the top-level node may be calculated using a Euclidean distance algorithm or the like. The present invention does not impose any limitation on the distance calculation method used.
[0082] After obtaining the HNSW index, the node connections can be optimized and adjusted. For example, by re-evaluating the distances between nodes, it may be found that the neighboring nodes of some nodes are not optimally selected, and some connections need to be added or deleted to improve the search efficiency of the index.
[0083] After obtaining the HNSW index, the hierarchical structure can be checked for balance to ensure that nodes are evenly distributed across the layers, avoiding situations where some layers have too many or too few nodes. If the hierarchical structure is found to be unbalanced, optimization can be performed by adjusting node distribution or reconnecting nodes.
[0084] In the above embodiment, by constructing the HNSW index with a multi-layer graph structure, each layer contains a sparse graph, with the top layer graph containing fewer nodes and the bottom layer graph containing more nodes. This hierarchical structure enables the search process to quickly narrow the scope, thereby improving efficiency.
[0085] In this embodiment, constructing the IVFFLAT index includes:
[0086] Configuring the number of clusters according to the number of vectors to be stored and data attributes;
[0087] Clustering the vectors to be stored according to the number of clusters to obtain the number of clusters;
[0088] Create an inverted index entry for each cluster; wherein the inverted index entry of each cluster is used to record the index information of the vectors contained in each cluster;
[0089] The vectors contained in each cluster are quantized to obtain the IVFFLAT index.
[0090] The inverted index item can quickly locate all vectors belonging to a cluster through an inverted table.
[0091] The index information can be used to record the location information of the corresponding vector. For example, in a text vector dataset such as financial transaction data, each cluster may represent a topic category, and the inverted index item will record the location or identifier of the text vector belonging to the topic category in the original dataset.
[0092] Uniform quantization can be used to divide the vector space into equally spaced intervals, mapping the vector to a discrete value based on the interval it falls within. Alternatively, non-uniform quantization can be used to divide the vector space into finer intervals in data-dense areas and coarser intervals in data-sparse areas, based on the data's distribution characteristics. For example, for medical imaging data, for color feature vectors, an appropriate quantization method can be selected based on the color distribution to map the color values to a finite set of discrete values.
[0093] Each dimension of each vector is quantized according to the selected quantization method, converting it into discrete quantized values. The quantized vector can be represented with less storage space and is also convenient for fast comparison and search in the index.
[0094] After obtaining the IVFFLAT index, the clustering quality can be checked to determine whether some clusters are too large or too small, or whether the boundaries between clusters are unclear. If problems are found, the clustering results can be optimized by adjusting the clustering algorithm parameters or re-clustering to make the clustering more reasonable and improve the index performance.
[0095] After obtaining the IVFFLAT index, parameters in the quantization process can be adjusted and optimized, such as the size of the quantization interval, the range of the quantization value, etc. Through experiments and evaluation, the parameter settings that can minimize the quantization error while ensuring the index accuracy are found.
[0096] In the above embodiment, by constructing the IVFFLAT index that combines quantization with an inverted file, the search space can be significantly reduced, improving search efficiency. Because the quantization process reduces data storage requirements, it is suitable for processing large-scale data. The IVFFLAT index can handle high-dimensional data and maintains good performance even as the data volume increases.
[0097] In this embodiment, after obtaining the mixed index of the vector to be stored, the method further includes:
[0098] When a change is detected in the vector to be stored, the change type is obtained; wherein the change type includes addition, deletion, and modification;
[0099] The mixed index of the vector to be stored is adjusted in real time according to the change type.
[0100] Through the above embodiments, the index structure can be adjusted in real time according to the dynamic changes of vector data (such as addition, deletion, modification, etc.) to ensure retrieval efficiency and accuracy.
[0101] S13: storing the mixed index and the vector to be stored in the centralized relational database.
[0102] Through the above embodiments, it is possible to support fast storage, indexing and retrieval of large-scale high-dimensional vector data while having good scalability and resource utilization.
[0103] This embodiment can be widely used in image retrieval, natural language processing, and recommendation systems in the fields of finance and healthcare.
[0104] For example, in the quantitative trading of stocks in the financial field, a large amount of high-dimensional time series data is generated. The price, trading volume, transaction amount and other data of each stock form a high-dimensional vector in chronological order. If data is collected on a daily basis, a stock's data for one year can constitute a sequence of approximately 240-250 data points, that is, a 250-dimensional vector; if recorded in minutes, the data dimension for one year will soar to 60,000. Through the hybrid index storage in this embodiment, the data points most similar to the current stock market data can be quickly and accurately located in the massive high-dimensional historical data vectors, providing an important reference for predicting market trends and formulating trading strategies.
[0105] For another example: For medical imaging data in the field of medical health, taking CT (Computed Tomography) images as an example, the grayscale value, spatial position (three-dimensional coordinates) and other information of each pixel constitute a high-dimensional vector. By using the hybrid index library in this embodiment to store the high-dimensional feature vectors of these images, doctors can input the feature vector of the image to be analyzed during diagnosis and quickly find similar historical imaging cases through efficient indexing and retrieval algorithms. The diagnostic results, treatment plans and other information of these similar cases can be used as a reference for doctors to assist them in more accurately judging the type and severity of the disease and formulating treatment plans.
[0106] It can be seen from the above technical solutions that the present invention calls the configuration interface to convert multimodal high-dimensional data into high-dimensional vectors in a configuration format. Through a unified vectorization interface, efficient conversion and interaction between different modal data can be achieved; high-dimensional vectors are compressed and reduced in dimension to obtain vectors to be stored in the database, which can effectively reduce storage overhead; based on the hybrid index strategy, HNSW indexes and IVFFLAT indexes of vectors to be stored in the database are constructed to obtain hybrid indexes of vectors to be stored in the database. The hybrid index structure can support the rapid indexing of large-scale high-dimensional vector data; the hybrid index and vectors to be stored in the database are stored in a centralized relational database. Since the high-dimensional data is compressed and reduced in dimension before being stored in the database, not only the storage overhead is reduced, but also the storage efficiency is improved. At the same time, combined with the hybrid index structure, the data retrieval efficiency can also be improved after storage.
[0107] like Figure 2 , which is a functional module diagram of a preferred embodiment of a high-dimensional data storage device based on a hybrid index according to the present invention. The high-dimensional data storage device 11 based on a hybrid index includes a conversion unit 110, a processing unit 111, a construction unit 112, and a storage unit 113. The modules / units referred to in the present invention refer to a series of computer program segments that can be executed by a processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0108] The conversion unit 110 is configured to, in response to a storage instruction for storing multimodal high-dimensional data in a centralized relational database, call a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format.
[0109] In this embodiment, the multimodal high-dimensional data may include, but is not limited to, transaction data and investment portfolio data in the financial sector, and electronic medical records and medical imaging data in the healthcare sector. This data may be in multiple modalities, such as text, images, and audio.
[0110] In this embodiment, the centralized relational database refers to a system in which all data is stored in a centralized system, which simplifies data management and maintenance and provides good consistency and security. For example, the centralized relational database may be a RASESQL (Reliability Availability Stability Enterprise SQL, an online relational database service based on a cloud computing platform) database.
[0111] In this embodiment, the storage instruction can be automatically triggered when it is detected that the multimodal high-dimensional data is uploaded to a designated page.
[0112] In this embodiment, the configuration interface is a unified interface for converting multimodal high-dimensional data into vectors in a specified format.
[0113] Specifically, before calling the configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format, data of different modalities are collected;
[0114] Preprocess the collected data to obtain data to be processed in multiple modes;
[0115] Call the corresponding feature extraction algorithm to extract features from the data to be processed of each modality to obtain feature data corresponding to each modality;
[0116] Obtaining a vectorization function required to convert characteristic data corresponding to each mode into the configuration format;
[0117] The obtained vectorized functions are encapsulated into the same interface class to obtain the configuration interface.
[0118] For example, data in various modalities such as text, images, and audio can be collected, and the collected data can be cleaned to remove noise, outliers, and duplicate data (for example, for text data, HTML (Hypertext Markup Language) tags, special characters, etc. can be removed; for image data, blurred or damaged images can be removed) to obtain the data to be processed.
[0119] Furthermore, the corresponding feature extraction algorithm is called to extract features from the data to be processed in each modality. For example, for text data, word embedding models (such as Word2Vec (Word to Vector, word vector), BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder model based on the Transformer architecture), etc.) can be used to convert text into vector representations; for image data, pre-trained convolutional neural networks (such as ResNet (Residual Neural Network, residual network), VGG (Visual Geometry Group, visual geometry group)), etc.) can be used to extract image feature vectors; for audio data, Mel (Mel Frequency Cepstrum Coefficient, Mel Frequency Cepstrum Coefficient, MFCC), deep neural networks (such as VGGish (Visual Geometry Groupish, audio feature extractor)), etc. can be used to extract audio features.
[0120] Furthermore, defining the configuration format can determine a unified vector data type and dimension, ensuring that data from different modalities have consistent representations after conversion to vectors. For example, vectors from all modalities can be converted to fixed-length floating-point vectors.
[0121] Furthermore, corresponding vectorization functions are implemented for data of different modalities to convert the raw data into vectors in a unified format. These functions can be encapsulated in a unified interface class for easy calling. For example, a Vectorizer class can be defined that contains functions such as text_to_vector, image_to_vector, and audio_to_vector.
[0122] Through the above embodiments, data in various modalities such as text, images, and audio can be uniformly converted into specified data types based on the configured unified interface, thereby assisting in achieving efficient interaction and retrieval between data in different modalities.
[0123] The processing unit 111 is used to compress and reduce the dimension of the high-dimensional vector to obtain a vector to be stored.
[0124] Whether in the financial field or the medical and health field, the amount of data is huge.
[0125] For example, transaction data in the financial sector generates a large number of orders daily, especially during special events like June 18th and Singles' Day, which sees a dramatic increase in transaction data volume. Another example is imaging data and electronic reports in the healthcare sector, which have a high volume due to the increasing adoption of paperless processes and increasingly high-definition images.
[0126] However, due to limited storage and computing resources, reducing data storage space is a priority.
[0127] In this embodiment, the processing unit 111 compresses and reduces the dimension of the high-dimensional vector to obtain the vector to be stored, including:
[0128] Obtaining a pre-configured discrete value set; wherein the discrete value set is used to store a preset number of data expression templates;
[0129] Performing clustering processing on the high-dimensional vector according to the preset number to obtain the preset number of cluster centers;
[0130] Obtaining each dimension value of the high-dimensional vector;
[0131] Calculate the distance between each dimension value and each cluster center;
[0132] For each dimension value, the data expression template corresponding to the cluster center closest to the dimension value is determined as the target template corresponding to the dimension value;
[0133] Map each dimension value to a discrete value according to the target template corresponding to each dimension value;
[0134] The discrete values are processed by principal component analysis to reduce the dimension, and the vector to be stored is obtained.
[0135] The data expression template is used to define the data expression form of each dimension, such as the data expression form of the size and color of image data.
[0136] Through the above embodiment, high-dimensional data is first quantized and compressed to reduce the amount of data, and then principal component analysis is performed on the compressed and quantized data to achieve further dimensionality reduction processing of the data, thereby retaining valid information in the data and speeding up the processing speed in a resource-constrained environment.
[0137] The construction unit 112 is configured to construct an HNSW (Hierarchical Navigable Small World) index and an IVFFLAT (Inverted File with Flat Quantization) index of the vector to be stored based on a hybrid indexing strategy to obtain a hybrid index of the vector to be stored.
[0138] In this embodiment, the construction unit 112 constructs the HNSW index and IVFFLAT index of the vector to be stored based on the hybrid index strategy, and obtains the hybrid index of the vector to be stored, including:
[0139] Constructing the HNSW index and constructing the IVFFLAT index;
[0140] Obtaining a mapping relationship between the vector to be stored in the HNSW index and the IVFFLAT index;
[0141] According to the mapping relationship, pointers pointing to each other are added in the HNSW index and the IVFFLAT index, and a unified index access interface of the HNSW index and the IVFFLAT index is constructed to obtain a mixed index of the vector to be stored in the database.
[0142] Wherein, pointers pointing to each other are added in the HNSW index and the IVFFLAT index according to the mapping relationship, so that an association relationship can be established between the HNSW index and the IVFFLAT index.
[0143] Among them, building a unified index access interface for the HNSW index and the IVFFLAT index can shield the internal implementation details of the HNSW index and the IVFFLAT index from upper-layer applications and support consistent operations such as insertion, query, and deletion.
[0144] The parameters of the HNSW index and the IVFFLAT index can also be adjusted jointly based on the actual query load and data characteristics. For example, if data is frequently updated, the HNSW index construction parameters can be appropriately reduced to speed up updates, while the IVFFLAT index clustering parameters can be adjusted to improve search accuracy.
[0145] Through the above embodiments, efficient vector retrieval can be achieved based on the hybrid index architecture.
[0146] In this embodiment, the construction unit 112 constructs the HNSW index including:
[0147] Randomly select one or more vectors from the vectors to be stored as top-level nodes;
[0148] Obtain the number and data attributes of the vectors to be placed in the library, and configure the number of layers and the number of nodes in each layer according to the number and data attributes of the vectors to be placed in the library; wherein each layer has a pyramid structure, and the layer closer to the top layer contains a lower number of nodes;
[0149] Calculating the distance between each vector in the vector to be stored and the top node;
[0150] In the order of the distance from near to far, each vector is sequentially assigned to the corresponding layer according to the number of layers and the number of nodes in each layer;
[0151] For each node in each layer, calculate the distance between every two nodes and connect the two nodes with the closest distance in sequence;
[0152] For each node between adjacent layers, calculate the distance between each two nodes and connect each two nodes with the closest distance in sequence;
[0153] The currently obtained multi-layer index structure is obtained as the HNSW index.
[0154] The distance between each vector in the vector to be stored and the top-level node may be calculated using a Euclidean distance algorithm or the like. The present invention does not impose any limitation on the distance calculation method used.
[0155] After obtaining the HNSW index, the node connections can be optimized and adjusted. For example, by re-evaluating the distances between nodes, it may be found that the neighboring nodes of some nodes are not optimally selected, and some connections need to be added or deleted to improve the search efficiency of the index.
[0156] After obtaining the HNSW index, the hierarchical structure can be checked for balance to ensure that nodes are evenly distributed across the layers, avoiding situations where some layers have too many or too few nodes. If the hierarchical structure is found to be unbalanced, optimization can be performed by adjusting node distribution or reconnecting nodes.
[0157] In the above embodiment, by constructing the HNSW index with a multi-layer graph structure, each layer contains a sparse graph, with the top layer graph containing fewer nodes and the bottom layer graph containing more nodes. This hierarchical structure enables the search process to quickly narrow the scope, thereby improving efficiency.
[0158] In this embodiment, the construction unit 112 constructs the IVFFLAT index including:
[0159] Configuring the number of clusters according to the number of vectors to be stored and data attributes;
[0160] Clustering the vectors to be stored according to the number of clusters to obtain the number of clusters;
[0161] Create an inverted index entry for each cluster; wherein the inverted index entry of each cluster is used to record the index information of the vectors contained in each cluster;
[0162] The vectors contained in each cluster are quantized to obtain the IVFFLAT index.
[0163] The inverted index item can quickly locate all vectors belonging to a cluster through an inverted table.
[0164] The index information can be used to record the location information of the corresponding vector. For example, in a text vector dataset such as financial transaction data, each cluster may represent a topic category, and the inverted index item will record the location or identifier of the text vector belonging to the topic category in the original dataset.
[0165] Uniform quantization can be used to divide the vector space into equally spaced intervals, mapping the vector to a discrete value based on the interval it falls within. Alternatively, non-uniform quantization can be used to divide the vector space into finer intervals in data-dense areas and coarser intervals in data-sparse areas, based on the data's distribution characteristics. For example, for medical imaging data, for color feature vectors, an appropriate quantization method can be selected based on the color distribution to map the color values to a finite set of discrete values.
[0166] Each dimension of each vector is quantized according to the selected quantization method, converting it into discrete quantized values. The quantized vector can be represented with less storage space and is also convenient for fast comparison and search in the index.
[0167] After obtaining the IVFFLAT index, the clustering quality can be checked to determine whether some clusters are too large or too small, or whether the boundaries between clusters are unclear. If problems are found, the clustering results can be optimized by adjusting the clustering algorithm parameters or re-clustering to make the clustering more reasonable and improve the index performance.
[0168] After obtaining the IVFFLAT index, parameters in the quantization process can be adjusted and optimized, such as the size of the quantization interval, the range of the quantization value, etc. Through experiments and evaluation, the parameter settings that can minimize the quantization error while ensuring the index accuracy are found.
[0169] In the above embodiment, by constructing the IVFFLAT index that combines quantization with an inverted file, the search space can be significantly reduced, improving search efficiency. Because the quantization process reduces data storage requirements, it is suitable for processing large-scale data. The IVFFLAT index can handle high-dimensional data and maintains good performance even as the data volume increases.
[0170] In this embodiment, after obtaining the mixed index of the vector to be placed in the library, when a change in the vector to be placed in the library is detected, the change type is obtained; wherein the change type includes addition, deletion, and modification;
[0171] The mixed index of the vector to be stored is adjusted in real time according to the change type.
[0172] Through the above embodiments, the index structure can be adjusted in real time according to the dynamic changes of vector data (such as addition, deletion, modification, etc.) to ensure retrieval efficiency and accuracy.
[0173] The storage unit 113 is used to store the mixed index and the vector to be stored in the centralized relational database.
[0174] Through the above embodiments, it is possible to support fast storage, indexing and retrieval of large-scale high-dimensional vector data while having good scalability and resource utilization.
[0175] This embodiment can be widely used in image retrieval, natural language processing, and recommendation systems in the fields of finance and healthcare.
[0176] For example, in the quantitative trading of stocks in the financial field, a large amount of high-dimensional time series data is generated. The price, trading volume, transaction amount and other data of each stock form a high-dimensional vector in chronological order. If data is collected on a daily basis, a stock's data for one year can constitute a sequence of approximately 240-250 data points, that is, a 250-dimensional vector; if recorded in minutes, the data dimension for one year will soar to 60,000. Through the hybrid index storage in this embodiment, the data points most similar to the current stock market data can be quickly and accurately located in the massive high-dimensional historical data vectors, providing an important reference for predicting market trends and formulating trading strategies.
[0177] For another example: For medical imaging data in the field of medical health, taking CT (Computed Tomography) images as an example, the grayscale value, spatial position (three-dimensional coordinates) and other information of each pixel constitute a high-dimensional vector. By using the hybrid index library in this embodiment to store the high-dimensional feature vectors of these images, doctors can input the feature vector of the image to be analyzed during diagnosis and quickly find similar historical imaging cases through efficient indexing and retrieval algorithms. The diagnostic results, treatment plans and other information of these similar cases can be used as a reference for doctors to assist them in more accurately judging the type and severity of the disease and formulating treatment plans.
[0178] It can be seen from the above technical solutions that the present invention calls the configuration interface to convert multimodal high-dimensional data into high-dimensional vectors in a configuration format. Through a unified vectorization interface, efficient conversion and interaction between different modal data can be achieved; high-dimensional vectors are compressed and reduced in dimension to obtain vectors to be stored in the database, which can effectively reduce storage overhead; based on the hybrid index strategy, HNSW indexes and IVFFLAT indexes of vectors to be stored in the database are constructed to obtain hybrid indexes of vectors to be stored in the database. The hybrid index structure can support the rapid indexing of large-scale high-dimensional vector data; the hybrid index and vectors to be stored in the database are stored in a centralized relational database. Since the high-dimensional data is compressed and reduced in dimension before being stored in the database, not only the storage overhead is reduced, but also the storage efficiency is improved. At the same time, combined with the hybrid index structure, the data retrieval efficiency can also be improved after storage.
[0179] like Figure 3 , which is a structural diagram of a computer device according to a preferred embodiment of the present invention for implementing a high-dimensional data storage method based on a hybrid index.
[0180] The computer device 1 may include a memory 12, a processor 13 and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a high-dimensional data storage program based on a hybrid index.
[0181] Those skilled in the art will understand that the schematic diagram is merely an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 may have either a bus structure or a star structure. The computer device 1 may also include more or less other hardware or software than shown in the figure, or a different arrangement of components. For example, the computer device 1 may also include input and output devices, network access devices, etc.
[0182] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that are suitable for the present invention should also be included in the scope of protection of the present invention and included here by reference.
[0183] The memory 12 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 12 may be an internal storage unit of the computer device 1, such as a mobile hard disk of the computer device 1. In other embodiments, the memory 12 may also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the computer device 1. Furthermore, the memory 12 may include both an internal storage unit of the computer device 1 and an external storage device. The memory 12 can be used not only to store application software installed on the computer device 1 and various types of data, such as the code of a high-dimensional data storage program based on a hybrid index, but can also be used to temporarily store data that has been output or is about to be output.
[0184] In some embodiments, the processor 13 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 13 is the control core (Control Unit) of the computer device 1, and utilizes various interfaces and lines to connect the various components of the entire computer device 1. It executes or executes programs or modules stored in the memory 12 (for example, executing a high-dimensional data retrieval program based on a hybrid index), and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0185] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-mentioned embodiments of the high-dimensional data storage method based on hybrid indexing, such as Figure 1 Steps shown.
[0186] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a conversion unit 110, a processing unit 111, a construction unit 112, and a storage unit 113.
[0187] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, computer device, or network device, etc.) or a processor to execute the portion of the high-dimensional data storage method based on hybrid indexing described in various embodiments of the present invention.
[0188] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the processes in the above-mentioned method embodiments by instructing relevant hardware devices through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments.
[0189] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory, etc.
[0190] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0191] Blockchain, as used in this article, refers to a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each block contains information about a batch of online transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0192] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The figure shows that only one straight line is used, but it does not mean that there is only one bus or one type of bus. The bus is configured to realize the connection and communication between the memory 12 and at least one processor 13.
[0193] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be detailed here.
[0194] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.
[0195] Optionally, the computer device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the computer device 1 and to display a visual user interface.
[0196] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0197] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0198] Combine Figure 1 The memory 12 in the computer device 1 stores a plurality of instructions to implement a high-dimensional data storage method based on a hybrid index, and the processor 13 can execute the plurality of instructions to implement:
[0199] In response to an instruction to store the multimodal high-dimensional data in a centralized relational database, calling a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format;
[0200] Compressing and reducing the dimension of the high-dimensional vector to obtain a vector to be stored;
[0201] Constructing an HNSW index and an IVFFLAT index of the vector to be stored based on a hybrid index strategy to obtain a hybrid index of the vector to be stored;
[0202] The mixed index and the vector to be stored are stored in the centralized relational database.
[0203] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0204] It should be noted that the data involved in this case were all obtained legally. The software tools or components not produced by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0205] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division, and actual implementation may employ other division methods.
[0206] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0207] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.
[0208] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0209] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0210] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0211] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the present invention may also be implemented by a single unit or device through software or hardware. Terms such as first and second are used to indicate names and do not imply any particular order.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A high-dimensional data storage method based on hybrid index, characterized in that: The high-dimensional data storage method based on hybrid index includes: In response to an instruction to store the multimodal high-dimensional data in a centralized relational database, calling a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format; Compressing and reducing the dimension of the high-dimensional vector to obtain a vector to be stored; Constructing an HNSW index and an IVFFLAT index of the vector to be stored based on a hybrid index strategy to obtain a hybrid index of the vector to be stored; The mixed index and the vector to be stored are stored in the centralized relational database.
2. The high-dimensional data storage method based on hybrid index according to claim 1 is characterized in that: Before calling the configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format, the method further includes: Collect data from different modalities; Preprocess the collected data to obtain data to be processed in multiple modes; Call the corresponding feature extraction algorithm to extract features from the data to be processed in each mode to obtain feature data corresponding to each mode; Obtaining a vectorization function required to convert characteristic data corresponding to each mode into the configuration format; The obtained vectorized functions are encapsulated into the same interface class to obtain the configuration interface.
3. The high-dimensional data storage method based on hybrid index according to claim 1 is characterized in that: The compressing and dimensionality reduction processing of the high-dimensional vector to obtain the vector to be stored in the database includes: Obtaining a pre-configured discrete value set; wherein the discrete value set is used to store a preset number of data expression templates; Performing clustering processing on the high-dimensional vector according to the preset number to obtain the preset number of cluster centers; Obtaining each dimension value of the high-dimensional vector; Calculate the distance between each dimension value and each cluster center; For each dimension value, the data expression template corresponding to the cluster center closest to the dimension value is determined as the target template corresponding to the dimension value; Map each dimension value to a discrete value according to the target template corresponding to each dimension value; The discrete values are processed by principal component analysis to reduce the dimension, and the vector to be stored is obtained.
4. The high-dimensional data storage method based on hybrid index according to claim 1 is characterized in that: The HNSW index and IVFFLAT index of the vector to be stored are constructed based on the hybrid index strategy to obtain the hybrid index of the vector to be stored, including: Constructing the HNSW index and constructing the IVFFLAT index; Obtaining a mapping relationship between the vector to be stored in the HNSW index and the IVFFLAT index; According to the mapping relationship, pointers pointing to each other are added in the HNSW index and the IVFFLAT index, and a unified index access interface of the HNSW index and the IVFFLAT index is constructed to obtain a mixed index of the vector to be stored in the database.
5. The high-dimensional data storage method based on hybrid index according to claim 4 is characterized in that: The constructing of the HNSW index includes: Randomly select one or more vectors from the vectors to be stored as top-level nodes; Obtain the number and data attributes of the vectors to be placed in the library, and configure the number of layers and the number of nodes in each layer according to the number and data attributes of the vectors to be placed in the library; wherein each layer has a pyramid structure, and the layer closer to the top layer contains a lower number of nodes; Calculating the distance between each vector in the vector to be stored and the top node; In the order of the distance from near to far, each vector is sequentially assigned to the corresponding layer according to the number of layers and the number of nodes in each layer; For each node in each layer, calculate the distance between every two nodes and connect the two nodes with the closest distance in sequence; For each node between adjacent layers, calculate the distance between each two nodes and connect each two nodes with the closest distance in sequence; The currently obtained multi-layer index structure is obtained as the HNSW index.
6. The high-dimensional data storage method based on hybrid index according to claim 5 is characterized in that: The constructing of the IVFFLAT index comprises: Configuring the number of clusters according to the number of vectors to be stored and data attributes; Clustering the vectors to be stored according to the number of clusters to obtain the number of clusters; Create an inverted index entry for each cluster; wherein the inverted index entry of each cluster is used to record the index information of the vectors contained in each cluster; The vectors contained in each cluster are quantized to obtain the IVFFLAT index.
7. The high-dimensional data storage method based on hybrid indexing according to claim 1 is characterized in that: After obtaining the mixed index of the vector to be stored, the method further includes: When a change is detected in the vector to be stored, the change type is obtained; wherein the change type includes addition, deletion, and modification; The mixed index of the vector to be stored is adjusted in real time according to the change type.
8. A high-dimensional data storage device based on hybrid index, characterized in that: The high-dimensional data storage device based on hybrid index includes: a conversion unit, configured to, in response to a storage instruction for storing multimodal high-dimensional data in a centralized relational database, call a configuration interface to convert the multimodal high-dimensional data into a high-dimensional vector in a configuration format; A processing unit, configured to compress and reduce the dimension of the high-dimensional vector to obtain a vector to be stored; A construction unit, configured to construct an HNSW index and an IVFFLAT index of the vector to be stored based on a hybrid index strategy, to obtain a hybrid index of the vector to be stored; A storage unit is used to store the mixed index and the vector to be stored in the centralized relational database.
9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the high-dimensional data storage method based on hybrid indexing as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the high-dimensional data storage method based on hybrid indexing as described in any one of claims 1 to 7.
Citation Information
Cited By
Fast intelligent voice assistant response method based on approximate nearest neighbor retrieval algorithm
CN120808784A
ViT model and Faiiss database-based image searching method and system
CN120950723A
Lightweight index method, system and device for high-dimensional data of edge device and medium
CN121636762A
Lightweight indexing method, system, device and medium for high-dimensional data of edge device
CN121636762B