Method, electronic device for mass data storage and fast retrieval

CN122862340APending Publication Date: 2026-10-02CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610919939.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

针对现有技术的不足,本发明提供了基于海量数据的存储和快速检索的方法,解决了现有技术中海量多模态数据处理过程中存在的并发写入导致索引阻塞、高维特征存储资源占用较高,以及冷热数据流转和冷区数据检索时产生的读写冲突与计算开销较大的技术问题

Benefits of technology

1、本发明引入日志结构合并树机制,将增量多模态数据顺序写入预写日志和活跃内存表,并在后台生成不可变磁盘文件的周期内同步构建增量索引片段。该机制将随机数据写入转换为顺序追加操作,缓解了海量数据并发写入时数据写入过程与多模态全量索引重算环节之间的锁竞争,降低了系统在高并发场景下的索引阻塞概率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122862340A_ABST
    Figure CN122862340A_ABST
Patent Text Reader

Abstract

This disclosure provides a method and electronic device for storing and rapidly retrieving massive amounts of data, relating to the field of data processing technology. The method involves writing multimodal data into a storage layer and constructing an index metadata model that includes text retrieval indexes, attribute retrieval indexes, and vector retrieval indexes. Based on access status information, a comprehensive popularity score for data objects is determined. When cold data migration conditions are met, the corresponding multimodal data is migrated to a cold data area. The vector retrieval index is compressed and encoded to generate a quantization encoding matrix, and the access route is updated. Multi-condition query instructions are received, and the access probability of cold data objects is determined. When prefetching conditions are met, the cold data objects and their quantization encoding matrices are loaded into an isolated circular prefetch buffer. Multi-condition query instructions are distributed to multiple index nodes. Distance calculation methods are selected according to the physical region where the data object is located. Local recall results are obtained and then globally merged, sorted, and re-scored before output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method and electronic device for storing and rapidly retrieving massive amounts of data. Background Technology

[0002] With the development of information technology, the amount of multimodal data generated continues to increase. To support subsequent data access, the system needs to extract features from received text, images, and other data and build corresponding index structures. When dealing with scenarios involving concurrent writing of massive amounts of data, traditional mechanisms typically trigger a complete index update after data is written. This approach leads to lock contention between the write operation and the index recalculation, causing system processing to become blocked.

[0003] Features extracted from multimodal data are typically represented as high-dimensional floating-point vectors, and storing them entirely in memory consumes significant physical storage resources. Over time, the access frequency of some data gradually decreases; without tiered scheduling and dimensionality reduction, this results in high hardware resource costs. Furthermore, during the transition to low-cost storage media, existing solutions are prone to concurrent read / write conflicts when modifying the underlying storage state, reducing the stability of data movement across physical locations.

[0004] During the retrieval phase, when a query request points to data stored in a cold region, physical reads across media incur significant time latency. Furthermore, traditional high-dimensional spatial distance calculations rely on floating-point multiply-accumulate operations. When faced with candidate datasets containing massive feature vectors, the lack of differentiated calculation strategies based on the physical region where the data resides results in high processor computational power consumption, increasing the system's computational resource overhead. Summary of the Invention

[0005] This disclosure provides a method and electronic device for storing and retrieving massive amounts of data. Addressing the shortcomings of existing technologies, this invention provides a method for storing and retrieving massive amounts of data, solving the technical problems in the processing of massive multimodal data, such as index blocking due to concurrent writes, high resource consumption for high-dimensional feature storage, and read / write conflicts and high computational overhead during the transfer of hot and cold data and retrieval of cold data.

[0006] According to a first aspect of this disclosure, a method for storing and quickly retrieving massive amounts of data is provided, comprising: The received multimodal data is written to the storage layer, and features are extracted from the multimodal data to construct an index metadata model containing text retrieval index, attribute retrieval index and vector retrieval index, forming a data object; The comprehensive popularity score is determined based on the access status information of the data object. When the comprehensive popularity score meets the cold data migration conditions, the multimodal data corresponding to the data object is migrated to the cold data area. The vector retrieval index is compressed and encoded to generate a quantization encoding matrix, and the access route of the data object is updated through the global metadata management node. Receive multi-condition query instructions, determine the access probability of cold zone data objects based on the multi-condition query instructions, and load the cold zone data objects and their associated quantization encoding matrices into the isolation prefetch buffer when the access probability meets the prefetch conditions. The multi-condition query command is distributed to multiple index nodes, so that each index node selects the corresponding distance calculation method based on the physical region where the data object is located to obtain local recall results. The local recall results are then globally merged, sorted, and re-scored for relevance, and the search response is output.

[0007] According to a second aspect of this disclosure, an apparatus for storing and rapidly retrieving massive amounts of data is provided, comprising: The extraction unit is used to write the received multimodal data into the storage layer and extract features from the multimodal data to construct an index metadata model containing text retrieval index, attribute retrieval index and vector retrieval index, forming a data object; The migration unit is used to determine the comprehensive popularity score based on the access status information of the data object. When the comprehensive popularity score meets the cold data migration conditions, the multimodal data corresponding to the data object is migrated to the cold data area. The vector retrieval index is compressed and encoded to generate a quantization encoding matrix, and the access route of the data object is updated through the global metadata management node. The loading unit is used to receive multi-condition query instructions, determine the access probability of cold zone data objects based on the multi-condition query instructions, and load the cold zone data objects and their associated quantization encoding matrices into the isolation prefetch buffer when the access probability meets the prefetch conditions. The output unit is used to distribute multi-condition query instructions to multiple index nodes, enabling each index node to select the corresponding distance calculation method based on the physical region where the data object is located to obtain local recall results, and to perform global merge sorting and relevance rescoring on the local recall results, and output the retrieval response.

[0008] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0009] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0010] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0011] The method and electronic device for storing and quickly retrieving massive amounts of data disclosed herein have the following beneficial effects: 1. This invention introduces a log structure merging tree mechanism, which sequentially writes incremental multimodal data into the write-ahead log and the active memory table, and synchronously builds incremental index fragments during the period of generating immutable disk files in the background. This mechanism transforms random data writing into sequential append operations, alleviating lock contention between the data writing process and the multimodal full index recalculation stage when massive data is written concurrently, and reducing the probability of index blocking in high-concurrency scenarios.

[0012] 2. This invention constructs a comprehensive popularity scoring mechanism to perform hot and cold tiered scheduling of data objects, and uses a product quantization mechanism to generate a quantization encoding matrix for high-precision floating-point vectors in the cold data area. Simultaneously, it employs a multi-version concurrency control mechanism to execute state transitions. This method converts high-precision floating-point numerical storage into low-bit-width indexed encoding storage, compressing the physical memory footprint of high-dimensional feature vectors, and achieving the transfer of data's physical storage location without triggering underlying concurrent read / write conflicts.

[0013] 3. This invention establishes a targeted prefetching mechanism based on a time-series prediction model. Cold zone data that meets access probability requirements is pre-loaded into an independent buffer, and asymmetric spatial distance coarse-sorting is performed on the quantization encoding matrix of the cold zone data during the distributed retrieval phase. This method replaces floating-point multiplication and addition operations in high-dimensional space by pre-extracting a query distance lookup table and performing cumulative summation, shortening the latency caused by long-tail queries reading across physical media and reducing computational resource overhead during cold zone data retrieval.

[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0015] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 A flowchart illustrating a method for storing and quickly retrieving massive amounts of data, provided in an embodiment of this disclosure; Figure 2 A schematic diagram of a system architecture provided for an embodiment of this disclosure; Figure 3 A comparison chart of write throughput under high-concurrency injection provided in this embodiment of the disclosure; Figure 4 This is a diagram illustrating the memory reduction and recall retention characteristics during a hot-to-cold transfer process, as provided in an embodiment of this disclosure. Figure 5 A comparison chart of P99 response latency under different prefetching strategies is provided for embodiments of this disclosure; Figure 6 This is a Pareto comparison chart of computational delay and accuracy under different retrieval operators provided in an embodiment of this disclosure.

[0016] Figure 7 A schematic diagram of the structure of a device for storing and quickly retrieving massive amounts of data, provided in an embodiment of this disclosure; Figure 8 A schematic block diagram of an example electronic device provided for embodiments of this disclosure. Detailed Implementation

[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0018] The embodiments disclosed herein, at every technical stage of the data lifecycle, including but not limited to data collection, transmission, storage, computation, use, disclosure, and destruction, are fundamentally based on strict adherence to and embedding of current laws, regulations, and regulatory requirements in their system architecture, protocols, and process controls. At the design level, the solution ensures, through systematic rules and strategies, that all processing activities automatically adhere to the principles of legality, legitimacy, necessity, and good faith, and technically implements core rules such as clear purpose, minimum necessity, transparency, and security.

[0019] For any data collection, processing, or other activities involved in the embodiments of this disclosure, corresponding verification, tracking, and constraint mechanisms are implemented at the system level to ensure that their execution has a clear legal basis or contractual foundation, and to automatically trigger and record the corresponding notification process. The processing purpose of related data is bound to its specific use at the metadata layer, and is strictly limited through the system's embedded flow strategy and access control model, thereby ensuring that data is accessed and used only within the scope necessary to achieve the initial collection purpose and as determined by technical criteria. The system has a multi-layered authorization management and compliance audit mechanism to ensure that related data will not be used for any other purpose without separate legal permission or valid separate consent from the information subject. This solution natively supports and protects the information subject's various legal rights to their data in its technical implementation, and provides standardized interfaces and automated processes to achieve efficient exercise of these rights.

[0020] The following describes, with reference to the accompanying drawings, a method and electronic device for storing and quickly retrieving massive amounts of data according to embodiments of the present disclosure.

[0021] Please see the appendix Figure 2 The present invention provides a system for storing and retrieving massive amounts of data. The system for executing this method mainly includes a storage layer, an index layer, and a retrieval engine layer.

[0022] The storage layer, index layer, and retrieval engine layer are physically and logically decoupled and isolated, and each layer communicates for data transmission and control signaling through network protocols.

[0023] The storage layer adopts a distributed object storage architecture for persistent storage of massive amounts of raw multimodal data.

[0024] The index layer consists of multiple distributed, multi-replica nodes, responsible for performing the construction, updating, and persistent storage operations of the metadata index.

[0025] The retrieval engine layer is deployed in a stateless node cluster and is responsible for receiving external query commands and scheduling the index layer to perform retrieval calculations.

[0026] Please see the appendix Figure 1 This invention provides a method for storing and quickly retrieving massive amounts of data, comprising the following steps: S10, Receive externally input multimodal data and write the multimodal data into the storage layer; S20, the index layer uses an incremental write mechanism to process multimodal data, writes incremental data to the memory table, and builds incremental index fragments during the flush cycle; S30, the index layer extracts attribute labels, text segmentation sequences and high-dimensional semantic feature vectors from the input multimodal data, and constructs an index metadata model that includes text retrieval index, attribute retrieval index and vector retrieval index; S40: The system background periodically performs lifecycle scheduling management and determines the comprehensive popularity score based on the access status information of data objects. S50, when the comprehensive popularity score meets the cold data migration conditions, triggers state transition and compression encoding processing; S60, during the execution process, migrates the multimodal data corresponding to the data object to the cold data area in the storage layer, and performs compression encoding on the vector retrieval index to generate a quantization encoding matrix; S70, after the cold data area file and quantization encoding matrix are persisted, the access route of the data object is updated through the global metadata management node, and the full-precision floating-point vector index resources occupied by the hot data area are released after the route update is completed; S80, the retrieval engine layer receives a multi-condition query instruction containing the feature vector and attribute conditions of the target to be queried, and determines the access probability of cold zone data objects; S90, when the access probability meets the prefetching condition, the retrieval engine layer triggers a directional prefetching instruction to load the corresponding cold zone data object and its quantization encoding matrix from the storage layer into the isolated circular prefetch buffer. S100, the retrieval engine layer distributes multi-condition query instructions to multiple index nodes in the index layer. Each node selects the corresponding distance calculation method based on the physical region where the data object is located to obtain local recall results, and then performs global merge sorting and relevance rescoring on the local recall results before outputting the retrieval response.

[0027] Multimodal, massive data streaming writes face high-concurrency data write demands. The index layer employs a log-structured merging tree mechanism to handle incremental data writing and index fragment construction, avoiding write blocking caused by rebuilding the global index. Specific incremental data persistence and index anti-blocking logic includes: The index layer receives the parsed incremental multimodal data, sequentially writes the incremental multimodal data into the write-ahead log to prevent data loss, and writes the incremental multimodal data into an active memory table located in memory. The active memory table uses a concurrent skip list data structure to sort and search data in memory. Regarding the specific disk write mechanism of the write-ahead log, those skilled in the art can choose synchronous or asynchronous disk flushing strategies according to actual availability requirements. The operating principles are well-known in the field and will not be elaborated here.

[0028] The monitoring thread within the index layer monitors the current capacity metrics of the active memory table in real time. When the current capacity metric reaches a preset capacity threshold, the monitoring thread changes the state of the active memory table to a frozen state to generate an immutable memory table, and initializes a new active memory table in memory to handle subsequent write requests. The preset capacity threshold specifically includes a memory byte size threshold and a data record count threshold. The memory byte size threshold is determined based on a preset percentage of the total physical memory allocated to the index layer by the running system, with this preset percentage ranging from 40% to 60%. The data record count threshold is calculated based on the expected retrieval response latency requirements and the average size of a single data record, with a value ranging from 500,000 to 2 million records.

[0029] The system wakes up a background asynchronous flushing thread to extract data from the immutable memory table and sequentially flush it to the persistent storage medium, generating a level-zero immutable disk file. The immutable disk file uses a key-value block sorting storage format. By configuring a separate background asynchronous flushing thread to perform input / output operations, the system achieves physical computational isolation between the data receiving thread and the disk write operations.

[0030] During the generation cycle of the zero-level immutable disk file, the background asynchronous flushing thread reads feature data from the immutable memory table and synchronously builds the corresponding incremental index fragments in memory. At the principle level, the log structure merging tree mechanism transforms random data writing into sequential data appending. During this merging cycle, the construction process of the incremental index fragment only utilizes the feature data currently residing in the immutable memory table. The system does not need to load or read existing historical disk data from the storage medium, thus avoiding resource contention caused by multimodal feature parsing and full index recalculation. The specific substructure of the incremental index fragment includes a local inverted index, a local attribute scalar index, and a local full-precision floating-point vector index for that batch of data.

[0031] After the incremental index fragment is constructed, the system physically binds the incremental index fragment to the corresponding level 0 immutable disk file and registers the addressing information of the incremental index fragment in the global metadata routing table. Under this incremental write mechanism, the query view of the global logical index at any given time is composed of the active index fragment in memory and the historical incremental index fragment on disk, and their logical relationship satisfies the following formula: ; In the formula, Represents a global logical index; This indicates the memory-state index fragment currently residing in the active memory table and the immutable memory table; Represents the logical union of indexed segments; Indicates the index number of an immutable disk file; This indicates the total number of immutable disk files in the current system; Indicates the number on the disk A disk-state incremental index fragment bound to an immutable disk file.

[0032] The above logic mechanism replaces the reconstruction and rearrangement of the global index tree with the appending of incremental index fragments. When the retrieval command is issued, it queries each independent incremental index fragment concurrently and merges the results at the retrieval engine layer, thereby decoupling the lock contention between massive data writing and multimodal index calculation and realizing the index anti-blocking function.

[0033] After receiving multimodal data, the system performs content parsing on the heterogeneous data and establishes logical relationships between different modalities. A multimodal feature parsing component is deployed in the index layer to perform differentiated feature extraction actions for different data formats.

[0034] The multimodal feature parsing component reads the system metadata and business-defined fields carried by the data object, extracts fixed-format information such as data creation time, data size, author identifier, and data source classification type, and formats it into standardized attribute tags.

[0035] The multimodal feature parsing component reads the text content from the data object, performs lexical segmentation using a natural language segmenter, filters stop words, extracts core entity words, and generates a text segmentation sequence. For the dictionary loading and part-of-speech tagging mechanism of the natural language segmenter, those skilled in the art can choose an open-source segmentation algorithm based on the actual business language; its text processing is well-known in the field and will not be elaborated upon here.

[0036] For keyframes of image data, audio and video data, and complex long text, the system invokes a deep learning coding model deployed on a graphics processing unit cluster to perform forward inference. The deep learning coding model maps unstructured data of different modalities to the same continuous numerical space, outputting corresponding high-dimensional semantic feature vectors. The specific data structure of the high-dimensional semantic feature vectors is a full-precision floating-point array containing multiple dimensions. This full-precision floating-point array is stored in 32-bit single-precision floating-point format, and its feature vector dimensions range from 128 to 2048 dimensions, with the specific feature vector dimensions determined according to the network structure of the deployed deep learning coding model. For the network layer design and tensor computation mechanism of the deep learning coding model, those skilled in the art can use a contrastive language image pre-trained model architecture; its model derivation and feature space mapping operations are well-known techniques in the field and will not be elaborated upon here.

[0037] After acquiring the three types of feature data mentioned above, the multimodal feature parsing component assigns a globally unique object identifier to the current data object. The object identifier is generated using a distributed snowflake algorithm to ensure that identifiers generated by different physical nodes do not collide. At the principle analysis level, multimodal retrieval scenarios involve a hybrid query requirement of structured conditional filtering and unstructured vector distance calculation. If various indexes are scattered across different database engines, data comparison and intersection operations between engines will incur network communication overhead. The index layer uses this object identifier as the primary key to associate heterogeneous feature data, constructing a unified index metadata model. The index metadata model integrates scattered features under the same logical computation view, and its structural relationship satisfies the following formula: ; In the formula, represents the index metadata model of the i-th data object; represents the identifier sequence number of the data object; represents the inverted index built based on the text segmentation sequence; represents the full-precision floating-point vector index built based on the high-dimensional semantic feature vector; and represents the attribute scalar index built based on the standardized attribute labels.

[0038] At the implementation level of sub-features, text retrieval indexes can use inverted indexes, which employ a trie structure with terms as keys and lists of object identifiers containing those terms as values. Attribute retrieval indexes can use attribute scalar indexes, which use a B+ tree structure to store numerical time tags with range retrieval requirements and a bitmap index structure to store discrete classification tags. Vector retrieval indexes can use full-precision floating-point vector indexes, which employ a hierarchical navigation small-world graph structure. By constructing long and short distance edges in a multi-layered connected graph, the connectivity of high-dimensional vectors in near-neighbor computation within space is ensured. This three-in-one fused index structure enables the retrieval engine layer to concurrently read the three types of index structures within the same physical node when processing cross-modal queries, avoiding network communication latency caused by cross-database calls.

[0039] In a tiered storage architecture, full-precision floating-point vector indexes need to reside in memory to ensure fast retrieval response times. To control memory usage, the system background executes a lifecycle scheduling management mechanism according to a preset time period. The preset time period ranges from 12 to 24 hours. This mechanism determines the physical location of the data object in the storage layer and the index structure in the index layer by assessing the access frequency of the data object.

[0040] A background scheduled task process within the index layer collects access logs for data objects. The logs include the object's creation timestamp and the time of each time the data object is recalled. For the technical implementation of distributed system log collection and aggregation, those skilled in the art can use open-source log collection components to construct data streams; the log aggregation principle is well-known in the field and will not be elaborated upon here.

[0041] Based on the collected access logs, the system constructs a comprehensive popularity scoring model for each data object. At the principle level, the access patterns of multimodal data conform to the principle of temporal locality, meaning that newly generated or recently frequently accessed data has a higher retrieval probability, while the access demand for data decreases over time. This model comprehensively considers historical access frequency, data time decay effect, and inherent business priority attributes. It combines historical frequency weight hyperparameters, recent activity weight hyperparameters, business prior weight hyperparameters, total number of visits, time period span, recent access frequency mapping value, and preset importance prior weights to perform a linear summation calculation. The system uses the following formula to calculate the comprehensive popularity score of the data object: ; In the formula, Indicates the first A comprehensive popularity score for each data object; This represents the historical frequency weight hyperparameter; Indicates the first The total number of times a data object has been retrieved in history; Indicates the first The time span from when a data object was first created; This indicates the hyperparameter representing the recent activity weight; This represents a recent access frequency mapping value calculated based on a time window; This represents the business prior weight hyperparameter; This indicates the pre-defined importance prior weight of the business domain to which the data belongs; The identifier number representing the data object.

[0042] The mechanism for implementing the sub-features of each calculation variable in the above formula is as follows: Time period span The quantification is achieved by subtracting the data object's creation timestamp from the current system timestamp. The formula adds 1 to the denominator to avoid division by zero errors. (Recent access frequency mapping value) Based on a preset sliding time window, the system counts the number of times the data object has been accessed within the past 7 to 30 days, and uses a max-min normalization algorithm to map this count to a numerical range of 0 to 1. During the data streaming write phase, the business system injects preset importance prior weights into the metadata. Preset importance prior weights. The value range is set according to the coreness of the business, and it ranges from 0.1 to 1.0. Three hyperparameters. , , The values ​​of all three are greater than 0 and less than 1, and satisfy the constraint that the sum of the three values ​​equals 1. The specific value ratios are determined by linear regression fitting based on the historical query load distribution during the initial stage of system operation.

[0043] After the index layer completes the comprehensive popularity score calculation for all data objects in the hot data area, it sorts them according to the numerical distribution of the comprehensive popularity scores. The system obtains a preset popularity threshold. The mechanism for determining the preset popularity threshold depends on the available physical memory of the current index layer. A monitoring process is deployed within the system to read the memory usage rate at the operating system level in real time. When the memory usage rate exceeds a preset safe usage ratio, which ranges from 75% to 85%, the system reads the comprehensive popularity scores of all data objects currently in the hot data area, sorts them in ascending order, and selects the score value in the bottom 20% as the current preset popularity threshold.

[0044] The system iterates through the calculation result set. When the overall popularity score of a data object falls below a preset popularity threshold, the system determines that the data object has transitioned from a hot to a cold state. Based on this, the system triggers a multi-version concurrency control mechanism for the data object to execute subsequent state transition and index dimensionality reduction logic.

[0045] After the system determines that a data object has undergone a hot-to-cold state transition, it triggers a multi-version concurrency control mechanism and executes the dimensionality reduction logic of the full-precision floating-point vector index.

[0046] Full-precision, high-dimensional floating-point feature vectors occupy a large amount of physical contiguous memory space. The product quantization mechanism divides the high-dimensional vector into multiple low-dimensional orthogonal subspaces, and within each orthogonal subspace, uses a finite number of cluster centers to approximate the original subvector. At the physical storage level, the system no longer stores the original 32-bit floating-point values, but instead stores integer index identifiers of the matched cluster centers, thus converting the storage of full-precision floating-point values ​​into low-bit-width indexed encoding storage, compressing the physical memory footprint.

[0047] The index layer extracts the full-precision floating-point feature vector of the data object to be dimensionality reduced. Let the feature vector dimension of this full-precision floating-point feature vector be . The system will adjust the feature vector dimension. Divided into equal parts There are *n* orthogonal subspaces, and the dimensions of the subvectors in each orthogonal subspace are respectively... Number of subspace partitions Based on feature vector dimension The spatial compression ratio is calculated and determined, and its value is an integer between 8 and 64, while ensuring the feature vector dimension. Number of subspaces that can be divided Divisible by.

[0048] For the segmented first For each orthogonal subspace, the system loads the corresponding pre-trained codebook. The pre-trained codebook contains multiple pre-calculated and allocated cluster centers. To adapt to the byte-aligned storage characteristics of the computer's underlying architecture, the total number of cluster centers in the pre-trained codebook corresponding to each orthogonal subspace is set to 256. Therefore, the cluster center index identifier for each orthogonal subspace only requires 8 bits, or 1 byte of storage space. The generation of the pre-trained codebook and the extraction mechanism of cluster center parameters can be calculated offline using the K-means clustering algorithm on massive amounts of business training samples by those skilled in the art. The model training and parameter convergence process is well-known in the field and will not be elaborated upon here.

[0049] The system in the Within each orthogonal subspace, the original subvector obtained from the segmentation is read, and the Euclidean distance between the original subvector and each cluster center in the pre-trained codebook is calculated. After completing the distance comparison, the system selects the cluster center with the smallest Euclidean distance value as the approximate replacement feature of the original subvector.

[0050] The system traverses all In each orthogonal subspace, the cluster center search and replacement operation is performed for each subvector. The indexes of all matched cluster centers are then concatenated to generate a low-bit-width quantization encoding matrix. This dimensionality reduction mapping relationship satisfies the following formula: ; In the formula, This represents the quantization encoding matrix generated after dimensionality reduction of the full-precision floating-point feature vectors. Indicates the first Full-precision floating-point feature vectors of each data object; The identifier number representing the data object; Denotes the closest vector to the original subvector in the first orthogonal subspace. One pre-trained cluster center; This represents the cluster center index identifier in the pre-trained codebook of the first orthogonal subspace; The second orthogonal subspace represents the closest vector to the original subvector. One pre-trained cluster center; This represents the cluster center index identifier in the pre-trained codebook of the second orthogonal subspace; Indicates the traversal sequence number of the data modality; Indicates the first The closest vector to the original subvector in the nth orthogonal subspace One pre-trained cluster center; Indicates the first Cluster center index identifiers in a pre-trained codebook of orthogonal subspaces.

[0051] After the quantization encoding matrix is ​​generated, the high-dimensional semantic features of the data objects are converted from a full-precision floating-point array into an integer array composed of the cluster center indices of each subspace. This quantization encoding matrix serves as the underlying data structure for coarse-ranking calculations in the cold data region and participates in subsequent asymmetric spatial distance coarse-ranking calculations.

[0052] During the dimensionality reduction transition of data physical location and index structure, the system faces read-write conflicts under high-concurrency queries. Directly modifying the underlying storage state will cause data read failures in retrieval requests. The system employs a multi-version concurrency control mechanism to handle concurrent reads and writes, and the transition process includes three operational stages: dual-write transition, atomic route switching, and asynchronous garbage collection.

[0053] During the dual-write transition phase, the system assigns a monotonically increasing new version number to the data object that triggers the hot-to-cold data transfer logic. The system maintains the old version number of this data object in an active and visible state in the global metadata management node. At the underlying physical media level, the hot data area of ​​the storage layer is constructed using high-performance solid-state drives, while the cold data area is constructed using large-capacity hard disk drives or low-cost object storage media. Query requests issued by the retrieval engine layer follow the routing addressing information corresponding to the old version number, continuously accessing the full-precision floating-point vector index residing in memory and the original data in the hot data area of ​​the storage layer. While the front-end query remains undisturbed, the system background initiates an asynchronous task process to copy the original multimodal data file to the cold data area within the storage layer and writes the dimensionality-reduced quantized encoding matrix into the index nodes of the cold data area of ​​the index layer.

[0054] After the file copies and quantization encoding matrix in the cold data area have completed disk persistence confirmation, the system enters the atomic routing switch phase. At the principle analysis level, the multi-version concurrency control mechanism relies on the global metadata state machine to record the physical location mapping of data. The routing table data structure inside the state machine includes the data object's identifier, current version number, and physical storage address. To avoid query thread blocking caused by traditional locking mechanisms, the system calls the underlying processor's compare and exchange instruction to perform an atomic commit operation on the global metadata state machine. This instruction ensures the indivisibility of the operation through hardware-level bus locking. This atomic routing switch logic satisfies the following state change formula: ; In the formula, Indicates the first A data object in The new routing status at any given moment; Indicates the time point at which the atomic operation is completed; Indicates the first A data object in Historical routing status at any given moment; Indicates the time point before the atomic operation is triggered; The identifier number representing the data object; This represents the function to compare and swap atomic operations; Indicates the first The old version number of the data object; Indicates the first The new version number after incrementing the data object; Indicates the first The new physical storage address corresponding to each data object in the cold data area.

[0055] By comparing and swapping atomic operation functions, the system redirects the access entry point of the data object to the cold data area, ensuring that the routing state has not been modified by other concurrent write threads. After this atomic operation completes, any new external query requests arriving are routed to the cold data area based on the new version number, in order to invoke the quantization encoding matrix to perform distance calculations.

[0056] After the atomic route switch is completed, the system enters the asynchronous garbage collection phase. The system does not immediately delete the original data resources in the hot data area. The global metadata management node maintains a read request reference counter and a preset safe time window for the old version number. The preset safe time window is set based on the maximum expected single query time of the system environment, and its value ranges from 500 milliseconds to 2000 milliseconds. When the retrieval engine layer initiates a read operation for the old version number, the read request reference counter is incremented; when the single read operation is completed and returns a result, the read request reference counter is decremented. When the read request reference counter reaches zero, and the current system time exceeds the preset safe time window, the system confirms that there are no pending read requests processing the data object within the engine. The system calls the operating system-level memory release interface to reclaim the full-precision floating-point vector index memory space occupied by the data object in the hot data area, and simultaneously deletes the original data file associated with the hot data area in the storage layer, completing the physical resource transfer of the data lifecycle.

[0057] In multimodal retrieval scenarios, some long-tail queries may point to data objects stored in the cold data area. Since the cold data area is constructed using high-capacity hard disk drives or low-cost object storage media, directly performing cross-media reads during the query phase will cause physical input / output latency. To shield the time consumed by physical device reads, the system adopts a confidence-based prefetching mechanism to load data with a high probability of being accessed into memory in advance.

[0058] The retrieval engine layer receives a multi-condition query instruction containing the feature vector of the target query and attribute conditions. The retrieval engine layer extracts the temporal query context from this multi-condition query instruction, specifically including the semantic features of the current query and the previous historical query sequence of the current user. Before executing the model's forward inference, the retrieval engine layer concatenates and aligns the semantic features of the current query and the previous historical query sequence to generate a temporal feature tensor that meets the model's input feature dimension requirements. The system inputs the temporal feature tensor into the temporal access prediction model. The temporal access prediction model employs a Long Short-Term Memory (LSTM) network architecture to capture long-term dependencies and potential intent shifts in continuous query behavior. For the internal hidden layer structure and gating update mechanism of the LTM network, those skilled in the art can deploy it using standard recurrent neural network code libraries; the principle of its model network forward inference is well-known in the field and will not be elaborated here.

[0059] The temporal access prediction model encodes the input temporal feature tensor and calculates the access confidence probability for data objects in the cold data region through the network output layer. This probability calculation process satisfies the following formula: ; In the formula, represents the access confidence probability for the i-th cold zone data object; represents the identifier number of the data object; represents the nonlinear activation function, specifically the Sigmoid activation function, used to map the linear calculation result to the probability value range of 0 to 1; represents the weight matrix of the output layer of the time-series access prediction model; represents the matrix dot product operation; represents the hidden layer state vector at the current query time obtained after encoding by the Long Short-Term Memory network, which is used to represent the current comprehensive access intent; and represents the bias vector of the output layer of the time-series access prediction model.

[0060] After obtaining the access confidence probability, the retrieval engine layer compares it with a preset probability threshold. The preset probability threshold is calculated based on the system's requirement to balance prefetch accuracy and hardware error tolerance, and its value ranges from 0.75 to 0.90. When the access confidence probability exceeds the preset probability threshold, the retrieval engine layer determines that the corresponding data object has a chance of being matched in the upcoming retrieval comparison, and then triggers a targeted prefetch instruction.

[0061] The targeted prefetch instruction drives the underlying data scheduling process, loading the corresponding cold data objects and their associated quantization encoding matrices from the storage layer into the isolated circular prefetch buffer within the retrieval engine layer. At the principle level, the prefetch mechanism of the isolated circular prefetch buffer transforms the cold data comparison operation, which originally relied on triggered disk reads, into a direct read operation from the memory cache, thereby eliminating the computation thread blocking phenomenon caused by cross-media reads. The isolated circular prefetch buffer allocates a fixed-length contiguous address space in physical memory and configures a head write pointer and a tail read pointer. The maximum capacity of this buffer is allocated proportionally based on the total available physical memory of the node where the retrieval engine layer resides, with a specific value ranging from 512 megabytes to 2048 megabytes.

[0062] When the total volume of the prefetched data objects reaches the maximum capacity of the buffer, the header write pointer will wrap back to the beginning of the contiguous address space, and the newly loaded prefetched data will directly overwrite the oldest historical data according to memory address order. By constructing this isolated circular prefetch buffer with circular overlay, the system can achieve seamless prefetching of cold data to reduce retrieval latency, while physically isolating the prefetch data cache from the resident main memory area of ​​the core retrieval engine, preventing large-scale erroneous prefetch operations from exhausting the system's main memory resources.

[0063] In the distributed storage and retrieval engine architecture, massive amounts of multimodal data are distributed across multiple physical nodes. The system employs a distribution and collection collaborative scheduling mechanism to handle multi-condition query requests across nodes concurrently. At the principle level, this mechanism distributes the global query pressure received from a single point to multiple physical worker nodes in the distributed cluster, leveraging the parallel computing resources of multiple nodes to reduce overall retrieval latency and address the computational bottleneck problem under massive data volumes.

[0064] The coordinating node within the retrieval engine layer receives multi-condition query commands containing textual, vector, and attribute conditions. The coordinating node reads the global metadata routing table to obtain the physical slice location of the target data set within the distributed cluster. At the implementation level of lower-level features, the system uses a consistent hashing algorithm to hash the identifier sequence number of the data object, mapping it to a virtual node on the hash ring, and determining the physical worker node to handle the query based on the mapping relationship. Based on the physical slice location, the coordinating node breaks down the multi-condition query command into multiple concurrent sub-query requests and distributes these sub-query requests to the corresponding physical worker nodes via a remote procedure call protocol.

[0065] After receiving a subquery request, each physical worker node performs multimodal feature retrieval locally. The physical worker nodes invoke differentiated computational logic based on the storage state of the data object. For data objects residing in the hot memory area, the physical worker nodes read the full-precision floating-point vector index, inverted index, and attribute scalar index to perform precise distance calculation and conditional filtering. For data objects residing in the cold data area, the physical worker nodes read the prefetched quantization encoding matrix in the isolated circular prefetch buffer to perform asymmetric spatial distance coarse-ranking calculation. To achieve adaptive coarse-fine calculation computational power allocation, the system sets a preset coarse-ranking truncation threshold for data objects in the cold data area. If the asymmetric spatial distance obtained from the coarse-ranking calculation is lower than this preset coarse-ranking truncation threshold, the physical worker node determines that the data object has high recall value and triggers the fine-ranking logic, reading the corresponding full-precision floating-point feature vector from the cold data area for precise distance calculation; if it does not exceed the threshold, it is directly discarded. The physical worker nodes integrate the local comparison results, generate a local candidate result set containing the data object's identifier sequence number and local similarity scores for each modality, and return it to the coordinating node.

[0066] The coordinating node receives the local candidate result sets returned by each physical worker node. Because the similarity calculation methods for different modalities differ—for example, text inverted index retrieval outputs term frequency inverse document frequency scores, while vector retrieval outputs cosine similarity distance—the two lack direct comparability in numerical dimensions. Before performing global ranking, the coordinating node performs global normalization on the local similarity scores across modalities. For the global normalization of scores across different modalities, those skilled in the art can use the max-min normalization algorithm to uniformly map the values ​​in different intervals to a standard interval of 0 to 1. This numerical mapping process is well-known in the field and will not be elaborated upon here.

[0067] After global normalization, the coordinating node performs weighted fusion calculations based on the importance of each modality in the current business scenario. It extracts the query weights of each modality issued by the front-end business system along with multi-condition query commands, performs matrix multiplication with the corresponding query weights, and sums the results to obtain the global fusion score for a single data object. This multimodal fusion scoring logic satisfies the following formula: ; In the formula, represents the global fusion total score of the i-th data object; represents the identifier sequence number of the data object; represents the summation operation; represents the total number of data modalities included in the query instruction; represents the traversal sequence number of the data modal; represents the preset query weight of the i-th modal, the value of which is issued by the front-end business system along with the multi-condition query instruction, and the sum of the query weights of each modal is equal to 1; represents the matrix multiplication operation; represents the global normalized local score of the i-th data object in the i-th modal.

[0068] Based on the calculated global fusion score, the coordinating node performs multi-way merge sorting on all collected candidate data objects. According to the descending order, the coordinating node extracts a preset recall number of data objects ranking at the top of the global fusion score. This preset recall number is determined based on the pagination display requirements of the front-end application or the truncation configuration of downstream services, and its value is an integer between 10 and 1000. The system generates a final multimodal retrieval response message from the extracted data objects and returns it to the query caller, completing the distributed collaborative retrieval task.

[0069] When performing retrieval on data objects residing in the cold data area, since the physical worker node reads the quantized encoding matrix, the system calls the asymmetric distance calculation mechanism to perform coarse-sorting calculations. At the principle level, the core of asymmetric distance calculation lies in replacing floating-point multiplication and addition operations in high-dimensional space with table lookup and accumulation operations. This reduces the computational time complexity from the product level of the feature vector dimension to the addition level of the number of orthogonal subspace partitions, accelerating local similarity evaluation while preserving some spatial representation accuracy.

[0070] After receiving a query command, the retrieval engine layer extracts the feature vector of the target object. Based on the aforementioned product quantization dimensionality reduction principle, the system proportionally divides the target feature vector into... The query subvector. For the nth... The system reads the corresponding pre-trained codebook from the orthogonal subspace and calculates the th . The system calculates the Euclidean distance between each query subvector and all cluster centers in the pre-trained codebook. The calculated distance values ​​are cached in the local memory of the physical worker node, generating a lookup table for the query distance corresponding to the current orthogonal subspace.

[0071] During coarse sorting, the physical worker node reads the quantization encoding matrix of the cold zone data objects. Based on the cluster center index identifiers of each orthogonal subspace recorded in the quantization encoding matrix, the system directly extracts the pre-calculated distance values ​​from the corresponding lookup table, and synchronously sums the distance values ​​of each subspace to output the asymmetric spatial distance. This underlying calculation mapping logic satisfies the following formula: ; In the formula, This indicates that the target feature vector to be queried is related to the first... The asymmetric spatial distance between the quantization encoding matrices of individual data objects; This represents the feature vector of the target to be queried; This represents the quantization encoding matrix generated after dimensionality reduction of the full-precision floating-point feature vectors. Indicates the first Full-precision floating-point feature vectors of each data object; The identifier number representing the data object; This represents the summation operation; Indicates the index of the orthogonal subspace; Indicates the traversal sequence number of the data modality; This represents the Euclidean distance extraction function; This indicates that the feature vector of the target to be queried is at the th position. The query subvector corresponding to each orthogonal subspace; Indicates the first The first pre-trained codebook in the orthogonal subspace Cluster centers; Indicates the first Cluster center index identifiers in a pre-trained codebook of orthogonal subspaces.

[0072] After calculating the asymmetric distance, the physical worker nodes sort the candidate results in the cold data area in ascending order based on the numerical value of the asymmetric spatial distance. The physical worker nodes extract the top-ranked data objects (within a preset coarse-ranking extraction count) as a coarse-ranking candidate set, and then perform fine-ranking flow judgment. This preset coarse-ranking extraction count is set based on the system's hardware memory capacity, and its value ranges from 10 to 50 times the pre-determined preset recall count. The system obtains a preset coarse-ranking truncation threshold. This preset coarse-ranking truncation threshold is set based on the business application's tolerance requirement for long-tail data recall rate, and its value ranges from 0.6 to 0.8.

[0073] The system iterates through and evaluates the results in the coarse-ranking candidate set. When the asymmetric spatial distance of a data object in the coarse-ranking candidate set is lower than the preset coarse-ranking truncation threshold, the physical worker node determines that the data object has passed the asymmetric distance coarse-ranking screening and then pushes the data object's identifier into its local fine-ranking processing queue. For data objects whose asymmetric spatial distance is greater than or equal to the preset coarse-ranking truncation threshold, the system removes them from the current retrieval context process.

[0074] The physical worker node, based on the identifier sequence number recorded in the fine-ranking processing queue, initiates a point-to-point lookup and load command to the underlying storage system, extracting the full-precision floating-point feature vector of the corresponding data object from the cold data area. The physical worker node invokes the hardware vector processing unit to perform precise Euclidean distance calculation using the extracted full-precision floating-point feature vector and the original target feature vector. This precise Euclidean distance is then reverse-mapped using the underlying exponential decay function to generate a precise similarity score. The system uses this precise similarity score as a local score for the vector modality, combining it with the data object's hit status in the text inverted index and attribute scalar index, and then summarizing and performing the aforementioned global normalization and weighted fusion calculation. By constructing a multi-stage re-ranking architecture that combines lookup-based cumulative coarse-ranking and full-precision vector fine-ranking, the system reduces the scale of physical data interaction and network transmission time for cold data media while maintaining the final retrieval accuracy of multimodal data.

[0075] Specific application examples: To verify the actual performance of the proposed hierarchical storage and fast retrieval system based on multimodal massive data, the applicant conducted multi-dimensional comparative experiments in a pre-defined distributed hardware cluster environment. The experimental environment consisted of a cluster of 10 physical nodes, with each node configured with a 128-core CPU, 512GB of memory, a 2TB NVMe SSD (for the hot data storage layer), and a 20TB HDD (for the cold data storage layer). The test dataset was a 1 billion-level multimodal dataset synthesized based on real-world business scenarios, containing high-dimensional semantic feature vectors (dimension D=768), text data, and structured attribute labels.

[0076] Incremental index anti-blocking write performance verification (corresponding appendix) Figure 3 ): Experimental objective: The validation index layer uses the Log Structure Merge Tree (LSM-Tree) mechanism to calculate the global logical index. At the same time, it has an anti-blocking effect on high-concurrency writes to the system.

[0077] Experimental Design: Comparing the traditional "Global Reconstruction Index Tree Algorithm" (Baseline) with the "LSM-tree Incremental Index Append Algorithm" (Ours) of this invention. In a continuous 120-minute streaming write test, peak concurrent write traffic was injected every 30 minutes.

[0078] Analysis of experimental results: like Figure 3As shown, the traditional algorithm (marked by light gray dashed squares in the figure) experiences a periodic precipitous drop in write throughput (QPS) when triggering index merging and rearrangement due to lock contention caused by multimodal feature parsing and full index recalculation, and exhibits continuous performance degradation in the later stages of testing. This invention, however, is based on the formula... By transforming random writes into incremental index fragment appends, the system does not need to load existing historical disk data throughout the entire process, as indicated by the solid black circle in the figure. The write QPS is stably maintained at over 25,000 times / second, completely eliminating the lock contention coupling between massive data writing and index calculation, and verifying the high availability of the mechanism of this invention.

[0079] Dynamic evaluation of data popularity and verification of memory compression performance (corresponding appendix) Figure 4 ): Experimental objective: Verification is based on comprehensive popularity score With Product Quantization (PQ) Dimensionality Reduction Formula The memory release effect.

[0080] Experimental Design: After system startup, all initial data is in a full-precision floating-point vector index state. Subsequently, lifecycle scheduling management is initiated, based on the formula... Identify cold data and map its vectors into low-bit-width quantization coding matrices.

[0081] Analysis of experimental results: like Figure 4 As shown in the dual Y-axis line graph, as the system ran for 1 to 15 days, the background scheduled task accurately identified long-tail cold data with decreasing access frequency. Under the protection of the Multi-Version Concurrency Control (MVCC) mechanism, a large number of 32-bit single-precision high-dimensional vectors were transformed into a quantized encoding matrix composed of 8-bit integer cluster center index identifiers. The solid triangle curve on the left axis indicates that the system's physical memory utilization rate smoothly decreased from an initial extremely high risk level (85%) and stabilized within a safe threshold (ultimately stabilizing at 28%). Meanwhile, the dashed diamond curve on the right axis shows that, because hot data remained resident in memory, the system's Top-100 recall rate (Recall@100) consistently remained at a minimally distorted 98.4%, achieving ultimate optimization of physical storage resource utilization.

[0082] Validation of the optimization of long-tail retrieval latency by the time-series confidence prefetching mechanism (corresponding appendix) Figure 5 ): Experimental objective: Verify the access confidence probability calculated based on the Long Short-Term Memory network and the shielding effect of the isolated ring prefetch strategy on physical I / O latency.

[0083] Experimental Design: For data queries in the cold data area (HDD medium), tests were conducted on "direct read without prefetching", "traditional LRU caching strategy", and the "confidence-isolated ring prefetching based on long short-term memory network" of the present invention.

[0084] Analysis of experimental results: like Figure 5 As shown in the bar chart, directly triggering cross-media reads causes the P99 latency for long-tail cold queries to spike to 1250ms. Traditional LRU strategies lack temporal intent awareness, resulting in low hit rates and latency still as high as 820ms. This invention, based on a formula, accurately predicts the user's subsequent access intent and loads the data into an isolated circular prefetch buffer. This mechanism transforms the original disk-triggered reads into direct reads from the memory cache, significantly reducing the P99 long-tail retrieval latency to 215ms and substantially improving the retrieval engine layer's resilience.

[0085] Verification of the efficiency of coarse-fine sorting collaborative retrieval using multimodal ADC (corresponding appendix) Figure 6 ): Experimental objective: Verify asymmetric spatial distance The coarse-sorting mechanism combined with the global fusion formula The ability to balance time consumption and accuracy in massive retrieval.

[0086] Experimental Design: When performing vector alignment in the cold data area, we compare the "full-precision brute-force exhaustive calculation (Flat L2)", "single-level coarse sorting calculation" and the multi-way merging strategy of "ADC lookup table accumulation coarse sorting + full-precision fine sorting and re-scoring" proposed in this invention.

[0087] Analysis of experimental results: like Figure 6 As shown in the scatter plot, although full-precision exhaustive search (light gray square dots) can achieve perfect mAP (0.985), the time taken for a single query is extremely long (850ms). Single-level coarse sorting (medium gray triangle points) is the fastest (45ms), but the accuracy loss (mAP drops to 0.720) cannot meet business requirements.

[0088] The physical working node of this invention is first based on the formula The multiplication-addition operation is reduced to a lookup table addition operation to quickly truncate irrelevant data. Then, full-precision ranking is performed on the very few candidate sets, and finally, the formula is combined... Multimodal fusion scoring was completed. This strategy (black pentagram) enabled the system to maintain a retrieval accuracy of 0.968 while taking only about 14% (120ms) of the time required for brute-force search.

[0089] The dashed line formed by connecting the three elements in the figure shows the Pareto Optimal Front, which intuitively proves that the strategy of this invention breaks through the performance bottleneck of traditional algorithms and achieves the best balance between computational latency and retrieval accuracy.

[0090] It should be noted that the embodiments of this disclosure may include multiple steps. For ease of description, these steps are numbered, but these numbers are not a limitation on the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of this disclosure do not limit this.

[0091] Corresponding to the aforementioned methods for storing and retrieving massive amounts of data, this disclosure also proposes an apparatus for storing and retrieving massive amounts of data. Since the apparatus embodiments of this disclosure correspond to the method embodiments described above, details not disclosed in the apparatus embodiments can be referred to the method embodiments described above, and will not be repeated here.

[0092] Figure 7 A schematic diagram of a device for storing and quickly retrieving massive amounts of data, provided in an embodiment of this disclosure, is shown below. Figure 7 As shown, it includes: Extraction unit 21 is used to write the received multimodal data into the storage layer and extract features from the multimodal data to construct an index metadata model containing text retrieval index, attribute retrieval index and vector retrieval index, forming a data object; Migration unit 22 is used to determine the comprehensive popularity score based on the access status information of the data object. When the comprehensive popularity score meets the cold data migration conditions, the multimodal data corresponding to the data object is migrated to the cold data area. The vector retrieval index is compressed and encoded to generate a quantization encoding matrix, and the access route of the data object is updated through the global metadata management node. The loading unit 23 is used to receive multi-condition query instructions, determine the access probability of cold zone data objects based on the multi-condition query instructions, and load the cold zone data objects and their associated quantization encoding matrices into the isolation prefetch buffer when the access probability meets the prefetch conditions. Output unit 24 is used to distribute multi-condition query instructions to multiple index nodes, so that each index node selects the corresponding distance calculation method based on the physical region where the data object is located to obtain local recall results, and performs global merge sorting and relevance rescoring on the local recall results, and outputs the retrieval response.

[0093] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and the principle is the same, so it is not limited in this embodiment.

[0094] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0095] Figure 8 A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0096] like Figure 8 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 302 or a computer program loaded from storage unit 308 into RAM (Random Access Memory) 303. The RAM 303 may also store various programs and data required for the operation of the electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An I / O (Input / Output) interface 305 is also connected to the bus 304.

[0097] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0098] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as methods for storing and retrieving large amounts of data. For example, in some embodiments, methods for storing and retrieving large amounts of data can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, computing unit 301 may be configured in any other suitable manner (e.g., by means of firmware) to perform the aforementioned method of storing and rapidly retrieving massive amounts of data.

[0099] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0100] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0103] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0104] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0105] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0106] The various numerical designations such as "first," "second," etc., used in this disclosure are merely for ease of description and are not intended to limit the scope of the embodiments of this disclosure, nor do they indicate a sequential order.

[0107] At least one of the features described in this disclosure can also be described as one or more, and multiple features can be two, three, four or more, and this disclosure does not impose any limitations. In the embodiments of this disclosure, for a technical feature, the technical features in that technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", etc., and there is no sequential order or size order among the technical features described by "first", "second", "third", "A", "B", "C" and "D".

[0108] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0109] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for storing and quickly retrieving massive amounts of data, characterized in that, include: The received multimodal data is written to the storage layer, and features are extracted from the multimodal data to construct an index metadata model containing text retrieval index, attribute retrieval index and vector retrieval index, forming a data object; Based on the access status information of the data object, a comprehensive popularity score is determined. When the comprehensive popularity score meets the cold data migration conditions, the multimodal data corresponding to the data object is migrated to the cold data area. The vector retrieval index is compressed and encoded to generate a quantization encoding matrix. The access route of the data object is updated through the global metadata management node. Receive a multi-condition query instruction, determine the access probability of a cold zone data object based on the multi-condition query instruction, and when the access probability meets the prefetch condition, load the cold zone data object and its associated quantization encoding matrix into the isolation prefetch buffer. The multi-condition query command is distributed to multiple index nodes, and each index node selects the corresponding distance calculation method based on the physical region where the data object is located to obtain local recall results. The local recall results are then globally merged, sorted, and re-scored for relevance, and a search response is output.

2. The method according to claim 1, characterized in that, The process of writing the received multimodal data into the storage layer and extracting features from the multimodal data to construct an index metadata model containing text retrieval indexes, attribute retrieval indexes, and vector retrieval indexes, forming a data object, includes: The parsed incremental multimodal data is sequentially written to the write-ahead log and then written to the active memory table in memory to form a memory-state dataset for receiving streaming write requests. Based on the capacity status of the active memory table, the active memory table is frozen to generate an immutable memory table, and a new active memory table is initialized to handle subsequent write requests. During the cycle of flushing the data in the immutable memory table to an immutable disk file, the feature data in the immutable memory table is read, and an incremental index fragment containing a local inverted index, a local attribute scalar index, and a local full-precision floating-point vector index is constructed. The constructed incremental index fragment is physically bound to the corresponding immutable disk file, and the addressing information of the incremental index fragment is registered in the global metadata routing table to obtain the index metadata model.

3. The method according to claim 1, characterized in that, The process of determining a comprehensive popularity score based on the access status information of the data object, migrating the multimodal data corresponding to the data object to the cold data area when the comprehensive popularity score meets the cold data migration conditions, compressing and encoding the vector retrieval index to generate a quantization encoding matrix, and updating the access route of the data object through the global metadata management node includes: The access status information is formed by obtaining the historical retrieval hit count, the time period span since the first creation, and the recent access frequency within the sliding time window of the data object based on the collected access logs. The recent access frequency is normalized, and the summation calculation is performed by combining the historical search hit count, the time period span, the business prior weight, and the corresponding weight coefficient to obtain the comprehensive popularity score. A heat threshold is determined based on the overall heat score ranking result of multiple data objects in the hot data area and the current resource occupancy status, and data objects whose overall heat score is lower than the heat threshold are identified as cold data migration objects; The multimodal data corresponding to the cold data migration object is copied to the cold data area, and the vector retrieval index corresponding to the cold data migration object is submitted to the compression encoding process to generate migration status information for the global metadata management node to update the access route.

4. The method according to claim 4, characterized in that, The step of compressing and encoding the vector retrieval index to generate a quantized encoding matrix includes: Extract full-precision floating-point feature vectors from the vector retrieval index corresponding to the cold data migration object, and divide the full-precision floating-point feature vectors into multiple orthogonal subspaces according to the dimension and spatial compression ratio of the full-precision floating-point feature vectors to obtain multiple original subvectors; For each orthogonal subspace, load the pre-trained codebook corresponding to the orthogonal subspace, compare the distance between the original subvector and the cluster center in the pre-trained codebook, and determine the cluster center index identifier that matches the original subvector. The cluster center index identifiers corresponding to all the orthogonal subspaces are concatenated in subspace order to generate the quantization encoding matrix used for coarse-rank retrieval in the cold data area.

5. The method according to claim 4, characterized in that, The step of updating the access route of the data object through the global metadata management node includes: When the data object is in a migration transition state, a new version number is generated for the data object, and the routing information corresponding to the old version number is kept visible in the global metadata management node; During the process of writing the multimodal data into the cold data area and writing the quantization encoding matrix into the cold data area index node, the arriving read request accesses the original data and full-precision floating-point vector index in the hot data area based on the old version number. After the file copy in the cold data area and the quantization encoding matrix have been persisted and confirmed, an atomic update is performed on the global metadata state machine to point the access entry of the data object to the cold data area. After confirming the end of the old read request based on the old version number's read request reference count and the safe time window, the full-precision floating-point vector index resources in the hot data area are released and the associated original data file is deleted.

6. The method according to claim 1, characterized in that, The process of receiving a multi-condition query instruction, determining the access probability of a cold zone data object based on the multi-condition query instruction, and loading the cold zone data object and its associated quantization encoding matrix into an isolation prefetch buffer when the access probability meets the prefetch conditions includes: Extract the current query semantic features and the previous historical query sequence of the current user from the multi-condition query instruction, and concatenate and align the current query semantic features and the historical query sequence into vectors to generate a temporal feature tensor that conforms to the model input dimension; The time-series feature tensor is input into the time-series access prediction model. The hidden state at the current query time output by the time-series access prediction model is used to represent the comprehensive access intention, and the access confidence probability of the cold zone data object is calculated. When the access confidence probability exceeds the probability threshold, a targeted prefetch instruction is triggered to write the cold zone data object and its associated quantization encoding matrix into the isolated circular prefetch buffer inside the retrieval engine layer.

7. The method according to claim 1, characterized in that, The step of distributing the multi-condition query instruction to multiple index nodes, so that each index node selects a corresponding distance calculation method based on the physical region where the data object is located to obtain a local recall result, includes: The coordinating node receives the multi-condition query instruction containing text conditions, vector conditions, and attribute conditions, and reads the global metadata routing table to determine the physical slice location where the target data set is located; Based on the physical slice location, the multi-condition query instruction is decomposed and encapsulated into multiple concurrent sub-query requests, and the concurrent sub-query requests are sent to the corresponding index nodes; For data objects residing in the hot data area, the index node reads the full-precision floating-point vector index, inverted index, and attribute scalar index to perform distance calculation and conditional filtering, and generates hot area candidate results; For data objects residing in the cold data area, the index node reads the prefetched quantization encoding matrix in the isolated circular prefetch buffer and performs asymmetric spatial distance coarse ranking calculation to generate a local candidate result set containing the identifier sequence number of the data object and the local similarity score of each modality as the local recall result.

8. The method according to claim 7, characterized in that, For data objects residing in the cold data area, the index node reads the prefetched quantization encoding matrix from the isolated circular prefetch buffer, performs asymmetric spatial distance coarse ranking calculation, and generates a local candidate result set containing the data object's identifier and local similarity scores for each modality as the local recall result, including: Extract the feature vector of the target to be queried, and divide the feature vector of the target to be queried into multiple sub-vectors to be queried according to the orthogonal subspace segmentation method corresponding to the quantization encoding matrix; Calculate the distance between each subvector to be queried and the cluster center in the corresponding pre-trained codebook, and cache the calculated distance values ​​in the local memory of the index node to generate a query distance lookup table for the corresponding orthogonal subspace; Read the cluster center index identifiers of each orthogonal subspace recorded in the quantization encoding matrix, extract the distance values ​​from the corresponding query distance lookup table based on the cluster center index identifiers, and accumulate the multiple distance values ​​to output the asymmetric spatial distance. Based on the asymmetric spatial distance, a coarse-rank candidate set is determined, and a full-precision floating-point feature vector is loaded onto the data objects that meet the fine-ranking conditions. The distance between the full-precision floating-point feature vector and the feature vector of the target to be queried is calculated to obtain the local similarity score of the vector modality.

9. The method according to claim 1, characterized in that, The step of performing global merge sorting and relevance rescoring on the local recall results and outputting the retrieval response includes: The coordination node receives the local candidate result set returned by each index node, and extracts the identifier number of the data object and the local similarity score of each modality from the local candidate result set; The local similarity scores across modalities are globally normalized to obtain globally normalized local scores under the same scoring scale. Based on the modal query weights carried by the multi-condition query instruction, the global normalized local scores are weighted and fused to obtain the global fusion total score for each data object; Based on the global fusion total score, perform multi-way merge sort on the candidate data objects, extract the data objects that meet the recall requirements, generate a multimodal retrieval response message, and use it as the retrieval response.

10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.