Unstructured data storage and calculation optimization system of distributed vector database

By converting unstructured data into vector data, calculating cosine similarity and performing distributed storage, extracting the theme of the set for search optimization, the problems of slow search speed and insufficient accuracy in unstructured data storage are solved, and efficient and accurate retrieval is achieved.

CN120277203AInactive Publication Date: 2025-07-08BEIJING XUNAO TECH

Patent Information

Application Number
CN202510734634.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing unstructured data vector storage technology has shortcomings in retrieval speed and accuracy, especially in intelligent AI Q&A. The traversal retrieval and approximate optimal solution methods lead to slow retrieval speed and insufficient accuracy.

Method used

Convert unstructured data into vector data, calculate the cosine similarity between vector data, divide the similar vector set, extract the theme for each set, perform distributed storage, search optimization based on the theme of the set, and use GPU acceleration and multi-node concurrent queries to improve computing performance.

Benefits of technology

The search time is greatly reduced, the retrieval efficiency and accuracy of unstructured data vector storage are improved, and the fast response needs of intelligent AI Q&A is met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277203A_ABST
    Figure CN120277203A_ABST
Patent Text Reader

Abstract

The invention discloses an unstructured data storage and calculation optimization system of a distributed vector database, which relates to the technical field of unstructured data vector storage, and comprises a data vectorization module, a vector data storage module and a retrieval calculation optimization module, the data vectorization module is used for converting unstructured data into vector data; the vector data storage module is used for dividing vector data into different similar vector sets, performing distributed storage on the similar vector sets, and extracting set subjects for the different similar vector sets; the retrieval calculation optimization module is used for retrieving vector data in a vector database on the basis of a set subject and performing performance optimization on a retrieval calculation process; the method is used for solving the problems that in an existing unstructured data vector storage technology, traversal retrieval and an approximate optimal solution mode are adopted for retrieval, so that the retrieval speed is low, and the retrieval precision is insufficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unstructured data vector storage, and specifically to an unstructured data storage and computing optimization system for a distributed vector database. Background Art

[0002] When the unstructured data vector storage technology processes unstructured data (such as text, images, audio, and video), it faces double challenges of data storage and computing performance. To achieve efficient unstructured data management and vector retrieval, it is necessary to optimize the storage architecture and computing model.

[0003] Existing unstructured data vector storage technologies are usually applied to intelligent AI question answering. Intelligent AI question answering encompasses a very large number of fields and knowledge. If the vector data in the database is traversed and retrieved, it will consume a large amount of retrieval time. However, intelligent AI question answering requires quick answers, so it is not applicable. Some unstructured data vector storage technologies improve the retrieval speed by retrieving approximate optimal solutions, but sacrifice accuracy. There are still quite a few gaps between the approximate optimal solution and the actual optimal solution. For example, in the Chinese patent with the application publication number: "CN119537532A", "A Quality Standard Retrieval Tool and Method Based on a Large Language Model" is disclosed. This solution directly performs similarity retrieval on the question vector data and each quality standard file knowledge base to retrieve the embedded vector data most similar to the question vector data. However, there is a large amount of data stored in the large prediction model, and the number of quality standard files is numerous. Traversing and analyzing during retrieval will cause a significant decrease in processing speed and require more time for retrieval. Existing unstructured data vector storage technologies also have problems of slow retrieval speed and insufficient retrieval accuracy due to using traversal retrieval and approximate optimal solution methods for retrieval. Summary of the Invention

[0004] The present invention aims to at least partly solve one of the technical problems in the prior art. By converting unstructured data into vector data, then calculating the cosine similarity between different vector data, dividing similar vector sets based on the cosine similarity, extracting the set theme for different similar vector sets, and performing distributed storage on the similar vector sets, then retrieving the vector data in the vector database based on the set theme, and optimizing the performance of the retrieval calculation process, to solve the problems of slow retrieval speed and insufficient retrieval accuracy existing in the existing unstructured data vector storage technology due to using traversal retrieval and approximate optimal solution methods for retrieval.

[0005] To achieve the above object, the present application provides an unstructured data storage and computing optimization system for a distributed vector database, including a data vectorization module, a vector data storage module, and a retrieval computing optimization module; the data vectorization module and the retrieval computing optimization module are respectively connected to the vector data storage module for data connection; The data vectorization module is used to convert unstructured data into vector data; The vector data storage module is used to divide vector data into different similar vector sets, and then perform distributed storage on the similar vector sets, and at the same time extract the set theme for different similar vector sets; The retrieval computing optimization module is used to retrieve vector data in the vector database based on the set theme and optimize the performance of the retrieval computing process.

[0006] Further, the data vectorization module is configured with a data vectorization strategy, and the data vectorization strategy includes: Obtain unstructured data, where the unstructured data includes text, images, audio, and video; Convert the unstructured data into vector data through a vector conversion model.

[0007] Further, the vector conversion model includes a Sentence-BERT model, a ResNet-50 model, a VGGish model, and a CLIP-ViT-B / 32 model, which are respectively used to process the vector conversion of text, images, audio, and video.

[0008] Further, the vector data storage module includes a similar set calculation unit, a theme extraction unit, and a distributed storage unit; The similar set calculation unit is used to calculate the cosine similarity between different vector data and divide similar vector sets based on the cosine similarity; The theme extraction unit is used to extract the set theme for different similar vector sets; The distributed storage unit is used to perform distributed storage on the similar vector sets.

[0009] Further, the similar set calculation unit is configured with a similar set calculation strategy, and the similar set calculation strategy includes: Number the vector data, represented by the symbol V d where d is a non-zero natural number and d is the serial number of V; Set the serial number D, where D is initially 1, starting from d = 2, and calculate the cosine similarity between V D and V d using the cosine similarity calculation method; In the cosine similarity calculation method, the V Dand V d Labeled as a and b respectively, a and b both represent a multi-dimensional vector; The formula for calculating the multidimensional space cosine function in the cosine similarity calculation method is: , where cos(θ) is the cosine similarity; Set a similarity threshold, and determine whether the pre-selected similarity is greater than or equal to the similarity threshold. If so, mark a and b as similar, and integrate a and b into the same similar vector set, and increase d by one and determine again; if not, mark a and b as dissimilar, increase d by one and determine again; repeat the process until d reaches the maximum value of d. When d reaches the maximum value of d, increase D by one to obtain a new D, and reset d to D+1. When resetting d, the reference D is the new D. Repeat the process until all vector data are divided into different similar vector sets; When D and d increase, it is necessary to judge V in real time. D or V d Does it already exist in the similar vector set? If V D Already exists in the similar vector set, then D and d are increased by one again. If V d If it already exists in the similar vector set, then d is increased by one again; For different unstructured data, similar vector sets are analyzed independently of each other and stored independently of each other.

[0010] Furthermore, the subject extraction unit is configured with a subject extraction strategy, and the subject extraction strategy includes: When analyzing any set of similar vectors, mark it as a pending set; Number any vector data in the collection to be mentioned, using the symbol G n Indicates, where n is a non-zero natural number and n is the sequence number of G; Set the sequence number N, N∈n, and calculate G N With G n The cosine similarity between them is denoted as H(N,n); By formula Calculate G N The subject matter is similar to that of N is the sum of similarity of the subject, max(n) is the maximum value of n; Find T N The maximum value is marked as the maximum similarity sum, and the vector data corresponding to the maximum similarity sum is marked as the subject of the collection.

[0011] Furthermore, the distributed storage unit is configured with a distributed storage strategy, and the distributed storage strategy includes: The set of similar vectors is stored in a vector database, which includes a text vector database, an image vector database, an audio vector database, and a video vector database; The text vector database, the image vector database, the audio vector database, and the video vector database are stored distributively.

[0012] Further, the retrieval calculation optimization module includes an optimization retrieval unit and a calculation optimization unit; The optimization retrieval unit is used to retrieve the vector data in the vector database based on the set theme; The calculation optimization unit is used to optimize the performance of the retrieval calculation process.

[0013] Further, the optimization retrieval unit is configured with an optimization retrieval strategy, which includes: When performing retrieval, obtain the input data of the retrieval, calculate the vector data of the input data, mark it as the input vector, and at the same time obtain the data type to be retrieved, named the retrieval type; Access the vector database of the retrieval type, marked as the target database; Calculate the cosine similarity between the input vector and the set theme of each similar vector set in the target database, named the retrieval similarity; Sort and number the retrieval similarities, and mark them as S i in descending order, where i is a non-zero natural number and i is the serial number of S; Set a preset number of sets, and name the top preset number of S i as the target set; Calculate the cosine similarity between the input vector and all vector data in the target set, find the maximum value among them, and mark the corresponding vector data as the target data; Output the unstructured data corresponding to the target data as the output value.

[0014] Further, the calculation optimization unit uses GPU acceleration and multi-node concurrent query to improve the calculation performance.

[0015] The beneficial effects of the present invention: By converting unstructured data into vector data, then calculating the cosine similarity between different vector data, dividing similar vector sets based on the cosine similarity, and storing the similar vector sets distributively, the advantage is that the vector data is initially divided, similar vector data is summarized into a set, and subsequent retrieval is performed in units of similar vector sets, which can greatly reduce the retrieval time and improve the efficiency of retrieval in the vector storage of unstructured data; The present invention extracts the set theme for different similar vector sets, then retrieves the vector data in the vector database based on the set theme, and optimizes the performance of the retrieval calculation process. The advantage is that the set theme reflects the general trend of the vector data in the similar vector set. When performing retrieval, the most similar vector set to the retrieved data is found through the set theme, and then searched from the vector data in the similar vector set, greatly reducing the retrieval time, improving the retrieval accuracy, and enhancing the efficiency and accuracy of retrieval in the storage of unstructured data vectors. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic block diagram of the principle of the system of the present invention; Figure 2 is a flowchart of the steps for converting unstructured data of the present invention into vector data; Figure 3 is a flowchart of the steps of the vector data storage module of the present invention; Figure 4 is a flowchart of the steps of the method of the present invention; Figure 5 is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0018] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0019] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0020] Embodiment 1. Refer to Figure 1 As shown, the present application provides an unstructured data storage and calculation optimization system for a distributed vector database, including a data vectorization module, a vector data storage module, and a retrieval calculation optimization module; the data vectorization module and the retrieval calculation optimization module are respectively connected to the vector data storage module for data connection; Please refer to Figure 2 As shown, the data vectorization module is used to convert unstructured data into vector data; the data vectorization module is configured with a data vectorization strategy, and the data vectorization strategy includes: Obtain unstructured data, where the unstructured data includes text, images, audio, and video; Convert unstructured data into vector data through a vector conversion model; The vector conversion model includes the Sentence - BERT model, ResNet - 50 model, VGGish model, and CLIP - ViT - B / 32 model, which are respectively used to process the vector conversion of text, images, audio, and video; In practical applications, the vector conversion model uses existing conversion algorithms. The processing processes of different unstructured data after vectorization are exactly the same. In this embodiment, only the processing process of text after vectorization by the Sentence - BERT model is taken as an example to illustrate the vector storage and retrieval of unstructured data. After the text is vectorized by the Sentence - BERT model, the obtained vector data is a 768 - dimensional vector, that is, a 1×768 vector, such as [0.314, - 0.258, 0.892, 0.157,..., - 0.117], which contains 768 numerical values in total, and the text needs to be converted into English before vectorization; The unstructured data vector storage technology is usually used in intelligent AI question - answering technologies, such as GPT. When training an intelligent AI, the training data needs to be vector - stored, and the efficiency and accuracy during retrieval need to be ensured.

[0021] Please refer to Figure 3 As shown, the vector data storage module is used to divide the vector data into different similar vector sets, then perform distributed storage on the similar vector sets, and at the same time extract the set theme for different similar vector sets; The vector data storage module includes a similar set calculation unit, a theme extraction unit, and a distributed storage unit; The similar set calculation unit is used to calculate the cosine similarity between different vector data and divide the similar vector sets based on the cosine similarity; The similar set calculation unit is configured with a similar set calculation strategy, and the similar set calculation strategy includes: Number the vector data, represented by the symbol V d where d is a non - zero natural number and d is the serial number of V; Set the serial number D, D is initially 1, starting from d = 2, calculate the cosine similarity between V D and V d by the cosine similarity calculation method; In the cosine similarity calculation method, the V D and V d used for calculation are respectively marked as a and b, and both a and b represent a multi - dimensional vector; The formula for calculating the cosine function in multi - dimensional space in the cosine similarity calculation method is , where cos(θ) is the cosine similarity; Set a similarity threshold, and determine whether the pre-selected similarity is greater than or equal to the similarity threshold. If so, mark a and b as similar, and integrate a and b into the same similar vector set, and increase d by one and determine again; if not, mark a and b as dissimilar, increase d by one and determine again; repeat the process until d reaches the maximum value of d. When d reaches the maximum value of d, increase D by one to obtain a new D, and reset d to D+1. When resetting d, the reference D is the new D. Repeat the process until all vector data are divided into different similar vector sets; When D and d increase, it is necessary to judge V in real time. D or V d Does it already exist in the similar vector set? If V D Already exists in the similar vector set, then D and d are increased by one again. If V d If it already exists in the similar vector set, then d is increased by one again; For different unstructured data, similar vector sets are analyzed independently and stored independently; In practical applications, vector data is numbered to obtain V d , 1≤d≤8976, set the serial number D, at this time D=1, start with d=2, at this time a is V1, b is V2, because the vector data is 768-dimensional data, it is inconvenient to show in the embodiment, so this embodiment only gives the final calculation result, in the formula for calculating the cosine function of multidimensional space, · is the dot product operator, || is the modulus operator, and the cosine similarity of V1 and V2 is 0.62 by calculation, and the calculation result is rounded to two decimal places; set the similarity threshold, which can be set by the developer. The smaller the similarity threshold, the better the training intelligence. The smaller the amount of calculation required for AI, the more time required for retrieval, and vice versa. In real life, people are accustomed to using 60% as the passing line to evaluate various types of data, and the range of cosine similarity is 0 to 1. Therefore, in this embodiment, the similarity threshold is set to 0.6. By comparing, the cosine similarity of V1 and V2 is greater than the similarity threshold, and V1 and V2 are marked as similar, and V1 and V2 are integrated into the same similarity vector set. At the same time, d is increased by one and judged again. At this time, it is necessary to judge whether V1 and V3 are similar. Repeat the execution until V1 and V2 are similar. 8976 After the judgment, D+1 is added, and D is now 2, but V2 is already in the same similar vector set as V1, so D is increased by one again, and D is now 3, and V3 does not exist in any similar vector set, and then d is reset to D+1, that is, d is reset to 3, and then analysis is performed again. After the analysis is completed, 1038 similar vector sets are obtained.

[0022] The subject extraction unit is used to extract the set subject for different similarity vector sets; The subject extraction unit is configured with a subject extraction strategy, which includes: When analyzing any set of similar vectors, mark it as a set to be extracted; Number any vector data in the set to be extracted, denoted by the symbol G n , where n is a non-zero natural number and n is the serial number of G; Set the serial number N, N ∈ n, and calculate the cosine similarity between G N and G n , marked as H(N, n); Calculate the main idea similarity sum of G through the formula N , where T N is the main idea similarity sum, and max(n) is the maximum value of n; Find the maximum value of T N , marked as the maximum similarity sum, and mark the vector data corresponding to the maximum similarity sum as the set main idea; In practical applications, taking a set of similar vectors as an example, mark it as a set to be extracted. There are 8 vector data in the set to be extracted, so numbering gives G n , 1 ≤ n ≤ 8. Set the serial number N, N ∈ [1, 8]. Taking N = 1 as an example, calculate the cosine similarity between G1 and G n , thus obtaining 8 data from H(1, 1) to H(1, 8). Each value of N corresponds to 8 H(N, n). Then calculate the main idea similarity sum of G N through the formula to obtain T N , obtaining a total of T1 to T8. Find that the maximum similarity sum is T3, and mark G3 as the set main idea. The set main idea can best represent the general trend of the vector data in the set of similar vectors.

[0023] The distributed storage unit is used for distributed storage of the set of similar vectors; The distributed storage unit is configured with a distributed storage strategy, and the distributed storage strategy includes: The set of similar vectors is stored in a vector database, and the vector database includes a text vector database, an image vector database, an audio vector database, and a video vector database; Perform distributed storage on the text vector database, the image vector database, the audio vector database, and the video vector database; In practical applications, use existing distributed storage technologies to store the text vector database, the image vector database, the audio vector database, and the video vector database.

[0024] The retrieval calculation optimization module is used to retrieve the vector data in the vector database based on the set main idea and optimize the performance of the retrieval calculation process; the retrieval calculation optimization module includes an optimized retrieval unit and a calculation optimization unit; The optimized retrieval unit is used to retrieve vector data in the vector database based on the set theme; The optimized retrieval unit is configured with an optimized retrieval strategy, and the optimized retrieval strategy includes: When performing retrieval, obtain the input data for retrieval, calculate the vector data of the input data, mark it as the input vector, and at the same time obtain the data type to be retrieved, named the retrieval type; Access the vector database of the retrieval type, marked as the target database; Calculate the cosine similarity between the input vector and the set theme of each similar vector set in the target database, named the retrieval similarity; Sort and number the retrieval similarities, and mark them as S in descending order i , where i is a non-zero natural number and i is the serial number of S; Set a preset number of sets, and set the top preset number of S i Named as the target set; Calculate the cosine similarity between the input vector and all vector data in the target set, find the maximum value among them, and mark the corresponding vector data as the target data; Output the unstructured data corresponding to the target data as the output value; In practical applications, when retrieving text, the text vector database is the target database. There are 1038 similar vector sets in the target database. Therefore, 1038 retrieval similarities are calculated and numbered as S in descending order i , 1 ≤ i ≤ 1038. The preset number of sets is set by the developer himself. In this embodiment, the preset number of sets is set to 5, that is, the top 5 S i The corresponding similar vector sets are named as the target sets, that is, S1 to S5. Then calculate the cosine similarity between the input vector and all vector data in S1 to S5, and find the vector data corresponding to the maximum value as the target data, and output the unstructured data corresponding to the target data as the output value.

[0025] The calculation optimization unit is used to optimize the performance of the retrieval calculation process; the calculation optimization unit uses GPU acceleration and multi-node concurrent query to improve the calculation performance.

[0026] Example 2, please refer to Figure 4 As shown, the present application provides a method for optimizing the storage and calculation of unstructured data in a distributed vector database, including the following steps: Step S1, convert unstructured data into vector data; Step S1 includes the following sub-steps: Step S101, obtain unstructured data, and the unstructured data includes text, image, audio, and video; Step S102, converting the unstructured data into vector data through a vector conversion model; Step S103, the vector conversion models include Sentence-BERT model, ResNet-50 model, VGGish model and CLIP-ViT-B / 32 model, which process the vector conversion of text, image, audio and video respectively; Step S2, dividing the vector data into different similar vector sets, and then distributing and storing the similar vector sets, and extracting set themes for the different similar vector sets; Step S2 includes the following sub-steps: Step S201, the similarity set calculation unit is used to calculate the cosine similarity between different vector data, and divide the similar vector set based on the cosine similarity; Step S201 includes the following sub-steps: Step S201.1, number the vector data, using the symbol V d Represents, where d is a non-zero natural number and d is the sequence number of V; Step S201.2, set the serial number D, D is initially 1, and starts with d=2, and calculate V by the cosine similarity calculation method D With V d The cosine similarity of Step S201.3, V calculated by the cosine similarity calculation method D and V d Labeled as a and b respectively, a and b both represent a multi-dimensional vector; Step S201.4, the formula for calculating the multidimensional space cosine function in the cosine similarity calculation method is: , where cos(θ) is the cosine similarity; Step S201.5, set a similarity threshold, determine whether the pre-selected similarity is greater than or equal to the similarity threshold, if so, mark a and b as similar, and integrate a and b into the same similarity vector set, and increase d by one and determine again; if not, mark a and b as dissimilar, increase d by one and determine again; repeat the process until d reaches the maximum value of d, when d reaches the maximum value of d, increase D by one to obtain a new D, and reset d to D+1, the reference D when resetting d is the new D, and repeat the process until all vector data are divided into different similarity vector sets; Step S201.6: When D and d increase, it is necessary to determine V in real time. D or V d Does it already exist in the similar vector set? If V D Already exists in the similar vector set, then D and d are increased by one again. If V d If it already exists in the similar vector set, then d is increased by one again; Step S201.7. For different unstructured data, the analysis of the similar vector sets is independent of each other, and the storage is also independent of each other. Step S202. The main idea extraction unit is used to extract the set main idea for different similar vector sets. Step S202 includes the following sub-steps: Step S202.1. When analyzing any similar vector set, mark it as the set to be extracted. Step S202.2. Number any vector data in the set to be extracted, denoted by the symbol G n where n is a non-zero natural number and n is the serial number of G. Step S202.3. Set the serial number N, N ∈ n, and calculate the cosine similarity between G N and G n , marked as H(N, n). Step S202.4. Calculate the main idea similarity sum of G through the formula N , where T N is the main idea similarity sum, and max(n) is the maximum value of n. Step S202.5. Find the maximum value of T N , marked as the maximum similarity sum, and mark the vector data corresponding to the maximum similarity sum as the set main idea. Step S203. The distributed storage unit is used to perform distributed storage on the similar vector sets. Step S203 includes the following sub-steps: Step S203.1. The similar vector sets are stored in the vector database, and the vector database includes a text vector database, an image vector database, an audio vector database, and a video vector database. Step S203.2. Perform distributed storage on the text vector database, the image vector database, the audio vector database, and the video vector database. Step S3. Retrieve the vector data in the vector database based on the set main idea, and optimize the performance of the retrieval calculation process; Step S3 includes the following sub-steps: Step S301. The optimization retrieval unit is used to retrieve the vector data in the vector database based on the set main idea. Step S301 includes the following sub-steps: Step S301.1. When performing retrieval, obtain the input data for retrieval, calculate the vector data of the input data, marked as the input vector, and at the same time obtain the data type to be retrieved, named the retrieval type. Step S301.2. Access the vector database of the retrieval type, marked as the target database. Step S301.3: Calculate the cosine similarity between the input vector and the set theme of each similar vector set in the target database, and name it the retrieval similarity; Step S301.4: Sort and number the retrieval similarities, and mark them as S in descending order i , where i is a non-zero natural number and i is the serial number of S; Step S301.5: Set a preset number of sets, and name the top S with the preset number of sets i as the target set; Step S301.6: Calculate the cosine similarity between the input vector and all vector data in the target set, find the maximum value among them, and mark the corresponding vector data as the target data; Step S301.7: Output the unstructured data corresponding to the target data as the output value; Step S302: The calculation optimization unit is used to optimize the performance of the retrieval calculation process; the calculation optimization unit uses GPU acceleration and multi-node concurrent query to improve the calculation performance.

[0027] Embodiment 3, please refer to Figure 5 as shown Figure 5 illustrates a schematic structural diagram of an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the method for storing and calculating the optimization of unstructured data in a distributed vector database are run to implement the following functions: convert unstructured data into vector data, then calculate the cosine similarity between different vector data, divide similar vector sets based on the cosine similarity, extract set themes for different similar vector sets, and perform distributed storage on the similar vector sets, and then retrieve the vector data in the vector database based on the set theme, and optimize the performance of the retrieval calculation process.

[0028] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0029] Embodiment 4, this application also provides a computer-readable storage medium. This application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, it runs the steps in the above-mentioned method for optimizing the storage and calculation of unstructured data in a distributed vector database to achieve the following functions: converting unstructured data into vector data, then calculating the cosine similarity between different vector data, dividing similar vector sets based on the cosine similarity, extracting the set main idea for different similar vector sets, and performing distributed storage on the similar vector sets, then retrieving the vector data in the vector database based on the set main idea, and optimizing the performance of the retrieval calculation process.

[0030] Through the description of the above embodiments, the embodiments of the present invention can be provided as a method, a system, or a computer program product. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0031] In the embodiments provided in the present application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division, and there may be other division methods in actual implementation. For another example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of systems, modules, and units can be in electrical, mechanical, or other forms.

[0032] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An unstructured data storage and computing optimization system for a distributed vector database, characterized in that It includes a data vectorization module, a vector data storage module and a retrieval calculation optimization module; the data vectorization module and the retrieval calculation optimization module are respectively connected to the vector data storage module; The data vectorization module is used to convert unstructured data into vector data; The vector data storage module is used to divide the vector data into different similar vector sets, and then perform distributed storage on the similar vector sets, and extract set themes for different similar vector sets; The retrieval calculation optimization module is used to retrieve the vector data in the vector database based on the collection subject and optimize the performance of the retrieval calculation process.

2. The unstructured data storage and computing optimization system of the distributed vector database according to claim 1, wherein The data vectorization module is configured with a data vectorization strategy, and the data vectorization strategy includes: Acquire unstructured data, wherein the unstructured data includes text, images, audio, and video; Convert unstructured data into vector data through vector transformation model.

3. The unstructured data storage and computing optimization system of the distributed vector database according to claim 2, characterized in that, The vector conversion models include Sentence-BERT model, ResNet-50 model, VGGish model and CLIP-ViT-B / 32 model, which process the vector conversion of text, image, audio and video respectively.

4. The unstructured data storage and computing optimization system of the distributed vector database according to claim 3, wherein The vector data storage module includes a similarity set calculation unit, a subject extraction unit and a distributed storage unit; The similarity set calculation unit is used to calculate the cosine similarity between different vector data, and divide the similar vector set based on the cosine similarity; The subject extraction unit is used to extract set subjects for different similarity vector sets; The distributed storage unit is used to perform distributed storage on a similar vector set.

5. The unstructured data storage and computing optimization system of the distributed vector database according to claim 4, characterized in that, The similarity set calculation unit is configured with a similarity set calculation strategy, and the similarity set calculation strategy includes: Number the vector data, denoted by the symbol V d where d is a non-zero natural number and d is the serial number of V; Set serial number D, where D is initially 1, starting from d = 2, and calculate V by the cosine similarity calculation method D with V d cosine similarity; Mark the V for calculation in the cosine similarity calculation method D and V d as a and b respectively, where both a and b represent a multi-dimensional vector; The formula for calculating the cosine function in multi-dimensional space in the cosine similarity calculation method is , where cos(θ) is the cosine similarity; Set a similarity threshold, and determine whether the pre-selected similarity is greater than or equal to the similarity threshold. If so, mark a and b as similar, and integrate a and b into the same similar vector set, and increase d by one and determine again; if not, mark a and b as dissimilar, increase d by one and determine again; repeat the process until d reaches the maximum value of d. When d reaches the maximum value of d, increase D by one to obtain a new D, and reset d to D+1. When resetting d, the reference D is the new D. Repeat the process until all vector data are divided into different similar vector sets; When D and d increase, it is necessary to judge V in real time D or V d already exists in the set of similar vectors. If V D already exists in the set of similar vectors, then increment D and d by one again. If V d already exists in the set of similar vectors, then increment d by one again; For different unstructured data, similar vector sets are analyzed independently of each other and stored independently of each other.

6. The unstructured data storage and computing optimization system of the distributed vector database according to claim 5, characterized in that The subject extraction unit is configured with a subject extraction strategy, and the subject extraction strategy includes: When analyzing any set of similar vectors, mark it as a pending set; Number any vector data in the set to be lifted, through the symbol G n is represented, where n is a non-zero natural number and n is the serial number of G; Set the serial number N, N ∈ n, and calculate G N The cosine similarity between G n is marked as H(N, n); Calculate G through the formula where the sum of the main idea similarities is calculated, and among them, T N is the sum of the main idea similarities, and max(n) is the maximum value of n; N ​ Find T N Find the maximum value, marked as the maximum similarity sum, and mark the vector data corresponding to the maximum similarity sum as the set theme.

7. The unstructured data storage and computing optimization system of the distributed vector database according to claim 6, characterized in that, The distributed storage unit is configured with a distributed storage strategy, and the distributed storage strategy includes: The similar vector set is stored in a vector database, wherein the vector database includes a text vector database, an image vector database, an audio vector database, and a video vector database; The text vector database, image vector database, audio vector database and video vector database are stored in a distributed manner.

8. The unstructured data storage and computing optimization system of the distributed vector database according to claim 7, characterized in that, The retrieval calculation optimization module includes an optimization retrieval unit and a calculation optimization unit; The optimization retrieval unit is used to retrieve vector data in the vector database based on the set theme; The calculation optimization unit is used to optimize the performance of the retrieval calculation process.

9. The unstructured data storage and computing optimization system of the distributed vector database according to claim 8, wherein The optimization retrieval unit is configured with an optimization retrieval strategy, and the optimization retrieval strategy includes: When performing retrieval, obtain the input data of the retrieval, calculate the vector data of the input data, mark it as the input vector, and at the same time obtain the data type to be retrieved, named the retrieval type; Access the vector database of the retrieval type, marked as the target database; Calculate the cosine similarity between the input vector and the set theme of each similar vector set in the target database, named the retrieval similarity; Sort and number the retrieval similarities, and mark them as S in descending order i , where i is a non-zero natural number and i is the serial number of S; Set a preset number of sets, and name the top S with the preset number of sets i as the target set; Calculate the cosine similarity between the input vector and all vector data in the target set, find the maximum value among them, and mark the corresponding vector data as the target data; Output the unstructured data corresponding to the target data as the output value.

10. The unstructured data storage and computing optimization system of the distributed vector database according to claim 9, characterized in that, The calculation optimization unit uses GPU acceleration and multi-node concurrent query to improve the calculation performance.

Citation Information

Patent Citations

  • Quality standard retrieval tool and method based on large language model

    CN119537532A

  • Retrieval method and device, electronic equipment, storage medium and program product

    CN114840666A

  • Vectorization data retrieval management method and device

    CN117453971A

  • RDF data storage and query method based on GPU acceleration

    CN117573664A

  • Search question and answer method based on large model

    CN117609444A

Cited By

  • Professional knowledge base retrieval optimization method based on AI deep semantic matching

    CN120994805A